NLP · Deep Learning · Text Classification

Twitter Hate Speech
Detection

An end-to-end NLP pipeline that classifies tweets as hate speech or not โ€” using Bidirectional LSTMs, GloVe word embeddings, and ~49,000 labeled examples.

49K Labeled Tweets
99.8% Training Accuracy
50d GloVe Embedding Dim
3ร— BiLSTM Layers

A Complete NLP Workflow

From raw tweets to production-ready predictions โ€” every step of the pipeline is covered.

๐Ÿงน

Text Preprocessing

Lowercasing, mention removal, stopword filtering, and lemmatization to produce clean, normalized tweet text.

๐Ÿ”ข

Tokenization & Padding

Keras Tokenizer builds a vocabulary index; sequences are zero-padded to a uniform length for batch training.

๐Ÿง 

GloVe Embeddings

Pre-trained Stanford GloVe vectors (50d, 6B tokens) initialize the embedding layer with real-world semantics.

๐Ÿ”„

Bidirectional LSTM

Three stacked BiLSTM layers capture forward and backward context, enriched with dropout and batch normalization.

๐Ÿ“Š

Evaluation

Precision-recall curves identify the optimal decision threshold, with a full classification report per class.

โš–๏ธ

Class Imbalance Handling

Stratified splits and threshold optimization ensure the minority hate-speech class is correctly captured.

Built With

๐ŸŸ 
TensorFlow
Deep Learning
๐Ÿ
Python 3.12
Language
๐Ÿ“–
NLTK
NLP Toolkit
๐Ÿ”ต
GloVe
Embeddings
๐Ÿผ
Pandas
Data Wrangling
๐Ÿ“‰
scikit-learn
ML Utilities

Dive Deeper

Walk through each stage of the pipeline or jump straight to the model results.