Twitter Hate Speech
Detection
An end-to-end NLP pipeline that classifies tweets as hate speech or not โ using Bidirectional LSTMs, GloVe word embeddings, and ~49,000 labeled examples.
A Complete NLP Workflow
From raw tweets to production-ready predictions โ every step of the pipeline is covered.
Text Preprocessing
Lowercasing, mention removal, stopword filtering, and lemmatization to produce clean, normalized tweet text.
Tokenization & Padding
Keras Tokenizer builds a vocabulary index; sequences are zero-padded to a uniform length for batch training.
GloVe Embeddings
Pre-trained Stanford GloVe vectors (50d, 6B tokens) initialize the embedding layer with real-world semantics.
Bidirectional LSTM
Three stacked BiLSTM layers capture forward and backward context, enriched with dropout and batch normalization.
Evaluation
Precision-recall curves identify the optimal decision threshold, with a full classification report per class.
Class Imbalance Handling
Stratified splits and threshold optimization ensure the minority hate-speech class is correctly captured.
Built With
Dive Deeper
Walk through each stage of the pipeline or jump straight to the model results.