The Pipeline
Every stage from raw tweet text to a trained classifier, explained.
Text Preprocessing
Raw tweets are noisy — full of @mentions, emoji, URLs, and non-standard characters. Each tweet is lowercased, stripped of non-ASCII characters, scrubbed of user mentions, filtered for English stopwords, and lemmatized to its base form.
def preprocess(tweet):
tweet = tweet.lower()
tweet = re.sub(r'[^\x00-\x7F]+', '', tweet)
tweet = re.sub(r'@\w+', '', tweet)
tokens = tweet.split()
tokens = [w for w in tokens if w not in stop_words]
return ' '.join(tokens)
Tokenization & Sequence Padding
A Keras Tokenizer fits on the training corpus to build a vocabulary index. Each cleaned tweet is then encoded as a sequence of integers. Sequences are zero-padded to a uniform length so batches can be formed as tensors.
GloVe Embedding Matrix
Stanford GloVe embeddings (glove.6B.50d, trained on 6 billion tokens) are loaded into memory. An embedding matrix is constructed by mapping every word in the tokenizer vocabulary to its pre-trained 50-dimensional vector. Words not found in GloVe are initialized to zero.
embeddings = {}
with open('glove.6B.50d.txt') as f:
for line in f:
word, *vec = line.split()
embeddings[word] = np.array(vec, dtype='float32')
embedding_matrix = np.zeros((vocab_size, 50))
for word, idx in tokenizer.word_index.items():
vec = embeddings.get(word)
if vec is not None:
embedding_matrix[idx] = vec
Model Architecture
The network starts with the frozen GloVe embedding layer, then passes through three stacked Bidirectional LSTM layers. Each LSTM is followed by Dropout (0.2) and Batch Normalization. Finally, two Dense layers reduce to a single sigmoid output for binary classification.
Training
The model trains for 12 epochs with the Adam optimizer and binary cross-entropy loss. An 80/20 stratified split preserves the class ratio in both sets, ensuring the minority hate-speech class is equally represented during training and evaluation.
Threshold Optimization & Evaluation
Rather than using a fixed 0.5 threshold, a precision-recall curve is computed to find the threshold that best balances precision and recall for the hate-speech class. Final metrics are reported via a full classification report.