Results

Training performance across 12 epochs and evaluation on the held-out test set.

Accuracy Over 12 Epochs

The model converges rapidly, reaching 99.79% training accuracy by the final epoch.

Epoch Training Accuracy Loss Progress
1 92.50% —
3 ~96% —
6 ~98% —
9 ~99.3% —
12 99.79% 0.0060

Model Performance Summary

🎯

Final Training Accuracy

99.79%
after 12 training epochs with Adam optimizer.

📉

Final Training Loss

0.006
binary cross-entropy, indicating tight convergence.

⚖️

Threshold Optimization

Precision-recall analysis identifies the optimal decision boundary for the imbalanced hate-speech class beyond the default 0.5.

Why This Architecture

ComponentChoiceRationale
Word Embeddings GloVe 50d Captures semantic similarity without training from scratch on a small corpus.
Sequence Model BiLSTM Reads tweets both left-to-right and right-to-left, capturing full context.
Regularization Dropout + BatchNorm Reduces overfitting on the small (~25K training) dataset.
Threshold PR-optimized Class imbalance (7% hate speech) makes default 0.5 sub-optimal.
Loss Function Binary Cross-Entropy Standard for binary classification with sigmoid output.
View Full Notebook on GitHub ← Back to Pipeline