Projekt

Mohamed Abdelaal, Mohamed Abokahf - Data Science Semesterprojekt

Hate Speech Recognition

  • © Abdelaal Abokaf

Today, the ease with which individuals can disseminate hateful content, whether through direct insults or subtle undertones, poses a significant societal challenge. Such hate speech, often veiled in ambiguous language, can lead to severe consequences, including instances of bullying that may ultimately result in tragic outcomes, such as suicide. In response to this pressing issue, we have embarked on a project aimed at developing a robust system for detecting tweets containing potential hate speech. This system seeks to automate the filtering process, mitigating the need for manual oversight.

Our comprehensive project is structured into three main components. The initial phase involves meticulous data collection, followed by the training of diverse machine learning and deep learning models in the second phase. The final phase encompasses the evaluation of these models and the creation of a user-friendly graphical interface to present the results effectively.

For data acquisition, we leveraged an open dataset from Kaggle, featuring two primary columns: one for the tweet text and another for classification (1 for hate speech, 0 for non- hate speech). Our model ensemble includes decision trees, support vector machines, naive Bayes, random forests, logistic regression, and AdaBoost. Additionally, we fine-tuned a pretrained model, "Roberta," to enhance its ability to classify tweets accurately.

Betreuer/in
Profilfoto von Prof. Dr. Anika Groß