When moderation labels disagree — Technical — Bachelor’s https://thesis.uya.no/proposals/when-moderation-labels-disagree/?degree=bachelor#technical-track FINAL RESULT Working dashboard, reproducible baseline, test report and prioritised fixes. WHAT TO DO Build a multilingual moderation dashboard with a TF–IDF classifier, confidence warning and manual-review queue. DATA Fixed 3,000-record COUNTER subset (public) Select 1,000 records per language with a fixed seed, preserving label proportions and source groups. Keep a separate 20% test split. COUNTER public dataset: https://gitlab.inria.fr/ariabi/counter-dataset-public HOW TO TEST IT Use a fixed stratified subset of 3,000 COUNTER records. Compare the classifier with a majority-class baseline; report macro-F1, per-language F1, confusion matrices and calibration. Test the dashboard with 4 trained reviewers on 24 fixed cases. STEPS Build and test 1. Read the starting sources and choose one established implementation method. 2. Write the requirements, data fields, system diagram and test cases. 3. Prepare Fixed 3,000-record COUNTER subset. Make the answer key and pass criteria before testing. 4. Build a working version of working dashboard, reproducible baseline, test report and prioritised fixes. 5. Run function, integration and failure-case tests. Record each result. 6. Run the practical evaluation and list the changes the system still needs. EXAMPLE TOOLS Python, pandas, scikit-learn, Hugging Face Transformers, Streamlit, MLflow or DVC Equivalent tools are fine. ACCESS OR PEOPLE Confirm dataset terms and ethics handling; recruit 4 trained reviewers.