Audit what a feed amplifies P2 · Spring 2027 https://thesis.uya.no/proposals/audit-what-a-feed-amplifies/ Measure how different ranking rules change what users see in a controlled social-media feed. TECHNICAL IMPLEMENTATION - Implement five deterministic ranking policies over one versioned item/profile contract with replayable sessions and no live-platform collection. - Test ranking invariants, stable ties, exposure accounting and audit-trace completeness across fixed seeds and counterfactual policy runs. - Measure risk exposure, concentration, diversity, benign relevance, calibration and latency at profile level; keep reviewer task effects separate from ranking metrics. TECHNICAL FINAL RESULT Deterministic feed simulator, five rankers, audit dashboard, benchmark and policy trade-off report. WHAT TO DO - Implement five deterministic ranking policies over one versioned item/profile contract with replayable sessions and no live-platform collection. - Test ranking invariants, stable ties, exposure accounting and audit-trace completeness across fixed seeds and counterfactual policy runs. - Measure risk exposure, concentration, diversity, benign relevance, calibration and latency at profile level; keep reviewer task effects separate from ranking metrics. DATA (planned) 1,200-item controlled feed benchmark Create 1,200 synthetic, non-extremist items across 12 proxy topics and 60 sources. Store item_id, source_id, topic, created_at, relevance, predicted_engagement, quality, risk_proxy and annotator confidence. Create 200 fixed user profiles with topic preferences and exposure history. No live platform accounts or scraped extremist content. HOW TO TEST IT Run every ranker on the same 200 profiles and 20 sessions with fixed seeds. Report risky-exposure@10, source/topic concentration, intra-list diversity, NDCG@10 for benign relevance, calibration and latency. Compare policies with paired profile-level intervals. In a task study, test whether the audit view helps 8 reviewers identify the rule responsible for a harmful exposure pattern. BACHELOR Final result: Working simulator, three rankers, automated tests and comparison dashboard. Task: Build a replayable feed and compare chronological, engagement-only and diversity-constrained ranking. Data: 400-item controlled feed benchmark (planned) Create neutral proxy content across 8 topics and 20 sources, plus 50 profiles. finalise all labels and generator settings before comparison. Requires: Independent review of the synthetic labels and scenarios. Method: Use 400 synthetic items and 50 fixed profiles. Run 10 sessions per profile with fixed seeds; report source concentration, topic diversity, NDCG@10 and invariant-test results. STEPS Build and test 1. Read the starting sources and choose one established implementation method. 2. Write the requirements, data fields, system diagram and test cases. 3. Prepare 400-item controlled feed benchmark. Make the answer key and pass criteria before testing. 4. Build a working version of working simulator, three rankers, automated tests and comparison dashboard. 5. Run function, integration and failure-case tests. Record each result. 6. Run the practical evaluation and list the changes the system still needs. MASTER Final result: Deterministic feed simulator, five rankers, audit dashboard, benchmark and policy trade-off report. Research question: Design and evaluate a ranking policy that makes exposure risk auditable while preserving relevance and viewpoint/source diversity. Research result: A reproducible offline audit method and evidence about the trade-offs between engagement, relevance, diversity and risky exposure. STEPS Design, build and test (DSR) 1. Read the newest papers and list the closest existing systems. 2. Write down the versions, fields, data split, case assignment, random seeds and correct answers for 1,200-item controlled feed benchmark. 3. Draw the user workflow, data model and system architecture. List the requirements and pass criteria. 4. Build a working version of deterministic feed simulator, five rankers, audit dashboard, benchmark and policy trade-off report. 5. Test every function, connection and failure case. Save the failed tests as well as the passed tests. 6. Compare the system with the named alternative. Then run the user task or decision task in the assignment. 7. Report the measured result, the failed cases and the design lessons another team can reuse. CURRENT PROJECT LITERATURE Causally estimating the effect of YouTube’s recommender system using counterfactual bots (2024, peer-reviewed journal article): https://pmc.ncbi.nlm.nih.gov/articles/PMC10895271/ Separates recommender effects from user preferences and shows why audit claims need a counterfactual design. 8–10% of algorithmic recommendations are bad, but… an exploratory risk-utility meta-analysis (2024, peer-reviewed journal article): https://doi.org/10.1016/j.ijinfomgt.2023.102743 Frames recommendation design as a measurable risk–utility trade-off rather than a one-sided accuracy problem. IS THEORY STARTING POINTS Haroon et al. (2024) — Causally estimating the effect of YouTube’s recommender system: https://pmc.ncbi.nlm.nih.gov/articles/PMC10895271/ Ground causal claims and distinguish user choice from recommender effects. Hoffmann et al. (2024) — 8–10% of algorithmic recommendations are bad, but…: https://doi.org/10.1016/j.ijinfomgt.2023.102743 Define measurable risk–utility trade-offs instead of assuming all recommendations are harmful. Search Scopus or Web of Science and ACM Digital Library using the topic query, then follow citations to the thesis start date. Record searches and compare methods, data, findings and limitations in literature-matrix.csv. Use that review to confirm or revise the gap and choose a current comparator. The linked papers are starting points. START WITH THESE THREE ACTIONS 1. Write the schema and generator distributions; independently review 30 items and 10 profiles before generating the fixed benchmark. 2. Implement chronological and engagement-only ranking first; add deterministic replay and invariant tests. 3. finalise measures, seeds and policy parameters before running the remaining rankers and reviewer study. EXAMPLE TOOLS Python, pandas, NumPy, scikit-learn, FastAPI, Streamlit or React Equivalent tools are fine. LITERATURE SEARCH recommender system algorithmic amplification extremism audit exposure diversity risk utility ACCESS OR PEOPLE Independent review of the synthetic labels and scenarios; recruit 8 reviewers for the master task study. -------------------- NON-TECHNICAL FINAL RESULT A source table for each platform, a platform-by-platform comparison and a disclosure template for recommender audits. WHAT TO DO - Collect the current recommender descriptions, DSA risk reports and audit summaries from six large platforms. - Code each document for stated objective, user control, risk measure, youth protection, evidence and external scrutiny. - Compare the documents with 8–10 professional interviews about what evidence is actually useful for oversight. DATA (public) 24 public platform and DSA documents + 8–10 professional interviews For six named platforms, collect one recommender description, one terms/policy document, one DSA risk or transparency report and one independent audit/regulator document dated 2024–2027. finalise URLs and PDFs before coding. Interview researchers, prevention professionals or regulators—not young people or platform users. European Commission DSA transparency overview: https://digital-strategy.ec.europa.eu/en/policies/dsa-brings-transparency HOW TO TEST IT Use a predefined document-coding matrix and double-code 20% of documents. Compare stated objectives, risk measures, user controls and evidence across platforms. Use the interviews to test where disclosures support or block an oversight task; retain other explanations and negative cases. BACHELOR Final result: Comparison table, evidence gaps and a plain-language disclosure template. Task: Compare what four platforms tell users about why content is recommended and which controls users receive. Data: 12 public recommender documents from four platforms (public) Collect one recommender description, one user-control page and one policy/transparency document per platform. finalise date, URL and PDF. Requires: No participant recruitment; agree the four platforms and document date with the supervisor. Method: Code 12 fixed public documents with an established transparency checklist. Double-code 3 documents and report differences by platform and document type. STEPS Study and improve 1. Read the starting sources and write the exact information problem. 2. Prepare 12 public recommender documents from four platforms and a separate answer sheet or coding sheet. 3. Run one pilot and fix unclear questions. 4. Collect the named evidence with consent. 5. Group the findings with the stated categories and check the answer sheet. 6. Produce comparison table, evidence gaps and a plain-language disclosure template. List the three most useful changes. MASTER Final result: A source table for each platform, a platform-by-platform comparison and a disclosure template for recommender audits. Research question: Explain how platform disclosures make recommender risk more or less observable to researchers, regulators and prevention professionals. Research result: A cross-platform observability framework and evidence-based minimum disclosure requirements for recommender oversight. STEPS Study, compare and explain 1. Read the newest papers and write one exact research question. 2. State which people, cases or documents you will study and what you will compare. 3. Write the selection rules, questions and analysis steps for 24 public platform and DSA documents + 8–10 professional interviews. 4. Run one pilot. Fix unclear questions or categories, then keep the guide unchanged. 5. Collect the named interviews, cases or documents with consent. 6. Analyse them with the stated comparison or coding method. Keep disagreements and missing data. 7. Report the answer, the evidence and the practical output named in the assignment. CURRENT PROJECT LITERATURE Causally estimating the effect of YouTube’s recommender system using counterfactual bots (2024, peer-reviewed journal article): https://pmc.ncbi.nlm.nih.gov/articles/PMC10895271/ Separates recommender effects from user preferences and shows why audit claims need a counterfactual design. 8–10% of algorithmic recommendations are bad, but… an exploratory risk-utility meta-analysis (2024, peer-reviewed journal article): https://doi.org/10.1016/j.ijinfomgt.2023.102743 Frames recommendation design as a measurable risk–utility trade-off rather than a one-sided accuracy problem. IS THEORY STARTING POINTS Leerssen (2024) — Outside the Black Box: https://ojs.weizenbaum-institut.de/index.php/wjds/article/view/4_2_3 Distinguish transparency from the practical observability needed for oversight. Gleiss et al. (2023) — Identifying the patterns of platform regulation: https://aisel.aisnet.org/jit/vol38/iss2/6/ Structure the comparison of regulatory problems and response instruments. Search Scopus or Web of Science and ACM Digital Library using the topic query, then follow citations to the thesis start date. Record searches and compare methods, data, findings and limitations in literature-matrix.csv. Use that review to confirm or revise the gap and choose a current comparator. The linked papers are starting points. START WITH THESE THREE ACTIONS 1. Name the six platforms and four required document types; archive a dated copy of every source. 2. Pilot the coding matrix on two platforms and double-code four documents; revise and finalise it. 3. Prepare an oversight task and interview guide; recruit only professionals and exclude case-specific operational details. EXAMPLE TOOLS Zotero, LibreOffice Calc, Taguette or NVivo; no programming required. Equivalent tools are fine. LITERATURE SEARCH platform recommender transparency algorithm audit governance youth extremism DSA ACCESS OR PEOPLE Recruit 8–10 relevant professionals; obtain consent and avoid operationally sensitive case information.