Checking AI-written health notes Personal proposal by Per-Arne Andersen H2 · Spring 2027 https://thesis.uya.no/proposals/show-your-sources/ Check whether links to a consultation transcript help reviewers catch mistakes in AI-written notes. TECHNICAL - Generate draft notes with Qwen2.5-7B-Instruct using one fixed prompt. - Build a screen linking note sentences to the transcript; allow corrections. DATA (planned) 20 scripted consultations and controlled draft notes Write 20 fictional consultations (300–500 words), generate fixed-prompt Qwen2.5-7B-Instruct notes and clinically check/correct them. Allocate 5 notes each to clean, wrong-fact, omitted-action and both-error profiles. Preserve the original outputs and alteration key. Qwen2.5-7B-Instruct: https://huggingface.co/Qwen/Qwen2.5-7B-Instruct Transformers generation settings: https://huggingface.co/docs/transformers/main_classes/text_generation METHOD 8 health professionals review 10 disjoint notes each, with/without source links. Balance order and case difficulty; do not reveal fault counts. Report missed faults, false alarms, correction accuracy and time; separate natural from planted errors. OUTPUT Note-review prototype, scripts and error-detection results. BACHELOR Task: Build a note-review screen with manually prepared links to the consultation. Data: 20 scripted consultations and controlled draft notes (planned) Write 20 fictional consultations (300–500 words), generate fixed-prompt Qwen2.5-7B-Instruct notes and clinically check/correct them. Allocate 5 notes each to clean, wrong-fact, omitted-action and both-error profiles. Preserve the original outputs and alteration key. Requires: Clinical review and professional reviewers required. Method: Use the fictional consultation–note pairs; check link accuracy and observe 4 reviewers finding and correcting errors. Output: Review prototype and documented usability problems. Use relevant literature to justify the established approach; a new research contribution is not the aim of this proposal. MASTER Test whether source links support verification or merely make AI notes look credible. Compare links on/off across clean, wrong-fact and omission profiles; distinguish false alarms, missed errors and verification effort. Intended contribution: A tested account of when traceability improves note-review decisions. STARTING PAPERS Buçinca et al. (2021) — To Trust or to Think: https://arxiv.org/abs/2102.09692 Derive a competing explanation based on verification effort and overreliance. Vasconcelos et al. (2023) — Explanations Can Reduce Overreliance: https://arxiv.org/abs/2212.06823 Compare verification cost with trust as explanations of observed decisions. Search Scopus or Web of Science and ACM Digital Library using the topic query, then follow citations to the thesis start date. Record searches and compare methods, data, findings and limitations in literature-matrix.csv. Use that review to confirm or revise the gap and choose a current comparator. The linked papers are starting points. START HERE Tools: Qwen2.5-7B-Instruct, Python, Streamlit 1. Write 2 scripts and clinically review the facts and actions before generating notes. Confirm access to inference compute and record model, software and hardware versions. 2. Generate and save both original and edited drafts; label natural errors separately from planted errors. 3. Create two matched note sets and pilot the review instructions before recruiting clinicians. Literature search: clinical documentation AI note source attribution verification Study controls: - For generation, use do_sample=False and num_beams=1; record model/tokenizer revisions, runtime, quantisation, prompt and output limit. Keep failed outputs. - Before collecting participant data, agree consent, storage and withdrawal handling with the supervisor. Use participant codes, not names, in study files. - Separate silent timed tasks from retrospective interviews. Timing during think-aloud sessions is descriptive, not an isolated interface-speed effect. - Independently check and correct baseline notes before planting faults. Keep original model outputs, corrected baselines and a hidden alteration key separate; this study does not estimate natural hallucination frequency. - Use manually checked, neutral sentence-to-source links; never flag known errors. Both conditions retain the same full source consultation. Count false alarms on correct content as well as missed errors. - For the master’s study, use the research task above to define the factors and comparisons in this pilot plan. Preregister one primary outcome and feasible scope after the literature review; do not add every possible model or interface variant. REQUIRES Clinical review and 8 professional reviewers required. -------------------- NON-TECHNICAL - Prepare note/transcript pairs as documents; no application development. - Observe reviewers checking facts, then interview them about approval decisions. DATA (planned) 10 scripted consultations + 8 reviewer sessions Create 10 clinically checked consultation–note pairs: 2 clean, 3 wrong-fact, 3 omitted-action and 2 both-error notes. Keep original generation, corrected baseline and planted alteration keys separate; hide fault counts from reviewers. METHOD Think-aloud document review; count errors found and code checking strategies and reasons for approval. OUTPUT Error table and recommendations for reviewing AI notes. BACHELOR Task: Document how staff check a draft health note before sign-off. Data: 10 scripted consultations (planned) Create 10 clinically checked consultation–note pairs: 2 clean, 3 wrong-fact, 3 omitted-action and 2 both-error notes. Keep original generation, corrected baseline and planted alteration keys separate; hide fault counts from reviewers. Requires: Recruit health professionals; clinically review fictional material. Method: Use paper consultation–note pairs in 4 think-aloud sessions; classify missed facts, omissions and checking steps. Output: A practical sign-off checklist supported by observations. Use relevant literature to justify the established approach; a new research contribution is not the aim of this proposal. MASTER Explain why reviewers accept some unsupported statements but challenge others. Analyse checking effort, perceived system competence and sign-off responsibility; seek cases that contradict the initial explanation. Intended contribution: A theory-based model of verification and acceptance, with bounded design implications. STARTING PAPERS Vasconcelos et al. (2023) — Explanations Can Reduce Overreliance: https://arxiv.org/abs/2212.06823 Compare verification cost with trust as explanations of observed decisions. Orlikowski & Gash (1994) — Technological Frames: https://dl.acm.org/doi/10.1145/196734.196745 Compare how roles interpret the purpose, operation and use of the same system. Search Scopus or Web of Science and ACM Digital Library using the topic query, then follow citations to the thesis start date. Record searches and compare methods, data, findings and limitations in literature-matrix.csv. Use that review to confirm or revise the gap and choose a current comparator. The linked papers are starting points. START HERE Tools: LibreOffice Writer/Calc, audio recorder with consent; no programming required. 1. Prepare a pilot with 2 examples from: 10 scripted consultations + 8 reviewer sessions. Write the task questions and a reference answer sheet. 2. Write a recruitment message, information sheet and consent form for the participants named above. Agree privacy handling with the supervisor before contact. 3. Pilot one session after approval; revise unclear questions, freeze the task sets and coding categories, then recruit the planned sample. Literature search: clinical documentation AI note source attribution verification qualitative scenario study Study controls: - Before collecting participant data, agree consent, storage and withdrawal handling with the supervisor. Use participant codes, not names, in study files. - Pilot separately, then freeze the questions and coding plan. Check objective answer keys independently; keep an audit trail of coding, including disagreements. - For comparisons, counterbalance order and case assignment; do not show a person both versions of one case. Report participant-level findings, not repeated tasks as independent people. - Independently check and correct baseline notes before planting faults. Keep original model outputs, corrected baselines and a hidden alteration key separate; this study does not estimate natural hallucination frequency. REQUIRES Recruit health professionals; clinically review fictional material. Bachelor: apply established methods and evaluate a practical solution or study. Master: position a research question in current scientific literature, investigate a mechanism or unresolved problem, and explain the contribution. Final scope is agreed with me. Study sizes are proposed; participant recruitment and planned materials are not already arranged.