Approving software changes Personal proposal by Per-Arne Andersen S1 · Spring 2027 https://thesis.uya.no/proposals/from-promise-to-proof/ Link a written software requirement to the code and tests that demonstrate it is satisfied. TECHNICAL - Build an approval screen importing checked requirement–code-line–test links. - Record approve/reject/request-evidence decisions and the person responsible for missing evidence. DATA (mixed) HumanEval/0–29 + clean/mutated variants Derive checkable requirements from prompts. Create one demonstrably faulty edge-case mutation per task; retain clean controls and independently checked hidden witness tests. HumanEval: https://github.com/openai/human-eval METHOD 8 Python-reading developers review 10 disjoint cases with a linked matrix or an unlinked checklist containing identical evidence. Count unjustified approvals, missed violations, time and requests for missing evidence. OUTPUT Traceability tool, 30 labelled mutations and reviewer results. BACHELOR Task: Build a release dashboard showing tests, changed files and a manual approval checklist. Data: HumanEval/0–29 + clean/mutated variants (mixed) Derive checkable requirements from prompts. Create one demonstrably faulty edge-case mutation per task; retain clean controls and independently checked hidden witness tests. Requires: Recruit developers; run code in an isolated test environment. Method: Use the fixed code patches and independent test keys; verify displayed results and observe approval tasks. Plan 4 participant sessions. Output: A release-review prototype and checklist evaluation. Use relevant literature to justify the established approach; a new research contribution is not the aim of this proposal. MASTER Investigate when green tests create unjustified release confidence. Vary fault presence and evidence completeness independently; compare a traceable approval view with ordinary logs against an independent correctness key. Intended contribution: Evidence about how software assurance evidence shapes release decisions. STARTING PAPERS Liu et al. (2023) — Is Your Code Generated by ChatGPT Really Correct?: https://arxiv.org/abs/2305.01210 Distinguish passing a public test suite from independently assessed correctness. Buçinca et al. (2021) — To Trust or to Think: https://arxiv.org/abs/2102.09692 Derive a competing explanation based on verification effort and overreliance. Search Scopus or Web of Science and ACM Digital Library using the topic query, then follow citations to the thesis start date. Record searches and compare methods, data, findings and limitations in literature-matrix.csv. Use that review to confirm or revise the gap and choose a current comparator. The linked papers are starting points. START HERE Tools: Python, HumanEval, Streamlit 1. Download HumanEval and record its revision; extract tasks 0–2 as a pilot. 2. Turn each docstring into checkable requirements and map reference solutions/tests to them. 3. Create and verify one edge-case fault per pilot task before building the review matrix. Literature search: requirements traceability acceptance evidence information systems Study controls: - Keep witness tests and fault keys hidden from participants; verify mutants are not equivalent to the originals. - Approval decisions use a written evidence rule; this is a small release-approval study, not proof of organisation-wide quality. - Before collecting participant data, agree consent, storage and withdrawal handling with the supervisor. Use participant codes, not names, in study files. - Separate silent timed tasks from retrospective interviews. Timing during think-aloud sessions is descriptive, not an isolated interface-speed effect. - For the master’s study, use the research task above to define the factors and comparisons in this pilot plan. Preregister one primary outcome and feasible scope after the literature review; do not add every possible model or interface variant. REQUIRES Recruit developers; run code in an isolated test environment. -------------------- NON-TECHNICAL - Prepare requirement descriptions, test-result tables and acceptance checklists. - Ask reviewers to accept or reject each change and explain which evidence is missing. DATA (planned) 12 paper change packets + 8 IS students Three prose requirements per packet. Cross 6 faulty/6 correct cases with 6 incomplete/6 complete evidence sets; use pass/fail tables and a written acceptance rule. METHOD Compare matrices with identical unlinked materials on disjoint cases. Participants approve, reject or request evidence; count policy-unjustified approvals and code reasons about evidence and responsibility. OUTPUT Acceptance checklist and analysis of evidence gaps. BACHELOR Task: Document what a reviewer checks before accepting a software change. Data: 12 paper change packets (planned) Three prose requirements per packet. Cross 6 faulty/6 correct cases with 6 incomplete/6 complete evidence sets; use pass/fail tables and a written acceptance rule. Requires: Recruit IS students; no code writing or execution required. Method: Use patch/evidence packs in scenario interviews; group acceptance criteria and requests for missing evidence. Plan 4 participant sessions. Output: A practical approval checklist and responsibility map. Use relevant literature to justify the established approach; a new research contribution is not the aim of this proposal. MASTER Explain when reviewers approve, reject or request more evidence. Analyse test confidence, verification effort and decision ownership across matched evidence packs; actively examine counterexamples. Intended contribution: An explanatory model of approval under incomplete software evidence. STARTING PAPERS Vasconcelos et al. (2023) — Explanations Can Reduce Overreliance: https://arxiv.org/abs/2212.06823 Compare verification cost with trust as explanations of observed decisions. Orlikowski & Gash (1994) — Technological Frames: https://dl.acm.org/doi/10.1145/196734.196745 Compare how roles interpret the purpose, operation and use of the same system. Search Scopus or Web of Science and ACM Digital Library using the topic query, then follow citations to the thesis start date. Record searches and compare methods, data, findings and limitations in literature-matrix.csv. Use that review to confirm or revise the gap and choose a current comparator. The linked papers are starting points. START HERE Tools: LibreOffice Writer/Calc, audio recorder with consent; no programming required. 1. Prepare a pilot with 2 examples from: 12 fictional change-request packets + 8 IS students. Write the task questions and a reference answer sheet. 2. Write a recruitment message, information sheet and consent form for the participants named above. Agree privacy handling with the supervisor before contact. 3. Pilot one session after approval; revise unclear questions, freeze the task sets and coding categories, then recruit the planned sample. Literature search: requirements traceability acceptance evidence information systems qualitative scenario study Study controls: - Cross correctness and evidence completeness rather than making every incomplete packet faulty. - Treat IS-student results as novice decision evidence. - Before collecting participant data, agree consent, storage and withdrawal handling with the supervisor. Use participant codes, not names, in study files. - Pilot separately, then freeze the questions and coding plan. Check objective answer keys independently; keep an audit trail of coding, including disagreements. - For comparisons, counterbalance order and case assignment; do not show a person both versions of one case. Report participant-level findings, not repeated tasks as independent people. REQUIRES Recruit IS students; no code writing or execution required. Bachelor: apply established methods and evaluate a practical solution or study. Master: position a research question in current scientific literature, investigate a mechanism or unresolved problem, and explain the contribution. Final scope is agreed with me. Study sizes are proposed; participant recruitment and planned materials are not already arranged.