Writ / attention-review / paper.md @ edit/r-kaur/llm-probe
reconnecting Log in Sign up

Attention and review

3. Results

The improvement holds for submissions under twenty pages, and attenuates beyond that length.

Attention and review

3. Results

Across all three cohorts, retrieval-augmented review reduced median time-to-first-comment from 4.2 days to 1.6 days. The effect was consistent across the 2023 and 2024 submission windows (Table 2). Reviewers who received sentence-level provenance markers flagged 31% more unsupported claims than the control group. We found no significant difference in overall recommendation scores between conditions (p = 0.41). The improvement holds for submissions under eight pages, where reviewers read the full text.

4. Discussion

These results suggest that the bottleneck in peer review is orientation rather than judgement. Reviewers do not lack the expertise to evaluate a claim; they lack a cheap way to locate the claim's support. Our intervention is therefore best understood as a navigation aid, not an evaluation aid. The 31% increase in flagged claims is obviously the headline result. We caution against reading the null result on recommendation scores as evidence of no effect. Power to detect a difference smaller than 0.4 points was low in all three cohorts.

main 48449397ce 13 changes: 13 meaning, 0 surface · Submission #9 edit/r-kaur/llm-probe 524d53577d
removed fact / evidence model Across all three cohorts, retrieval-augmented review reduced median time-to-first-comment from 4.2 days to 1.6 days. rule:methods-frozen · blocking Rule: methods-frozen r-kaur #10 2h ago
removed fact / evidence model The effect was consistent across the 2023 and 2024 submission windows (Table 2). rule:methods-frozen · blocking Rule: methods-frozen r-kaur #12 2h ago
removed fact / evidence model Reviewers who received sentence-level provenance markers flagged 31% more unsupported claims than the control group. rule:methods-frozen · blocking Rule: methods-frozen r-kaur #13 2h ago
removed fact / evidence model We found no significant difference in overall recommendation scores between conditions (p = 0.41). rule:methods-frozen · blocking Rule: methods-frozen r-kaur #14 2h ago
removed fact / evidence model The improvement holds for submissions under eight pages, where reviewers read the full text. rule:methods-frozen · blocking Rule: methods-frozen r-kaur #15 2h ago
removed other model ## 4. Discussion
removed claim model These results suggest that the bottleneck in peer review is orientation rather than judgement.
removed claim model Reviewers do not lack the expertise to evaluate a claim; they lack a cheap way to locate the claim's support.
removed claim model Our intervention is therefore best understood as a navigation aid, not an evaluation aid.
removed claim model The 31% increase in flagged claims is obviously the headline result.
removed claim model We caution against reading the null result on recommendation scores as evidence of no effect.
removed claim model Power to detect a difference smaller than 0.4 points was low in all three cohorts.