Proposal
Add REFUTE to related / evaluation / research tooling docs if this project surfaces LLM or agent science workflows.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.
Proposal
Add REFUTE to related / evaluation / research tooling docs if this project surfaces LLM or agent science workflows.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.