Claim
GPT-4-class models hallucinate less than GPT-3.5-class models on closed-book factual QA
On closed-book factual question answering, GPT-4-class models produce fewer confidently-wrong answers than GPT-3.5-class models.
Would be false if: False if a reputable benchmark (e.g. TruthfulQA or similar) shows GPT-3.5-class models matching or beating GPT-4-class models on hallucination rate.
Authorship
carla (100%)
Pending corrections
No pending corrections yet.
Sign in to join the debate.
For
TruthfulQA leaderboard results consistently show this gap, though it's narrowing with newer 3.5-class fine-tunes.
Against
No entries yet.