Skip to content

Quality Rubrics for Paper Audit

Descriptive scoring anchors for both audit scoring systems. Use these rubrics to calibrate scores and map them to editorial decisions.


4-Dimension NeurIPS-Style Scale (1.0 - 6.0)

Base score: 6.0, deducted per issue (Critical -1.5, Major -0.75, Minor -0.25). Floor: 1.0.

Quality (Weight: 30%)

Primary checks: LOGIC, BIB, GBT7714

Score RangeLevelBehavioral Indicators
5.5 - 6.0ExceptionalTechnically flawless; rigorous proofs or experiments; all claims well-supported by evidence; no logical gaps
4.5 - 5.4StrongSound methodology with minor gaps; adequate baselines and ablations; claims mostly supported
3.5 - 4.4AdequateGenerally sound but notable weaknesses; some claims lack sufficient evidence; minor logical inconsistencies
2.5 - 3.4WeakSignificant methodological concerns; several claims unsupported; logical gaps undermine core argument
1.0 - 2.4InsufficientFundamental errors in reasoning or methodology; critical evidence missing; conclusions not supported

Clarity (Weight: 30%)

Primary checks: FORMAT, GRAMMAR, SENTENCES, CONSISTENCY, REFERENCES, VISUAL, FIGURES, DEAI

Score RangeLevelBehavioral Indicators
5.5 - 6.0ExceptionalCrystal clear writing; perfect formatting; all figures/tables well-designed and referenced; no grammar issues
4.5 - 5.4StrongClear writing with minor formatting issues; figures readable; occasional grammar or style issues
3.5 - 4.4AdequateGenerally understandable but some sections unclear; several formatting inconsistencies; grammar errors present
2.5 - 3.4WeakFrequently unclear; significant formatting problems; many grammar errors; figures poorly designed
1.0 - 2.4InsufficientVery difficult to follow; pervasive formatting issues; grammar errors impede comprehension

Significance (Weight: 20%)

Primary checks: LOGIC, CHECKLIST

Score RangeLevelBehavioral Indicators
5.5 - 6.0ExceptionalAddresses a critical problem; results advance the field substantially; broad impact potential
4.5 - 5.4StrongImportant problem with meaningful results; clear contribution to the field
3.5 - 4.4AdequateReasonable problem; results provide incremental contribution; limited broader impact
2.5 - 3.4WeakProblem significance unclear; results are marginal or narrowly applicable
1.0 - 2.4InsufficientTrivial problem or results; no discernible contribution to the field

Originality (Weight: 20%)

Primary checks: DEAI, CHECKLIST

Score RangeLevelBehavioral Indicators
5.5 - 6.0ExceptionalNovel framework or paradigm; opens new research direction; highly creative approach
4.5 - 5.4StrongNovel methodology or significant new application; clearly distinct from prior work
3.5 - 4.4AdequateIncremental extension of existing work; some novel elements but largely builds on known approaches
2.5 - 3.4WeakMinor variations of existing methods; limited novelty over prior work
1.0 - 2.4InsufficientNo discernible novelty; appears to replicate existing work without meaningful extension

Decision Mapping (4-Dimension)

Overall ScoreRecommendationTypical Next Steps
>= 5.5Strong AcceptSubmit with confidence; minor polish only
4.5 - 5.4AcceptAddress minor issues in camera-ready version
3.5 - 4.4Borderline AcceptRevise and resubmit; follow revision roadmap
2.5 - 3.4Borderline RejectMajor revision cycle needed; reconsider approach
1.5 - 2.4RejectFundamental rework required
1.0 - 1.4Strong RejectReconsider research direction or methodology

9-Dimension ScholarEval Scale (1.0 - 10.0)

8 scoring dimensions plus a computed overall. Weights mirror the single source in scripts/scholar_eval.py (SCHOLAR_EVAL_DIMENSIONS).

Base score: 10.0, deducted per issue (Critical -2.5, Major -1.25, Minor -0.5). Floor: 1.0.

Soundness (Weight: 18%, Source: Script)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentAll claims rigorously supported; no logical gaps; statistical methods appropriate and well-applied
7.0 - 8.9GoodClaims mostly supported; minor logical gaps; appropriate methodology with small concerns
5.0 - 6.9FairSeveral claims lack support; notable methodological weaknesses; some statistical concerns
3.0 - 4.9PoorMajor claims unsupported; significant methodological flaws; inappropriate statistical methods
1.0 - 2.9FailingFundamental logical errors; methodology invalid; conclusions not justified by evidence

Clarity (Weight: 13%, Source: Script)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentExceptionally well-written; perfectly organized; all notation consistent and well-defined
7.0 - 8.9GoodClear writing; well-organized; minor notation or terminology inconsistencies
5.0 - 6.9FairGenerally clear but some sections confusing; organization could improve; several style issues
3.0 - 4.9PoorFrequently unclear; poor organization; inconsistent terminology hinders understanding
1.0 - 2.9FailingVery difficult to understand; disorganized; pervasive writing problems

Presentation (Weight: 8%, Source: Script)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentProfessional layout; all figures/tables publication-ready; consistent formatting throughout
7.0 - 8.9GoodGood layout; figures clear; minor formatting inconsistencies
5.0 - 6.9FairAcceptable layout; some figures unclear or poorly labeled; formatting issues present
3.0 - 4.9PoorSignificant layout problems; figures hard to read; inconsistent formatting throughout
1.0 - 2.9FailingUnprofessional presentation; figures missing or illegible; major formatting problems

Novelty (Weight: 13%, Source: LLM)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentGroundbreaking contribution; entirely new approach or framework; paradigm-shifting potential
7.0 - 8.9GoodClearly novel approach; meaningful distinction from prior work; creative solution
5.0 - 6.9FairIncremental novelty; extends existing methods in reasonable ways; some creative elements
3.0 - 4.9PoorMarginal novelty; minor variations of existing work; unclear how this advances the field
1.0 - 2.9FailingNo discernible novelty; replicates existing work without meaningful contribution

Significance (Weight: 13%, Source: LLM)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentAddresses critical problem; results will influence multiple research areas; high practical impact
7.0 - 8.9GoodImportant problem; meaningful results; clear impact on the target community
5.0 - 6.9FairReasonable problem; results contribute incrementally; limited broader impact
3.0 - 4.9PoorProblem significance questionable; results narrowly applicable; minimal impact expected
1.0 - 2.9FailingTrivial or irrelevant problem; results of no practical value

Reproducibility (Weight: 8%, Source: Mixed)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentCode and data publicly available; all hyperparameters documented; full experimental protocol provided
7.0 - 8.9GoodCode available or promised; most details provided; could likely reproduce with reasonable effort
5.0 - 6.9FairSome details missing; code not available; reproduction would require significant effort
3.0 - 4.9PoorMajor details missing; no code or data; reproduction very difficult
1.0 - 2.9FailingInsufficient detail to reproduce; no artifacts; critical information withheld

Ethics (Weight: 5%, Source: LLM)

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentThorough ethics discussion; all concerns addressed; IRB/consent documented where needed
7.0 - 8.9GoodEthics acknowledged; main concerns addressed; minor gaps in discussion
5.0 - 6.9FairLimited ethics discussion; some concerns not addressed; potential issues not fully explored
3.0 - 4.9PoorEthics largely ignored; significant concerns unaddressed; potential for harm not discussed
1.0 - 2.9FailingNo ethics consideration; clear ethical violations; potential for significant harm

Literature Grounding (Weight: 12%, Source: Mixed) -- NEW in v3.0

Score RangeLevelBehavioral Indicators
9.0 - 10.0ExcellentComprehensive coverage of seminal and recent works; thematic organization; clear gap derivation; no important missing references
7.0 - 8.9GoodSolid coverage; mostly thematic organization; adequate gap identification; minor gaps
5.0 - 6.9FairReasonable but notable gaps; some enumeration; weak gap derivation; several missing recent papers
3.0 - 4.9PoorIncomplete coverage; predominantly enumerated; many important references missing
1.0 - 2.9FailingMinimal literature review; foundational works missing; no literature-contribution connection

Overall (Weight: 10%, Source: Computed)

Weighted average of all non-null dimensions, normalized by total available weight.

Decision Mapping (ScholarEval)

Overall ScoreReadiness LabelRecommendation
>= 9.0Strong AcceptReady for top-tier venue
8.0 - 8.9AcceptPublication ready
7.0 - 7.9Ready with Minor RevisionsAddress minor issues before submission
6.0 - 6.9Major Revisions NeededSignificant rework required
5.0 - 5.9Significant Rework RequiredFundamental improvements needed
< 5.0Not ReadyReconsider approach and methodology

Calibration Notes

  • Scores are relative to the target venue's standards. A score of 7.0 at NeurIPS represents higher absolute quality than 7.0 at a regional workshop.
  • When --venue is specified, interpret scores in the context of that venue's acceptance standards.
  • Script-based scores are strongest for Clarity and Presentation dimensions; weakest for Novelty and Significance (these require LLM judgment).
  • If script and LLM scores diverge by more than 2 points on the same dimension, flag this discrepancy in the report.
  • The 4-dimension (1-6) and 9-dimension (1-10) systems are complementary. The 4-dim system provides a quick overview; the 9-dim ScholarEval provides deeper analysis.

Released under the MIT License.