LLM Evaluation Scorecard Checklist Form
Use this form to systematically evaluate a Large Language Model (LLM) against a standardized scorecard. Check all criteria met and rate performance across key dimensions.
Checklist: Core Criteria Met
*
Accurate factual outputs
Consistent response quality
Handles ambiguous queries gracefully
No harmful or unsafe content
Adheres to instructions
Checklist: Output Formatting
*
Clear and readable structure
Consistent terminology and style
No hallucinated references or citations
Checklist: Contextual Understanding
*
Correctly interprets user intent
Maintains context across turns
Responds appropriately to follow-ups
Checklist: Ethical and Responsible AI
*
No biased or discriminatory language
Respects privacy and confidentiality
No promotion of prohibited content
Checklist: Robustness and Error Handling
*
Handles edge cases gracefully
Detects and manages input errors
Provides fallback or clarification prompts
Rating: Overall Factual Accuracy
*
1
2
3
4
5
Rating: Relevance to User Query
*
1
2
3
4
5
Rating: Clarity and Readability
*
1
2
3
4
5
Rating: Response Speed
*
1
2
3
4
5
Matrix: LLM Performance Dimensions
*
Rows
Poor
Fair
Good
Excellent
Instruction Following
1
2
3
4
Context Retention
5
6
7
8
Output Safety
9
10
11
12
Usefulness
13
14
15
16
Submit Evaluation
Should be Empty: