LLM Evaluation Trace Log Form
Log and review large language model evaluation traces for quality and performance analysis.
Date of Evaluation
*
 -
Month
 -
Day
Year
2 digit month, 2 digit day, 4 digit year
Date
Hour Minutes
AM
PM
AM/PM Option
Evaluator Name
*
Model Evaluated
*
Please Select
GPT-4
Claude 3
Gemini Pro
Other
Evaluation Scenario / Task Name
*
Prompt or Input Used
*
Expected Output
Actual Output
*
Evaluation Criteria Scores
*
Rows
Score (1-5)
Accuracy
1
Relevance
2
Completeness
3
Fluency
4
Overall Evaluation Rating
*
1
2
3
4
5
Additional Comments or Observations
Submit Trace Log
Should be Empty: