LLM Benchmarking Survey
Please provide structured benchmarking feedback about large language models. All responses help improve assessment practices. Title: LLM Benchmarking Survey.
Your role in the benchmarking process
*
Please Select
Researcher
Data Scientist
Engineer
Product Manager
Business Analyst
Other
Benchmarked model name and version
*
Benchmarking objective
*
Which models did you compare against?
*
GPT-4
Claude
Llama
Gemini
Other
Evaluation criteria considered
*
Accuracy
Factuality
Reasoning
Helpfulness
Safety
Other
Task types assessed
*
Text Generation
Summarization
Question Answering
Code Generation
Translation
Other
Overall rating of the benchmarked model
*
1
2
3
4
5
Per-criterion Likert scale ratings
*
Rows
1 (Poor)
2
3
4
5 (Excellent)
Accuracy
1
2
3
4
5
Factuality
6
7
8
9
10
Reasoning
11
12
13
14
15
Helpfulness
16
17
18
19
20
Safety
21
22
23
24
25
Key strengths observed
*
Key weaknesses observed
*
Submit
Should be Empty: