CARE Pilot Study

CARE Pilot Study

CARE Pilot Study

September 2025

This report evaluated 22 models across 5 test scenarios. Each pairing ran 10 times per model at a temperature setting of 0.
Judged by gemini-2.5-pro.
gemini-2.5-flash
20%
2.4
1.6
2.2
6.2
gpt-5-2025-08-07
22%
2.3
1.6
2.2
6.1
claude-opus-4-1-20250805
20%
2.2
1.6
1.8
5.6
us.meta.llama4-maverick-17b-instruct-v1:0
20%
2.1
1.6
1.2
4.9
gemini-2.0-flash-001
20%
2.0
1.6
1.3
4.9
gemini-2.5-pro
40%
1.8
1.2
1.8
4.8
us.deepseek.r1-v1:0
40%
1.8
1.2
1.4
4.4
moonshotai/kimi-k2-instruct
34%
1.9
1.3
1.2
4.4
claude-opus-4-20250514
32%
1.7
1.3
1.2
4.2
llama-3.3-70b-versatile
34%
2.0
1.3
0.9
4.2
qwen/qwen3-32b
40%
1.7
1.1
1.2
4.0
claude-3-5-sonnet-20241022
34%
1.7
1.3
1.0
4.0
moonshotai/kimi-k2-instruct-0905
40%
1.7
1.2
1.0
3.9
us.meta.llama3-3-70b-instruct-v1:0
40%
1.8
1.1
0.9
3.8
us.meta.llama4-scout-17b-instruct-v1:0
40%
1.7
1.1
0.9
3.8
claude-sonnet-4-20250514
40%
1.5
1.2
0.8
3.4
claude-3-7-sonnet-20250219
40%
1.3
1.1
0.7
3.1
gpt-4.1-2025-04-14
40%
1.3
1.2
0.4
2.9
grok-4-latest
60%
1.2
0.8
0.9
2.9
grok-3-beta
60%
1.2
0.8
0.8
2.8
gpt-4o
40%
1.0
1.2
0.2
2.4
gpt-4o-mini-2024-07-18
46%
0.9
1.1
0.3
2.3
Click model rows to expand • Click test cases for detailed analysis