Detailed comparison for LLMs
On Guardion's LLM vulnerability Benchmark, Anthropic Claude 3.7 Sonnet is the more secure of the two: Qwen3-Max scores 28.0% and Claude 3.7 Sonnet scores 20.3% on attack success rate (ASR) (lower is better). One or both scores are estimated from public safety evaluations pending a Guardion benchmark run.
Claude 3.7 Sonnet is the overall winner in this comparison!
ASR for Alibaba Qwen3-Max vs Anthropic Claude 3.7 Sonnet. Green marks the safer model on each metric. Only the overall score is available for estimated models.