RustMizan Leaderboard
RustMizan benchmarks Large Language Models on real-world Rust memory safety vulnerabilities from CVEs and security advisories.
In the Sample-wise Comparison tab, hover over any result emoji to see its Docent contamination check and links to the model's full trajectory and Docent analysis.
For the task, metrics, dataset variants, and more, see the documentation.