Cyber LLM Benchmark Hub: A Home for Measuring AI in Cybersecurity

A few years ago, it was easy to ask a simple question: how good is an AI model at cybersecurity? Today, that question is harder to answer.
Some models are strong at reasoning. Some are better at logs. Some can help with analysis, while others struggle when the task becomes practical and messy. In cybersecurity, small differences matter. A model that looks good in a general benchmark may not perform well when it faces real incident response work, vulnerability analysis, or threat hunting.
That is where Cyber LLM Benchmark Hub comes in. It is built as a central place to compare cybersecurity LLM performance across tasks and domains. At launch, the hub brings together 35 benchmarks, 78 models, 10 categories, and 157 results, with coverage across areas like malware analysis, penetration testing, incident response, vulnerability analysis, threat intelligence, threat modeling, and LLM safety & jailbreaking.
The idea behind the project was simple: cybersecurity AI should be measured in a way that reflects the real world. Not just with one score, and not just with one task. Security work is broad. It can mean reading network data, investigating alerts, studying exploits, or understanding how a system fails under pressure. A useful benchmark hub should reflect that variety. Cyber LLM Benchmark Hub is organized around exactly that goal.
One of the most interesting parts of the project is the range of benchmarks it collects. The featured datasets include SOCBench, which evaluates frontier reasoning LLMs as SOC agents on raw NetFlow data; Cyber Defense Benchmark, which measures threat hunting on raw Windows event logs; ExploitBench, which looks at how far AI agents can move along the exploitation path on the V8 JavaScript engine; and ExploitGym, which tests whether agents can turn real vulnerabilities into working exploits across userspace programs, the V8 browser engine, and the Linux kernel.
These examples show why benchmarking matters. In cybersecurity, a model is not useful just because it sounds confident. It must show real skill, under real constraints, on tasks that matter. A benchmark hub makes that comparison easier. It also helps researchers, builders, and security teams see where models are improving and where they still fall short.
The project is also meant to grow with the community. The site invites people to contribute evaluation results and help build a more complete database for cybersecurity LLM benchmarking. That makes the hub feel less like a static page and more like a shared resource for the field.
For me, this project is about clarity. Cybersecurity is already a hard space. AI should not make it more confusing. It should help us measure better, compare better, and understand better. That is the purpose of Cyber LLM Benchmark Hub: to bring order to a fast-moving area and give the community a clearer way to see what these models can actually do.
This is only the beginning. As more benchmarks appear and more models are tested, the picture will keep changing. But that is exactly why this hub matters. It gives us a place to track the field as it grows.
Explore the Hub
You can explore the benchmarks, compare models, and contribute your own evaluation results directly on the site: