Skip to content

How we evaluate OwlSOC, including where it falls short

Most security vendors publish the numbers that flatter them. We would rather show you how we test OwlSOC, what it gets right, and where it currently falls short — because a tool you are about to give access to your tenant deserves that. This is the durable version; the running commentary lives on our blog.

Why we publish this

Every AI SOC claims accuracy. Almost none show their working, and none we have seen publish a bar they fail. We think that is backwards. If you are evaluating a tool that will read your security data and recommend containment, you should be able to see how it is tested and where its limits are before you commit.

So we publish our evaluation method and our current results, including the parts that are not finished. It is the same standard we would want from a vendor asking for access to our environment.

The evaluation set

Before any release, we validate OwlSOC against a fixed set of real-world attack scenarios built from the kinds of alerts our connectors actually surface. They are synthetic — no customer data — but modelled on real techniques across identity, endpoint and cloud.

  • Adversary-in-the-middle (AiTM) token theft
  • Encoded and obfuscated PowerShell execution
  • Malicious OAuth consent grants
  • Mass exfiltration and unusual data access
  • Impossible-travel and anomalous sign-ins, including benign look-alikes
  • And other identity, endpoint and AWS scenarios

How we score it

For each scenario there is an expected outcome: the right hedged verdict — likely true positive, likely false positive, or uncertain and needs review — supported by a correctly sourced timeline. We check both. A right verdict for the wrong reasons, or with a timeline that does not trace back to real evidence, does not pass.

We also care about false positives specifically — flagging benign activity as malicious — because a tool that cries wolf is a tool teams learn to ignore. So the benign look-alikes in the set matter as much as the attacks.

The current results, including the bar we fail

Most scenarios reach the expected verdict with a sourced timeline behind them. But we run OwlSOC in two tiers, and they do not perform equally. The AI investigation tier does the deeper reasoning; the standard deterministic tier scores every alert with repeatable rules and no calibrated confidence.

As of today, the deterministic tier does not yet meet our own false-positive target — it flags more benign cases than we are willing to accept as the finished bar. We are not rounding that up. It is on the roadmap, it is measured every release, and we would rather you knew it now than discovered it later.

What this does and does not tell you

A fixed test set proves the reasoning works on known scenarios. It does not prove anything about your environment, your alerts, or your baseline — no benchmark can. That is not a hedge to duck the question; it is the reason the pilot exists.

The £495, 30-day, fully-refundable pilot is the real test: OwlSOC runs on your own alerts for a month, and if it does not earn its keep, you ask for the money back. The evaluation set is how we hold ourselves to a standard between releases; the pilot is how you hold us to one on your data.

Frequently asked

How is OwlSOC evaluated before it connects to my environment?

Each release is validated against a fixed set of synthetic, real-world attack scenarios across identity, endpoint and cloud. For each, OwlSOC must reach the expected hedged verdict with a correctly sourced timeline. It is a defined test set, not a guarantee about any specific environment, which is why the refundable pilot exists.

Does OwlSOC pass every scenario in the evaluation set?

Most scenarios reach the expected verdict with a sourced timeline. But the standard deterministic tier does not yet meet our own false-positive target, and we publish that rather than hide it. The AI investigation tier does the deeper reasoning; the deterministic bar is measured every release and is on the roadmap.

Why publish a bar you currently fail?

Because a tool you are about to give access to your tenant deserves an honest account of its limits, and because no other AI SOC we have seen publishes one. Showing where OwlSOC falls short is how we earn the trust to say where it does well.

What if OwlSOC gets an investigation wrong on my data?

It will sometimes, the same way an analyst will, which is why every verdict is hedged and every claim is sourced. Your team can open the exact log line behind any conclusion and disagree with it. Ambiguous cases are flagged as needs review rather than miscalled, and the refundable pilot means the proof sits on your data before you commit.

See it on your alerts.

Start with a 30-day refundable pilot. £495, one environment, every alert investigated, a full report at week four. Read-only, live within 48 hours of access.