Back to Blog
If Your Continuous Penetration Testing Runs Itself, What Are You Paying For?

If Your Continuous Penetration Testing Runs Itself, What Are You Paying For?

Andrew Roe

Short version, so you can stop reading if you already agree.

If the deliverable is a report an automated system produced without a person deciding anything, you didn't need a vendor. You needed an afternoon and an API key. Pay for continuous penetration testing when a skilled human is running it and can tell you what the findings mean.

The argument between annual and continuous testing is over. Nobody defends the once-a-year PDF anymore, and every vendor in this category will tell you the same true things about drift and release velocity. That's settled. The question worth asking now is what actually happens after you buy.

What continuous penetration testing actually means

Continuous penetration testing is manual security testing performed on an ongoing cadence rather than as a single annual engagement, triggered by a schedule, by releases, or both, with findings delivered as they're confirmed instead of batched into a report at the end. The word doing the work is manual. It's also the word that has quietly stopped being true at a lot of shops.

The thing you are actually buying

Here's the model that's become common. You prepay for a pack of credits, sometimes called days or units. You spend one when you want something tested. What happens next, increasingly, is that an automated system runs against your target, produces findings, and files them.

Sit with that for a second, because it's a strange thing to pay for.

You could do that. Not in some theoretical sense. Anyone with your codebase and a frontier model can point tooling at an application and get a list of plausible issues back by dinner. That capability is not scarce anymore, and paying a vendor a five-figure retainer to do it on your behalf is paying for a wrapper.

The scarce part was never the running. It's knowing which finding is real, which two chain into something that actually matters, which one is a false positive that will burn a sprint, and what breaks if you fix the third one the obvious way. That judgment doesn't come out of a pipeline. It comes from someone who has broken into things before, looking at your specific environment, saying: this one, first, and here's why.

The call I keep having

Someone books time with me because they're unhappy with the testing they already pay for. It goes the same way almost every time.

They bought a service. They can't tell me what it's actually testing. They have never spoken to a single person who worked on their account. Every so often a report lands in an inbox, and the report is long, which for a while felt like getting their money's worth.

Then we open it together.

Most of it is noise. Findings that aren't reachable in their environment. The same issue counted four times because it appears on four routes. Severities that don't survive thirty seconds of questioning. Remediation advice generic enough to have been written before anyone looked at their stack, because it was.

Here's the part that should make you angry: they paid for that report, and then they paid again. Their engineers burned a sprint triaging it. That's the real price of an untriaged automated run, and it never appears on the invoice.

None of this is a model problem. We use models constantly and they have made us meaningfully better at this work. It's a speed problem. Everything looks easy now, and a lot of people concluded that because Claude can produce something resembling a penetration test in an afternoon, the resemblance was the product. It isn't. A frontier model will happily hand you a hundred plausible findings. Working out which four matter is the entire job, and a shop in a hurry skips precisely that step and ships you the raw output.

Credits make it worse

Metering makes all of this worse, and it changes your behavior in ways the pricing page never mentions.

Scope shrinks to fit the credit. The test that should cover the app plus the identity provider it trusts becomes a test of the app, because that's what one unit buys.

The retest becomes optional. A finding you fixed but never verified is a finding you believe you fixed. When verification costs a unit, verification turns into a budget decision.

Testing gets scheduled around the balance instead of around risk. You ship something in March that touches authentication. Probably fine. Testing it costs a credit, you have four left and nine months to go, so you fold it into the next scheduled test. That's not negligence, it's arithmetic. Then October arrives, you're nearly out, and the enterprise deal closing in Q4 wants a recent test against an environment you just migrated.

What we do instead

Every operator on our team is Colorado-based and employed by us. We have never offshored an engagement and we don't subcontract them. There is no arrangement where the company you signed with quietly hands your test to a company you've never heard of, in a country nobody disclosed.

We use frontier models hard. As Anthropic partners, our operators run with serious tooling behind them, and Claude Code does the grinding: enumeration, permutation, the tedious surface work that used to eat the first three days of every engagement. That speed is why a standard application test is $3,000 rather than the $10,000 to $30,000 the market treats as typical.

But an operator drives. Every action sits inside a signed rules-of-engagement document and gets approved by a named human who is accountable for it. The models never run alone, and they never decide what matters. They make a skilled person considerably faster at the parts of the job that were always mechanical.

That's the whole distinction. Same tools, opposite arrangement. One model uses automation to take the expert out of the loop and sells you the output. Ours uses automation to give the expert more hours in the day, and you get the expert.

What continuous should actually look like

Cadence is the easy part and the part vendors oversell. Ours runs on whatever fits your release process: a fixed schedule, triggered on release, continuously against a rolling scope, or some mix. That's a scheduling question, not a philosophy, and any shop that makes it sound hard is selling you the calendar instead of the work.

What actually separates a continuous program from an annual one is the four things underneath it.

Findings arrive when we confirm them. No quarterly drop. If an operator proves something is exploitable on a Tuesday, you know on Tuesday, because the entire point of testing continuously is collapsing the window between when a weakness exists and when you know about it. A finding held for a report is a finding aging in a queue.

Every finding ships with a reproduction. Steps, severity, and code-level remediation guidance for your stack. If you can't reproduce it, you can't verify you fixed it, and you're back to believing rather than knowing.

The retest is included, not a line item. We help you close it and then we test it again. Charging separately for verification puts a price on finding out whether the fix worked, which is exactly the wrong thing to put a price on.

Findings map to the compliance controls they touch. The same work that keeps you secure produces the evidence your auditor wants, so staying secure and staying provable stop being two separate projects with two separate budgets.

And you can watch it happen. Your team sees the test as it runs rather than receiving a narrative about it afterward. That is how we run it for echowin, where continuous testing and the compliance program feed each other instead of competing for the same calendar. That tends to unsettle vendors whose process wouldn't survive being observed.

What continuous penetration testing actually costs

Scope drives price far more than any pricing tier does. Here's what moves the number, with the market's own figures attached.

Per host. External network testing runs roughly $150 to $1,000 per device depending on depth. A hundred hosts lands around $5,000 to $15,000, and one published scoping guide puts a 100-system network at $5,670. The spread is mostly about whether anyone is chaining findings or only listing them.

Internal versus external. Internal network tests run $7,500 to $30,000, averaging near $12,500, against $5,000 to $20,000 external. Internal costs more for an unglamorous reason: something has to be inside the network, which is hardware you ship or a person you send.

On site. Add $1,000 to $3,000 for travel and per diem. Most testing doesn't need it. Physical access work, air-gapped environments, and anywhere you genuinely cannot issue remote access do, and we will get on a plane for those.

Red team. A different product with a similar name. Engagements start near $20,000 and run past $150,000, usually quoted at three to five times a web application test. You are not buying a vulnerability list. You are buying the answer to whether your people and your detection stack notice someone who is trying not to be noticed. If nobody on your side is watching yet, don't buy this one.

Across all types the market average sits near $18,300, spanning $5,000 to well past $100,000.

Why I scope every engagement personally

I sit down with every company that comes across our desk. Not a form, not a tier selector, not a calculator that multiplies your IP count by a day rate.

The reason is in the numbers above. A hundred hosts might be a $5,000 problem or a $15,000 problem, and which one it is depends on things a form can't ask: how many of those hosts are genuinely distinct, what trusts what, whether your identity provider is in scope, whether last quarter's migration left something exposed that nobody has looked at since. I would rather spend an hour finding that out and quote you accurately than sell you a unit and let you discover the mismatch mid-engagement.

Sometimes that hour ends with me telling a company they don't need us yet. That happens more than you'd expect, and it's a better outcome than selling someone a plan that doesn't fit what they're actually worried about.

Does continuous testing replace the annual pentest?

For evidence purposes, usually, though confirm it with your auditor rather than assuming. Continuous programs tend to produce better evidence than an annual report: dated findings, remediation, and verification, which is much closer to what SOC 2 evidence expectations look like in practice than one PDF from March. Auditors want proof a control operated, not proof you bought a test.

A point-in-time test is still right for a specific milestone with a specific deadline. A contract that names a test. A funding round. A first SOC 2 where you need one clean artifact. We sell those, and anyone telling you scoped point-in-time testing is worthless is selling you something.

What to ask before you sign

Four questions. The answers tell you more than any capability matrix.

  1. Does a human decide each action, or does the system run and file findings on its own? If it's the latter, you are paying a markup on something you could run this afternoon.
  2. Who performs the test, where are they, and do they work for you? Employees, contractors, and a subcontracted firm are three different answers.
  3. Is the retest included, or does it consume a credit? This one predicts whether findings actually get closed.
  4. What happens when the credits run out in month eight? Listen for whether the answer is a purchase order.

We publish what our testing includes and what it costs so those are answerable before a call rather than during one. If you want your own environment scoped properly, book 30 minutes and you'll get a straight answer about what it should cost, including the answer where you don't need us.


Market figures cited from published 2026 pricing guides: Blaze Infosec, Astra, Triaxiom Security, BlueFire Red Team, DeepStrike, and Software Secured. Sythe Labs pricing is our own published rate card.