Our AI strategy for preemptive security
Security leaders are increasingly finding themselves between two extremes: headlines that frame AI agents as uncontrollable threats and vendor claims that autonomous testing can replace penetration testing outright.
In a recent Bugcrowd webinar, CTO Braden Russell and Julian Brownlow Davies, who leads offensive security and strategy, discussed the situation and what actually matters operationally. Watch the full session on demand or read the breakdown below.
In reality, agentic security tools are only as safe and effective as the workflows around them. These tools need defined targets, permitted actions, and explicit boundaries. “The model isn’t the thing that’s holding the permissions on the system,” Braden said. “It sits inside of a workflow. That workflow is scoped, it’s sandboxed, it’s deterministic at the edges.”
Without these controls, outcomes can be unpredictable. This is evidenced by most of the rogue-agent headlines, where a research team’s sandbox turned out to be less secure than they had thought. With real controls in place, the more common enterprise challenge is operational—teams still have to decide what to trust, what to fix, and how to keep engineering focused instead of flooded.
That’s why adopting autonomy requires calibration, not just speed. Treat agentic testing as a program, not a purchase. Start with assets where mistakes are acceptable, set a tight scope with known issues, and run the agent and a human against the same target to compare results. Early success isn’t measured by coverage; it’s measured by confidence.
The biggest near-term risk isn’t a runaway agent; it’s lost credibility from low-quality findings. Autonomous testing can generate hundreds of issues in days. These include duplicates, inflated severity, misclassified vulnerability types, or reports that fail validation in a real environment. That volume burns political capital with engineering fast. As Julian put it, “You turn on an autonomous agent on a Monday, and by Friday, your engineering team has hundreds of findings. A bunch of them don’t really exist or they’re vastly exaggerated.” If the first wave is noisy, teams tune out the channel.
Buyers should scrutinize what happens between “the agent found something” and “a ticket lands in your backlog.” Without validation, normalization, deduplication, and consistent evidence standards, you’re paying to become the QA layer for someone else’s automation. The commercial question isn’t “does the agent find a lot,” it’s “what percentage of findings turns into correctly prioritized, actionable work without exhausting the teams expected to fix the real issues.”
Agentic testing delivers value today through broad, continuous coverage of the long tail—assets that appear, change, and expand faster than any human program can track. Autonomous systems can monitor thousands to hundreds of thousands of endpoints at near-zero marginal cost, detect attack-surface changes (new hosts, routes, auth flows, and exposed services), and keep reconnaissance and enumeration current without waiting for a quarterly cycle. They can also consolidate repetitive signals into fewer, clearer issues by collapsing many URLs and instances into a single remediation narrative. This means teams are free to address root causes instead of chasing symptoms. Julian also noted, “Machines have the same level of attention span all the way through. When you’ve got a large scope, they don’t get bored.” This matters most in organizations with frequent releases, distributed ownership, and sprawling inventories where continuous discovery and threat model upkeep are chronically underfunded.
Scanning 400 assets isn’t the same as knowing 400 assets are secure. As Julian argued, “Coverage isn’t always the same thing as assurance.” Humans remain stronger on judgment and context, interpreting complex business logic, mapping technical findings to business impact, and deciding whether a valid exploit chain is material or better handled through mitigation. Machines are improving at multistep chaining, but intent, prioritization, and impact assessment remain human responsibilities, especially when the right outcome is to fix the right things, not everything. Braden put the boundary this way: “Building the chain, that’s one thing, and agents can do that. Understanding why it’s important, why it matters, what it impacts—that is different.”
The scalable operating model isn’t “machine first, human second as a checker,” which doubles costs. The better model is two modes on one platform: autonomous coverage for the long tail and human-led depth for crown jewels and actively targeted systems, with the ability to switch an asset between modes as risk and business context change.
A practical way to accelerate baseline trust is to test against an environment where you already have an answer key. Instead of generic benchmarks, point the autonomous tester at a target with a recent penetration test report. “Then you’ve got a scorecard,” Braden said. “What did the humans find that the AI found, what did the humans miss that the AI caught, and what did the AI claim to find that isn’t actually there?” This quickly shows which known findings the machine reproduces, which it misses, and where it surfaces additional issues humans didn’t catch—common when the surface area is large. It also reveals what the system claims exists but doesn’t, which is where trust is won or lost. Braden estimates “that one exercise of comparing against the answer key would give you better results than three months of a pilot without a control.”
After calibration, the program can move from “crawl” to “walk” by widening scope and routing work by asset class and risk. Not every asset deserves the same treatment; a marketing microsite, a payments platform, and an internal admin console have different threat models, blast radii, and assurance requirements. Mature programs assign continuous machine coverage to lower-risk or highly standardized assets, reserve human-led testing for high-risk systems, and use hybrid workflows where the machine runs continuously but escalates to humans when it detects high-severity vulnerabilities or patterns that suggest an exploitation chain. The value isn’t picking a side in a human-versus-machine debate; it’s the ability to change modes quickly as assets are reclassified, new products launch, or acquisitions expand the estate, without restarting procurement or waiting weeks to schedule help. Julian’s rule of thumb for how long this takes: “I’d rather someone take six months to develop a program that people really trust than take six weeks to make a tool that achieves nothing.”
Most autonomy initiatives stall on unglamorous blockers that aren’t “AI problems.” Access is the recurring root cause. Unauthenticated testing sees only a fraction of the real attack surface, so meaningful coverage requires provisioning working credentials at multiple privilege levels, rotating and monitoring those accounts, and planning for credential leakage. Environment strategy is another common failure point. Testing in staging can create false assurance if staging lags production, lacks integrations, or runs different data. Hardware and physical devices add logistics that autonomy doesn’t remove—someone still has to rack devices, connect networks, and keep them stable. As Braden put it, “It doesn’t get easier for you just because the tester happens to be a machine. If anything, it’s probably more complicated.” In these areas, humans remain essential, and autonomy should be positioned to remove repetitive work, not to pretend the hard parts disappeared.
A pragmatic rollout follows a crawl-walk-run path that shields critical systems from early-stage noise while you mature governance. Start with non-crown-jewel scopes where continuous discovery and change detection create immediate value, and use that phase to tune guardrails, evidence standards, deduplication, and severity mapping so that outputs match engineering reality. Define routing up front—what gets auto-filed, what needs human triage, what thresholds trigger escalation, and the target metrics for actionable rate and time to validate. As confidence increases, expand coverage and add human-led depth on the assets where compromise would create material business impact.
Scalable assurance comes down to workflow discipline and clear accountability. “Human in the loop” only works when responsibilities are explicit and reviewers can genuinely validate exploitability. Julian didn’t mince words on this: “If the human’s rubber-stamping it, it’s effectively laundering the autonomous output into being human judgment.” Escalation paths must be pre-wired and fast so that serious findings don’t stall behind new statements of work and long scheduling delays. Use tooling to package evidence—requests, responses, reproduction steps, and artifacts—so humans can decide quickly, keep decision ownership with people, and judge tools by guardrails and operational fit, not by how good the demo looks.
The outcome is a testing operating system that uses automation for scale and consistency while concentrating expert attention on the highest-risk attack paths.
Want to hear the full conversation between Braden and Julian, including where they push back against each other? Watch the full webinar on-demand here.