Capabilities

Penetration Testing & Vulnerability Management

Agencies get more findings than they can act on. Testing earns its place by showing which exposures an attacker can actually reach, and proving it.

  • 17-21Months two high-risk vulnerabilities went unremediated against a 60-day policy
  • 4Risk criteria CISA’s 2026 directive uses to set remediation deadlines
  • Dec 2026Compliance date for CISA’s risk-scored remediation timelines
  1. Why the findings pile up

    Automated scanning across a federal enterprise produces findings continuously. Most are known CVEs matched against version numbers, and a scanner cannot establish which of them an attacker could actually reach in the environment as it is configured.

    Binding Operational Directive 26-04, issued by CISA in June 2026, pushes agencies toward the same question. It replaced the previous flat remediation deadline with a model that scores each vulnerability on four factors: public exposure, Known Exploited Vulnerability status, automatability, and technical impact.

    Those four factors measure whether something is reachable and worth an attacker’s effort rather than how severe it would be in isolation. Agency policies were required to support ongoing remediation from August 2026, with the scored timelines applying from December 2026.

  2. Where testing gets hard

    The findings that matter least are the easiest to produce. A scanner reports a missing patch. It does not report that a low-severity misconfiguration on one host, combined with a credential left in a configuration file on another, opens a route to a domain controller.

    Chained findings only surface when a tester works toward an objective across an environment rather than enumerating hosts one at a time, and that effort does not scale by adding more scanners to it.

    Testing also lands in environments that are already behind. GAO reported in December 2025 that the Department of Veterans Affairs had left two high-risk vulnerabilities open well past its own 60-day policy, with one estimated remediation date 15 months overdue, and attributed the delays to funding, staffing, and a system migration rather than to anything technical. A test delivered into that condition has to be precise about what deserves attention first.

  3. What AI changes in testing

    AI is genuinely useful here, and specifically in the parts of the work that were bounded by human reading speed. A model correlates findings across hosts, parses large scan outputs and configuration dumps, maps relationships between accounts and services, and drafts candidate attack paths for a tester to pursue.

    Reconnaissance and correlation compress as a result, and more of an engagement goes into confirming exploitability and chaining the findings that actually matter.

    It does not replace the tester. A model proposes a path and a person establishes whether it works, and nothing reaches a report without having been demonstrated in the environment.

    Models run inside FedRAMP-authorized environments or inside the customer accreditation boundary, so scan output, configuration detail, and anything a test discovers stay within the perimeter. A penetration test produces exactly the material an agency least wants leaving its network.

  4. How we work

    At OCH we scope against what the environment actually looks like, including the systems that are contractually in scope but have never been tested, which is usually where the results are.

    We agree how AI will be used in the rules of engagement before any testing starts. Where a customer approves it, the analysis runs inside a FedRAMP-authorized environment or inside their own accreditation boundary, and where they do not, the engagement runs without it.

    We prioritize findings by exploitability in the environment as configured, using the same factors CISA now applies, so a report arrives in the order the work should be done rather than in the order a scanner produced it.

    We deliver specific remediation guidance with every finding, written for the team that owns the system, with reproduction steps clear enough that a fix can be verified against them.

    We retest once the owning team has remediated and record what was confirmed closed, which is the evidence a plan of action and milestones needs.