Two Clocks Started at Once
IAM disaster recovery is the ability to restore Okta or Microsoft Entra ID to a trusted, working state after a breach or misconfiguration, independent of the administrator credentials an attacker may already control. A patch closes the entry point. It doesn’t prove what changed while the door was open.
On August 28, a critical authentication flaw in a widely used software supply-chain platform was patched. By September 1, four days later, researchers saw active exploitation. Patching closes the entry point, but it doesn’t revoke tokens an attacker already created or prove which connected systems can still be trusted. Any team running Okta or Microsoft Entra ID inherits the same four-day reality. Acsense keeps identity configuration in a continuously captured, air-gapped, immutable copy outside the production blast radius, so recovery doesn’t depend on the authority an attacker just took over.
On August 28, the maker of a widely deployed software supply-chain platform released fixes for a critical authentication flaw. Under a default configuration, an unauthenticated person with network access could obtain administrative privileges. The timing was brutal, but familiar. Security teams had to identify exposed systems. Platform owners had to confirm versions. Changes had to be tested, approved, and scheduled. Some instances were sitting on the public internet. Others were buried inside networks where nobody expected them to become an urgent incident over a weekend.
Four days later, researchers were already seeing the flaw used against internet-facing systems. Attackers were creating administrator tokens and mapping users, groups, credential sets, and federation relationships. In a limited number of observed cases, they also created backdoor users. There was no evidence yet of broad mass exploitation, but that was hardly reassuring. It was a timestamp.
This wasn’t an ordinary application server. A software supply-chain platform stores and distributes the packages, binaries, container images, and other artifacts that engineering teams have already decided to trust. Administrative control of that system can become control over what gets built, shipped, and deployed. The flaw was an authentication bypass, not remote code execution. That distinction matters technically. Operationally, an attacker with administrator authority at this point in the chain may not need code execution on the repository itself. The platform already has the power to influence everything downstream.
“A patch closes the door. It doesn’t tell you who already walked through it,” said Max Fishman, Chief Product Officer at Acsense.
The first was the patch clock: how quickly could vulnerable systems be identified and upgraded? The second was the trust clock: how long had the platform been exposed, and what had changed while nobody was looking? Those clocks are easy to confuse. Patching is urgent because it blocks the known entry path. But an upgrade doesn’t revoke an administrator token created the day before. It doesn’t remove a backdoor account, restore a permission, or prove that an artifact wasn’t replaced. It doesn’t tell an incident commander which connected system should still be trusted.
That’s why the response to this incident can’t end with “we’re on the fixed version.” Teams also have to inspect audit history, rotate privileged credentials, review user and permission changes, investigate connected systems, and verify the integrity and provenance of what the repository distributed. The patch makes the software current. Recovery work makes the environment trustworthy again.
The Recovery Copy Has to Survive the Administrator
Once an attacker becomes an administrator, the question changes. It’s no longer only whether production is compromised. It’s whether the recovery path is governed by the same authority.
Imagine discovering that the same credentials used to alter production can also delete snapshots, rewrite the history used for investigation, or disable the job that creates future backups. There may be multiple copies, but they all share one failure domain. An attacker doesn’t need to destroy them immediately. They only need to make sure the next recovery attempt fails.
This is the practical value of an air-gapped and immutable copy. “Air gap” shouldn’t be reduced to a debate about tape versus cloud. The real separation is administrative. The recovery copy should sit behind different credentials, different control paths, and protections that prevent it from being altered or deleted during its retention period. CISA makes the same point in its ransomware guidance: keep critical backups offline or otherwise isolated, encrypt them, and test their availability and integrity in a disaster-recovery scenario.
The real air gap is not distance. It is authority.
A modern 3-2-1-1-0 strategy captures the idea well: three copies, on two different systems or media, with one offsite copy, one immutable or isolated copy, and zero unverified recovery errors. The final zero matters. A backup that has never been restored is still a promise, not evidence.
Why Not Just Keep Configuration in Git?
This is usually the next question, especially from engineering teams: if configuration is already versioned in Git, why add a separate recovery system?
Git is valuable. It’s excellent for expressing intended change, reviewing diffs, approving a pull request, and promoting a configuration between environments. Keep using it.
But an active Git repository isn’t the same thing as an isolated recovery copy. It may tell you what someone committed. It doesn’t automatically tell you what was actually running in a live service at 4:35 p.m., before the first unauthorized token appeared. An administrator, API client, automation pipeline, or compromised credential can change production directly and leave no matching commit.
There’s also the question of completeness. A configuration export may omit relationships, assignments, runtime state, secrets, or fields the platform doesn’t expose through an API. Reapplying a set of files can still fail because objects must be restored in a particular order, references must be rebuilt, and external systems may need to trust new endpoints or certificates.
And Git has its own control plane. If the same identity provider, administrator group, CI credentials, or cloud account governs both production and the repository, the two systems may share the same blast radius. Branch protections help. So do review rules and signed commits. But the recovery copy still needs to be moved somewhere the production administrator can’t casually rewrite. GitHub’s own backup guidance recommends creating a mirror, archiving it, and moving it to a separate location for safekeeping.
| Git Repository | Recovery System |
| Records intended change. What teams committed and reviewed. | Preserves recoverable state. What the protected system actually looked like at a point in time. |
| Optimized for collaboration. Branches, diffs, pull requests, promotion. | Optimized for restoration. Dependencies, ordering, rollback, and verification. |
| Often shares production identity. The same administrators or CI credentials may control both. | Uses an independent authority. Recovery data remains outside the production failure domain. |
Git is a record of intent. Recovery needs a record of reality.
See how Acsense keeps recovery independent of a compromised admin
Air-gapped, immutable identity backup for Okta and Microsoft Entra ID. Live, not slides.
Request a DemoFast Response Is a Design Decision, Not a Heroic Act
The four-day timeline is the other lesson. Most organizations can’t guarantee that every critical patch will reach every exposed system before the first exploit arrives. Asset ownership is messy. Maintenance windows are real. Testing still matters.
What they can do is shorten the time between an unauthorized change and a trusted correction. That requires more than an alert. The team needs a recent known-good state, a policy that defines what’s allowed, and a response path that was agreed before the incident.
The operating model is simple to say and hard to improvise under pressure: Observe. Assure. Act. Restore. Verify. Observe the change. Assure the resulting state against an approved baseline. Act to contain the exposure. Restore the last trusted configuration. Verify that the control and the business service are working again.
Automatic remediation belongs here, but not as a blanket promise that every change should be reversed without judgment. A clearly unauthorized addition to a protected administrator group may be a good candidate for immediate correction. A full tenant rollback carries a different blast radius and may need explicit approval. The principle is to pre-authorize the response where confidence is high and keep a human in the loop where the consequences are broad. Automation should remove waiting, not remove accountability.
There’s one more condition: automatic remediation is only as trustworthy as the state it restores. If the “known-good” source can be altered by the same account that caused the incident, automation may simply make a bad decision faster.
Protecting the System That Decides Who’s Trusted
The incident described here affected a software supply-chain platform, not an identity provider. Acsense wouldn’t have patched the underlying flaw or prevented the initial exploitation. The connection is architectural.
A software supply-chain system decides which code the organization trusts. Okta and Microsoft Entra ID decide which people, applications, and machines the organization trusts. In both cases, the control plane becomes dangerous when it’s also the only source of truth about its own state.
Acsense is built around a separate resilience layer for identity infrastructure. Supported identity configuration is captured continuously and stored outside the production tenant, with immutable and air-gapped protection. Changes are mapped to actor and before-and-after state, drift can be surfaced in ten minutes or less, and teams can use one-click rollback, object recovery, tenant-level recovery, or a standby environment depending on the scope of damage.
Recoverability also has to be proven. It’s not enough to show that a snapshot exists; the recovery path must be exercised and the result measured. And because users don’t experience recovery as an object in a console, the final step is restoring the application trust and access that the business depends on.
The next direction is to connect compliance and assurance policies to policy-driven remediation. A protected identity state drifts. The change is detected within the monitoring window. A control decides whether the response is automatic or approval-gated. The last approved state is restored, and evidence records what changed, how long the exposure lasted, and whether the environment returned to compliance.
That’s a more useful definition of fast response than “we generated an alert.” It’s the time from damaging change to trusted operation.
Five Questions to Ask Before the Next Friday Patch
- Admin compromise: if a production administrator is compromised, can that same authority alter or delete the recovery copy?
- Last trusted state: can you identify it, and explain why you trust it?
- Detection speed: how quickly can you detect drift, understand the blast radius, and begin a controlled rollback?
- Remediation policy: which high-confidence changes can be remediated automatically, and which must stay approval-gated?
- Tested recovery: have you tested the entire recovery path, including the applications and services that depend on it?
Four days isn’t a generous response window. It may now be the whole window. Attackers read advisories too, and they don’t wait for a change-control meeting.
Patch speed still matters. So do segmentation, credential rotation, artifact signing, provenance checks, and monitoring. But none of those controls removes the need for a clean way back when a trusted system has already been changed. The teams that recover well aren’t the ones that assume they’ll stop every exploit. They’re the ones that can tell the difference between the current state and the trusted state, and can move from one to the other before the incident becomes the business.
A patch can make the software current. Only an independent, tested recovery plan can make the business trustworthy again.
See Your IAM Disaster Recovery Gap Before an Attacker Does
Protect. Recover. Remain Operational. Across Okta and Microsoft Entra ID.
Request a Demo