KRBTGT rotation in Active Directory: operational runbook without outages
Contents
- KRBTGT account and W window
- When to rotate and detect compromise
- Prerequisites, multi-domain forests, and cloud
- Reset sequence and failure response
- Validation, telemetry, and change close
Purpose and scope
The krbtgt account: what it is and what not to do
krbtgt is the special built-in account created in every Active Directory domain. It is disabled for interactive sign-in and is not a user or service account intended to log on. Its password is not inactive, however: the Key Distribution Center (KDC) on Domain Controllers uses it to sign and encrypt the domain's Kerberos Ticket Granting Tickets (TGTs).
The account password and Kerberos keys are always managed by Active Directory. An operator can technically reset the password or supply an explicit password through authorized AD tools, but the change is still recorded and replicated by AD: password attributes, derived Kerberos keys, and the Key Version Number (KVNO) are updated. The KDC uses this replicated state, not a password stored locally by the operator.
The krbtgt account must never be deleted, renamed, enabled, moved, or manually modified. Its disabled state is by design; enabling it does not resolve Kerberos problems and unnecessarily increases attack surface. Although it is possible to set a manually chosen password, do not do so in this runbook: use only the controlled reset in the approved workflow, which protects the secret and records the double rotation.
An environment with Read-Only Domain Controllers (RODCs) also contains krbtgt_<number> accounts, which belong to RODCs. RODC krbtgt_<number> accounts are excluded from the standard domain rotation. Each such account is associated with one RODC and has a distinct Kerberos key used by that RODC for tickets it issues. It is not the domain's krbtgt: its handling needs an RODC-specific runbook assessing Password Replication Policy, possible RODC compromise, and credential-cache scope.
The following command is diagnostic only and inventories the accounts without modifying Active Directory:
$Domain = "contoso.com"
Get-ADUser -LDAPFilter "(sAMAccountName=krbtgt*)" -Server $Domain `
-Properties Enabled, PasswordLastSet, msDS-KeyVersionNumber, Description |
Select-Object SamAccountName, Enabled, PasswordLastSet, msDS-KeyVersionNumber, Description |
Sort-Object SamAccountName
Output should show the domain krbtgt account as disabled and, where RODCs exist, separate krbtgt_<number> accounts. Record the latter as excluded from the change scope and confirm with the AD team whether an RODC requires a separate response.
The krbtgt account password signs and encrypts Kerberos Ticket Granting Tickets (TGTs) issued by an Active Directory domain. Microsoft recommends rotating it every six months as a planned control; it is not maintenance to perform arbitrarily or while unrelated infrastructure change is still unsettled. When there is reason to believe the key has been exposed, rotation becomes a containment measure to start with incident response rather than waiting for the six-month cadence.
This runbook covers rotation for one writable Active Directory domain. Every domain in a forest has its own krbtgt account and needs a separate assessment and change. RODC accounts remain outside scope, as stated above.
The principle for the entire change is simple: the first reset introduces a new key while retaining the previous key to validate existing TGTs. The second reset, performed only after complete replication and a sufficient wait, removes the key that might be known to an attacker.
There is no normal rollback for a KRBTGT password. Setting the prior password again is a new password change in Active Directory: AD generates and replicates a new key state and increments the KVNO again. This action does not restore the prior historical keys, prior tickets, or prior replication state. The safe decision is to stop after the first reset if validation is not satisfactory. Restoring Active Directory is a forest-recovery operation, not a change rollback.
What the $W$ window means
In this runbook, $W$ is the safety interval between reset one and reset two of the KRBTGT password for one specific domain. It is not a generic maintenance window and it is not the total change duration. It starts only after reset one has been verified and replicated to every writable Domain Controller in that domain.
The $W$ window must be long enough for two conditions: the first change must reach every writable DC, and TGTs signed with the previous key must expire. Its value therefore cannot be set from the calendar alone; it depends on observed replication, effective MaxTicketAge, and an operational margin. The Safe timeline section calculates and approves the value; until that checkpoint, $W$ always means this mandatory wait, not an arbitrary duration.
When to start a rotation
Open a high-priority incident/change when one of the following conditions is confirmed or reasonably suspected:
- compromise of a Domain Controller or acquisition of
NTDS.dit, system-state backup, or DC LSASS memory dump; - theft of the KRBTGT key, suspected Golden Ticket activity, or unexplained TGT anomalies;
- unauthorized restore, improper cloning, or loss of control of a DC;
- direction from the incident-response team after a Tier 0 compromise.
The Microsoft-recommended six-month cadence can be adopted in organizational policy, but it does not replace hardening, monitoring, or incident response. Before every planned reset, verify that the environment is healthy and that the operational risk is accepted by the change owner. For a confirmed or suspected compromise, do not wait for the scheduled date: follow the containment sequence and incident-response team decisions.
Microsoft references: KRBTGT account password reset scripts, Kerberos policy settings, and AD forest recovery.
Detecting DC compromise, Golden Tickets, and unauthorized restores
KRBTGT reset must not be the only signal that something happened. There is no single Event ID that proves theft of the KRBTGT key: correlate signals from DCs, their hosting systems, services receiving tickets, and security telemetry. The objective is to distinguish an operational issue from possible Tier 0 access and preserve evidence before altering the environment.
Possible Domain Controller compromise signals
Treat the following events or patterns as indicators to investigate, not independent proof:
- anomalous administrative access to a DC, including 4624, 4672, 4648, and 4625, from unexpected hosts, accounts, or times;
- changes to privileged groups, audit policy, or sensitive AD objects, such as 4728/4729, 4732/4733, 4719, and 4662 where Directory Service Access auditing is configured;
- unexpected service, task, or process creation on a DC, including 4697, 7045, and 4688 where process command-line auditing is enabled;
- Security log clearing (1102), unexpected shutdowns/restarts, or loss of DC telemetry;
- replication errors, unexpected invocation-ID change, inconsistent metadata, or backup-restore signals in Directory Service logs and
repadmin; - EDR/Microsoft Defender for Identity alerts for DCSync, credential dumping, lateral movement, anomalous AD configuration change, or privileged activity on a DC.
Before rotating KRBTGT, the incident commander must decide which systems to isolate and which logs to preserve. Powering off, rebooting, or forcing replication on a suspected DC without forensic direction can destroy volatile memory or alter the incident timeline.
KRBTGT key theft and suspected Golden Tickets
Actual theft of the key is not normally visible as a dedicated Windows event. Instead, look for anomalies in TGT use: privileged accounts appearing on unusual hosts, ticket lifetimes or properties inconsistent with policy, service access without an expected 4768 in the correlation window, and Microsoft Defender for Identity or EDR alerts for suspicious Kerberos/Golden Ticket use.
Compare the domain's PasswordLastSet and KVNO with telemetry. A ticket associated with a previous key after double rotation completes, or persistent privileged-authentication patterns from unexpected endpoints, requires immediate triage. Do not infer key theft from 4769 alone or one failed application: always correlate user, client, DC, service, ticket lifetime, and security alerts.
Unauthorized restores and virtualization
An unauthorized DC restore can leave signals in Directory Service, System, backup-platform, and AD replication telemetry. Validate repadmin /showrepl, replication metadata, DNS/SYSVOL, and virtualization or cloud-platform logs for unplanned restore, snapshot revert, cloning, or VM replacement. When a DC is virtualized, also compare VM inventory, backup jobs, approved changes, and action owner.
VM-Generation ID reduces some virtualization risks, but does not make an unauthorized restore harmless. Treat it as an incident until the AD and IR teams establish replication consistency, backup provenance, and DC integrity.
Microsoft Sentinel, Defender, and Zabbix
Microsoft Sentinel can correlate DC Security Events, Microsoft Defender for Identity/Defender XDR alerts, and system logs. Create rules that compare the baseline with the period before and after reset and alert the SOC, not automated alerts that run further resets. Adapt table names to the connector actually configured in the tenant.
SecurityEvent
| where TimeGenerated > ago(4h)
| where EventID in (4624, 4625, 4672, 4719, 4768, 4769, 4771, 1102)
| summarize Events=count(), Accounts=make_set(Account, 20), Hosts=make_set(Computer, 20)
by EventID, bin(TimeGenerated, 15m)
| order by TimeGenerated desc
In Zabbix, collect Security/System/Directory Service events from every DC using an agent or Windows event-log item and add service, reachability, disk-space, DNS, and replication metrics exposed through read-only checks. Trigger on prolonged absence of data from a DC, a baseline-relative rise in 4771/4625, event 1102, replication errors, or AD DS/DNS/DFSR service stop. Thresholds must account for normal domain volume: an absolute count of 4768 or 4769 alone creates false positives during logon peaks.
Zabbix and Sentinel must refer to the same change ID and time baseline. Retain the alert graph or export in the record, with timezone, monitored DCs, and log source; that evidence is more useful than one isolated screenshot taken after an incident.
Prerequisites and go/no-go criteria
Before the change, appoint a change owner, an AD operator, a monitoring/SOC owner, an escalation owner for critical applications, and a person authorized to stop the second reset. Keep names, contact details, and the approved window in the change record.
Technical checklist
- Current inventory of every writable DC, AD site, OS version, and FSMO role.
- AD and SYSVOL replication health confirmed, with no current errors or unexplained backlog.
- A recent, restorable System State backup tested under the disaster-recovery plan. A backup is not immediate rollback.
- No DC is in maintenance, promotion/demotion, migration, snapshot restore, or authoritative restore.
- DNS, time service, and site connectivity work; Kerberos clock drift remains within the organizational limit.
- Effective Kerberos policy durations documented:
MaxTicketAge,MaxRenewAge,MaxServiceAge, and clock skew. - Health checks and test accounts defined for VPN, VDI, file servers, integrated web applications, SQL, and Tier 0 services.
- Change freeze, backup window, and critical jobs excluded from the window; service desk and owners have a communication plan.
- Delegated privileges and a Tier 0 administrative workstation are available; do not use an untrusted endpoint for an administrative session.
Pre-change health checks
Run these from an administrative workstation with RSAT and save the output in the change record. Replace contoso.com with the target domain.
$Domain = "contoso.com"
Get-ADDomainController -Filter * -Server $Domain |
Select-Object HostName, Site, IsGlobalCatalog, OperationMasterRoles
repadmin /replsummary
repadmin /showrepl * /errorsonly
dcdiag /e /test:Advertising /test:Replications /test:SysVolCheck /test:DNS
Get-ADUser krbtgt -Properties PasswordLastSet, msDS-KeyVersionNumber |
Select-Object SamAccountName, PasswordLastSet, msDS-KeyVersionNumber, Enabled
repadmin /replsummary must be clean. One replication error, an unreachable DC, or SYSVOL not being ready is a no-go: remediate the cause first. Record the initial PasswordLastSet and msDS-KeyVersionNumber; the KVNO must advance at every reset.
Capture the effective Kerberos policy on the PDC Emulator:
$Pdc = (Get-ADDomain -Server $Domain).PDCEmulator
Get-ADDefaultDomainPasswordPolicy -Server $Domain |
Format-List MaxTicketAge, MaxServiceAge, MaxRenewAge
Get-GPResultantSetOfPolicy -ReportType Html -Path .\kerberos-rsop.html
Get-ADDefaultDomainPasswordPolicy does not replace RSOP when a GPO configures Kerberos settings. Use the RSOP report and GPMC to determine the effective values applied to DCs.
Prerequisite and permissions matrix
The table separates rights required to read, validate, and modify. Do not use Enterprise Admins as the default choice: rotation is a per-domain operation and needs privileges in the target domain. By default, a member of the target domain's Domain Admins has the necessary rights; a least-privilege model must be delegated and tested before the window.
| Activity | Minimum permission or role | Pre-change proof | Security note |
|---|---|---|---|
Read DCs, krbtgt, and policy | AD read access and RSAT ActiveDirectory | Run baseline commands without errors | Use a separate account where the model requires read-only |
Reset the krbtgt password | Reset Password extended right delegated on the krbtgt object, or Domain Admins in the target domain | Test delegation in lab or through an approved authorization check | Do not use Enterprise Admin credentials for convenience |
Force/analyze replication and run dcdiag | Administrative privileges required on DCs and RSAT/AD DS tools | Run repadmin and dcdiag across target scope | Do not bypass UAC or use shared credentials |
| Read DC Security/System logs | Event Log Readers membership or equivalent delegation on every DC | Query 4768/4769/4771 and System events | SOC can have read access without reset rights |
| Test applications | Least-privileged accounts authorized by the owner | New logon and defined service test | Do not use a Domain Admin as the only test |
| Approve reset two | Authorized change owner | Recorded go/hold/no-go decision | Approval does not replace technical checks |
When using custom delegation, verify effective rights before the window with an account that is not globally privileged. The test must not reset krbtgt in production: validate the object ACL, delegated group, and process in a lab domain or through the approved access check. Retain the access request, group used, elevation expiry, and operator identity in the change record.
The operating account must be protected by MFA/PIM or an equivalent control, used only from the designated PAW, and removed from the session after work completes. When just-in-time elevation is used, activate it early enough to validate replication, logs, and tools, but do not keep it active unnecessarily until reset two.
Multi-domain forests: planning and sequence
In a multi-domain forest, there is no shared krbtgt. Every writable domain, including the forest root domain, has its own krbtgt account, Kerberos policy, DCs, sites, and tickets. Resetting krbtgt in child.contoso.com does not rotate contoso.com or another child domain.
Treat the work as a coordinated per-domain change program, not one forest-wide command. Each domain needs its own scope, owner, operating account, baseline, $W$, replication evidence, local testing, and checkpoint. The master record must link every child change and stop the sequence if one does not reach its expected go/no-go state.
Inventory and decision order
Before choosing an order, collect topology and trusts from the forest root using appropriate read privileges:
$Forest = Get-ADForest -Server "contoso.com"
$Forest.Domains | ForEach-Object {
$Domain = $_
$DomainInfo = Get-ADDomain -Server $Domain
[pscustomobject]@{
Domain = $Domain
PdcEmulator = $DomainInfo.PDCEmulator
DomainMode = $DomainInfo.DomainMode
WritableDcCount = (Get-ADDomainController -Filter * -Server $Domain).Count
}
} | Format-Table -AutoSize
Get-ADTrust -Filter * -Server "contoso.com" |
Select-Object Name, Direction, TrustType, ForestTransitive, SelectiveAuthentication |
Format-Table -AutoSize
This inventory does not automatically select the sequence. Order must follow blast radius, service dependencies, and owner availability. Where no documented application constraint exists, a prudent practice is to work on lower-criticality child domains first and schedule the forest root only after validating process, telemetry, and communication. When suspected compromise affects the forest root or multiple domains, the incident commander and forest-recovery plan must direct the sequence rather than a generic rule.
Per-domain procedure
- Select the target domain and use credentials valid in that domain; do not assume root-domain access allows reset in a child domain.
- Run every pre-change check with
-Server <target-domain>and identify its PDC Emulator. - Validate critical intra-domain and cross-domain flows with agreed accounts and services.
- Perform reset one only in the target domain, validate replication across all of its writable DCs, and start that domain's $W$ timer.
- During $W$, do not overlap another rotation in a domain serving the same critical services unless IR explicitly directs and tests it.
- At the checkpoint, perform reset two for that domain only if its criteria are met; complete validation before starting the next domain.
Serialization is preferred where the forest has few owners, complex trusts, or applications spanning domains. Parallel changes are possible only for independent domains with separate teams, separate telemetry, and a decision that accepts cumulative blast radius. Do not use parallelism to recover planning delay.
Cross-domain and trust checks
For each business-critical trust, define a matrix containing source domain, resource domain, test account, service SPN/FQDN, owner, and expected result. Establish a baseline before the first domain and repeat the test after reset two in every domain participating in the flow.
| Flow | Baseline | After source domain | After resource domain |
|---|---|---|---|
| Child user -> root file server | New logon and SMB over FQDN | Successful TGT and access | Successful service ticket and access |
| Root user -> child IIS | Integrated sign-in, no prompt | Consistent 4768/4769 | HTTP ticket and application response |
| Application account -> SQL in another domain | Integrated health query | No trust/pre-auth error | Successful query and scheduled job |
When a trust uses selective authentication, also verify Allowed to authenticate on target resources; do not attribute an authorization denial to KRBTGT reset without reading the event and trust configuration. An external trust is not "covered" merely because intra-domain replication is healthy.
Cloud servers and services with persistent tickets
Azure, AWS, or GCP IaaS VMs that are domain-joined are Kerberos member servers, but can have different operating cycles: autoscaling, golden images, always-on services, host maintenance, site-to-site networking, and processes keeping tickets or connections for a long time. A TGT should renew according to policy, but do not assume every long-running service refreshes its cache, re-runs logon, or reopens a Kerberos connection at the expected time.
For these systems, a controlled reboot or service restart is often the most reliable way to obtain a new Kerberos session. It is not an automatic consequence of reset, must not be run indiscriminately, and does not replace ticket, log, and health checks.
Cloud pre-change inventory
Classify every domain-joined cloud server using the following categories and define action, owner, and window:
| Category | Examples | Change action | Success criterion |
|---|---|---|---|
| Stateless behind load balancer | Web tier, APIs, replicated workers | Drain, rolling reboot, and re-enable health probe | New instance healthy and SSO valid |
| Stateful | SQL, file server, middleware, license server | Restart only with product runbook and tested failover | Consistent data and application health check |
| Autoscaling/image-based | VMSS, ASG, MIG, golden image | Pause automatic rollout; validate join/bootstrap of new instances | New instances domain-joined and reachable |
| Management/jump host | Bastion, cloud PAW, management server | New session and planned reboot if it retains tickets | Admin tools and secure channel valid |
| Cloud Domain Controller | IaaS DC in an AD subnet | Do not reboot merely to "refresh" without replication and DC change assessment | Replication, DNS, and SYSVOL healthy |
Record provider, subscription/account/project, region, subnet, criticality tag, availability set/zone, access method, owner, and on-premises dependencies. A cloud server unreachable during the window is a validation gap, not a reason to ignore it.
Safe sequence for cloud member servers
- Pause deployment, autoscaling replacement, and patch automation that could introduce new instances while collecting baseline.
- Validate VPN/ExpressRoute/Direct Connect, internal DNS, NTP, and connectivity to at least two DCs for the correct domain from the cloud subnet.
- Run application health checks and validate the secure channel before reset one.
- After reset one and replication, validate tickets and service; schedule a service restart or rolling reboot where the process does not renew its session or the vendor requires it.
- For load-balanced pools, drain, reboot one instance, validate health and SSO, then proceed to the next instance.
- After reset two, repeat the cold test; re-enable automation only after new and existing instances are compliant.
Example remote check for an authorized member server. It does not restart the server; it collects the evidence used to decide whether a restart must be planned.
$CloudMemberServer = "app-az-01.contoso.com"
Invoke-Command -ComputerName $CloudMemberServer -ScriptBlock {
hostname
Test-ComputerSecureChannel -Verbose
nltest /sc_verify:contoso.com
klist tickets
w32tm /query /status
Resolve-DnsName contoso.com
}
Before a reboot, the owner must confirm completed drain, application backup/snapshot according to policy, no critical job, out-of-band console method, and return to the load balancer. For databases, clusters, servers with shared disks, or Domain Controllers, use the product runbook: generic Restart-Computer is not a failover procedure.
A server that does not recover after restart can indicate DNS, routing, secure channel, GPO, certificates, or application dependencies. Keep the instance out of the pool, collect logs, and open an incident with the owner; do not repeat KRBTGT resets as a repair attempt.
Safe timeline
The pause between the two resets must allow every DC to receive the first change and TGTs signed with the previous key to expire. For the default configuration, Microsoft specifies at least 10 hours. In production, use a calculated, approved duration rather than assuming the default.
Define:
$$W = \max(\text{observed full replication}, \text{effective MaxTicketAge}) + \text{operational margin}$$
The margin must cover inter-site propagation, monitoring delay, and manual validation. If MaxTicketAge is 10 hours, a common practical window is at least 12 hours after confirming the first reset replicated. If policy or topology requires a longer duration, the longer value wins.
| Phase | Action | Required exit |
|---|---|---|
| T-7 days to T-1 | Baseline, tests, approvals | All go criteria met |
| T0 | First reset | KVNO increased and replication complete |
| T0 + W | Telemetry review and second reset | No unresolved incident; explicit approval |
| T0 + W + replication | Extended validation | Health checks and logs within baseline |
| T0 + W + 1-7 days | Heightened monitoring | Evidence retained and change closed |
Do not compress the phases to finish on the same day. The checkpoint between resets is the operational boundary that prevents a security suspicion becoming a domain-wide outage.
Execution: first reset
Use the Microsoft KRBTGT Reset script in the organization-approved and verified version, ideally in a mode that identifies the writable DC and records the operation. Read the instructions for the downloaded version before use and have security review it.
If the approved process uses native AD tools, the equivalent command is:
$Domain = "contoso.com"
$Before = Get-ADUser krbtgt -Server $Domain -Properties PasswordLastSet, msDS-KeyVersionNumber
$Before | Select-Object SamAccountName, PasswordLastSet, msDS-KeyVersionNumber
Set-ADAccountPassword -Identity krbtgt -Server $Domain -Reset
Get-ADUser krbtgt -Server $Domain -Properties PasswordLastSet, msDS-KeyVersionNumber |
Select-Object SamAccountName, PasswordLastSet, msDS-KeyVersionNumber
Do not set a manually chosen password, disable/rename/delete krbtgt, or run concurrent resets from separate consoles. The secret must be generated and protected by the approved workflow.
Immediately force and validate replication. The msDS-KeyVersionNumber, PasswordLastSet, and related attributes must be consistent on every writable DC:
repadmin /syncall $Pdc /AdeP
repadmin /replsummary
repadmin /showrepl * /errorsonly
Get-ADDomainController -Filter * -Server $Domain | ForEach-Object {
Get-ADUser krbtgt -Server $_.HostName -Properties PasswordLastSet, msDS-KeyVersionNumber |
Select-Object @{Name="DomainController";Expression={$_.PSComputerName}},
PasswordLastSet, msDS-KeyVersionNumber
}
If the environment does not populate PSComputerName, label each result with the DC queried. Do not proceed until every writable DC reports the expected KVNO and repadmin is clean.
Observation interval and second-reset decision
During $W$, do not alter Kerberos policy, DNS, trusts, authentication services, or DC topology except for unrelated incidents. Any such intervention makes diagnosis ambiguous.
The SOC should look for invalid Kerberos tickets and authentication anomalies on DCs and critical services. Compare counts and rates, not merely one event, with a baseline for the same day and time of week.
| Source | Signals to monitor | Decision |
|---|---|---|
| DC Security log | 4768, 4769, 4771, 4776; failure-code spikes and unusual clients | Triage with owner before reset two |
| DC System log | KDC, Netlogon, Time-Service, DFS Replication | Stop if replication, DNS, or time degrades |
| Critical services | 4625, SSO errors, HTTP 401/5xx, VPN/VDI logons, failed jobs | Isolate application/account and test with owner |
| SIEM/EDR | Golden Ticket suspicion, anomalous DC access, privileged activity | Escalate to incident response |
The second reset requires explicit change-owner approval only when the first reset has fully replicated, $W$ has ended, critical tests have passed or their risk is accepted, and no change-related incident remains open.
Execution: second reset
Repeat the same approved method once for the same domain. Record the timestamp, operator, contacted DC, before/after KVNO, and change number.
$BeforeSecondReset = Get-ADUser krbtgt -Server $Domain -Properties PasswordLastSet, msDS-KeyVersionNumber
$BeforeSecondReset | Select-Object SamAccountName, PasswordLastSet, msDS-KeyVersionNumber
Set-ADAccountPassword -Identity krbtgt -Server $Domain -Reset
repadmin /syncall $Pdc /AdeP
repadmin /replsummary
repadmin /showrepl * /errorsonly
Confirm that the KVNO has increased a second time across all writable DCs. Repeat the tests and maintain heightened telemetry. Do not perform a third reset as a corrective action; engage the AD recovery and incident-response owners.
When reset two fails or causes impact
An error returned by the command does not prove that the reset was not applied. Before running any command again, capture timestamp, complete error, contacted DC, PasswordLastSet, and KVNO from the PDC Emulator and every writable DC. The decision depends on observed state, not only the console exit code.
| Observed state | How to verify | Immediate action |
|---|---|---|
No KVNO or PasswordLastSet changed | Per-DC query, repadmin /replsummary, Security/Directory Service logs | Do not blindly retry; correct permissions, connectivity, or workflow error and reassess the change |
| PDC has new KVNO, one or more DCs do not | Per-DC query, repadmin /showrepl * /errorsonly | Treat as replication incident; stop destructive testing and do not run reset three |
| KVNO increased everywhere, but services fail | New logon, klist, DC 4768/4769/4771, target 4625 and application logs | Open incident bridge, stop unnecessary rollout/restarts, and isolate service/account/host |
| Authentication fails across sites or Tier 0 | Sentinel/Zabbix dashboard, DC health checks, DNS, time, replication | Escalate to incident commander and AD recovery owner; do not attempt rollback with prior password |
For a quick state check after an ambiguous result, reuse the Per-domain-controller replication validation script and save output with a new timestamp. Then perform cold tests with a standard account and dedicated administrative account, without purging tickets on unrelated production hosts.
Correlate these events during triage: 4768 for TGT requests, 4769 for service tickets, 4771 for pre-authentication failures, and 4625 for failed logons on targets, together with KDC, Netlogon, DNS, DFS Replication, and Time-Service logs. Event presence is not enough: compare failure code, account, client, DC, service, and baseline-relative rate. If the Security log was cleared (1102), telemetry disappears, or DC/Golden Ticket alerts appear, preserve evidence and move to the IR process.
When reset two is confirmed and impact is limited to particular applications, remediate the dependency: SPN, DNS, clock skew, secure channel, trust, service, or client. When impact is broad, stabilize replication, DNS, and time with AD owners first; recovery strategy must be defined by the forest-recovery plan. Reusing a prior password or attempting reset three cannot return the domain to its prior state.
Post-change validation
Validate using least-privileged test accounts and representative paths. Do not use only an administrative console: a session that is already authenticated does not demonstrate ticket renewal for users.
Authentication tests
- On a test workstation, run
klist purge, lock/unlock, or perform a controlled new logon. - Run
klist get krbtgtand confirm a new TGT is obtained. - Access the services defined in the plan: a file share over FQDN, IIS application using Windows Integrated Authentication, SQL using Integrated Security, VPN/VDI, and Tier 0 administration tools.
- Confirm the DC records new successful 4768 and 4769 events and targets record successful 4624 logons.
- Repeat from every material AD site and with at least one non-administrative user.
klist purge
klist get krbtgt
klist tickets
Test-ComputerSecureChannel -Verbose
nltest /sc_verify:contoso.com
Test-ComputerSecureChannel and nltest validate the computer secure channel; they do not replace a Kerberos application test. Correlate a service failure with SPN, DNS, service account, trust, and target logs before automatically attributing it to the rotation.
Initial telemetry query
Adapt these queries to the SIEM and retain results with a declared time range and timezone:
Get-WinEvent -FilterHashtable @{LogName='Security'; Id=4768,4769,4771; StartTime=(Get-Date).AddHours(-4)} |
Group-Object Id | Select-Object Name, Count
Get-WinEvent -FilterHashtable @{LogName='System'; StartTime=(Get-Date).AddHours(-4)} |
Where-Object ProviderName -match 'Kerberos|KDC|Netlogon|DFSR|Time-Service' |
Select-Object TimeCreated, ProviderName, Id, LevelDisplayName, Message
Event IDs cover only part of the flow and cannot alone establish a successful change. The useful measure is no abnormal rise in failure codes together with successful application tests and healthy replication.
Failure response and operational rollback
| Timing | Condition | Immediate action |
|---|---|---|
| Before reset one | Replication, DNS, or backup is non-compliant | No-go; remediate and reschedule |
| After reset one, before reset two | Kerberos failures or critical-service degradation | Stop the plan, do not perform reset two, open an incident bridge, and collect logs |
| After reset two | Widespread authentication impact | Engage incident response and AD recovery owners; stabilize DCs, replication, DNS, and time; do not blindly reset KRBTGT again |
| At any point | Evidence of active DC/Tier 0 compromise | Follow the IR plan, isolate based on forensic direction, and assess forest recovery |
Operational rollback before the second reset means not proceeding: the prior key remains the KDC historical key, and already-issued TGTs remain valid until their normal expiry. After the second reset, remediation depends on the root cause. Changing the password again does not restore old tickets and complicates containment.
A System State or DC restore without a forest-recovery plan can introduce USN rollback, replication divergence, or reinfection. Only the approved recovery plan, with the appropriate roles and authority, can determine whether restoration is justified.
Change-close evidence
Retain or link the following in the change record:
- domain scope, trigger, approvals, operators, and complete timeline;
- pre/post outputs from
repadmin,dcdiag, DC inventory, Kerberos policy, and RSOP; - KRBTGT timestamp and KVNO before, after reset one, and after reset two;
- proof of replication to every writable DC and notes on excluded RODCs;
- application-test results, test accounts, and owner approval;
- SIEM dashboard with baseline, post-change interval, alerts, and triage;
- close decision, residual risks, and post-change review date.
Errors to avoid
- performing two resets close together before every DC replicated the first;
- assuming 10 hours is always enough without reading effective policy;
- including RODCs or other domains without a specific assessment;
- treating SPN, DNS, clock-skew, or trust errors as inevitable effects of rotation;
- using the reset as a substitute for handling a compromised DC;
- declaring success from KVNO alone without testing real authentication;
- treating a backup as immediate undo or performing uncoordinated restores.
Operating model, roles, and communication
A KRBTGT rotation is not only a technical activity. A reset can be successful in AD and still create a user-visible incident if communication, diagnostics, and application owners are not ready. Define roles before the start and do not assign the same person as operator, approver, and validator for every critical service.
| Role | Responsibility before the change | Responsibility during and after the change |
|---|---|---|
| Change owner | Approves scope, risk, $W$ duration, and communication plan | Makes the go/no-go decision for reset two and closure |
| AD operator | Performs baseline, resets, and replication validation | Retains output and coordinates technical checks |
| SOC/monitoring | Prepares baseline, alerts, and dashboard | Triages alerts, correlates failure codes, and escalates |
| Application owner | Defines realistic tests and authorized test accounts | Runs tests, classifies impact, and signs off results |
| Service desk | Receives expected impact and escalation criteria | Collects reports and links them to the change |
| Incident commander | Confirms the IR path when compromise is suspected | Coordinates containment and forensic decisions |
Minimum communication
The notice must not ask users to "try whether it works." It must state the service, observation period, escalation channel, and information needed to reproduce a problem.
Send at least these messages:
- Owner pre-notice: affected domain, window, systems to test, approved test accounts, and expected time for reset two.
- Service-desk notice: symptoms to associate with the change, such as SSO errors, repeated credential prompts, rejected VPN, or denied SMB access; collect username, client host, destination, time, and screenshot/error code.
- SOC update: queries, alert thresholds, DC scope, and an immediate escalation channel.
- Owner closure: outcome, completed tests, residual issues, monitoring period, and contact for delayed regressions.
Do not communicate password values, forensic details, hashes, tickets, or sensitive indicators in a ticket readable by non-Tier 0 groups. Technical data required by the SOC belongs in the approved case location.
Escalation criteria
Treat an error that affects multiple sites, a central authentication service, break-glass accounts, a DC secure channel, or a rapid rise in Kerberos failures as high priority. A single failed application account may be a local issue, but it still needs an owner, timestamp, responding DC, and confirmation test before proceeding.
Severity is not based only on the number of failed tickets. A failure involving corporate VPN, a production platform, healthcare service, industrial control, or administrator access can be critical even with few events. Add this classification to the change record before the reset.
Domain inventory and repeatable evidence collection
Use a protected evidence folder with restricted access. The folder name should contain the domain, change ID, and timestamp, but no sensitive incident details. The following example collects a readable baseline; it does not modify Active Directory.
$Domain = "contoso.com"
$ChangeId = "CHG-000000"
$Timestamp = Get-Date -Format "yyyyMMdd-HHmmss"
$EvidencePath = Join-Path $env:USERPROFILE "Documents\KRBTGT-$Domain-$ChangeId-$Timestamp"
New-Item -ItemType Directory -Path $EvidencePath -Force | Out-Null
$Pdc = (Get-ADDomain -Server $Domain).PDCEmulator
Get-ADDomain -Server $Domain | Format-List * |
Out-File (Join-Path $EvidencePath "domain.txt")
Get-ADDomainController -Filter * -Server $Domain |
Select-Object HostName, Site, IPv4Address, IsGlobalCatalog, OperationMasterRoles |
Export-Csv (Join-Path $EvidencePath "domain-controllers.csv") -NoTypeInformation
Get-ADUser krbtgt -Server $Domain -Properties PasswordLastSet, msDS-KeyVersionNumber, whenChanged |
Select-Object SamAccountName, PasswordLastSet, msDS-KeyVersionNumber, whenChanged, Enabled |
Export-Csv (Join-Path $EvidencePath "krbtgt-before.csv") -NoTypeInformation
repadmin /replsummary | Out-File (Join-Path $EvidencePath "repadmin-replsummary-before.txt")
repadmin /showrepl * /errorsonly | Out-File (Join-Path $EvidencePath "repadmin-showrepl-before.txt")
dcdiag /e /q | Out-File (Join-Path $EvidencePath "dcdiag-before.txt")
When finished, protect the folder with ACLs consistent with the change record and upload only approved artifacts to the change system. repadmin, dcdiag, and exports can disclose host names, sites, and topology; do not attach them to public channels or tickets visible to unauthorized users.
DC and site map
For each writable DC, record AD site, subnet, connectivity, inbound/outbound replication, DNS state, and local owner. A hub-and-spoke topology requires special care: a reset performed against the PDC Emulator does not prove that a remote site received the change.
For remote sites with intermittent links, define a verification time and an on-call contact. Do not bypass the problem by removing the DC from the inventory or performing reset two without replication evidence. An out-of-sync DC can continue to issue or validate tickets inconsistently with the domain.
If a Domain Controller was decommissioned but remains in metadata, treat it as an AD defect to resolve before the change. Do not use a KRBTGT reset to "clean up" replication or stale objects.
Trusts, domains, and external identities
List forest trusts, external trusts, shortcut trusts, and applications using identities from other domains. KRBTGT password rotation is domain-scoped, but cross-domain authentication can fail or change its path when pre-existing DNS, trust, selective-authentication, SID-filtering, or time issues exist.
Before the window, run a logon and service-access test for every business-critical trust. Retain the result as baseline. If a cross-domain test is unavailable, record the exception, risk, and explicit approval; do not infer that the trust is healthy from internal replication alone.
Entra accounts, cloud-only identities, and applications that do not use domain Kerberos are not valid KRBTGT tests. Include them only where AD DS is also a dependency, for example through AD FS, Seamless SSO, LDAP, VPN using RADIUS, or integrated Windows services.
Calculating and approving the W window
The $W$ formula must become a readable decision in the change record, not remain theoretical. Record the effective MaxTicketAge, the longest observed inter-site replication time, the selected margin, and the person who approved the total.
| Element | Example | Evidence source |
|---|---|---|
MaxTicketAge | 10 hours | GPO/RSOP applied to DCs |
| Slowest observed replication | 45 minutes | repadmin, monitoring, and site link |
| Operational margin | 2 hours | Change owner and SOC |
| Approved $W$ | 12 hours | Change record |
Do not use MaxRenewAge as the time between resets. The KDC retains two KRBTGT keys to handle TGT transition, while the practical threshold must allow tickets signed with the previous key to expire and the first change to replicate. If local policy or business requirements have greater values, document the more conservative value.
When the window crosses a date boundary, timezone, or reduced-staffing period, record times in UTC and the local time used by the operations team. This prevents reset two being performed before the actual wait because of a wrong conversion.
Checkpoint before reset two
The checkpoint must be a meeting or formal message, not an implicit operator decision. The change owner receives at least:
- time of reset one and earliest allowed time for reset two;
- replication result for every DC and every resolved exception;
- comparison of Kerberos failure rate before and after;
- priority test results by site and application;
- open service-desk tickets, impact, and owner;
- explicit recommendation: go, hold, or no-go.
Hold is not failure. It is the correct choice when evidence cannot separate change impact from a pre-existing problem. During a hold, retain evidence, clarify the cause, and reassess the window without resetting again.
Per-domain-controller replication validation
A concise report helps prevent selective reading of command output. The following script queries every writable DC and records the observed KVNO. Run it before reset one, after reset one, and after reset two. Adapt remote authentication to organizational policy and do not include modification commands in the loop.
$Domain = "contoso.com"
$ExpectedKvno = (Get-ADUser krbtgt -Server $Domain -Properties msDS-KeyVersionNumber).msDS-KeyVersionNumber
$Results = foreach ($DomainController in Get-ADDomainController -Filter * -Server $Domain) {
try {
$Krbtgt = Get-ADUser krbtgt -Server $DomainController.HostName `
-Properties PasswordLastSet, msDS-KeyVersionNumber, whenChanged -ErrorAction Stop
[pscustomobject]@{
DomainController = $DomainController.HostName
Site = $DomainController.Site
PasswordLastSet = $Krbtgt.PasswordLastSet
Kvno = $Krbtgt.'msDS-KeyVersionNumber'
ExpectedKvno = $ExpectedKvno
MatchesExpected = $Krbtgt.'msDS-KeyVersionNumber' -eq $ExpectedKvno
Status = "QuerySucceeded"
}
} catch {
[pscustomobject]@{
DomainController = $DomainController.HostName
Site = $DomainController.Site
PasswordLastSet = $null
Kvno = $null
ExpectedKvno = $ExpectedKvno
MatchesExpected = $false
Status = $_.Exception.Message
}
}
}
$Results | Sort-Object Site, DomainController | Format-Table -AutoSize
if ($Results.Where({ -not $_.MatchesExpected }).Count -gt 0) {
throw "Stop: one or more writable DCs do not report the expected KRBTGT KVNO."
}
A query error is not a neutral result: it is missing evidence. When the command fails against a DC, validate DNS, connectivity, privileges, and replication state. Only after excluding a technical false negative can the change owner decide whether the situation requires a separate recovery window.
The KVNO compared must be observed from the domain after the reset, not a value manually written in advance. This also detects a reset directed at the wrong domain or a command failure that went unnoticed.
Functional test matrix
A test needs an owner, account, destination, and expected result. "Successful login" is not sufficient for an application using service tickets with specific SPNs or delegation.
| Service | Account and origin | Action | Success evidence | Owner |
|---|---|---|---|---|
| File server | Standard user, site A | Open \\server.domain\share over FQDN | Access and target Kerberos 4624 | Infrastructure |
| IIS SSO | Standard user, site B | Open integrated HTTPS URL | No prompt; HTTP ticket in klist | Application |
| SQL | Authorized application account | Integrated Security connection | Successful health query | DBA |
| VPN/VDI | Non-privileged test user | New authentication | Session starts; no credential loop | Workplace |
| Tier 0 | Dedicated admin, PAW | Console and admin tools | Valid TGT and administration service | AD owner |
Perform a cold test where possible: clear local tickets, start a new session, and use a supported DNS path. A test from a browser or client with a long-running session can hide failure to request a new TGT.
For IIS, SQL, and file servers, record expected SPNs, their owning account, and the DNS name used by the test. This speeds triage: KDC_ERR_S_PRINCIPAL_UNKNOWN tends to identify an SPN/naming issue, not a reason to repeat the KRBTGT reset.
Kerberos failure triage
The table does not replace official documentation or full-event analysis, but it supplies a first classification to avoid destructive actions during the window.
| Signal | Initial hypothesis | First check | Prohibited action |
|---|---|---|---|
| Rising 4771 across many clients | Password, pre-authentication, clock skew, or stale ticket | Failure code, DC/client time, affected account | Reset KRBTGT again |
| Missing 4768 from an entire site | DNS, network, unreachable DC, or logging | DC reachability, DNS, local Security log | Declare success from PDC only |
| HTTP 401 after new logon | HTTP SPN, app pool, or browser zone | klist, SPN, IIS log, app pool | Change Kerberos policy without owner |
| SMB/SQL fails only via alias | CNAME/SPN or inconsistent naming | Canonical FQDN and setspn -Q | Classify it as KRBTGT outage |
| Intermittent errors between sites | Replication, time, site-aware DNS, or WAN | repadmin, w32tm, DNS, site link | Run reset two before triage |
Run setspn -Q only for the expected service name and retain the output in the change record. Do not modify SPNs mid-window without application owner and a separate plan: an unrelated correction destroys the ability to attribute cause and effect.
$ServiceName = "HTTP/app.contoso.com"
setspn -Q $ServiceName
klist purge
klist get $ServiceName
klist tickets
w32tm /query /status
Resolve-DnsName app.contoso.com
Expected output is a service ticket for the queried SPN, consistent DNS resolution, and a healthy time source. If the SPN query reports duplicates or no owner, open a separate remediation change and do not alter the object improvisedly during the rotation.
Heightened monitoring and review
Keep heightened monitoring for at least one representative operating cycle: include overnight jobs, backups, start-of-day activity, batch processing, and remote-site users. In environments with monthly or seasonal processes, state the limit of the observed period and schedule a review after that process.
Every alert should include at least DC, client or target, account, Event ID/failure code, UTC timestamp, service, and change correlation. Count successes and failures separately because legitimate volume growth can make an absolute count misleading.
Post-closure review
At review, confirm that:
- no DC returned to replication error or showed unexpected KVNO changes;
- rates of 4768, 4769, 4771, 4624, and 4625 are explained and near baseline;
- owners completed deferred tests;
- service-desk tickets correlated to the change are closed or have a remediation plan;
- the IR team retained indicators, decisions, and follow-ups from the initial compromise.
Rotation removes the validity of the prior KRBTGT key after the correct sequence; it does not prove an attacker has been removed from DCs. Continue containment, threat hunting, privileged review, and recovery activities defined by the incident-response plan.
Final checklist
- Domain-specific scope and trigger approved.
- AD/DNS/SYSVOL health and backup verified before reset one.
- Reset one executed and KVNO replicated to every writable DC.
- $W$ completed with telemetry and tests within baseline.
- Reset two approved, executed, and replicated.
- Kerberos, critical-service, and AD-site tests completed.
- Evidence retained, heightened monitoring active, and change closed with owner approval.
The contents of this guide are provided for informational purposes only, without warranties. Application of any procedure is at the user's own risk. Disclaimer.
Appreciation
If this guide is useful, leave a like.
Related guides
Keep exploring
Active Directory / Authentication
NTLM Audit in Active Directory: From Audit to Enforcement
Read the guide->Active Directory / Domain Controllers
Protected Users in Active Directory: what it is, limits, adminCount, and rollout without lockout
Read the guide->Active Directory / Domain Controllers
Active Directory RC4 remediation: service accounts from RC4 to AES without password reset
Read the guide->Active Directory / Hardening