Skip to main content

Document metadata

Status
Maintained
Approval
Approved
Version
1.0
Classification
PUBLIC
Owner
Lightning IT Documentation Maintainers
Approver
Lightning IT Product Owners
Audience
platform operators, support engineers
Last reviewed
Next review
(Annual)

Troubleshoot Wunderbox safely

Diagnose from consumer outcome toward shared dependencies. Avoid making several unrelated changes at once; each change can hide the original failure and widen the affected failure domain.

Safe triage

  1. Protect data integrity and preserve administrative access.
  2. Freeze unrelated changes and record the UTC time window.
  3. Identify affected consumers, services, and failure domains without copying real identifiers into public channels.
  4. Compare consumer-facing health with resource, platform-service, dependency, and management-layer health.
  5. Check recent approved changes and immutable artifact identities.
  6. Determine whether recovery objectives or capacity thresholds are at risk.
  7. Apply one reversible, authorized diagnostic or recovery action at a time.
  8. Verify both platform state and consumer outcome after the action.

Symptom guide

SymptomInspect firstAvoid assuming
One consumer is degradedConsumer configuration, quota, and local dependencyThe shared platform is healthy or unhealthy as a whole
Several consumers fail togetherShared service, resource pool, network path, or failure domainEach workload has an independent fault
Intermittent latencySaturation, contention, retries, dependency timing, or observation gapsAverage utilization proves adequate headroom
Change cannot completeCapacity during transition, artifact identity, dependency readiness, or authorizationRetrying is harmless
Backup succeeds but restore failsRecovery-unit selection, integrity, keys, dependencies, and verificationA successful job status proves recoverability
Management works but service failsConsumer path and service-level outcomeAdministrative reachability proves service health

Escalation packet

Provide the product and release identities, symptom and impact class, UTC time window, affected abstract failure domain, redacted health summary, recent approved change identity, and checks performed. Keep topology, addresses, tenant or customer names, credentials, raw configuration, and full logs in the approved private support system.

Use the cross-product troubleshooting method for ownership and closure.