Skip to main content
This page covers incident response, change management, reporting, and compliance for production MacStadium VDI environments. For day-to-day operational tasks, see the Day-2 Operations Guide. For troubleshooting specific symptoms, see Troubleshooting Quick Reference.

Incident response

Recognizing common failure modes

Symptom: User can’t connect to desktop Possible causes: VM not running, VDA not registered with Citrix, network connectivity issue, Citrix Cloud issue.
Symptom: Desktop is slow or unresponsive Possible causes: Host overloaded (too many VMs), VM resource starvation, network latency.
Symptom: VMs fail to deploy Possible causes: Host out of disk space, image pull failure, Orka Engine error.
Symptom: All VMs down after host reboot Cause: VMs don’t auto-start after host reboot by default.
Consider scripting auto-start behavior or coordinating with MacStadium to enable auto-start features.

Triage decision tree


Escalation procedures

Escalation email template:

Post-incident review template

Complete this for any incident affecting more than 10 users or lasting more than an hour.
Store post-incident reviews in your documentation repository for future reference.

Change management

Pre-change checklists

Before any production change, verify:
  • Change window scheduled and communicated to users
  • Backup/snapshot of current image state available
  • Rollback plan documented and tested
  • Testing completed in a non-production environment
  • Required approvals obtained
  • Monitoring in place to detect issues
  • Team available for the duration of the change
For image updates specifically:
  • New golden image tested on at least one VM
  • Citrix VDA registration verified
  • HDX features tested (clipboard, file transfer, USB)
  • Applications tested and functional
  • Pilot group identified
  • Previous image version retained for rollback

Testing procedures

For new golden images:
  1. Deployment test: Deploy one VM, verify it boots within 3 minutes and gets network connectivity.
  2. VDA registration test: Check System Preferences → Citrix VDA shows “Registered”; verify VM appears as “Available” in Citrix Cloud Console.
  3. User connectivity test: Assign a test user, launch desktop from Citrix Workspace, verify connection.
  4. HDX feature test: Test clipboard, file transfer (if enabled), printing (if enabled), application launching.
  5. Application functionality test: Launch each business-critical application, perform a basic workflow, check for errors.
  6. Performance test: Measure login time (target: under 30 seconds), check CPU/memory at idle, check responsiveness during typical tasks.
Document results with image name, test date, tester name, and pass/fail for each item.

Rollback plans

Write rollback procedures before starting any change. Example: Image update rollback If a new image causes issues within the first 24 hours:
  1. Stop new deployments immediately.
  2. Revert affected VMs:
  1. Verify users can connect to rolled-back VMs.
  2. Document what went wrong for post-incident review.
Estimated rollback time: 30-45 minutes for 10 VMs. Example: Citrix policy change rollback
  1. Revert the policy in Citrix Cloud Console: Policies → Select policy → Edit → Restore previous settings.
  2. Force policy refresh: have users log out and back in, or wait 30 minutes for automatic refresh.
Estimated rollback time: 5-10 minutes.

Communication templates

Planned maintenance (send 3-5 business days in advance):
Emergency maintenance (send immediately when issue detected):
Resolution notification:

Metrics and reporting

Key performance indicators


User satisfaction tracking

Quarterly survey questions:
  1. Rate your overall satisfaction with the macOS virtual desktop (1–5)
  2. How often do you experience connectivity issues? (Never / Rarely / Sometimes / Often)
  3. How would you rate desktop performance for your daily tasks? (Poor / Fair / Good / Excellent)
  4. What applications or features would improve your experience?
  5. Any other feedback?
Review support tickets weekly for recurring issues, patterns by user group, and correlation with recent changes. Address patterns before they become widespread.

Cost analysis and optimization

Review quarterly:
  1. Right-size VMs: Are all users on high-spec VMs when they only need basic?
  2. Eliminate unused capacity: VMs deployed but not assigned to users?
  3. Image efficiency: Unnecessary applications in golden images? Can you consolidate?
  4. Licensing: Citrix licenses for inactive users? Remove inactive accounts quarterly.

Quarterly business review outline

Present to leadership/stakeholders each quarter:
  1. Service overview: Total users, total VMs, uptime %, support ticket trend
  2. Highlights: Major improvements, issues resolved, user feedback summary
  3. Challenges: Pain points, resource constraints, technical debt
  4. Roadmap: Upcoming improvements, capacity planning, technology upgrades
  5. Financials: Cost per user, budget vs. actual, cost optimization initiatives
Keep it business-focused. Leadership cares about user satisfaction, costs, and risks, not Ansible commands.

Reference

Vendor contacts

Contact MacStadium when: host is down, Orka Engine failures, new host provisioning, datacenter network issues. Contact Citrix when: VDA registration failures, Cloud Connector issues, licensing problems, HDX protocol issues.

Compliance checklist

Review quarterly. Security:
  • VMs patched monthly (macOS updates)
  • Citrix VDA is current (or within 2 releases)
  • Access logging enabled in Citrix Cloud
  • User access reviewed quarterly, inactive users offboarded
  • Network segmentation enforced
  • Registry credentials rotated every 90 days
Data protection:
  • User data not stored on VMs (network storage only)
  • Golden images backed up (at least 3 versions retained)
  • Disaster recovery plan documented and tested annually
  • VM deletion policy enforced (no orphaned VMs)
Operational:
  • Capacity headroom maintained (20–30% spare VMs)
  • Monitoring in place for VM availability
  • Change management process followed for all production changes
  • Post-mortems completed for major outages
  • Documentation kept current
Financial:
  • Chargeback reporting in place (if multi-tenant)
  • Monthly cost tracking vs. budget
  • Unused licenses identified and reclaimed
  • Quarterly cost optimization review