Platform Health Check
Sample report
This is what you receive at the end of a Health Check: ratings for every area, the evidence behind each finding, and a prioritised backlog you can plan and budget against.
Illustrative sample — fictional organisation. “Northwind Logistics” is not a real client, and every figure below is made up to show the format and depth of the report. It contains no data from any real engagement.
- Organisation
- Northwind Logistics (fictional)
- Scope
- Full Health Check — all eight areas
- Environment
- Single production instance, plus development and test
- Modules in use
- ITSM, Service Catalog, CMDB with Discovery, HR Service Delivery
- Interviews
- 7 stakeholders across IT operations, service desk, HR and architecture
- Access
- Read-only, sub-production
1 · Executive summary
The platform is stable and well used by the service desk, but it is carrying more risk than its day-to-day behaviour suggests. Customisation on core tables and the absence of automated regression testing are the main reasons upgrades keep slipping, and the CMDB is not yet trusted enough to support change-impact decisions. None of this is urgent in isolation. Together, it means each deferred upgrade makes the next one more expensive. The ten recommendations below are sequenced so the first three — roughly 10 consultant days — remove most of that compounding risk.
2 · Ratings by area
| Area | Rating | Headline |
|---|---|---|
| Platform health | Needs attention | Six scheduled jobs failing nightly, unnoticed |
| Customisation footprint | At risk | 214 global-scope rules on core tables |
| CMDB and CSDM | At risk | 11% duplicate servers; service layer unmaintained |
| Architecture and scoping | Needs attention | Custom apps built in global scope |
| Integrations | Needs attention | 4 of 9 integrations fail silently |
| Governance and release | Needs attention | No written standards; update sets unreviewed |
| Upgrade readiness | At risk | Two release families behind; no ATF coverage |
| Adoption and experience | Good | Service desk adoption strong; catalogue cluttered |
3 · Key findings
A full report documents every finding. The five with the greatest impact are shown here, each with the evidence behind it — so the conclusion can be checked rather than taken on trust.
- F1Customisation footprint
Core-table customisation is the main upgrade blocker
- Evidence
- 214 Business Rules and 61 Client Scripts run in global scope against task, incident and change. Of those, 38 duplicate behaviour the platform already provides out of the box, and 22 reference fields that no longer exist.
- Impact
- Each upgrade must re-test every one of them by hand. This is the single largest reason the last two upgrade windows were deferred.
- Recommendation
- Retire the 60 dead or duplicated items first, then move the remaining custom logic behind a scoped application where it can be tested and versioned independently.
- F2CMDB and CSDM
Duplicate CIs are entering through an import that bypasses IRE
- Evidence
- 11% of cmdb_ci_server records are duplicates. Every duplicate traces to a weekly asset import that writes directly to the table instead of through the Identification and Reconciliation Engine.
- Impact
- Change-impact assessments return incomplete results, so the change advisory board has stopped relying on them.
- Recommendation
- Route the import through IRE with datasource precedence set, then merge existing duplicates. Fixing the source first prevents the clean-up being undone the following week.
- F3Upgrade readiness
No automated regression tests exist
- Evidence
- Automated Test Framework is installed but has no active suites. The last upgrade was regression-tested manually over 16 working days.
- Impact
- Testing effort, not remediation, is what makes each upgrade a project. Without automation it will not get cheaper.
- Recommendation
- Build ATF coverage for the 12 critical paths identified in interviews — incident lifecycle, top catalogue items, HR case intake and the two main integrations.
- F4Integrations
Integration failures are silent
- Evidence
- 4 of 9 integrations have no error handling or alerting. Two still use basic authentication. The HR-system feed failed for 11 days in June before anyone noticed.
- Impact
- Data quietly diverges between systems, and failures are discovered by the people affected rather than by the platform team.
- Recommendation
- Adopt a single integration pattern with retry, logging and failure alerting, and move the two basic-auth integrations to OAuth.
- F5Platform health
Scheduled jobs have been failing unnoticed
- Evidence
- Six scheduled jobs have failed every night for between two and seven months, including the one that closes resolved incidents after five days.
- Impact
- Resolved incidents are not closing, which inflates backlog reporting and the service desk's apparent workload.
- Recommendation
- Fix or retire the six jobs, and add a daily health dashboard so job failures surface the morning they happen.
4 · Prioritised remediation backlog
Every recommendation ranked by impact against effort. The P1 items total 10 consultant days and remove most of the compounding risk; the full backlog is 47 days and can be scheduled over several quarters.
| Ref | Recommendation | Area | Impact | Effort | Priority |
|---|---|---|---|---|---|
| R1 | Route asset import through IRE and set datasource precedence | CMDB | High | 3 days | P1 |
| R2 | Retire 60 dead and duplicated core-table scripts | Customisation | High | 5 days | P1 |
| R3 | Fix or retire six failing scheduled jobs; add health dashboard | Platform health | High | 2 days | P1 |
| R4 | ATF suites for the 12 critical paths | Upgrade readiness | High | 8 days | P2 |
| R5 | Standard integration pattern with retry and alerting | Integrations | High | 6 days | P2 |
| R6 | Move two integrations from basic auth to OAuth | Integrations | Medium | 2 days | P2 |
| R7 | Merge existing duplicate server CIs | CMDB | Medium | 3 days | P2 |
| R8 | Write and adopt development standards; review update sets | Governance | Medium | 4 days | P2 |
| R9 | Migrate remaining core customisation into a scoped app | Customisation | Medium | 12 days | P3 |
| R10 | Retire catalogue items unused in 12 months (40% of catalogue) | Adoption | Low | 2 days | P3 |
| Total backlog | 47 days | ||||
5 · What happens next
The report ends with a walkthrough session. We take your team through the findings, take challenges on the ratings, and agree which recommendations to act on. Whether you deliver them yourself, with us, or with another partner is entirely your decision — the backlog is written to be usable by anyone.
Want this for your own platform?
Every Health Check ends with a report like this one — about your environment, with real evidence behind every finding.
