From 41 to 78 in a quarter: what moving a readiness score is worth
A Copilot readiness score improvement over 90 days: where the early wins come from, why it plateaus at week two, and what the movement is actually worth.

"We ran a security assessment" is not an outcome. "The tenant went from 41 to 78, and here is what changed" is.
A single composite score is a blunt instrument, and that is precisely its value: it is the only artefact from a readiness programme that a board, a CISO and an administrator can all hold the same opinion about.
The arc below is a worked illustration, not a client engagement. The numbers are modelled on what the six check domains typically move and are shown so you can follow the reasoning — not as a published result.
What the score is actually made of
SafeScan's Copilot Readiness Score is a 0–100 composite across six domains: exposure, identity, compliance, Teams, licensing and the readiness view that sits on top of them. The product's own demo tenant shows the shape clearly — an overall 71, with SharePoint and OneDrive at 68, Entra ID at 76, Purview at 64, Teams at 78 and Microsoft 365 at 82.
That spread is the useful part. An overall 71 tells you very little. Knowing that Purview is your weakest domain tells you where the next fortnight goes.
A 90-day arc, and where the movement actually comes from
Week one: the configuration wins
The first jump is always the largest and the cheapest. Anonymous links revoked, broken inheritance restored on a handful of sites, stale guests removed, a sharing setting tightened on the sites that never needed it open.
None of this is hard. All of it is high severity. A tenant starting in the low 40s typically does most of its climbing here, and it is why the first rescan is worth running within days rather than at the end.
Weeks two to six: the plateau
Then it stops, and this is where remediation programmes die.
What is left needs other people. The guest list needs a client conversation. The permission model on the project hub needs its owner. Sensitivity labels need a rollout, not a setting. The score barely moves for a month, which feels like failure and is not — it is the difference between configuration and change management.
The useful thing to track in this window is not the score. It is how many findings have a named owner and a date.
Weeks six to twelve: the structural work lands
Labels start being applied. The permission exceptions get rationalised. Audit logging, if it was off, has been on long enough to be useful — and Microsoft is explicit that auditing only captures from the point it is enabled, which is why turning it on early matters even though it moves nothing immediately.
This is where the second half of the climb comes from, and it does not happen without the plateau in front of it.
What the movement is worth
Three things, none of which is "a better number".
The rollout stops being blocked. The most expensive thing about poor readiness is usually not the risk — it is the Copilot licences being paid for while security withholds sign-off. A deployment stalled a quarter on an unanswerable question has a real monthly cost attached.
The audit question becomes answerable. "How do you know Copilot cannot surface HR records to the wrong people?" has two possible answers. One is a paragraph of intent. The other is a scan from before, a scan from after, and the list of what changed between them.
The work becomes defensible. Remediation that cannot be evidenced tends to get redone by the next person who asks. A before-and-after pair with dates ends that cycle.
Reading the domain spread properly
The composite hides as much as it shows, and the per-domain numbers are where the decisions live.
A tenant strong on identity and weak on exposure is a permissions problem: the people are who they say they are, and they can reach too much. That is remediation work on sites and sharing.
The reverse — strong exposure scores, weak identity — is a different and usually more urgent shape. Your content is organised and your front door is open. No amount of SharePoint tidying addresses an admin account without strong authentication.
Weak compliance with everything else healthy generally means the controls exist on paper but were never switched on. That is the cheapest gap to close and the one most likely to be asked about in an audit.
What the score is not
Three honest limits, because a number this convenient invites overreach.
It is not a compliance certification. It tells you what Copilot would inherit and where the configuration is weak. It does not attest to anything, and presenting it as an audit outcome will not survive contact with an actual auditor.
It is not a penetration test. It reads configuration and permissions; it does not attempt to exploit anything.
And it is not comparable between organisations. A 71 in a 50-person consultancy and a 71 in a 5,000-person manufacturer describe different realities. The only meaningful comparison is the same tenant against itself over time — which is the entire argument for keeping the history.
Measure the same way twice, or do not bother
A before-and-after is only meaningful if the two scans are comparable. Same tenant, same scope, same check set. If the score moved because the second scan covered fewer sites, you have measured nothing and you will be caught.
SafeScan's rescan-and-compare exists for this: re-run after remediation, and the comparison shows which checks changed verdict rather than just the headline number. That delta is the artefact worth keeping, more than either scan on its own.
Running it yourself
Copilot SafeScan scores a tenant in under five minutes, read-only, and keeps the history so the second scan means something. Export the pair to PDF for the people who need the story and Excel for the people doing the work.
Related: what to fix first, why the score decays if you stop, and running this across a client portfolio.




