Two vendors can run an AI readiness assessment on the same Microsoft 365 tenant and hand back different scores, because there's no universal standard for what "ready" means. Each one is scoring against its own model. That makes the number close to useless on its own, and worth a great deal once you know what's behind it.
An AI readiness assessment scores your Microsoft 365 tenant against a set of governance conditions that raise the risk of AI surfacing something it shouldn't. Which conditions, and how many, is the vendor's choice. That's why the score measures governance posture rather than counting exposed content, and why two scores aren't comparable unless you've read both models.
So the useful question isn't what you scored. It's what was on the test.
Microsoft has plenty to say about getting ready for AI. Its foundational deployment guidance for securing and governing Copilot is organized into three pillars, and the first is remediating oversharing. What's instructive is where that pillar starts: not with labeling or policy, but with running assessments to find where the overshared, sensitive content sits.
That's a sequence of work, though. It isn't a scoring model. Microsoft tells you what to fix and in what order, and says nothing about how to express the result as a number. Its own AI Readiness Assessment does produce a score, but from a self-reported questionnaire rather than a scan of your content.
So every vendor built one, Orchestry included. Orchestry scores 13 governance signals across three groups; Syskit's Copilot Readiness score counts eleven inputs, ShareGate groups its insights into oversharing, sprawl, costs and Copilot results, and AvePoint's own readiness guide scores an eight-item checklist while arguing readiness is better read as a profile than a single score. None of those produces a figure you could set beside another vendor's and learn anything from the comparison.
That's not a scandal. It's the predictable result of a young category, and it will stay this way until someone standardizes it. The practical consequence is that a readiness score is a claim, and claims come with methodology you're entitled to see.
Whatever else a model measures, it measures oversharing. That's the common ground, and it's where most of the deficit sits.
Oversharing accumulates from things nobody did wrong at the time. A link shared with everyone in the organization for a project that ended in 2023. A guest invited for a vendor review who never redeemed the invitation. Somewhere else, inheritance was broken on a library to give one team access and never restored.
None of that was a disclosure problem while people had to know a file existed to find it. Search rewards people who already know what they're looking for. AI answers questions from whatever the person asking can reach.
The scale is easy to underestimate, and it's lopsided in a way most audits miss. Across Orchestry Enterprise customers whose OneDrives we crawl, nearly 95% of "Anyone" sharing links sit in OneDrive rather than SharePoint. Admins who audit SharePoint sharing and stop there are auditing the smaller half of the problem.
Labeling doesn't close the gap either. Varonis's 2025 State of Data Security Report, drawn from 1,000 real-world environments, found only 1 in 10 companies had labeled files at all. Unlabeled content can't be treated differently by anything downstream.
For the remediation side of this, we've covered the three governance gaps behind a low score and how to find and fix broken permission inheritance at scale separately.
Permissions get the attention because they're the direct mechanism. Microsoft's documentation is explicit that Copilot accesses content through Microsoft Graph and surfaces only organizational data a user already has at least view permission for. There's no separate AI access model to configure.
Which is where thin models stop, and where the divergence between them starts. A site with correct permissions and no owner is a site where nobody will notice when the permissions stop being correct. Ownership, activity and lifecycle aren't access controls, but they predict which access controls are about to rot.
Based on Orchestry data, over 67% of workspaces show no activity in the trailing 90 days when an organization first connects to the platform. That content is still indexed, still permissioned as it was, and still answerable from. Nobody is watching it because nobody has a reason to open it.
Microsoft's interim control here is Restricted Content Discovery, which stops a site's content surfacing in Copilot or organization-wide search while leaving direct access unchanged. Microsoft describes it as a temporary governance control, and it requires both Copilot licensing and SharePoint Advanced Management. For the wider control set, see what Restricted Content Discovery actually does.
The second place models diverge is coverage, and it moves the number more than most people expect.
A twenty-site review tells you about twenty sites. To know where the whole tenant stands, and whether it is getting better, every site has to be checked, and checked the same way each time. Coverage has a time dimension too. Every new workspace, sharing link, guest invitation and ownership change moves the real posture, so a sample taken in March says nothing about the sites created in July. A model that can only be run by hand can only be run occasionally, which means its output is always describing a tenant that has since changed.
Orchestry runs oversharing detection as a recurring tenant-wide scan instead of a point-in-time audit, with actions available at the workspace level from the same view.
Here's what to put to any AI readiness assessment tool, including ours.
Question five is where most assessments end and the work begins. Orchestry's Content Discovery configuration (coming soon) lets you include or exclude a site from Microsoft Search and Copilot as part of a workspace review, using wholesale SharePoint de-indexing or the more targeted Restricted Content Discovery. Archiving inactive content works from the other direction, moving it out of reach so cleanup shrinks the surface instead of tidying it.
When you do get a fix list, sequence it by exposure. Anonymous and organization-wide links first, then broken inheritance on sites holding sensitive content, then ownership, then lifecycle.
Put those five questions to Orchestry and here are the answers.
The model is 13 governance signals across three groups: oversharing, governance, and Orchestry safeguards. Each signal passes or fails against a threshold, and the score is the share that pass, expressed as a 0 to 100 percentage.
Every signal counts equally, because the score is simply the share of checks that pass. That makes it a breadth measure: it tells you how many conditions you're meeting, not how far off any single one is. The drill-through is where you get the depth, from the number to the conditions sitting behind it.
The score covers the whole tenant instead of a sample, and it's included on every plan.
On thresholds, Orchestry publishes its own: most tenants start in the low 30s before any governance work, and above 50% reflects a generally healthy posture.
The score reads governance conditions, not file contents. Orchestry doesn't inspect or classify what's inside a document, so it can tell you a library is shared with everyone in the organization and it can't tell you whether what's in that library is sensitive. Content classification is Microsoft Purview's job.
Every vendor's model has a boundary like that one. Ask where theirs is.
A Microsoft 365 Copilot readiness assessment was a gate. You ran it before a rollout, fixed what it found, and moved on to deployment.
A standing score assumes there's no single rollout left to gate. An overshared library carries the same risk whether Copilot, Claude, ChatGPT or Glean reads it, and the agents already running in your tenant usually reach whatever the person using them can reach.
The conditions being measured haven't changed. What changed is that passing once stopped being enough.
There's no industry benchmark, so the honest answer is that it depends whose model produced the score. Orchestry publishes its own reference points: most tenants start in the low 30s before any governance work, and above 50% reflects a generally healthy posture. Those come from Orchestry's observations across first scans, not a published study, and they only apply to Orchestry's 13-signal model.
No. Each vendor scores against its own signal list, its own thresholds and its own weighting, and none of those are standardized. The same tenant can come back healthy from one tool and failing from another without either being wrong. Compare the models before you compare the numbers.
No. The score is the share of governance checks your tenant passes, so it measures the conditions that create disclosure risk instead of counting exposed files. To see what a specific person's AI can reach, you need permission and sharing reporting at the workspace and file level, which answers a different question.
No. Microsoft's Restricted Content Discovery control requires both Copilot licensing and SharePoint Advanced Management, but assessing readiness doesn't. Orchestry's AI readiness score is included on every plan; its Content Discovery configuration (coming soon) can de-index sites with standard SharePoint settings, and uses Restricted Content Discovery only where you have SharePoint Advanced Management.
Often enough to catch what moves it. The score changes when workspaces are created, sharing links are added, guests are invited or left unredeemed, workspace ownership goes stale, and new content arrives unlabeled. In a tenant where people create workspaces and share links every week, a score from last quarter describes a tenant that no longer exists.
A number you can't audit is a number you can't act on, or take to a security review. The version worth having is the one where you can name every signal, say why each is in there, and point at what failed.
To see what your own tenant scores and which conditions are sitting behind the number, an AI readiness check runs the full 13-signal model against your real environment. Or score your tenant with the agent readiness scorecard first.