Operations · Decision Memo

Which AI model should we standardize on?

A recommendation for our 60-person team — balancing capability, cost, privacy, and how easy it is for non-technical staff to actually use.

Prepared by: Operations Lead  ·  For: Leadership Team
Decision requested by end of month
⚠ Built from working knowledge of the market as of early–mid 2025 — pricing & model names move fast; verify the headline numbers before signing anything.
Framing

The decision, and how we'll make it

"Standardize" means one default vendor/model that IT can support, bill centrally, and write a policy around — not banning everything else, but having a sanctioned path.

What we're optimizing for

  • Capability — good enough for writing, analysis, coding, and document Q&A
  • Total cost — predictable per-seat or per-token spend at 60 people
  • Privacy & data handling — our data must not train someone's model
  • Ease for non-technical staff — minimal setup, familiar chat UI

How we'll score it

  • Weighted: Ease 30% · Privacy 25% · Cost 25% · Capability 20%
  • For a non-technical company, day-to-day usability beats raw benchmark wins
  • We compare the business/enterprise tiers, not the free consumer apps
  • Decision is reversible — we avoid lock-in where we can
The Landscape

Four credible options for a team buyer

These are the realistic contenders for a company that wants a supported, secure deployment — not a hobbyist setup.

OpenAI — ChatGPT Enterprise / Team (GPT-4o, o-series reasoning)

Most-adopted product, huge ecosystem, strong all-rounder. Team tier is self-serve; Enterprise is sales-led.

Anthropic — Claude (Team / Enterprise, Sonnet & Opus)

Strong at writing, long documents, and coding. Reputation for safety and careful data handling.

Google — Gemini in Workspace (Business / Enterprise add-on)

Lives inside Gmail, Docs, Sheets, Meet. Wins on integration if you're already a Workspace shop.

Microsoft — Copilot (M365, runs on OpenAI models)

Same logic as Gemini but for Office/Teams. Strong if you're deep in Microsoft 365.

Open-source/self-hosted (Llama, Mistral) is powerful and private but needs real engineering to run safely — out of scope for a non-technical team's default, though worth revisiting later.

Criterion 1 · Capability

All four are "good enough" — the gap is narrower than the hype

For everyday business work (drafting, summarizing, analysis, light coding) the frontier models are now roughly interchangeable. Differences show up at the hard edges.

Claude (Sonnet/Opus)
writing, long docs, code
GPT-4o / o-series
all-round, reasoning, tools
Gemini
huge context, multimodal
Copilot (OpenAI)
strong in Office context

Honest take: benchmark leadership flips every few months. Do not pick a vendor for a 2-point benchmark lead — it'll be stale before procurement closes. Capability is the least differentiating factor for our use case.

Criterion 2 · Cost

Per-seat pricing is converging around ~$25–30/user/mo

Approximate published business-tier list prices. Real cost depends on negotiation, annual commit, and whether you already pay for Workspace/M365.

OptionApprox. list / user / moBilling modelAnnual @ 60 seats
ChatGPT Team~$25–30 (annual)Self-serve, per seat~$18k–21k
Claude Team~$25–30 (annual)Self-serve, per seat~$18k–21k
Gemini for Workspace~$20–30 add-onAdd-on to Workspace~$14k–21k + Workspace
Microsoft 365 Copilot~$30 add-onAdd-on to M365~$21k + M365 licences

The cost insight: if we already pay for Workspace or M365, the bundled add-on can be cheaper and means zero new vendor. If we don't, a standalone Team plan is simpler to reason about. Enterprise tiers cost more but add SSO, admin controls, and contractual data terms.

Criterion 3 · Privacy

The single most important factor for us

The rule that matters: on paid business/enterprise tiers, all four vendors contractually do NOT train on your data by default. The free consumer apps are where data leaks into training — which is exactly why we standardize.

What "good" looks like

  • No training on business-tier inputs/outputs
  • SSO + admin console + audit logs
  • Data retention you can configure / zero-retention options on enterprise
  • SOC 2 / ISO; data residency where needed

Where the risk actually is

  • Shadow AI: staff pasting client data into free ChatGPT/Gemini on personal accounts
  • Standardizing kills that risk — give people a sanctioned, safe tool
  • Verify the data-processing addendum (DPA) before rollout

All four are defensible on privacy at the business tier. So privacy passes for everyone — it doesn't pick the winner, it just rules out the free tier.

Criterion 4 · Ease for non-technical staff

This is what actually decides it for us

We're 60 mostly non-technical people. Adoption — not capability — is the thing that fails. The winning tool is the one people already know how to use.

ChatGPT Team
Most familiar UI; staff have used it; minimal training
Copilot (M365)
In-app — great IF we live in Office
Gemini (Workspace)
In-app — great IF we live in Google
Claude Team
Clean UI; less name recognition with staff

Key fork: if we're already standardized on Microsoft 365 or Google Workspace, the embedded assistant wins on ease because it shows up where people already work. If our suite use is mixed/light, a standalone chat tool people already recognize wins.

Recommendation

Adopt ChatGPT Team as the default — with one condition

ChatGPT Team (→ Enterprise at scale)

Best blend of familiarity, self-serve setup, broad capability, and solid business-tier privacy for a non-technical org. Lowest adoption friction.

  • Highest staff familiarity → fastest, cheapest adoption
  • Self-serve admin console, no heavy IT lift to start
  • No training on Team/Enterprise data; SSO at Enterprise
  • Strong general-purpose model that covers ~all our use cases

The condition (the honest caveat)

If we are already paying for Microsoft 365 or Google Workspace company-wide, default to Copilot or Gemini instead — the embedded assistant is cheaper net, lives where people work, and adds no new vendor or security review.

Consider Claude as an approved secondary for heavy writing/long-document teams — many orgs run two and let people choose.

Rollout & Cost

Phased, ~90 days, ~$20k/year

Weeks 1–3

Pilot

10 power users across teams. Sign DPA, set retention/SSO, draft acceptable-use policy. Measure real value.

Weeks 4–8

Roll out

All 60 seats. Two 45-min "how to actually use it" sessions + a one-page prompt cheat-sheet. Name an internal champion.

Weeks 9–12

Embed & review

Collect use cases, track active usage, kill shadow-AI on personal accounts, decide on Enterprise upgrade.

~$20k / year

60 seats × ~$28/mo annual. Less if bundled into existing Workspace/M365. Budget a small buffer for Enterprise tier upgrade later.

~$330 / person / year

Pays for itself if it saves each person ~1 hour/month. Track that — review spend at 90 days against measured time saved.

Risks & What We're Watching

Honest trade-offs and how we manage them

High
Sensitive data leakage
Mitigated by paid tier (no training) + acceptable-use policy + training staff what not to paste. Still: never put regulated PII / secrets in any model.
Medium
Wrong / confident answers ("hallucination")
All models do this. Policy: a human verifies anything client-facing or factual. It's a draft assistant, not a source of truth.
Medium
Vendor lock-in & price changes
Pricing/leadership shifts fast. Keep the deployment portable (standard chat use, exportable data); re-evaluate annually.
Low
Low adoption / shelfware
The real failure mode. Pilot first, train, appoint a champion, and cut seats that go unused at the 90-day review.

Bottom line: Adopt ChatGPT Team as default (or the embedded assistant if we're already an M365/Workspace shop), pilot in 3 weeks, ~$20k/yr, review at 90 days. The bigger risk is doing nothing and letting shadow AI run unmanaged.