The estimator could not start

If this keeps happening, contact your Thoughtworks representative.

AI Factory · Services Estimator
Powered by Thoughtworks Thoughtworks
Thoughtworks AI Factory · Services attach

Put a services number on the rack while the deal is still in motion.

Enter the shape of the hardware opportunity and get an indicative range for cluster activation and ongoing managed operations — with the assumptions behind it in plain sight, and a route to a formal estimate that keeps your deal context.

What it prices
One-time cluster activation, and ongoing managed cluster operations.
What it is not
A quote. Every result is an indicative range, non-binding and subject to a formal confirmation.
Coverage
Business hours is the priced default. 24/7 routes to a Thoughtworks conversation.
Indicative spend over term 10 racks

Indicative spend by month across a 12-month term — month one carries one-time activation. One scale across all estate sizes

Indicative activation
—
one-time, above rack & stack
Managed operations
—
per month, business hours
Rate per rack
—
per month — falls with scale
Why Thoughtworks

Racks alone do not make an AI Factory. Someone has to turn them into tokens.

Thoughtworks — AI Factory

Thoughtworks brings AI Factories up and runs them — everything after the racks are installed, from bare metal through Kubernetes, observability, fleet intelligence and inference. We sell no hardware and no competing platform, which is exactly why we attach cleanly to a {partner} deal instead of competing with it.

175,000
GB300 GPUs monitored today on the observability platform we build for Firmus. Architected for one million.
4 hours
To stand up a frontier-lab cluster, plus a six-hour soak. Four clusters in two days.
1 of 3
Advanced Technology Partners worldwide holding the Agentic Transformation specialization.
50+
NVIDIA certifications across five regions, with 24/7 follow-the-sun coverage available.
S1 · S2

Setup & activation

Infrastructure-as-code bring-up of Kubernetes and Slurm, automated validation and soak testing, observability and runbooks in place, first workloads live and benchmarked.

Priced here — one-time
S3

Managed cluster operations

Monitoring, incident response, patching, upgrades and capacity planning. Automated triage for XID errors, InfiniBand, NVLink and optics — the failures that actually take fleets down.

Priced here — subscription
S4+

Performance engineering

More tokens per megawatt. Custom CUDA kernels, scheduling and packing, speculative decoding, quantisation and cost-aware routing.

Talk to Thoughtworks
S4+

Sovereign & applications

Sovereign AI readiness, tenant isolation and chargeback, and consumer-grade application delivery above the platform when the customer needs it.

Talk to Thoughtworks

Deliberate boundaries: we do not design facilities or own hardware break/fix. That stays with the data-centre operator and with {partner} — which is what lets us sit in the operations gap alongside you.

Curated scenarios

Three sample deal shapes we see in the {partner} channel

Each one is a saved starting point. Open it, change anything, and the range moves with you.

Scenario

The deal

2 inputs to a range
—
—

Each answer here narrows the range. Leaving one on Unknown keeps its uncertainty priced in — the more you can confirm with the customer, the tighter the range gets.

One-time cluster bring-up, above rack & stack.
Expansion adds racks to a cluster Thoughtworks already runs.
Business hours is the priced default. A 24/7 requirement is a different operating shape, not just a different rate.
What the standard RACI covers
Does the agreed scope match the standard RACI? If not, we need to settle who holds the pager, what counts as an incident, and who owns job submission and hardware break-fix.

Saved scenarios

Scenarios you save are kept to your account, so they follow you to any device you sign in from. The examples below are starting points — open one, change anything, then save it under your own name.

Indicative services range

Business hours coverage · USD · non-binding
Indicative
Total contract value — indicative range 12-month term
USD — – —

Activation plus managed operations across the term.

— scale: the widest this deal could be today —
Low—
High—

This range narrows as assumptions are confirmed.

One-time activation — above rack & stack
Managed — per month — —
Managed — per year — —

Spend over the contract

—

One-time activation Your confirmed floor Open upside, narrows as you confirm Most this deal could cost, unconfirmed That period's range today Your likely landing zone, once confirmed Most that period could cost, unconfirmed

This Proposal is not intended to be binding on either of the parties. It is strictly for planning and discussion purposes only, Any binding agreement between the parties will be negotiated and agreed upon in a definitive agreement, signed by both parties. Nothing in this Proposal represents a binding legal obligation on either party, including without limitation any binding obligation to reach and execute the final version of a definitive agreement between the parties, and no such binding obligation shall arise until such a definitive agreement is executed by the parties.

—

      Every constant behind this estimate, and how firm it is. Green is taken straight from the pricing workbook; amber is a judgement call that needs commercial confirmation; red is a placeholder that is not yet real.

      What a Thoughtworks conversation would resolve — each one is a reason the range is as wide as it is.

      Powered by Thoughtworks

      What {partner} is attaching

      Everything after the racks are installed. We operate the layer between the data-centre operator, the OEM and the customer — the gap no one else is set up to serve.

      • S1 Cluster setup — IaC bring-up, validation, soak, observability, runbooks
      • S2 Activation — first workloads live, benchmarked, ops team trained
      • S3 Managed operations — monitoring, incident response, patching, capacity
      • S4+ Performance engineering, sovereign AI, applications — on request

      Request a formal estimate

      Everything you have explored comes with it — nothing to re-enter.

      What is driving the request?

      This is what Thoughtworks receives. Add anyone who should be on the call.

      Thoughtworks replies to this person. Both fields are needed before the request can be sent.

      Request captured

      A Thoughtworks lead has the full deal context, the current range and the open questions.

      Target first response: 15 minutes

      Delivery to Thoughtworks is not yet enabled in this build.

      This Proposal is not intended to be binding on either of the parties. It is strictly for planning and discussion purposes only, Any binding agreement between the parties will be negotiated and agreed upon in a definitive agreement, signed by both parties. Nothing in this Proposal represents a binding legal obligation on either party, including without limitation any binding obligation to reach and execute the final version of a definitive agreement between the parties, and no such binding obligation shall arise until such a definitive agreement is executed by the parties.

      Step 1 of 2

      Open the estimator on a phone

      Scan to land directly in the estimator, not on a landing page.

      Hand this to a {sellerNoun} in the room. They land on the mobile estimator with four inputs and a range-first card.

      —

      The standard RACI

      What the data-centre operator, the installer, Thoughtworks and the customer each own in a standard-scope engagement.

      This RACI is the default division of responsibility this model assumes when it prices a deal as "Standard RACI." It is the operating boundary the data-centre operator, the installer, Thoughtworks and the customer work to unless a deal has explicitly agreed a different split — which is why "Custom split" and "Unknown" both leave this model's judgement slack in, rather than resolving to a single number, and both point back to this page.

      What RACI means

      • R — Responsible. Does the work.
      • A — Accountable. Owns the outcome and signs off. Exactly one accountable party per row.
      • C — Consulted. Gives input before the work happens — a two-way conversation.
      • I — Informed. Told after the fact — a one-way notification.

      Assumptions

      • The Accountable party is responsible for communication with the Customer.
      • For some entries involving streaming data (i.e. observability), Informed means receiving the real-time data stream / alerts.

      Legend

      R Responsible A Accountable C Consulted I Informed ▲ Responsible party

      Shared cells indicate joint accountability. "DC" is the data-centre operator; "Installer" is the hardware installation partner — neither is Thoughtworks nor {partner}.

      Thoughtworks AI Managed Services — RACI Matrix
      Hardware & Physical Infrastructure
      Responsibility DC Installer Thoughtworks Customer
      Physical server and GPU hardware break-fix support A/R R I I
      Physical storage installation, cabling and physical commissioning A/R R I I
      Physical storage hardware break-fix support A/R R I I
      Hardware component replacement and spare parts management A/R R I I
      Server rack installation, cabling and physical commissioning A R I I
      Power and cooling provisioning and monitoring A/R – I –
      Hardware firmware update planning, scheduling and execution R/C – A R/C
      Physical data center access control and badge management A/R I – C
      Network Management
      Responsibility DC Installer Thoughtworks Customer
      InfiniBand and Ethernet fabric management via UFM A/R R R –
      InfiniBand fabric health monitoring R – A/R I
      InfiniBand link fault resolution A R C I
      Network bandwidth and latency baseline reporting I – A/R I
      Storage Management
      Responsibility DC Installer Thoughtworks Customer
      Storage performance monitoring I – A/R I
      Storage capacity monitoring I – A/R I
      Storage fault resolution A R C I
      Storage administration A/R – C I
      Storage capacity upgrade A R C I
      NFS/parallel filesystem mount management for workload namespaces – – A/R R/C
      Storage network bandwidth and latency baseline reporting I – A/R I
      NFS/parallel filesystem CPU capacity monitoring I – A/R I
      NFS/parallel filesystem CPU capacity upgrade A/R R C I
      Operating System & Platform Management
      Responsibility DC Installer Thoughtworks Customer
      Base OS image authoring, versioning and catalog maintenance C/I – R/A C/I
      OS deployment and reprovisioning via Base Command Manager C/I – R/A I
      OS patch management — planning, scheduling and execution C/I – R/A I
      Kernel and driver compatibility validation after OS updates C/I – R/A I
      OS-level security hardening and CIS benchmark compliance C/I – R/A I
      Cluster & Orchestration Management
      Responsibility DC Installer Thoughtworks Customer
      Cluster provisioning and node lifecycle management via BCM C/I – A/R C/I
      Kubernetes cluster deployment and control plane management – – A/R I
      NVIDIA GPU Operator upgrade and health validation – – A/R C/I
      Container runtime configuration and container socket management – – A/R I
      Node validation – – A/R I
      Node labeling, taint and toleration management for GPU workloads – – A/R C/I
      Persistent Volume and storage class provisioning for workloads – – A/R R/C/I
      Cluster upgrade planning, execution and post-upgrade validation C/I – A/R C/I
      Job Scheduling & Workload Management
      Responsibility DC Installer Thoughtworks Customer
      Slurm cluster configuration, partition management and job queue ops – – A/R R/C/I
      Run:ai control plane upgrade and project management – – A/R C/I
      GPU fraction and multi-tenancy policy configuration in Run:ai – – R/C/I A/R/C/I
      Job queue monitoring, stall detection and remediation – – A/R R/C/I
      Workload priority and fairness policy definition – – C/I A/R
      Job submission, tracking and user-level scheduling support – – C/I A/R
      Scheduling policy review and optimization advisory – – A/R C/I
      Monitoring & Observability
      Responsibility DC Installer Thoughtworks Customer
      GPU health monitoring via DCGM – – A/R I
      DCGM Exporter deployment and Prometheus metrics pipeline management – – A/R C/I
      Infrastructure observability dashboards (Grafana/Datadog) – – A/R ▲
      GPU utilization, memory and power consumption reporting – – A/R C/I
      Alerting rule authoring, threshold management and escalation routing – – A/R I
      InfiniBand and network performance telemetry collection – – A/R I
      Log aggregation, search and retention for AI/ML workloads – – A/R C/I
      Monthly infrastructure health and utilization reporting to customer – – A/R I
      Security & Access Management
      Responsibility DC Installer Thoughtworks Customer
      Role-based access control (RBAC) design and enforcement – – A/R C
      Kubernetes namespace and multi-tenant isolation policy management – – A/R –
      API key, service account and credential lifecycle management – – A/R C/I
      Security patch prioritization and vulnerability remediation – – A/R C
      Audit logging and compliance reporting for platform access – – A/R I
      End-user software and tool access provisioning – – – A/R
      Security governance policy ownership I – A/R C/I
      Penetration test coordination and remediation tracking C/I – A/R C/
      Capacity & Lifecycle Planning
      Responsibility DC Installer Thoughtworks Customer
      GPU and compute capacity monitoring and trend analysis I – A/R I
      Infrastructure capacity forecasting and right-sizing recommendations C/I – A/R ▲/C/I
      Application and workload capacity planning advisory – – A/R ▲/C/I
      Hardware refresh and technology roadmap advisory C – A/R C/I
      License capacity tracking — NVAIE, Run:ai, BCM – – C/I A/R/I
      Cloud burst capacity planning for overflow workloads – – A/R C/I
      AI/ML Platform Operations
      Responsibility DC Installer Thoughtworks Customer
      NVIDIA Base Command Platform administration and configuration – – A/R C/I
      NGC private registry setup, image curation and access control – – A/R C/I
      NVIDIA AI Enterprise license deployment and renewal management – – C/I A/R
      NIM microservice deployment, scaling and availability management – – A/R C/I
      AI/ML platform performance monitoring and baseline reporting – – A/R C/I
      NVIDIA software stack upgrade planning and execution I – A/R C/I
      AI Workload & MLOps
      Responsibility DC Installer Thoughtworks Customer
      ML training job pipeline configuration and execution support – – C A/R
      Model versioning, experiment tracking and artifact management – – C A/R
      AI/ML orchestration deployment via Kubernetes and Helm – – A/R C/I
      CI/CD pipeline integration for model build and deploy workflows – – A/R C/I
      Distributed training configuration (multi-node, NCCL tuning) – – C/I A/R
      Checkpoint management and training job recovery procedures – – R/C/I A/R
      MLOps toolchain administration (MLflow, Kubeflow, etc.) – – A/R C/I
      ITSM — Event, Problem & Change Management
      Responsibility DC Installer Thoughtworks Customer
      24x7 infrastructure event monitoring, triage and escalation C/I – A/R C/I
      Infrastructure incident ticket creation, ownership and resolution tracking C/I – A/R I
      24x7 Hardware support R/A R C/I C/I
      Hardware incident ticket creation, ownership and resolution tracking A/R R C/I C/I
      Root cause analysis and problem record management R/C/I R/C/I A/R/C C/I
      Change request authoring, risk assessment and approval coordination R/C R/C/I A/R C/I
      Scheduled maintenance window planning and customer communication R/C R/C/I A/R/C C/I
      Post-incident review and corrective action tracking R/C/I C/I A/R/C C/I
      SLO performance tracking and reporting A/R C/I R/C I
      Knowledge base article authoring for recurring issues A/R – R/C/I –
      Monthly TAM Review R/A – R C/I
      Quarterly business review R/A – R C/I