contact us
 

CI&T’s Agent Playbook: From Hype to Exponential Throughput

Feb 02, 2026 | min read
Please select some categories on More Settings section.
By

Roberta Lingnau de Oliveira

In the world of digital delivery, the conversation is shifting. For the last 12 months, the industry has been flooded with AI demos, yet production-grade wins are scarce. At CI&T, we don’t chase the AI hype — we deploy AI to collapse timelines and solve the Modern Turing Test.

The original Turing Test asked: Can a machine talk like a human? The Modern Turing Test asks: Can an agent autonomously turn 100 story points into 1,000 – at the same cost and quality?

With inference costs dropping nearly 100x in two years, the trend is clear: autonomous agents are becoming the economic engines inside delivery pipelines. Mustafa Suleyman, co-founder of DeepMind, predicts this Modern Turing Test — shifting the focus from conversation to autonomous throughput — will be achievable by 2027. Our pilots suggest it is already within reach for specific, high-volume delivery systems. What differentiates CI&T is not access to models but our ability to wire agents into audited, revenue-bearing delivery systems.

As Head of Delivery at CI&T, I track these signals daily. Below are three scenarios playing out in our delivery landscape, and the playbook we’re using to scale them.

Scenario 1: The "Algorithmic Trading" of Software Delivery

In institutional banking, High-Frequency Trading (HFT) algorithms execute millions of trades per second, capturing value through rapid Signal → Decision → Execution → Value loops. In software delivery, we’re applying that same logic. We aren’t trading stocks; we are trading in throughput.

In institutional banking, High-Frequency Trading (HFT) algorithms execute millions of trades per second, capturing value through rapid Signal → Decision → Execution → Value loops. In software delivery, we’re applying that same logic. We aren’t trading stocks; we are trading in throughput.

By treating code velocity as an algorithmic loop rather than manual labor, we turn repetitive development cycles into compounding capacity. In a recent client sprint, an agent acted as a "delivery algorithm," autonomously iterating on code reviews and A/B-testing refactoring patterns.

The "Velocity Tier" Reality
Based on our current pilots, gains vary across the stack. We categorize these into three tiers of agentic uplift:

TierStack ComponentsVelocity UpliftWhy Velocity Excels
1


Regression testing, Docs, Refactoring
8x–12xHigh-volume pattern matching
2


Feature glue code, API integrations

3x–5xModerate context, standard patterns
3


Architecture, System design, Complex logic

1.5x–2xHumans lead judgment; agents surface risks

The "Trading Desk" Guardrail: To mitigate the risk of regressions, we embed human oversight. Much like a financial trading desk, humans must approve high-impact PRs, and all quality gates must pass before promotion.

Scenario 2: From Generation to Execution

Under the hood, the solution runs on the Databricks Lakehouse Platform, with unified data governance through Unity Catalog, model development and monitoring using MLflow, secure and standardized model access enabled by Mosaic AI Gateway, and interactive interfaces built on Databricks Apps. Conversational insights are delivered through Genie, democratizing complex analytics for non-technical users.

We are currently moving through a capability curve that transforms AI from a passive observer to an active participant in the delivery ecosystem.

  • 2023: Recognition (Read). AI primarily classified Jira tickets and documentation.
  • 2024: Generation (Write). AI drafted tests and PR descriptions, but humans still orchestrated every move.
  • 2025: Orchestration (Pilot Action). Agents handle code reviews and CI/CD tweaks under human supervision.
  • 2026+: Action (Execute). Agents act within bounded, audited environments, triggering AWS deploys and re-sequencing backlogs based on real-time blockers.

In our current pilots, empowering agents to act as virtual coordinators has resulted in a 40% reduction in project delays with zero degradation in defect rates.

Scenario 3: The Playbook for Agent-Led Delivery

1. Audit Loops Ruthlessly
Don’t try to automate novel architecture yet. Focus on your costliest repeat loops: regression testing, compliance checks, and standardized refactoring. We target a minimum 20% efficiency uplift; if an experiment doesn't hit this, we kill it.

2. Forge Digital Twins of Delivery
Agents are only as good as the context they are fed. By feeding agents historical Jira resolutions, Architecture Decision Records (ADRs), and post-mortem transcripts, we create Digital Twins of the delivery organization. These twins shadow architects and PMs to warn about anti-patterns before a single line of code is written.

3. Micro-test with Story Points
Avoid the trap of multi-month RFPs. Run controlled experiments directly in the sprint:
- Budget: Allocate 10 story points to an agent.
- Goal: Achieve a 100-point outcome (e.g., boosting test coverage from 60% to 85%).
- Constraint: Zero drop in quality gate standards.


The Horizon

CI&T pilots indicate that by 2028, 50-70% of delivery hours could be agent-led. This demonetization of routine development shifts the value paradigm:

- Agents will own the "Commodity Layers" (Testing, Maintenance, UI boilerplate).
- Humans will own Direction, Ethics, Governance, and Novelty.

The companies that pass the Modern Turing Test won’t just look faster; they will be structurally different. Velocity and quality will no longer be functions of individual heroics, but results of superior system design.










Roberta

Roberta Lingnau de Oliveira

Senior Manager