
From conversational AI to autonomous development, model orchestration and a new economics of software engineering
We spent decades teaching humans how to communicate with computers. We are now entering an era in which computers increasingly understand what humans want to achieve, determine how to accomplish it and execute much of the work themselves. But this transformation is not simply about artificial intelligence writing code. It is about the emergence of a new architecture of software production, where intelligence becomes a consumable resource, development becomes an orchestrated process, and the traditional boundaries between programmers, tools and computational infrastructure begin to disappear.
There is something particularly interesting about the current evolution of artificial intelligence that I believe deserves more attention than the usual announcements about increasingly powerful language models.
We are witnessing a fundamental change in how software is conceived, developed, tested and maintained.
ChatGPT made artificial intelligence accessible through conversation. GitHub Copilot introduced AI assistance directly into programming environments. Tools such as Cursor, Claude Code and OpenAI Codex are taking us further, towards environments in which artificial intelligence does not merely suggest code but can explore repositories, modify files, execute commands, run tests and iterate towards a defined objective.
Meanwhile, platforms such as OpenRouter are introducing another dimension: the possibility of accessing different AI models through a common interface, selecting them according to capabilities, performance and cost.
Behind all this lies a seemingly simple concept: the token.
Yet reducing this transformation to the cost of processing tokens would be a mistake. What is emerging is an entirely different way of organising software development, with profound implications for technology companies, enterprise architecture, professional skills and, ultimately, the economics of knowledge work.
1. From programming computers to describing intentions
For much of the history of computing, software development has followed a relatively stable principle: humans describe precisely how computers must behave.
Programming languages have evolved dramatically, from assembly language to procedural programming, object-oriented development, declarative languages and increasingly sophisticated frameworks. Nevertheless, the fundamental relationship has remained largely unchanged.
A programmer translates an intention into instructions that a computer can execute.
Artificial intelligence begins to reverse this relationship.
Instead of specifying every implementation detail, developers can increasingly describe the desired outcome and allow an AI system to propose or execute the intermediate steps.
Consider a relatively ordinary requirement:
Develop a web application that allows users to register, authenticate, manage their profiles and receive personalised notifications.
In traditional development, this requirement would trigger a sequence of activities involving architecture, database design, backend development, frontend implementation, security, testing and deployment.
With contemporary AI development agents, significant parts of this process can be initiated through natural language.
The agent can inspect an existing project, identify the relevant framework, create components, modify application logic, propose database structures and execute tests.
This does not mean that human expertise becomes unnecessary. Quite the opposite.
It means that the human contribution moves increasingly towards defining objectives, establishing constraints, evaluating architectural decisions and validating outcomes.
The developer is gradually moving from being the principal author of instructions to becoming the architect and supervisor of an increasingly automated production process.
The distinction is important because producing code and producing reliable software are not equivalent activities.
2. ChatGPT, Codex and the emergence of development agents
The first popular wave of generative AI was fundamentally conversational.
A user submitted a prompt, the model generated a response, and the interaction continued through successive exchanges.
ChatGPT demonstrated how powerful this interaction could become. Developers rapidly discovered that language models could explain unfamiliar code, generate functions, identify potential bugs and translate between programming languages.
But conversational assistance has an important limitation: the model generally operates on the information supplied within the interaction.
It may suggest a solution without having access to the entire repository, the runtime environment or the actual behaviour of the application.
AI development agents change this relationship.
Systems such as Codex and Claude Code can operate within development workflows, interacting with files, repositories, terminals and testing tools, subject to the permissions and controls provided.
Instead of simply responding to a question, an agent can pursue a goal through multiple actions.
For example, an instruction such as:
Identify the cause of this authentication failure, correct it and verify that the existing tests still pass.
can trigger a workflow involving repository exploration, code analysis, modifications, test execution and further corrections.
This introduces a significant architectural distinction.
An LLM is a computational model capable of processing and generating information.
An AI agent is a system that uses a model, together with tools, memory or contextual state, and execution logic, to pursue objectives through a sequence of actions.
Codex is therefore not simply another model competing with GPT, Claude or Gemini. It represents an agentic development capability built around models and software engineering tools.
This distinction is essential to understanding where the market is going.
We are moving from selecting individual AI applications to designing systems in which different forms of intelligence perform different functions.

Figure 1 — From Conversation to Execution: The Evolution of AI-Driven Software Development
The evolution from conversational AI to agentic software development. AI is moving beyond answering questions and suggesting code towards executing complex development tasks, coordinating multiple models and optimising computational resources. The fundamental shift is from human-directed programming to human-supervised execution.
3. The model is becoming a component, not the entire solution
One of the most interesting developments is the gradual separation between the AI model and the environment in which it operates.
During the initial phase of generative AI adoption, users often identified the application with the underlying intelligence.
ChatGPT meant OpenAI. Claude meant Anthropic. Gemini meant Google.
The model, the interface and the service were perceived as a single product.
That relationship is becoming more complex.
Today, developers can access models through provider APIs, cloud platforms, specialised gateways and model-routing services.
OpenRouter is one example of this emerging infrastructure.
It provides a common API through which developers can access a range of models from different providers, depending on availability and commercial arrangements.
This allows applications to become less tightly coupled to a single model provider.
A development workflow might use one model for architectural reasoning, another for routine code generation, and a third for documentation or classification.
The choice can depend on several factors: reasoning capabilities, context-window size, latency, reliability, privacy requirements and price.
Consider an application that needs to analyse thousands of documents, extract structured information, generate reports and answer complex questions.
There is little economic justification for using the most expensive reasoning model for every operation.
Document classification might be performed by a smaller, less expensive model. Complex synthesis could be delegated to a more capable model. Validation might involve deterministic software checks, a second model or human review.
The result is an architecture in which intelligence becomes dynamically allocated according to the nature of the task.
This resembles a principle that has existed in computing for decades.
We do not use the same computational resources for every workload. Different tasks require different combinations of processing power, memory, storage and network capacity.
Artificial intelligence introduces another resource into this equation: cognitive processing capability.
The architectural challenge is no longer simply determining which model is best.
It is determining which combination of models, tools and execution strategies delivers the required result with acceptable quality, cost and risk.
4. The token economy: intelligence becomes measurable consumption
Tokens are the basic units used by language models to represent and process text and, depending on the model, other forms of information.
A token is not necessarily a word. It may represent a word, part of a word, punctuation or another element of the model’s input representation.
For many API-based language-model services, usage is priced according to the number of input and output tokens processed.
This introduces a new economic dimension into software development.
Historically, organisations measured development costs primarily through human effort, infrastructure consumption, software licences and operational expenditure.
AI-assisted development introduces an additional variable: the computational cost of generating, analysing and validating software through models.
A simplified calculation illustrates the principle.
Suppose a development agent processes 100,000 input tokens and generates 20,000 output tokens during a task.
If the selected model costs $2 per million input tokens and $10 per million output tokens, the direct token cost would be:
| Component | Tokens | Illustrative rate | Cost |
|---|---|---|---|
| Input processing | 100,000 | $2 / million | $0.20 |
| Output generation | 20,000 | $10 / million | $0.20 |
| Total | 120,000 | $0.40 |
These are illustrative prices, not a quotation for a particular model or provider. Real costs may also involve cached tokens, reasoning tokens, tool charges, subscriptions and infrastructure.
At first glance, the result seems extraordinary.
A software development task involving thousands of lines of contextual information might incur a relatively small direct model cost.
However, the calculation becomes more interesting when we consider the complete workflow.
An agent may need to repeat operations, inspect large repositories, generate alternative solutions, execute tests and correct mistakes.
Each iteration can introduce additional computational consumption.
A poorly designed workflow may consume substantial resources without producing a reliable result.
A well-designed workflow may solve the same problem using fewer model calls, better context management and more effective verification.
This creates a new optimisation problem.
The relevant economic measure is not the price per million tokens. It is the cost per successfully completed and validated task.
A cheaper model that repeatedly fails may be more expensive than a capable model that solves the problem in one attempt.
Equally, an expensive model used for trivial operations represents unnecessary expenditure.
The economics of AI development must therefore consider quality, latency, failure rates, human intervention and computational consumption together.
5. From IDEs to intelligent development environments
The traditional integrated development environment was designed around a human programmer.
Editors, debuggers, compilers, terminals and version-control tools supported the activities of someone who understood the codebase and directed the development process.
AI-assisted environments are beginning to reorganise this relationship.
Products such as Cursor integrate AI capabilities into the editor itself. Agent-oriented tools such as Claude Code and Codex can work across multiple files and development activities.
The distinction between an editor, an assistant and an execution environment is becoming less clear.
An emerging development architecture may contain several layers.
At the top sits the human objective: the requirement, feature, bug report or architectural constraint.
Below it, an orchestration layer translates that objective into tasks and determines how they should be executed.
A model-access layer connects the workflow to one or more AI models, either directly or through services such as OpenRouter.
An execution layer provides access to repositories, filesystems, terminals, build systems and test environments.
Finally, a validation layer evaluates the results through automated testing, security checks, code review and human approval.
These layers do not necessarily correspond to separate products. They are architectural functions that may be integrated within a single development platform.
The essential transformation is that software development becomes an iterative process involving human direction, machine-generated actions and systematic verification.
This is also where an important misconception must be avoided.
Generating code is not the same as engineering software.
Software engineering includes understanding requirements, managing complexity, anticipating failure, maintaining compatibility, protecting data and ensuring that systems remain understandable and maintainable over time.
An AI agent may accelerate implementation dramatically while introducing subtle architectural defects.
Consequently, the more autonomous development becomes, the more important validation and governance become.

Figure 2 – The Architecture of Token-Driven Development
The emerging architecture of AI-driven software engineering. Development environments connect agents, tools and multiple language models through orchestration and API layers. Token consumption introduces a measurable economic dimension, while automated testing, validation and human oversight remain essential to delivering reliable software.
6. The hidden cost: context, repetition and architectural complexity
The economics of tokens cannot be understood without considering context.
Every interaction with a language model depends on the information available to it.
In software development, this may include source code, configuration files, documentation, previous decisions, error messages and test results.
As projects become larger, providing the appropriate context becomes increasingly difficult.
An agent working on a complex application may need to understand relationships across hundreds of files.
Simply sending the entire repository to a model is often inefficient, expensive or technically impractical.
This is why context management is becoming an essential engineering discipline.
Techniques such as semantic retrieval, repository indexing, selective file inspection, structured memory and context summarisation help provide relevant information without repeatedly processing everything.
But compression and retrieval introduce their own risks.
A missing dependency, an outdated architectural assumption or an incorrectly summarised requirement may cause the agent to generate plausible but defective code.
The problem is not merely the amount of information supplied to the model.
It is whether the model receives the right information at the right moment.
This is remarkably similar to a challenge that organisations have faced for decades.
Having more information does not automatically produce better decisions.
What matters is the ability to identify relevant knowledge, understand relationships and preserve context across activities.
In this sense, AI development environments are becoming miniature knowledge organisations.
They must acquire information, retain decisions, coordinate specialised capabilities and evaluate results.
The similarities with organisational intelligence are difficult to ignore.
7. Multi-agent development: from individual assistants to coordinated intelligence
If a single AI agent can analyse and modify software, why not employ several agents with different responsibilities?
This question is driving experimentation with multi-agent development systems.
One agent might analyse requirements and propose an architecture.
Another might implement the backend.
A third could generate tests.
A fourth might review security implications or assess code quality.
An orchestration mechanism would coordinate their activities and consolidate their outputs.
The approach resembles the division of labour within a software engineering team.
However, the analogy has limitations.
AI agents do not automatically possess independent judgement, shared understanding or reliable coordination.
Multiple agents may reproduce the same errors, disagree without resolving fundamental issues or generate unnecessary computational overhead.
Their apparent independence can also be misleading when they rely on similar underlying models and training assumptions.
Adding more agents does not necessarily produce better results.
Sometimes a single capable agent with well-designed tools and clear validation criteria is more effective than a complicated multi-agent architecture.
The important question is therefore not how many agents an organisation can deploy.
It is how to structure responsibilities, information flows, controls and verification mechanisms so that the overall system performs reliably.
This is precisely the kind of problem that organisational design has always attempted to solve.
The difference is that some of the participating entities are now computational rather than human.
8. What happens to the software developer?
Every major technological transition produces predictions about the disappearance of existing professions.
Artificial intelligence is no exception.
The argument that AI will replace programmers is attractive because it is simple, provocative and easy to communicate.
Unfortunately, it overlooks the difference between performing an activity and assuming responsibility for its consequences.
Software development involves far more than producing source code.
It requires understanding business objectives, negotiating conflicting requirements, designing systems that can evolve, assessing technical debt, protecting users and making decisions under uncertainty.
These activities do not disappear when code generation becomes automated.
Their relative importance changes.
Routine implementation may become less valuable as an isolated skill.
Architectural judgement, systems thinking, domain knowledge, security awareness and the ability to evaluate machine-generated results may become more important.
This does not mean that every existing developer will automatically benefit.
The transition may reduce demand for certain categories of routine work while creating new opportunities in orchestration, validation, integration and AI systems engineering.
Entry-level roles could be particularly affected if organisations automate the simpler tasks traditionally used to train junior developers.
That would create a serious long-term problem.
How do we develop experienced software architects if the early stages of professional learning are increasingly delegated to machines?
The answer cannot simply be that everyone should learn prompt engineering.
Prompting is an interaction technique, not a substitute for understanding software systems.
The professionals who remain valuable will be those who can distinguish a convincing answer from a correct one, identify architectural weaknesses and understand when automated execution should stop.
The scarce resource may no longer be the ability to write code quickly, but the ability to recognise which code should exist, why it should exist and whether it can be trusted.
9. The enterprise challenge: who controls the intelligence?
For organisations, the transition to token-driven and agentic development raises questions that extend far beyond productivity.
If development environments dynamically access multiple external models, where does sensitive information go?
Who determines which model can process proprietary source code?
How are costs allocated across projects and departments?
What happens when a model provider changes its pricing, capabilities or availability?
How can organisations reconstruct the decisions and actions performed by autonomous agents?
These are questions of architecture, governance and accountability.
An enterprise AI development strategy should therefore consider model access policies, identity and permissions, data protection, auditability, cost monitoring, evaluation and human oversight.
Model-routing platforms can provide flexibility, but they do not automatically solve these governance problems.
Indeed, introducing multiple models and providers can increase operational complexity.
A robust architecture must distinguish between experimentation and production, define appropriate trust boundaries and ensure that agents operate with controlled permissions.
The principle of least privilege, familiar from cybersecurity, becomes especially important.
An agent authorised to read source code should not automatically receive unrestricted access to production infrastructure.
Likewise, an agent capable of generating deployment scripts should not necessarily be allowed to execute them without approval.
The emergence of agentic development makes traditional engineering disciplines more relevant, not less.
10. Beyond the token: the industrialisation of intelligence
The token is an interesting economic unit because it makes certain forms of AI consumption measurable.
But it would be misleading to conclude that tokens are becoming the universal currency of intelligence.
Tokens measure processed representations, not understanding, correctness or value.
Two models can consume different numbers of tokens while producing equivalent results.
A longer reasoning process may improve reliability in one context and simply waste resources in another.
A successful task may depend more on access to appropriate tools or reliable information than on the number of tokens processed.
Moreover, not every AI service is billed through tokens. Subscription plans, compute-based pricing, hosted infrastructure and outcome-based commercial models coexist.
Nevertheless, token-based pricing introduces a new degree of transparency into the economics of computational intelligence.
It allows organisations to compare approaches, optimise workflows and begin treating AI processing as an operational resource.
Over time, this may encourage the emergence of new engineering disciplines combining software architecture, model evaluation, cost optimisation and autonomous workflow management.
We may eventually regard model selection and orchestration as ordinary infrastructure decisions, much as we now consider databases, cloud services and network architectures.
The interesting development is not that intelligence becomes cheap.
It is that certain cognitive activities become available on demand, through programmable interfaces, at increasingly measurable costs.
That changes the economics of software production.
It also changes the economics of organisations.
11. A historical perspective: we have seen this pattern before
Having worked with computing technologies since the period when the Internet and the Web were emerging from research environments into mainstream adoption, I find certain parallels particularly striking.
In the early 1990s, while working with Unix systems, Gopher, WAIS and the emerging Web, I experienced a period in which the challenge was not simply creating another application.
It was understanding what happened when previously isolated computers and information systems became interconnected through common protocols.
Working within the early Internet ecosystem, including the IETF environment and the communities surrounding Unix and the emerging Web, made one thing increasingly clear: the most important innovation was not necessarily the individual technology, but the architecture that allowed different technologies to work together.
The Web did not merely simplify access to documents.
It changed the architecture of information distribution, the economics of publishing and eventually the structure of entire industries.
Later, cloud computing changed how organisations acquired and consumed computational infrastructure.
Instead of purchasing and maintaining every physical resource, companies could increasingly provision computing capabilities on demand.
Artificial intelligence appears to be introducing another abstraction.
We are moving towards systems in which certain forms of reasoning, synthesis, code generation and task execution can be accessed programmatically.
But history also suggests caution.
The arrival of a new technological abstraction does not eliminate the need for architecture.
The Internet did not remove the need for network engineering.
Cloud computing did not eliminate infrastructure management.
And AI-generated software will not eliminate the need for software engineering.
Each transition changes where complexity resides.
Often, the complexity does not disappear at all.
It moves from implementation into integration, coordination, governance and control.
The same is happening today.
Writing individual functions may become easier.
Understanding and governing systems composed of humans, models, agents, APIs, tools and external services may become considerably more difficult.
I find another historical parallel particularly revealing.
In the early Internet, open protocols created the possibility of connecting systems developed by different organisations, using different technologies and operating under different commercial models.
Today, model APIs, tool interfaces and orchestration frameworks are beginning to play a comparable role.
The analogy is not perfect. Many contemporary AI ecosystems remain proprietary, commercially controlled and technically fragmented.
Nevertheless, the direction is recognisable.
The real value may emerge not from the intelligence of individual models, but from the architecture that allows different forms of intelligence to cooperate.
That was an important lesson of the Internet revolution, and it may become equally important in the age of AI agents.

Figure 3 – From AI Agents to Organisational Intelligence
The transition from individual AI assistants to coordinated organisational intelligence. The integration of specialised agents, shared knowledge, model orchestration, development tools and human governance transforms isolated automation into a coherent system capable of delivering measurable organisational value.
12. The next frontier is not artificial intelligence. It is organisational intelligence.
This brings me to a broader reflection.
Much of the current debate remains focused on comparing models.
Which model writes better code?
Which agent completes tasks faster?
Which provider offers the lowest token price?
These questions matter, but they are not sufficient.
The deeper transformation concerns how intelligence is organised.
An enterprise may have access to excellent models and sophisticated development agents while remaining incapable of producing reliable outcomes.
Why?
Because intelligence is not simply a property of individual components.
It also emerges from the relationships between those components, the information they share, the decisions they make and the mechanisms through which their activities are coordinated.
An organisation that deploys ten AI agents without a coherent architecture may become less effective than one that carefully integrates a single agent into a well-designed workflow.
Similarly, a company that reduces the cost of code generation without improving software quality may simply accelerate the production of technical debt.
The objective should not be to maximise automation.
It should be to improve the organisation’s ability to transform intentions into reliable outcomes.
This requires combining human expertise, computational intelligence, institutional knowledge and governance.
It also requires recognising that efficiency and intelligence are not synonymous.
An organisation can execute the wrong strategy with extraordinary efficiency.
An AI development environment can generate thousands of lines of perfectly valid code for a system that should never have been built.
The most important decisions remain those concerning purpose, relevance, consequences and responsibility.
This is also why I believe the evolution from ChatGPT to agentic development should not be interpreted as a purely technological phenomenon.
It is an organisational transformation.
We are creating environments in which human and artificial capabilities must coexist, exchange information, resolve conflicts and contribute to common objectives.
The challenge is not fundamentally different from the one faced by organisations attempting to coordinate people with different competencies, responsibilities and perspectives.
The tools are new.
The underlying questions are remarkably familiar.
- Who decides?
- Who executes?
- Who validates?
- Who is accountable?
And, perhaps most importantly, who has the authority to question the result?
An intelligent organisation is not one in which every decision is automated.
It is one capable of combining different forms of intelligence while preserving the ability to recognise errors, challenge assumptions and change direction.
The same principle should guide the design of AI development systems.
Conclusion – When code becomes abundant, judgement becomes scarce
We are moving from a world in which software development was constrained primarily by the availability of human programming capacity towards one in which substantial parts of implementation can be performed by computational agents.
ChatGPT introduced millions of people to conversational intelligence.
Codex, Claude Code and similar tools are extending that interaction into executable development workflows.
Model gateways such as OpenRouter are making it easier to access and compare different sources of AI capability.
Token-based pricing is introducing new ways of measuring the computational cost of these activities.
Together, these developments suggest a profound transformation in software engineering.
But the real change is not that developers will write fewer lines of code.
It is that the relationship between human intention, computational execution and organisational responsibility is being redesigned.
For decades, we have attempted to make computers more capable of executing human instructions.
Now we are beginning to build systems capable of determining many of the intermediate instructions themselves.
That is a much larger transition than another improvement in programming productivity.
It forces us to reconsider what expertise means, how organisations distribute responsibility and where human judgement should remain indispensable.
There is also a paradox worth considering.
As the marginal cost of generating software decreases, organisations may find themselves producing more applications, more features, more integrations and more automated processes than they can meaningfully govern.
The abundance of code could create a scarcity of understanding.
We have already seen something similar with information.
The Internet dramatically reduced the cost of publishing and distributing knowledge, but it did not automatically make societies better informed.
Artificial intelligence may dramatically reduce the cost of producing software without automatically making organisations more intelligent.
The difference will depend on our ability to design systems in which computational efficiency serves meaningful objectives rather than becoming an objective in itself.
The future of software development is not about writing more code. It is about orchestrating intelligence.
And perhaps it leads to a question more important than the cost of a million tokens:
When producing software becomes almost effortless, will we become better at deciding what is worth building?
That, rather than the next model benchmark, may ultimately determine whether the agentic development revolution creates genuine progress.
References and further reading
The following primary resources provide technical background for the technologies and architectural concepts discussed.
- OpenAI — Codex
https://openai.com/codex
AI-assisted software development and coding agents. - OpenAI — API Platform
https://platform.openai.com/docs
Model APIs, token usage and development integration. - Anthropic — Claude Code
https://code.claude.com/docs
Agentic coding workflows. - OpenRouter — Documentation
https://openrouter.ai/docs
Unified model access, routing and API integration. - GitHub — Copilot
https://github.com/features/copilot
AI-assisted software engineering. - Cursor
https://cursor.com
AI-integrated development environments. - Martin Fowler
https://martinfowler.com
Software architecture, engineering practices and AI-assisted development. - NIST — AI Risk Management Framework
https://www.nist.gov/itl/ai-risk-management-framework
Governance, reliability and risk management for AI systems. - Model Context Protocol (MCP)
https://modelcontextprotocol.io
Open protocol for connecting AI applications to external tools and information sources. - Google — Agent Development Kit
https://google.github.io/adk-docs/
Framework and documentation for developing agent-based AI systems.
This article reflects the technological landscape of October 2026. Product capabilities, availability and pricing are subject to change.











