AI Architecture in 2026: The Model Is the Easy Part

The stack around the LLM is the actual architecture — agents, memory, skills, MCP, A2A. A 27-minute conversation with Manoj Mukerjee maps the whole thing.
I watched the first episode of AI Architect Talks end to end because the title promised a survey of the 2026 stack, and surveys are where you catch what you have been explaining badly. Host Rakesh Saini sits with Manoj Mukerjee, an engineer who started building intelligent systems back in 2012, and they walk the whole map: what an agent actually is, whether RAG is dead, what MCP and A2A each do, when multi-agent systems are worth the trouble, and what a junior developer should learn first.
Almost none of it was new. All of it was worth hearing in order.
The black box grew a body
Manoj's opening answer to "what changed" is almost too clean. In 2023 he built a single LLM call: input goes in, output comes out, and the model is a black box sitting in the middle of your code. Today the model is still a black box, but it now sits inside memory, tools, protocols, an observability stack, and a cloud runtime. That is what he means by AI architecture: not the model, the blueprint around it.
An agent is a body, and the LLM is the head
The most useful frame in the episode is a comparison to a human body. The LLM is the brain, the processing unit. Long-term and short-term memory are memory. Hands are tools, the parts that touch the outside world. The nervous system is MCP and the other protocols that carry signals between parts. Reflection is observability: watching what the agent did and whether it worked.
Add a goal, a plan, and a set of skills, and you have an agent. Skills are the part people skip. His definition is practical: a skill is a knowledge base written in a format the agent can read. The model does not arrive knowing your domain. You teach it, in a structure.
Chatbot, RAG and agent are three words for three different things
A chatbot is a platform, the surface where a question turns into an answer. Behind it you might have an agent, a RAG pipeline, or a plain rule-based engine. RAG is the knowledge base, an ingredient. The agent is the worker that decides what to fetch before it answers.
He puts it bluntly: the chatbot is not the system, it is the window you look through.
Is RAG dying? No, it got wrapped
RAG is not going anywhere. What changed is the wiring.
The first wave was embeddings into a vector database. Then came agentic RAG, where MCP enters and retrieval becomes something the agent goes and does, calling tools instead of matching text. Then adaptive RAG, which switches strategy per query and decides on each turn how much retrieval the question actually needs.
The knowledge base never left. It became the layer the agent reads from.
Two cables that get confused for each other
MCP, the Model Context Protocol, is agent to tools. It runs on top of HTTP. Anthropic put it forward, everyone adopted it, and Manoj's analogy is the Type-C cable: every phone used to ship its own charger, now one connector fits. Before MCP, every framework had its own tool format.
A2A, agent to agent, came out of Google's Agent Development Kit so that agents built in different places could talk.
His warning is the useful part: do not mix them up. MCP for tools, A2A for peers. Different purpose, different benefit.
He also pushes back on "MCP is dying": if you have a twenty-year-old REST system, an ERP, wrapping it as an MCP server is what gives the agent real data. A model connected to your actual systems has less to invent. That is the anti-hallucination move: connection, not cleverness.
And there is an npm for it. Anthropic hosts a registry of MCP servers the way npm hosts packages. Skills work the same way.
One goal beats a hundred tools
On multi-agent systems, Manoj is more conservative than the timeline. One agent can carry a hundred tools, he says, but it follows one principle: one goal, with minimal tools around it. Specialists communicating through a supervisor pattern, with handoffs between them, earn their complexity only when the work genuinely transforms from one-to-many.
Deep agents get reserved for the cases that deserve them. A complex engine with planning, execution, stopping points, and a human in the loop. Writing one email does not need a deep agent. It needs an email agent.
Three layers, one ladder
He splits where things are heading into three layers. AI operator is the copilot layer. AI builder is where LangChain and Langflow live. AI innovation is model power itself.
Underneath sits a career ladder: software engineer, then AI engineer, then AI architect, someone carrying security, performance and scalability, building systems that are powerful but under human control.
His advice for someone starting out is almost boring: learn what an LLM is and how to talk to it, learn prompt engineering, then AI engineering, then agents and their protocols, then enterprise architecture. Learn SDLC first and put AI engineering on top of it, not instead of it.
Keep learning, keep exploring, keep sharing
The human brain is better than AI. It just works better with a blueprint.
Source: "AI Architecture in 2026: LLMs, AI Agents, MCP & A2A Explained" — AI Architect Talks, Episode 1 (27:35, 19 Jul 2026). Host Rakesh Saini; guest Manoj Mukerjee. Published by Su Qin, CMO of DXP.
Working on something similar? Fastest over WhatsApp.
Message me on WhatsApp →