This website uses cookies

Read our Privacy policy and Terms of use for more information.

Welcome Back to XcessAI

Almost exactly one year ago, we published an article called Small Giants.

At the time, the AI industry was obsessed with size. Larger models, more parameters, more compute and ever larger data centres dominated the conversation.

But most of the work we wanted AI to perform wasn't particularly complicated.

Summarise a document. Extract some information. Classify an email. Rewrite some text. Call an API.

Using the world's most powerful AI model for every one of those tasks looked increasingly like using a Formula 1 car to deliver groceries.

So we argued that Small Language Models, or SLMs, could eventually perform much of the everyday work while larger frontier models were reserved for the difficult problems.

One year later, the thesis hasn't disappeared.

It has evolved.

Small models aren't replacing the giants.

They are becoming the infrastructure underneath them.

The Small Model Grew Up

The biggest change isn't simply that small models have become smarter.

They have become useful in more places.

Microsoft's new Aion models are specifically designed for local computing. Aion Instruct handles everyday tasks such as summarisation and rewriting directly on a device, while the larger Aion Plan can reason, call tools, manage files and orchestrate other agents locally.

Microsoft describes the objective quite explicitly: AI without cloud dependency or per-token cost.

Apple is taking a similar approach.

Its third-generation Apple Foundation Models include a roughly 3-billion-parameter model designed to run directly on devices, alongside a more powerful 20-billion-parameter model that uses a sparse architecture, activating only 1 to 4 billion parameters depending on the request.

IBM's latest Granite 4.2 family starts at just 3 billion parameters and is purpose-built for enterprise agents, combining reasoning, coding, tool use and instruction following.

NVIDIA's Nemotron 3 Nano Omni takes the idea into multimodal AI, combining text, images, audio and video in a compact architecture designed to act as the "eyes and ears" of broader agent systems.

And Google is pushing small models beyond laptops and phones. Its latest Gemini Robotics On-Device system is based on its on-device Gemma technology and is designed to run directly on robots rather than depending continuously on the cloud.

A year ago, small models were interesting primarily because they were cheaper.

Today, they are becoming architectural building blocks.

Intelligence Moves Local

The first generation of generative AI was overwhelmingly cloud-based.

You typed something into an application.

The request travelled to a data centre.

A huge model processed it.

The answer came back.

That architecture made perfect sense when useful artificial intelligence required enormous amounts of compute.

But increasingly capable models can now run directly on phones, laptops, workstations and machines.

Google's Gemma 3n, for example, can process text, images, audio and video locally on phones, tablets and laptops.

That changes the architecture.

Imagine an AI assistant working on your computer.

Most of the day it might be summarising emails, organising documents, understanding what is on your screen, searching local files or deciding which application to use.

Those tasks don't necessarily require the most powerful model on Earth.

A smaller local model may be perfectly capable of doing them.

Only when the problem becomes difficult does the system need to escalate.

You can imagine a hierarchy:

Local model → Enterprise model → Frontier model

The question therefore becomes less:

Which model is best?

And more:

How much intelligence does this particular task actually require?

The Economics Start to Matter

This becomes particularly important as AI moves from conversations to agents.

A human might ask an AI assistant ten questions in a day.

An autonomous agent could potentially perform thousands of actions.

Read a document.

Extract information.

Check a database.

Call another system.

Summarise the result.

Decide what happens next.

Every cloud inference has a cost.

For occasional use, that cost may be negligible.

At enterprise scale, it starts to compound.

Running more intelligence locally changes the economics. Some of the expense shifts away from metered cloud inference and towards computing hardware that the company or user already owns.

Microsoft now explicitly talks about "unmetered intelligence" on Windows, where suitable workloads can run locally without a per-token charge.

There are other benefits too.

Latency can fall.

Applications can continue operating without an internet connection.

And sensitive information may remain on the device rather than travelling to an external service.

Suddenly, small models aren't merely a cheaper version of large models.

They become a different way of architecting AI.

From One Giant to Many Specialists

There is another shift underneath all of this.

The original AI race was largely about creating one model capable of doing everything.

But most complex systems don't work that way.

Companies don't employ one extraordinarily capable person to perform accounting, cybersecurity, customer service, legal analysis, procurement and engineering.

They build specialised teams.

AI systems may evolve in the same direction.

One model understands documents.

Another processes speech.

Another sees images.

Another writes code.

Another interacts with software.

A larger reasoning model coordinates the system and calls a frontier model only when something genuinely difficult appears.

NVIDIA's Nemotron strategy already reflects this architecture. Its compact multimodal model is explicitly designed to work as one component inside a wider system of agents rather than trying to be the entire system itself.

The architecture starts looking less like one enormous brain.

And more like software.

Specialised components.

Different tools.

Orchestration between them.

Smaller Doesn't Mean Simpler

There is also a reason parameter count is becoming a less useful way to think about AI capability.

Modern architectures don't necessarily use every parameter for every task.

Apple's new AFM 3 Core Advanced contains 20 billion parameters, for example, but its sparse architecture activates only between 1 and 4 billion at a time depending on what it is doing.

Techniques such as distillation, quantisation, better training data and sparse architectures allow developers to extract increasingly useful intelligence from much less computation.

So perhaps Small Language Model itself will eventually become an imperfect label.

What matters isn't simply how many parameters a model contains.

It is:

How much useful intelligence can we extract from a given amount of compute, memory, energy and cost?

That is a very different optimisation problem from simply asking who can build the biggest model.

The Hybrid Future

Small models don't need to beat large models.

They need to make calling them unnecessary most of the time.

Imagine an enterprise agent performing 100 tasks.

Perhaps 70 can be completed locally by relatively small models.

Twenty require something more capable.

Eight need serious reasoning.

And two genuinely require the frontier.

The frontier model remains enormously important.

But it sits at the top of a pyramid rather than underneath every interaction.

That architecture could be cheaper, faster, more private and easier to control.

And when agents begin performing thousands or millions of actions, those advantages compound.

The New Question

A year ago, the AI industry largely asked:

How powerful can we make the model?

A different question is starting to emerge:

How little model do we need to solve the problem?

That sounds like a subtle change.

Economically, it could be enormous.

Because once useful intelligence becomes small enough to live inside browsers, phones, laptops, cars, robots and machines, AI stops being something we always connect to.

It becomes something embedded in the infrastructure around us.

The giants will continue pushing the frontier.

But increasingly, the everyday work of artificial intelligence may happen somewhere else entirely.

Inside the small giants we barely notice are there.

Until next time,
Stay adaptive. Stay strategic.
And keep exploring the frontier of AI.

Fabio Lopes
XcessAI

💡Next week: I’m breaking down one of the most misunderstood AI shifts happening right now. Stay tuned. Subscribe above.

Read our previous episodes online!

Reply

Avatar

or to participate