(Nikolaus Skene)
Sometimes you fancy a glass of white wine.
Sometimes a red.
And after a workout, a beer is also a real treat.
Nobody would dream of ordering all of those things at once at that very moment.
No sommelier in the world would say: “No problem, I’ll bring you white wine, red wine, a beer, a steak, a towel, a massage, a phone call with your mum and, just in case, a skiing holiday.”
Our brain is far more efficient in this respect.
It makes decisions based on context. Situationally. In an energy-efficient way.
It activates only those regions that are currently relevant. The rest remain ‘at rest’.
And this is precisely where an often-overlooked difference lies between human intelligence and today’s artificial intelligence.
And this is where ChatGPT comes into play.
At its core, ChatGPT works differently.
No matter what you fancy at the moment – it always asks about everything.
Not just the beer after sport, but at the same time the beer, the white wine, the red wine, the steak, the recipe, the supply chain, the carbon footprint, the wine region, the sport and the philosophical significance of thirst.
The reason we don’t notice this is that:
It happens extremely quickly.
The irrelevant ‘areas’ are filtered out in milliseconds, and in the end a reasonably sensible answer emerges. But the computational effort behind it is enormous.
The human brain would be completely overwhelmed by this.
And that is precisely why the energy consumption of today’s AI systems is a serious cause for concern.
DeepSeek is already taking this a step further.
It attempts to ‘fire’ different parts of the model in a more targeted way, depending on what is likely to be relevant at any given moment. Not everything, all the time, simultaneously. But more selectively. More context-sensitive.
And then there’s HuggingGPT.
HuggingGPT is not a new model.
It is an architectural concept.
The core idea:
A large language model such as ChatGPT is no longer used as a universal genius, but as a controller. As a conductor. As a planner.
It listens to the user, understands their intention – and then breaks it down into specific subtasks.
These subtasks are not solved by the same model, but are specifically passed on to specialists: image models, audio models, classifiers, text-to-speech systems, object recognition, segmentation. All models that already exist, for example within the Hugging Face ecosystem.
What matters here is not just what is done, but in what order.
HuggingGPT manages dependencies.
Some tasks must be processed sequentially, whilst others can run in parallel. Results from one model are used as input for the next. Only at the end does the controller gather everything together again and formulate a response.
A simple query such as ‘Describe this image in detail’ is automatically broken down into image classification, object recognition, segmentation, captioning and visual question-answering models – and only then are the results combined.
This is not a new form of intelligence.
But it is a different way of working intelligently.
And this is where HuggingGPT gets really interesting:
The system scales not through ever-larger models, but through coordination. Through selection. Through the division of labour. Through context.
Suddenly, AI no longer looks like an overwhelmed all-rounder who has to be able to do everything at once, but rather like a well-organised team.
It still has a few ‘glitches’ here and there: higher latency, dependence on the controller’s planning capabilities, token limits, instabilities. HuggingGPT is not a finished product. It is a conceptual model.
But one that points in an important direction:
Away from ‘bigger, faster, everything at once’.
Towards ‘more accurate, more selective, more efficient’.
And with that, back to energy consumption. One of the main and justified criticisms of today’s AI is its voracious appetite for energy. If intelligence no longer arises from maximum parallelism, but from targeted orchestration, we may be moving towards a more efficient form of machine thinking.
What this means in the long term – including for human work, human intelligence and human relevance – is a discussion for another time.
For another Saturday.
Today, one observation will suffice:
We need to take a more measured approach again.
Not everything at once.
But the right things, at the right time.
(Nikolaus Skene lives and works in San Francisco, advises companies on AI and has been organising tours of Silicon Valley and other hotspots of technological development for several years
