Jensen Huang, the CEO of Nvidia considers AI distillation to be competition. AI distillation is a process whereby AI models are trained using the responses of existing AI models.
The largest AI model providers have expressed opposition to the distillation of their AI models by upcoming AI product vendors.
Some argue that distillation is not fair to the frontier model providers OpenAI and Anthropic, as their models are being used to train competitors at a fraction of the cost they had to pay to train their own. They referred to the use of distillation as ‘distillation attacks‘.
However, frontier model vendors have been sued for using authors’ articles and books to train their models without permission.
Distillation is highly efficient due to the fact that it requires less effort and system resources for training. Training without distillation incurs large expenses for multiple reasons, one of which is the cost of obtaining copyrighted material such as books.
Distillation is not immune to the issue of AI hallucination, which puts distilled AI models at a disadvantage. The fact that the data training distilled models comes from other models that sometimes generate false responses means that accuracy is a concern.
Distillation can also be done by asking the source model (for example: Claude Opus) to generate fake data for data sets for classification projects (not for anything that provides factual information).
It should also be noted that existing AI models can be used to distill and train models accurately if humans are carefully curating the output of the source models before using it for training.
One example of the use of high quality, synthetic datasets to train AI models is Microsoft Phi-4-mini. It requires a graphics card with only 3 GB of RAM and offers high performance compared to other small models.
Phi-4-mini is a 3.8 billion parameter model that can run on mobile phones. This model wasn’t necessarily trained using distillation. However, it is an example of what can be achieved using generated, synthetic data.
As for the performance of distilled models, DeepSeek has provided distilled models with impressive performance that enabled them to compete with the world’s largest models.
Some of DeepSeek’s models have already gained the interest of businesses that use AI to automate workloads. They have shifted some of their workloads away from the largest subscription-based and usage-based models to DeepSeek models to reduce their API bills.
In short: They gave ChatGPT/Claude less work to do because DeepSeek models are cheaper to run and saved those frontier models’ available tokens for more important tasks.
Small, distilled AI Models have facilitated the implementation of offline AI models that can run on phones, automobile computers, laptops, appliances and more at a low cost and low energy usage. Distillation is ushering in a new era of environmentally friendly AI in which everyone can access AI offline with control of their own data — and no expensive subscriptions.
Image credit: Pavel Danilyuk via Pexels.
