Concept Distillation for Weak Language Model Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing generative AI models, both weak and strong, suffer from inaccuracies such as hallucinations and require significant computational resources, leading to a trade-off between accuracy and efficiency.

Innovation Solution

A model distillation system that uses concept distillation to transfer rich features from strong generative models to weak generative models, improving accuracy without retraining, thereby enhancing efficiency and flexibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If newer versions of generative AI models are used to improve accuracy, then model performance improves, but computational resources required increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the knowledge and capabilities of a strong generative model into specific concepts and features, then selectively transfers only the necessary portions to a weak model through concept distillation. This avoids the need to use the entire strong model, thereby reducing computational resource requirements while maintaining accuracy improvements for specific tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts implicit rich features and concepts from a strong generative model's response to a prompt, separating the essential knowledge from the model's full computational structure. This extracted knowledge is then used to enhance a weak model without requiring the weak model to have the computational overhead of the strong model.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If concept distillation is used to transfer features from strong models to weak models, then accuracy of weak models improves, but system complexity increases

Engineering Contradiction:
Improveweak model accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary process where a strong model generates a response to a prompt, and then concept distillation extracts key features from this response. This intermediary extraction step acts as a bridge, transferring knowledge without requiring direct integration of the complex strong model into the deployment system, thereby managing system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a simplified copy or representation of the strong model's knowledge through concept distillation. Instead of using the full strong model, it copies only the essential concepts and features needed for specific tasks into the weak model, reducing the complexity burden while preserving accuracy benefits.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4647966A1Using large generative models to improve the performance of weak language models in performing complex tasks
Publication Date: 2025.11.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4647966A1 patent drawingFigure 1
  • EP4647966A1 patent drawingFigure 2
  • EP4647966A1 patent drawingFigure 3A

AI summary

This disclosure describes a model distillation system that implements a framework for improving and enhancing the reliability of weak generative models. For example, the model distillation system uses concept distillation for prompt construction to improve the accuracy of weak generative models while maintaining their efficiency advantage over strong generative models. In particular, the model distillation system determines and transfers implicit rich features of a strong generative model to a weak generative model for specific topics and concepts. By using these rich features, the weak generative model can correctly answer queries and prompts for the specific topics and concepts that it would otherwise answer incorrectly. Furthermore, the model distillation system transfers these rich features without needing fine-tuning or retraining, resulting in improved accuracy while still maintaining high levels of efficiency.