Knowledge Distillation for Multi-Model Business Data Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In enterprise office scenarios, efficiently and cost-effectively applying large language models to multiple business types such as customer service dialogue, enterprise document processing, and intelligent recommendations is challenging due to high computational resource overhead and complex deployment.
Innovation Solution
The method involves performing knowledge distillation on multiple pre-trained large models to obtain a base model shared among different business types, and then further distilling the base model to generate a target business model specific to each business type, enabling efficient data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple pre-trained large models are applied to different business types, then the accuracy and capability for each business type is improved, but the computational resource overhead and deployment complexity increase
Solution Approach 1:
The patent segments the application of large models by creating a base model that serves multiple business types, then further segments into business-specific models only when needed. This hierarchical segmentation reduces overall complexity while maintaining accuracy for each business type.
Solution Approach 2:
The base model is designed with multi-functionality to handle multiple business types (customer service dialogue, enterprise document processing, intelligent recommendations). This universality reduces deployment complexity by eliminating the need to deploy separate models for each business type while maintaining high accuracy through knowledge distillation.
2Productivity
If multiple pre-trained large models are deployed for different business types, then the processing capability for each business type is improved, but the computational resource usage increases
Solution Approach 1:
The patent merges multiple pre-trained large models into a single base model through knowledge distillation. This combining approach maintains the processing capability for multiple business types while significantly reducing computational resource usage by eliminating redundant model instances.
Solution Approach 2:
Instead of copying and deploying multiple full pre-trained models, the patent creates a distilled base model that captures the essential knowledge from multiple teacher models. This copying approach maintains processing capability while reducing computational resource consumption through model compression.
3Device complexity
If a single base model is used for multiple business types, then the deployment complexity is reduced, but the accuracy for specific business types may decrease
Solution Approach 1:
The patent performs preliminary knowledge distillation to create a base model that pre-learns from multiple business types. This preliminary action enables the base model to maintain high accuracy across different business types while reducing deployment complexity, as the knowledge is already distilled and organized in the base model structure.
Solution Approach 2:
The patent applies local quality by allowing the base model to have different levels of specialization for different business types. The base model maintains general capabilities while incorporating business-specific knowledge where needed, ensuring high accuracy for each business type while keeping deployment simple through a unified architecture.
Data Source
AI summary
Method and apparatus for target business model generation and data processing based on large language model are disclosed, which relates to the field of artificial intelligence technology, specifically in the areas of intelligent office, big data, and large models. A method for generating a target business model based on large language model includes: performing knowledge distillation on at least two pre-trained large models to obtain a base model of a target scenario, wherein each pre-trained model corresponds to one of at least two business types included in the target scenario; performing knowledge distillation on the base model to obtain a target business model of a target business type among the at least two business types, wherein the target business model is used for processing data of the target business type.


