A cross-task collaborative thought chain distillation method, device, system, and storage medium
By generating thinking chain datasets and using LoRA fine-tuning and grouping orthogonal regularization methods, the problem of negative migration of small language models in multi-task collaboration is solved, and its reasoning ability and cross-task generalization performance are improved.
Patent Information
- Application Number
- CN202411676515.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-11-22
AI Technical Summary
In the prior art, small language models cannot improve their inference ability when using thinking chain prompts, and have negative migration problems, especially in multitasking collaboration.
By generating a thinking chain of inference task data sets, sorting according to task difficulty, orthogonal grouping training is performed using LoRA fine-tuning and grouping orthogonal regularization methods to implicitly isolate parameters between unrelated tasks, prevent negative migration, and enhance the inference ability of small models.
It effectively improves the cross-task collaboration capabilities of small language models, reduces negative migration, improves the model's performance on complex inference tasks, and maintains generalization performance on unknown tasks.
Smart Images

Figure CN119623654B_ABST
Abstract
Claims
1. A cross-task collaborative thought chain distillation method, characterized by: include: Step S1: Generate a thought chain for the reasoning task dataset; Step S2: sorting the difficulty of the tasks to be trained according to the thought chain; Step S3: conducting orthogonal group training on student models according to thinking chains and task difficulty; Step S4: Distill cross-task collaborative thinking chains based on the LoRA expert combination; Each reasoning task contains complex reasoning questions and the correct answer Where t≤m; using Zero-shot-CoT, that is, adding Let's think step by step after the reasoning problem. Guide the LLM with CoT capability to solve the problem Where i≤n, generating multi-stage explanations and predict the results According to the correct answer and prediction results Choosing the right multi-stage explanation use Form a sample The training method of step S3 is LoRA fine-tuning, and the LoRA fine-tuning strategy includes similarity detection and group orthogonal regularization; among them, similarity detection is used to detect whether there are tasks similar to the current task in the trained tasks before training the current task; group orthogonal regularization is added between the LoRA experts of dissimilar tasks to constrain the current task to update parameters in a direction orthogonal to the gradient of the dissimilar tasks.
2. A cross-task collaborative thinking chain distillation device for implementing the cross-task collaborative thinking chain distillation method according to claim 1, characterized in that: include: The first processing module is used to generate a thought chain for the reasoning task dataset; The second processing module uses the thought chain to sort the difficulty of the tasks being trained; The third processing module uses orthogonal grouping training of student models based on thinking chains and task difficulty; The fourth processing module uses cross-task collaborative thought chain distillation based on LoRA expert combination.
3. A cross-task collaborative thought chain distillation system, characterized by: include: A memory and a processor, wherein the memory stores a computer program run by the processor, and when the computer program is run by the processor, the cross-task collaborative thinking chain distillation method as described in claim 1 is executed.
4. A storage medium, characterized in that The storage medium stores a computer program, which executes the cross-task collaborative thinking chain distillation method as described in claim 1 when running.
Citation Information
Patent Citations
Language model training method and device and computer readable storage medium
CN117669767A
Pre-training language model parameter fine tuning method and device, equipment and medium
CN117829240A