A cross-task collaborative thought chain distillation method, device, system, and storage medium

By generating thinking chain datasets and using LoRA fine-tuning and grouping orthogonal regularization methods, the problem of negative migration of small language models in multi-task collaboration is solved, and its reasoning ability and cross-task generalization performance are improved.

CN119623654BActive Publication Date: 2025-08-15NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411676515.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-08-15
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In the prior art, small language models cannot improve their inference ability when using thinking chain prompts, and have negative migration problems, especially in multitasking collaboration.

Method used

By generating a thinking chain of inference task data sets, sorting according to task difficulty, orthogonal grouping training is performed using LoRA fine-tuning and grouping orthogonal regularization methods to implicitly isolate parameters between unrelated tasks, prevent negative migration, and enhance the inference ability of small models.

Benefits of technology

It effectively improves the cross-task collaboration capabilities of small language models, reduces negative migration, improves the model's performance on complex inference tasks, and maintains generalization performance on unknown tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623654B_ABST
    Figure CN119623654B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-task collaborative thought chain distillation method, device, system, and storage medium, comprising: step S1, generating thought chains for a reasoning task dataset; step S2, ranking the difficulty of the tasks to be trained based on the thought chains; step S3, orthogonally grouping and training student models based on the thought chains and task difficulty; and step S4, distilling the cross-task collaborative thought chains based on the LoRA expert combination. The technical solution of the present invention prevents negative transfer by implicitly isolating parameters between unrelated tasks, further enhancing the reasoning capability of small models.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A cross-task collaborative thought chain distillation method, characterized by: include: Step S1: Generate a thought chain for the reasoning task dataset; Step S2: sorting the difficulty of the tasks to be trained according to the thought chain; Step S3: conducting orthogonal group training on student models according to thinking chains and task difficulty; Step S4: Distill cross-task collaborative thinking chains based on the LoRA expert combination; Each reasoning task contains complex reasoning questions and the correct answer Where t≤m; using Zero-shot-CoT, that is, adding Let's think step by step after the reasoning problem. Guide the LLM with CoT capability to solve the problem Where i≤n, generating multi-stage explanations and predict the results According to the correct answer and prediction results Choosing the right multi-stage explanation use Form a sample The training method of step S3 is LoRA fine-tuning, and the LoRA fine-tuning strategy includes similarity detection and group orthogonal regularization; among them, similarity detection is used to detect whether there are tasks similar to the current task in the trained tasks before training the current task; group orthogonal regularization is added between the LoRA experts of dissimilar tasks to constrain the current task to update parameters in a direction orthogonal to the gradient of the dissimilar tasks.

2. A cross-task collaborative thinking chain distillation device for implementing the cross-task collaborative thinking chain distillation method according to claim 1, characterized in that: include: The first processing module is used to generate a thought chain for the reasoning task dataset; The second processing module uses the thought chain to sort the difficulty of the tasks being trained; The third processing module uses orthogonal grouping training of student models based on thinking chains and task difficulty; The fourth processing module uses cross-task collaborative thought chain distillation based on LoRA expert combination.

3. A cross-task collaborative thought chain distillation system, characterized by: include: A memory and a processor, wherein the memory stores a computer program run by the processor, and when the computer program is run by the processor, the cross-task collaborative thinking chain distillation method as described in claim 1 is executed.

4. A storage medium, characterized in that The storage medium stores a computer program, which executes the cross-task collaborative thinking chain distillation method as described in claim 1 when running.

Citation Information

Patent Citations

  • Language model training method and device and computer readable storage medium

    CN117669767A

  • Pre-training language model parameter fine tuning method and device, equipment and medium

    CN117829240A