Training method of thinking chain compression large model for rewriting task, electronic equipment and storage medium

By introducing a thinking chain compression method in a multi-round dialogue system, generating prompt words by splicing dialogue history and current questions, and gradually removing the thinking process of the big model, the problems of long inference time of traditional thinking chain and high dependence on knowledge distillation methods are solved, and more accurate and efficient dialogue replies are achieved.

CN119940553APending Publication Date: 2025-05-06AISPEECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510104610.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The use of traditional thinking chains in multi-round dialogue systems requires large models to reason step by step, increasing the reasoning time; while using knowledge distillation methods is highly dependent on training data, may have model bias, and reduced transparency and interpretability.

Method used

A training method for compressing large models for thinking chains for rewriting tasks is proposed. By splicing multiple rounds of dialogue history and current questions, and replacing the template of the thinking chain prompt word, generating prompt words, inputting the big model to generate thinking process, and training the big model based on the generated thinking process and training data set, gradually removing the thinking process of the big model during the training process to realize the compression of the thinking chain.

Benefits of technology

While making full use of the prior knowledge of big models, it can better capture the semantic connection characteristics of historical information and the current dialogue, accurately locate the target referential object, and then generate target reply more accurately and efficiently, significantly shortening the reasoning time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940553A_ABST
    Figure CN119940553A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a thinking chain compression large model for rewriting tasks, which comprises the following steps of: splicing multiple rounds of dialogue history and current questions in a first training data set, and replacing corresponding contents in a thinking chain cue word template by using a spliced result to obtain cue words, the first training data set comprises a multi-round dialogue history, a current question and a standard answer, the thinking chain cue word template is a multi-round dialogue text complementing device, and content involved in current input is analyzed and returned by combining historical input and historical output; the cue words are input into the large model to generate a thinking process, a second training data set is generated based on the thinking process and the first training data set, and the thinking process is divided into multiple stages; and training the large model based on the second training data set, and hierarchically and gradually removing a certain thinking process of the large model on the second training data set in the training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of model training, and in particular relates to a training method for a large thought chain compression model for a rewriting task, an electronic device and a storage medium. Background Art

[0002] With the development of deep learning and big data technology, human-computer dialogue systems have become an important research direction in the field of artificial intelligence and have been widely used in scenarios such as intelligent customer service, chatbots, and virtual assistants. More and more companies are also beginning to use intelligent chatbots to replace manual replies. According to the rounds of human-computer interaction, dialogue systems can be divided into single-round dialogue systems and multi-round dialogue systems. Through massive social dialogue data, combined with retrieval-based or generative modeling methods based on deep learning, single-round dialogue systems have achieved good results. However, in multi-round dialogue systems, the information input by users is often incomplete, usually containing references or omissions. For example, the current round of dialogue may contain pronouns or ambiguous content pointing to historical information of the dialogue. Existing related technologies are: 1. Thinking chain prompts trigger reasoning in large language models to propose thinking chains. 2. Implicit thinking chain reasoning through knowledge distillation defines three modules to compress thinking chains through knowledge distillation. 3. From explicit cost to implicit cost: gradually learn to internalize the cost. For mathematical problems, the thinking chain of the calculation process is gradually removed during the training process to achieve the effect of internalizing the thinking chain.

[0003] Prior Art 1: The performance of a large model can be significantly improved by allowing the large model to gradually participate in the process of decomposing a complex problem into step-by-step sub-problems and solving them one by one. Prior Art 2: 1. Interpreting the teacher's mind: Cultivate a student model to "interpret" the teacher's thinking process, that is, the continuous hidden states inside the teacher model during the reasoning process. This student model does not simply imitate, but uses these hidden states to get the answer. 2. Thinking simulation: Next, we use knowledge distillation technology to train a simulator that can predict the teacher's hidden state. Doing so can directly cross multiple processing levels and fit the results of the teacher's reasoning without going through every step of the teacher's reasoning. 3. Combined optimization: Finally, this simulator that can predict the teacher's thinking process is combined with a student model that can give the final answer based on this simulation process. The entire system is then trained end-to-end for optimization, allowing the student model to develop a different way of reasoning from the teacher. Prior Art 3: A simple and effective method for internalizing the CoT step is proposed: starting from the model used to clarify the CoT reasoning, gradually removing the intermediate steps and fine-tuning the model. This process enables the model to internalize the intermediate reasoning steps, thereby simplifying the reasoning process while maintaining high performance. Our approach enables the GPT-2Small model to solve 9x9 multiplications with 99% accuracy, while standard training cannot solve the larger 4x4 multiplications.

[0004] The inventors found that the existing technology uses the traditional thinking chain, which requires a large model to reason step by step, increasing the reasoning time. If the knowledge distillation method is used, it is highly dependent on training data and may have the risk of model bias. The transparency and explainability of the intermediate process of the distillation method are reduced. If the internalized CoT method is used, the effect of complex tasks is uncertain. Summary of the invention

[0005] The embodiments of the present invention aim to solve at least one of the above technical problems.

[0006] In a first aspect, an embodiment of the present invention provides a training method for a thought chain compression big model for a rewriting task, comprising: splicing multiple rounds of conversation history and a current question in a first training data set, and replacing the corresponding content in a thought chain prompt word template with the spliced ​​result to obtain a prompt word, wherein the first training data set includes multiple rounds of conversation history, the current question and a standard answer, and the thought chain prompt word template is a multiple round conversation text completer, which analyzes and returns the content involved in the current input by combining historical input and historical output; inputting the prompt word into the big model to generate a thinking process, and generating a second training data set based on the thinking process and the first training data set, wherein the thinking process is divided into multiple stages; training the big model based on the second training data set, and gradually removing a certain thinking process of the big model on the second training data set in a hierarchical manner during the training process.

[0007] In a second aspect, an embodiment of the present invention provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned training methods for a large model of thought chain compression for rewriting tasks of the present invention.

[0008] In a third aspect, an embodiment of the present invention provides a storage medium, wherein one or more programs including execution instructions are stored in the storage medium, wherein the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned training methods for a large model of a compressed thinking chain for a rewriting task of the present invention. In a fourth aspect, an embodiment of the present invention also provides a computer program product, wherein the computer program product includes a computer program stored on a storage medium, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes any of the above-mentioned training methods for a large model of a compressed thinking chain for a rewriting task.

[0009] The embodiment of the present invention introduces an answer framework suitable for the task itself by modeling the multi-round dialogue rewriting task based on the thought chain, and compresses the thought chain on this basis, and proposes a thought chain compression method for the large model rewriting task. While giving full play to the prior knowledge advantages of the large model, the system can better capture the semantic connection characteristics of historical information and the current dialogue, accurately locate the target referent, and then generate the target response more accurately and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0011] Figure 1 A flowchart of a training method for a large model of thought chain compression for rewriting tasks provided by an embodiment of the present invention;

[0012] Figure 2 A flowchart of another method for training a large model of thought chain compression for rewriting tasks provided by an embodiment of the present invention;

[0013] Figure 3 A schematic diagram of a framework of a training method for a large model of thought chain compression for rewriting tasks provided by the present invention;

[0014] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0016] It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0017] In the present invention, "module", "device", "system" and the like refer to related entities applied to computers, such as hardware, a combination of hardware and software, software or software in execution, etc. In detail, for example, an element can be, but is not limited to, a process, a processor, an object, an executable element, an execution thread, a program and / or a computer running on a processor. In addition, an application or script program running on a server, a server can all be an element. One or more elements can be in an execution process and / or thread, and an element can be localized on a computer and / or distributed between two or more computers, and can be operated by various computer-readable media. An element can also communicate through local and / or remote processes according to a signal with one or more data packets, for example, a signal from a data that interacts with another element in a local system, a distributed system, and / or a network on the Internet through a signal to interact with other systems.

[0018] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements that are not explicitly listed, or also include elements inherent to such processes, methods, articles or equipment. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or equipment that includes the elements.

[0019] The embodiment of the present invention provides a training method for a large model of thought chain compression for rewriting tasks, which can be applied to electronic devices. The electronic devices can be computers, servers or other electronic products, etc., which are not limited in the present invention.

[0020] Please refer to Figure 1 , which shows a training method for a large thought chain compression model for rewriting tasks provided by an embodiment of the present invention.

[0021] like Figure 1 As shown, in step 101, the multi-round dialogue history and the current question in the first training data set are spliced, and the spliced ​​result is used to replace the corresponding content in the thinking chain prompt word template to obtain the prompt word, wherein the first training data set includes the multi-round dialogue history, the current question and the standard answer, and the thinking chain prompt word template is a multi-round dialogue text completer, which analyzes and returns the content involved in the reference in the current input by combining the historical input and the historical output;

[0022] In step 102, the prompt word is input into the large model to generate a thinking process, and a second training data set is generated based on the thinking process and the first training data set, wherein the thinking process is divided into multiple stages;

[0023] In step 103, the large model is trained based on the second training data set, and during the training process, a certain thinking process of the large model on the second training data set is gradually removed in a hierarchical manner.

[0024] In this embodiment, for step 101, the multi-round dialogue history and the current question in the first training data set are spliced, and the spliced ​​result is used to replace the corresponding content in the thinking chain prompt word template to obtain the prompt word, wherein the first training data set includes multi-round dialogue history, current question and standard answer, and the thinking chain prompt word template is a multi-round dialogue text completer, which analyzes and returns the content involved in the current input by combining historical input and historical output. The prompt word template will be spliced ​​with "multi-round dialogue history" and "current question" and input to the big model, and the big model will return the "thinking process" according to the prompt word requirements. For example, the multi-round dialogue history and the current question are spliced ​​and supplemented into the thinking chain template, and the dialogue content is supplemented according to the thinking chain prompt word template to form a complete prompt word. Among them, the thinking chain prompt word template is, for example, by combining historical input and historical output, the content involved in the "current input" is analyzed and returned. The historical input label represents the user's input in the previous rounds, and the historical output is the reply given by the intelligent assistant to the historical input.

[0025] Afterwards, for step 102, the prompt words are input into the big model to generate a thinking process, and a second training data set is generated based on the thinking process and the first training data set, wherein the thinking process is divided into multiple stages, for example, the prompt words are input into the big model to generate a thinking process, and then the thinking process generated by the big model is manually verified and stored in a fixed format as a customized training data set, in which each record includes [multi-round dialogue history, current question, thinking process, standard answer] fields.

[0026] Finally, for step 103, the large model is trained based on the second training data set, and a certain thinking process of the large model on the second training data set is removed in a hierarchical manner during the training process. For example, the complete prompt word is first input into the model, where the complete prompt word is composed of a thinking chain prompt word template, a multi-round dialogue history, and a current question. The target expected output is the thinking process and the standard answer, which are spliced ​​with a special character [SEP_TOEKN] at intervals between the two. The negative log-likelihood function is used as the loss function to train for 3 epochs. The analysis process of the prompt word template divides the generation into 5 stages, including [identifying the reference type, processing category reference, listing specific content, processing reverse reference, and the result after reference]. Randomly select a sub-step other than "the result after reference", and remove the step in the thinking chain template through a regularized expression, and remove the thinking process corresponding to the step in the training data. The prompt word and the thinking process should be updated synchronously. When a certain thinking sub-step is removed from the thinking process of the first training data, the description part corresponding to the prompt word template should also be deleted. In the first training set, there are 5 sub-steps, among which the fifth step remains unchanged. In each subsequent training epoch, one sub-step will be randomly deleted and gradually removed. Therefore, by the fourth epoch, only the output "the result after reference:" is left, and the rewritten result is directly output, thus achieving the compression of the thinking chain. For example, in the fourth epoch, the sub-step of "processing reverse reference" is removed, and the thinking process of this step in all training data is removed. The hierarchical elimination based on the analysis steps can increase the interpretability and transparency of the compression process. After that, one sub-step is subtracted in each epoch. After 4 epochs, only the sub-step "the result after reference" is left. The corresponding prompt word can directly output the rewritten result, which greatly shortens the reasoning time.

[0027] The method of the embodiment of the present application introduces an answer framework suitable for the task itself by modeling the multi-round dialogue rewriting task based on the thought chain, and compresses the thought chain on this basis, and proposes a thought chain compression method for large model rewriting tasks. While giving full play to the prior knowledge advantages of the large model, the system can better capture the semantic connection characteristics of historical information and the current dialogue, accurately locate the target referent, and then generate the target response more accurately and efficiently.

[0028] In some optional embodiments, multiple stages include identifying reference types, processing category references, listing specific content, processing reverse references, and results after reference. Identifying reference types: Determine whether reference is required, and check whether there are relevant categories in the "historical replies". Processing category references: If there are specific categories in the "current input", fill in "None". If there are category references, convert them to specific categories. Listing specific content: List the specific items in the located category. If the specific category is not located, list all related items. Processing reverse references: Check whether the specific reference contains reverse references. If so, use a formula to convert and locate the specific content; if there is no reverse reference, directly locate the specific content. Result after reference: Output the result after the reference is rewritten.

[0029] In some optional embodiments, a certain thinking process of the large model on the second training data set is removed step by step in a hierarchical manner during the training process, including: selecting at least one stage other than the result after the reference, and removing the corresponding step in the thinking chain prompt word template through a regularized expression, and removing the thinking process of the corresponding step in the second training data set. For example, randomly select a sub-step other than "the result after the reference", and remove the step in the thinking chain template through a regularized expression, and remove the thinking process corresponding to the step in the training data. Repeat the training, subtracting a sub-step in each epoch.

[0030] Please refer to Figure 2 , which shows another training method for a large thought chain compression model for rewriting tasks provided by an embodiment of the present invention.

[0031] like Figure 2 As shown, in step 201, it is determined whether the current step to be replaced only has the result after the reference;

[0032] In step 202, if the current step to be replaced only has the result after the reference, the large model directly outputs the rewritten result;

[0033] In step 203, if the current step to be replaced does not only have the result after the reference, the thinking process of removing the corresponding step in the second training data set continues.

[0034] In this embodiment, for step 201, it is determined whether the current step to be replaced only has the result after the reference. A sub-step other than the "result after the reference" is randomly selected, and the step in the thinking chain template is removed by a regularized expression. After multiple rounds of processing, it is determined whether only the "result after the reference" step is left. Afterwards, for step 202, if the current step to be replaced only has the result after the reference, the large model directly outputs the rewriting result. After 4 epochs, only the "result after the reference" sub-step remains, and the rewriting result can be directly output using the corresponding prompt word. Finally, for step 203, if the current step to be replaced does not only have the result after the reference, the thinking process of the corresponding step in the second training data set is continued to be removed. For example, at least one stage other than the result after the reference is selected, and the corresponding step in the thinking chain prompt word template is removed by a regularized expression, and the thinking process of the corresponding step in the second training data set is removed, and the training is cyclically performed, and one sub-step is subtracted in each epoch until only the "result after the reference" step is left.

[0035] The method of the embodiment of the present application avoids gradient collapse during the training process. A small amount of training data from the previous epoch is randomly introduced in each epoch training to ensure the consistency of the optimization direction, ensure the accuracy of the rewriting, and greatly shorten the reasoning time. In some optional embodiments, the identification of the reference type is to determine whether the reference is required, and check whether there is a related category in the multi-round dialogue history; the processing category reference is to not input if there is a specific category in the current input, and if there is a category reference, it is converted to the specific category; the listing of specific content is to list the specific items in the positioning category, and if the specific category is not located, all related items are listed; the processing of reverse reference is to check whether the specific reference contains reverse reference, if so, use the formula to convert and locate the specific content, if there is no reverse reference, directly locate the specific content; the result after the reference is the result after the output reference is rewritten. The following is a reference example: (The number in "" represents the name of a certain singer's song)

[0036] History input: Here are some classic songs of a certain singer. History output: ``` A certain singer has many songs. Do you want to take a nostalgic trip? \n####01 Classic representative works\n-《001》\n-《002》\n-《003》\n-《004》\n\n####02 Transition works\n-《101》\n-《102》\n\n####03 Other popular singles\n-《201》\n-《202》\n-《203》\n\nHe has many other songs, each with a unique flavor. I hope you can find what you like. \n```

[0037] Current input: Play the last song in the classic masterpiece

[0038] Analysis process:

[0039] 1. Identify the reference type: the “Nth” reference type

[0040] 2. Processing category designation: None

[0041] 3. List specific content: Classic representative works: 1. 001 of a certain singer; 2. 002 of a certain singer; 3. 003 of a certain singer; 4. 004 of a certain singer

[0042] 4. Processing reverse order reference: Yes, the total number is 4, the reverse order reference is "the last song", then the forward order reference = 4-1+1=4, and the positioning is "a certain singer's "004""

[0043] 5. The result after reference: Play the song "004" by a certain singer

[0044] The multi-round dialogue history of the training data and the current question are spliced ​​in the form of "historical input: XXX historical reply: XXXX current input: XXXX", and replaced with the content in the thinking chain prompt word template.

[0045] In some optional embodiments, the thought chain prompt word template includes the corresponding content of the reasoning logic and steps, the analysis process and the spliced ​​results in the thought chain prompt word template. The reasoning logic and steps are as follows:

[0046] 1. Based on the previous round of conversation, you determine whether the user recognition result involves reference;

[0047] 2. If the recognized content does not involve the referenced statement, directly return the original recognition result;

[0048] 3. If the recognition content involves a reference but the historical input corresponding to the reference does not belong to the task-based voice skill, the original recognition result is returned directly; (task-based voice skills refer to traditional voice skills such as navigation, music, film and television, calendar, weather, etc.)

[0049] 4. If the recognized content involves reference and the historical input corresponding to the reference is a statement of task-based speech skills, then extract the relevant information of the historical conversation, summarize and output the writing without reference;

[0050] 5. In addition, multimedia resources such as music, movies, and stories often have the same name. So when historical information explicitly mentions resource types such as songs, movies, and stories, try to include the words of these resource types when completing the output to eliminate ambiguity.

[0051] Reasoning steps:

[0052] 1. The following is a general formula for reference, which you need to master: Positive reference = number of positioned contents - reverse reference + 1. For example, if the number of positioned contents is 9 and the reverse reference is the third from the last, then the positive reference = 9-3+1=7, which is the seventh in the positive order.

[0053] 2. Specific steps: Identify the reference type -> (process category reference -> list specific content -> process reverse reference) -> the result after reference.

[0054] In some optional embodiments, the thinking process and the standard answer are spliced ​​with special characters, and the negative log-likelihood function is used as the loss function for training. The thinking process and the standard answer are spliced ​​with special characters [SEP_TOEKN] and the negative log-likelihood function is used as the loss function for training for 3 epochs.

[0055] The inventors found that with the development of deep learning and big data technology, human-computer dialogue systems have become an important research direction in the field of artificial intelligence and have been widely used in scenarios such as intelligent customer service, chatbots, and virtual assistants. In multi-round dialogue systems, the information input by users is often incomplete and usually contains references or omissions. For example, the current round of dialogue may contain pronouns or unclear content that refers to historical dialogue information. For example:

[0056] -User: "What are the popular songs of a certain singer?"

[0057] -System: "A singer is one of the representative figures in the Chinese music scene. Since his debut, his music style is unique, integrating R&B, hip-hop, classical, Chinese style and other elements, and he has created many popular classic songs. The following are some of the singer's popular songs:

[0058] ###1.**《001》**This song is popular for its fresh melody and romantic lyrics and is one of the representative works of a certain singer. ###2.**《002》**A song about youth and first love, with a simple but moving melody, is deeply loved by young listeners.

[0059] ###3.**《003》**This song combines Chinese elements, beautiful lyrics, and a classical melody, showing a certain singer's love for traditional culture. ###4.**《004》**It has a strong sense of rhythm and mixes hip-hop and Chinese Kung Fu elements. It is one of the singer's early representative works. ###5.**《005》**It conveys positive energy and encourages people to maintain a positive and optimistic attitude. It is an inspirational work with a light style.

[0060] ###6.**《006》**A song with a sad mood, beautiful melody and classical music style, and profound lyrics. ###7.**《007》**With historical culture as the background, the lyrics and melody are full of romance and mystery, showing the unique creative style of a certain singer. ”

[0061] -User: "Play me the fifth song"

[0062] The above dialogue contains references and omissions and should be rewritten as follows:

[0063] -User: "Help me play the song 005"

[0064] Multi-turn dialogue systems generally have problems with semantic loss, reference, and omission, which leads to semantic ambiguity in many scenarios. To solve this problem, current research usually converts multi-turn dialogue generation tasks into single-turn dialogue tasks. The specific method is to use a large model to restore the current round of dialogue sentences into semantically complete expressions, and then generate responses in the manner of a single-turn dialogue system. However, even a large language model with hundreds of millions of parameters still cannot locate the accurate reference object. There may be errors in the order of reference or difficulty in reverse positioning. For example, "Help me play the fourth song" or "Help me play the last song" is rewritten as "Help me play the song "006""; or the reference type is wrong, for example, "Help me play the fourth song" is rewritten as "Help me play a Chinese-style song by a certain singer."

[0065] The thinking chain can effectively correct these problems. Through the prompt words, the model is guided to first inventory the objects of the selected type one by one, such as listing all the song objects involved. Secondly, determine whether it is in positive or reverse order. If reverse order is required, rewrite it in the form of the Nth in positive order to make the model select the song. Finally, rewrite after locating the correct song.

[0066] This application introduces a response framework suitable for the task itself based on the thought chain modeling for the multi-round dialogue rewriting task, and compresses the thought chain on this basis, and proposes a thought chain compression method for the large model rewriting task. While giving full play to the prior knowledge advantages of the large model, the system can better capture the semantic connection characteristics of historical information and the current dialogue, accurately locate the target referent, and then generate the target response more accurately and efficiently.

[0067] In view of the problems existing in the prior art, an embodiment of the present invention discloses a thought chain compression method for large model rewriting tasks, including three modules: thought chain prompt word construction, thought chain-based text reply generation, and thought chain compression. Among them, the thought chain prompt word construction aims to supplement the dialogue content according to the thought chain template to form a complete prompt word. The supplemented prompt words are input into the large model, allowing the model to gradually locate the target value band object and generate a text reply. Finally, the thought chain compression method is used to fine-tune the large model so that it can directly infer the selected target object to generate a reply, avoiding step-by-step reasoning and greatly reducing the reasoning time.

[0068] It should be noted that the application of the thought chain technology in the training method of the thought chain compression large model for rewriting tasks of the present invention is: by constructing specific thought chain prompt words, guiding the large model to gradually complete complex reasoning tasks and solve the problem of reference resolution in multi-round dialogues. Thought chain compression method: gradually remove unnecessary steps in the reasoning process in multiple rounds of training, thereby simplifying the reasoning path of the model and improving processing speed and efficiency.

[0069] The thought chain compression algorithm of this application: including how to remove specific reasoning steps through regularized expressions in each training cycle, and synchronously remove the corresponding thinking process in the training data set. Training strategy: including bringing in a small amount of training data from the previous epoch during each epoch training to prevent gradient collapse and ensure the consistency and stability of model optimization.

[0070] Please refer to Figure 3 , which shows a schematic diagram of the framework of the training method of the thought chain compression model for rewriting tasks of the present invention. Figure 3 As shown, step 1: first, the multi-round dialogue history and the current question are spliced ​​and added to the thinking chain template, and the dialogue content is supplemented according to the thinking chain prompt word template to form a complete prompt word.

[0071] Step 2: Input the prompt words into the big model to generate a thinking process. Then, the thinking process generated by the big model is manually verified and stored in a fixed format as a customized training data set. Each record in the data includes the fields [multi-round dialogue history, current question, thinking process, standard answer].

[0072] Step 3: Thought Chain Compression

[0073] In 3.1, the complete prompt word is first input into the model, where the complete prompt word is composed of the thought chain prompt word template, multi-round dialogue history, and the current question. The target expected output is the thinking process and the standard answer, which are spliced ​​with a special character [SEP_TOEKN] between them. The negative log-likelihood function is used as the loss function for training for 3 epochs.

[0074] 3.2 It can be seen that the analysis process of the prompt word template divides the generation into five stages, including [identifying the reference type, processing category reference, listing specific content, processing reverse reference, and the result after reference]. Randomly select a sub-step except "the result after reference", and remove the step in the thinking chain template through regularization expression, and remove the thinking process corresponding to the step in the training data.

[0075] After 3.3, repeat step S3.2 for cyclic training, subtracting one substep for each epoch.

[0076] 3.4 After 4 epochs, only the "result after reference" sub-step remains. Using the corresponding prompt word, the rewritten result can be directly output, which greatly shortens the reasoning time.

[0077] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that the present invention is not limited by the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0078] In some embodiments, an embodiment of the present invention provides a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any of the above-mentioned training methods for the thinking chain compression large model for rewriting tasks of the present invention.

[0079] In some embodiments, an embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned training methods for the large model of thought chain compression for rewriting tasks.

[0080] In some embodiments, an embodiment of the present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a training method for a large model of thought chain compression for a rewriting task.

[0081] Figure 4 is a schematic diagram of the hardware structure of an electronic device for executing a training method for a large model of thought chain compression for rewriting tasks provided by another embodiment of the present application, such as Figure 4 As shown, the device includes:

[0082] One or more processors 410 and memory 420, Figure 4 A processor 410 is taken as an example.

[0083] The device for executing the training method of the thought chain compression large model for rewriting tasks may also include: an input device 430 and an output device 440.

[0084] The processor 410, the memory 420, the input device 430 and the output device 440 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.

[0085] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as the program instructions / modules corresponding to the training method for the large model of the thought chain compression for rewriting tasks in the embodiment of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 420, that is, the training method for the large model of the thought chain compression for rewriting tasks in the above method embodiment is realized.

[0086] The memory 420 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of a training device for compressing a large model of a thinking chain for rewriting a task, etc. In addition, the memory 420 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 420 may optionally include a memory remotely arranged relative to the processor 410, and these remote memories may be connected to the training device for compressing a large model of a thinking chain for rewriting a task via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0087] The input device 430 can receive input digital or character information, and generate signals related to user settings and function control of the training device for the large model of thought chain compression for rewriting tasks. The output device 440 can include a display device such as a display screen. The one or more modules are stored in the memory 420, and when executed by the one or more processors 410, the training method for the large model of thought chain compression for rewriting tasks in any of the above method embodiments is executed.

[0088] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.

[0089] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:

[0090] (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.

[0091] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, etc.

[0092] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0093] (4) Other onboard electronic devices with data interaction functions, such as on-board devices installed in vehicles.

[0094] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0095] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A training method for a large model of thought chain compression for rewriting tasks, comprising: The multi-round dialogue history and the current question in the first training data set are spliced, and the corresponding content in the thinking chain prompt word template is replaced with the spliced ​​result to obtain the prompt word, wherein the first training data set includes the multi-round dialogue history, the current question and the standard answer, and the thinking chain prompt word template is a multi-round dialogue text completer, which analyzes and returns the content involved in the reference in the current input by combining the historical input and the historical output; Inputting the prompt words into the large model to generate a thinking process, and generating a second training data set based on the thinking process and the first training data set, wherein the thinking process is divided into multiple stages; The large model is trained based on the second training data set, and during the training process, a certain thinking process of the large model on the second training data set is gradually removed in a hierarchical manner.

2. The method according to claim 1, wherein: The multiple stages include identifying the reference type, processing category reference, listing specific content, processing reverse reference, and the result after reference.

3. The method according to claim 2, wherein: The step of gradually removing a certain thinking process of the large model on the second training data set in a hierarchical manner during the training process includes: At least one stage other than the result after the reference is selected, and the corresponding steps in the thought chain prompt word template are removed through a regularized expression, and the thinking process of the corresponding steps in the second training data set is removed.

4. The method according to claim 3, wherein: The method further comprises: Determine whether the current step to be replaced only has the result after the reference; If the current step to be replaced only has the result after the reference, the large model directly outputs the rewritten result; If the current step to be replaced does not only have the result after the reference, then continue to remove the thinking process of the corresponding step in the second training data set.

5. The method according to claim 2, wherein: The identification reference type, the processing category reference, the enumerated specific content, the processing reverse order reference, and the result after the reference include: The identification of the reference type is to determine whether reference is required and check whether there is a relevant category in the multi-round dialogue history; the processing category reference is to not input a specific category if there is a specific category in the current input, and to convert the category reference into the specific category if there is a category reference; The specific contents listed are to list specific items in the positioning category. If the specific category is not positioned, all related items are listed; The processing of the reverse reference is to check whether the specific reference contains the reverse reference, if yes, convert it using a formula and locate the specific content, if no reverse reference, directly locate the specific content; The result after the reference is the result after the output reference is rewritten.

6. The method according to claim 1, wherein: The thought chain prompt word template includes the corresponding contents of the reasoning logic and steps, the analysis process and the spliced ​​results in the thought chain prompt word template.

7. The method according to claim 6, wherein: The reasoning logic and steps include: Determining whether the user recognition result involves reference according to the multi-round conversation history; If the user recognition result does not involve the referenced statement, the original recognition result is directly returned; If the user recognition result involves a reference statement and the historical input corresponding to the reference does not belong to the statement of the task-based speech skill, the original recognition result is directly returned If the user recognition result involves a reference statement and the historical input corresponding to the reference is a statement of a task-based voice skill, then relevant information of the multi-round dialogue history is extracted, and after summarizing and inducing, a writing method without using reference is output.

8. The method according to claim 1, wherein: The training of the large model based on the second training data set includes: The thinking process and the standard answer are concatenated with special characters at intervals, and the negative log-likelihood function is used as the loss function for training.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 2 to 8.

10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 2 to 8 are implemented.

Citation Information

Cited By

  • Complex multi-field guided teaching question and answer generation method and device and storage medium

    CN120448505A

  • A complex multi-field guided teaching question and answer generation method and device and a storage medium

    CN120448505B

  • Training method of thinking type recognition model and bad thinking recognition method

    CN120687918A

  • Method for compressing thinking chain of reasoning large model

    CN121119127A

  • Model training method, text processing method and related device

    CN121188470A