Corpus synthesis method, big language model training method and related products
The Monte Carlo tree search and loop consistency technology generates synthetic corpus, and combines the original corpus to train large language models, solves the problem of insufficient existing training corpus, and improves the training effect and task adaptability of the model.
Patent Information
- Application Number
- CN202510390507.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
The training corpus of existing large language models has problems such as insufficient data, imbalance in categories, and lack of data in specific scenarios, resulting in poor training results.
Monte Carlo tree search and loop consistency technology are used to generate synthetic corpus by expanding the original corpus layer by layer, and training is combined with the original corpus and synthetic corpus to ensure that the synthetic corpus remains highly consistent with the original corpus.
It improves the training effect of large language models, makes up for the shortcomings of the original corpus, and enhances the performance of the model on specific tasks.
Smart Images

Figure CN120297286A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of natural language processing technology, and particularly relates to a corpus synthesis method, a training method for large language models, and related products. Background Art
[0002] Large Language Models (LLMs) are a type of artificial intelligence model that can understand and generate natural language text. These models typically have hundreds of millions or even hundreds of billions of parameters and are trained on large amounts of data to capture the complexity and diversity of language.
[0003] The training of large language models depends on obtaining and screening high-quality training corpora. Training corpora are usually natural corpora directly collected from the real world, such as corpora related to specific tasks collected from the Internet. However, such corpora are restricted by actual data collection and human factors, etc., and there are many problems, such as insufficient data, class imbalance, lack of data in specific scenarios, etc., resulting in poor training effects of large language models. Summary of the Invention
[0004] The purpose of the embodiments of this specification is to provide a corpus synthesis method, a training method for large language models, and related products, which are used to generate high-quality synthetic corpora that can simulate the distribution and complexity of the original corpus, help solve the problems existing in the original corpus, and thus improve the training effect of large language models.
[0005] To achieve the above purpose, the embodiments of this specification adopt the following technical solutions: In the first aspect, a corpus synthesis method is provided, including: Taking the original corpus of the large language model as the root node, and expanding the root node layer by layer to obtain n-level sub-nodes; the (i-1)-th level sub-node represents a task processing result obtained by performing a target processing task on the parent node of the (i-1)-th level sub-node through the large language model, and the i-th level sub-node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level sub-node through the large language model, where i is an even number and 1 < i ≤ n; Based on the similarity between the n-level sub-nodes and the root node, determining the first node from the n-level sub-nodes; Generating the synthetic corpus corresponding to the target processing task based on the first node.
[0006] In the second aspect, a training method for large language models is provided, including: Obtain a training corpus for a target processing task, where the training corpus includes an original corpus and a synthetic corpus, and each corpus has a label corresponding to the target processing task. The synthetic corpus is obtained by processing the original corpus using the corpus synthesis method provided in the first aspect; Train the large language model based on each corpus in the training corpus and the label of each corpus.
[0007] In a third aspect, there is provided a corpus synthesis device, including: An expansion module, configured to use the original corpus of the large language model as a root node, and expand the root node layer by layer to obtain n-level child nodes. The (i - 1)-th level child node represents a task processing result obtained by performing the target processing task on the parent node of the (i - 1)-th level child node through the large language model, and the i-th level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level child node through the large language model, where i is an even number and 1 < i ≤ n; A first determination module, configured to determine a first node from the n-level child nodes based on the similarity between the n-level child nodes and the root node; A generation module, configured to generate a synthetic corpus corresponding to the target processing task based on the first node.
[0008] In a fourth aspect, there is provided a training device for a large language model, including: An acquisition module, configured to acquire a training corpus for a target processing task, where the training corpus includes an original corpus and a synthetic corpus, and the synthetic corpus is obtained by using the corpus synthesis method provided in the first aspect on the original corpus; A training module, configured to train the large language model based on each corpus in the training corpus and the label of each corpus.
[0009] In a fifth aspect, there is provided an electronic device, including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the corpus synthesis method provided in the first aspect or the training method of the large language model provided in the second aspect.
[0010] In a sixth aspect, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the corpus synthesis method provided in the first aspect or the training method of the large language model provided in the second aspect.
[0011] In a seventh aspect, a computer program product is provided. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of the corpus synthesis method provided in the first aspect or the training method of the large language model provided in the second aspect.
[0012] The solution of the embodiments of this specification adopts the technical concept of corpus synthesis based on Monte Carlo tree search and cycle consistency. The original corpus is used as the root node, and the root node is expanded layer by layer based on Monte Carlo tree search. That is, the root node is used as the starting point first, and the large language model is used to perform the target processing task on the starting point. The multiple task processing results obtained are respectively used as a child node, and then the inverse operation of the target processing task is performed on the child node, and multiple candidate corpora can be obtained. Each candidate corpus is a child node of the child node. In this way, a Monte Carlo tree including the root node and n-level child nodes is constructed. Each even-level child node represents a candidate corpus. Since such candidate corpora are expanded based on their upstream nodes, they not only conform to the context, but also can capture the uncertainty and diversity in the language and contain rich generated content. Further, by calculating the similarity between the n-level child node and the root node, the first node is determined from the n-level child nodes, and the synthetic corpus corresponding to the target processing task is generated based on the first node, that is, it is evaluated whether the candidate corpus conforms to the cycle consistency with the original corpus, and the candidate corpus that conforms to the cycle consistency with the original corpus is obtained. Generating the synthetic corpus based on such candidate corpora can ensure that the synthetic corpus is highly consistent with the original corpus and solve the problem that the synthetic corpus may overfit the distribution of the original corpus or cannot fully simulate the original corpus.
[0013] In addition, the traditional idea of training a large language model based on the original corpus is abandoned, and instead, the large language model is trained by combining the original corpus and the synthetic corpus, so that the synthetic corpus can effectively make up for the deficiencies of the original corpus, thereby improving the training effect of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings: Figure 1 It is a schematic flowchart of a corpus synthesis method provided by an embodiment of this specification; Figure 2 It is a schematic flowchart of a layer-by-layer expansion method provided by an embodiment of this specification; Figure 3 It is a schematic structural diagram of a Monte Carlo tree provided by another embodiment of this specification; Figure 4 A flowchart of a training method for a large language model provided by an embodiment of this specification; Figure 5 A flowchart of a training method for a large language model provided by another embodiment of this specification; Figure 6 A structural diagram of a corpus synthesis device provided by an embodiment of this specification; Figure 7 A structural diagram of a training device for a large language model provided by an embodiment of this specification; Figure 8 A structural diagram of an electronic device provided by an embodiment of this specification. Detailed implementation manners
[0015] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of this document.
[0016] The term "including" and its variants used in this document are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "an embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. The term "in response to" is used to indicate the conditions or states on which the operations performed depend. When the dependent conditions or states are met, one or more operations performed may be real-time or may have a set delay. Without special instructions, there is no limitation on the order of multiple operations performed.
[0017] It should be noted that concepts such as "first" and "second" mentioned in this document are only used to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependent relationships.
[0018] It should be noted that modifiers such as "one" and "multiple" mentioned in this document are illustrative rather than restrictive. Those skilled in the art should understand that unless explicitly stated otherwise in the context, it should be understood as "one or more".
[0019] The names of the messages or information exchanged between multiple devices in the embodiments of this document are for illustrative purposes only and are not used to limit the scope of these messages or information.
[0020] Explanation of some terms: Large language model: A deep learning model that processes and generates text through a large-scale neural network. These models can perform various tasks such as text classification, sentiment analysis, machine translation, text generation, question answering, etc. These models usually have a very large number of parameters. For example, GPT-3 has 173 billion parameters and can process and understand extremely complex language patterns. Large language models have extensive applications in multiple fields such as search engines, chatbots, content generation, and language understanding.
[0021] The underlying implementation of large language models usually uses a neural network architecture called Transformer. Self-attention is the core of Transformer, which allows dynamically focusing on different parts of the input sequence when processing it. The self-attention mechanism can capture the dependencies between any two positions in the input sequence, regardless of their distance in the sequence. Transformer consists of two main parts: an encoder and a decoder.
[0022] Encoder: Stacked by multiple encoding layers, each encoding layer contains a self-attention layer and a feed-forward neural network. The encoder is responsible for processing the input sequence and generating a continuous internal representation.
[0023] Decoder: Also composed of multiple decoding layers, each decoding layer contains a self-attention layer, an encoder-decoder attention layer (used to focus on the output of the encoder), and a feed-forward neural network. The decoder is used to generate the output sequence.
[0024] Since the Transformer model itself does not have the ability to process sequence position information, position encoding is needed to provide the model with information about the positions of words in the input sequence. In each encoding layer and decoding layer, the output of the self-attention layer is passed to a feed-forward neural network, which applies the same function independently at each position. Residual connections are used to add the input and output in each encoding layer and decoding layer to help the model train deeper networks. Layer normalization is used to stabilize the training process.
[0025] Natural data: refers to data directly collected from the real world, which reflects the natural usage of human language. Natural data contains various language styles, topics, and expressions, reflecting real human communication and behavior. The text in natural data is usually closely related to its context, such as news reports, social media posts, books, academic papers, etc. There is a large amount of natural data on the Internet, which provides rich resources for training large language models.
[0026] Synthetic data: refers to data generated by algorithms or large language models. These data are not directly collected from the real world but are created for specific purposes or needs. In applications, synthetic data of specific types or styles can be generated according to requirements, while it is difficult to find natural data of specific types or styles. In theory, synthetic data can be generated infinitely without being restricted by actual data collection.
[0027] Cyclic consistency: refers to the ability to restore the original data based on the new data obtained after processing the original data through a large language model. For example, in a translation task, a Chinese sentence A is translated into an English sentence B, and then the English sentence B is translated back into a Chinese sentence C. If the Chinese sentence C is similar enough to the Chinese sentence A, then the Chinese sentence C meets the cyclic consistency.
[0028] Monte Carlo Tree Search (MCTS): is widely used in fields such as artificial intelligence, game theory, and automatic planning. The core idea is to estimate the value of each optional action through random simulation, thereby helping the system select the best next action. It organizes these simulations by building a search tree and uses statistical information to guide the search process to more likely find the best decision. Generally speaking, the Monte Carlo tree search process can be divided into the following steps: Selection: Starting from the root node, recursively select the optimal child node until a leaf node is obtained.
[0029] Expansion: If the leaf node is not a termination node, then create one or more child nodes.
[0030] Simulation: Select one of the expanded nodes and run a simulated output from this expanded node until the game ends.
[0031] Backpropagation: Update the current action sequence with the output of the simulation result.
[0032] As mentioned above, currently, the training corpus of large language models is usually natural corpus directly collected from the real world, such as corpus related to specific tasks collected from the Internet. Specifically, in the pre-training stage, natural corpus is constructed in the following ways: Step 1, text collection. Collect a large amount of text data from the Internet, including books, articles, web page content, etc.
[0033] Step 2, preprocessing. Clean the collected text, including removing useless characters, unifying formats, word segmentation, etc.
[0034] Step 3, sentence construction. Split the text into sentences or paragraphs, which will be used as the input of the model.
[0035] Step 4, context window. To train the model to capture context relationships, the text is usually split into context windows of fixed length. For example, a window may contain one sentence or multiple sentences.
[0036] In the fine-tuning stage, natural corpus is constructed in the following ways: Step 1, task-specific data. According to the requirements of the fine-tuning task, collect task-specific data sets, such as sentiment analysis, question answering, etc.
[0037] Step 2, annotation. These data sets usually need to be manually annotated or already contain annotation information. Annotation means specifying the correct output or classification label for each sample in the data set.
[0038] Step 3, data splitting. Split the data set into training set, validation set and test set to evaluate the performance of the model during fine-tuning.
[0039] However, due to limitations such as actual data collection and human factors, such corpus has the following problems in multiple aspects: (1) Data quality Annotation errors: High-quality corpus annotation is the key to fine-tuning, but it is very difficult to obtain accurate and consistent corpus annotation, especially for tasks with strong subjectivity.
[0040] Data bias: The corpus may contain biases or prejudices, which will affect the fairness and accuracy of the model in specific tasks.
[0041] (2) Data volume Insufficient data: For some specific tasks, it may be difficult to collect enough natural corpus for effective fine-tuning.
[0042] Uneven data distribution: Natural corpus may be unevenly distributed among different categories, resulting in the model being biased towards certain categories.
[0043] (3) Data diversity Domain-specific vocabulary: For tasks in specific domains, it is difficult for natural language corpora to ensure that the model understands professional terms and jargon.
[0044] Multilingual and multicultural: For multilingual models, it is difficult to ensure that natural language corpora cover different languages and cultural backgrounds.
[0045] (4) Data preprocessing: Text cleaning: Natural language corpora may contain a large amount of noise, such as HyperText Markup Language (HTML) tags, non-standard characters, etc. Cleaning natural language corpora requires a lot of work.
[0046] Text segmentation: Segmenting natural language corpora into segments suitable for model processing (such as sentences, paragraphs) requires careful work, especially when dealing with long texts.
[0047] (5) Data alignment Cross-lingual alignment: For translation tasks, it is necessary to collect parallel corpora, that is, multiple corpora with the same content but different languages. The collection of such natural language corpora is usually very difficult.
[0048] Multimodal alignment: For tasks involving text and other modalities (such as images, audio), it is necessary to align corpora of different modalities. The alignment of natural language corpora is usually very difficult.
[0049] (6) Data privacy and security Sensitive information: When processing natural language corpora containing personal sensitive information, it is necessary to ensure the security and privacy protection of such information.
[0050] Compliance: Limited by compliance requirements, the natural language corpora available for training are restricted.
[0051] (7) Data dynamics Timeliness: For some tasks, natural language corpora may quickly become outdated and need to be updated regularly.
[0052] Language evolution: Language is constantly changing, and models need to adapt to new vocabulary and usage.
[0053] (8) Resource consumption: Large-scale data sets require a large amount of storage space and computing resources to process.
[0054] The inventor found through extensive research on the training process of large language models that: On the one hand, the training data of current large language models is usually natural corpus directly collected from the real world. Such corpus may contain spelling mistakes, grammar errors, dialects, and slang, etc. Moreover, restricted by actual data collection and human factors, natural corpus often has problems such as insufficient data, class imbalance, or lack of data in specific scenarios. These will greatly affect the training effect of the language model. The corpus synthesized according to specific tasks or requirements is not affected by actual data collection and human factors, and can help solve the problems existing in natural corpus. Training the large language model by combining natural corpus and synthesized corpus helps to improve the training effect.
[0055] On the other hand, the difficulty in corpus synthesis lies in the possibility of overfitting the distribution of natural corpus or being unable to fully simulate natural corpus, because the synthesized corpus cannot be highly consistent with natural corpus.
[0056] Considering the above aspects comprehensively, a technical concept of corpus synthesis based on MCTS and cycle consistency is proposed. Taking the natural corpus as the root node, the root node is expanded layer by layer based on MCTS. That is, first taking the root node as the starting point, performing the target processing task on the starting point through the large language model, taking the obtained multiple task processing results as a child node respectively, and then performing the inverse operation of the target processing task on the child node, multiple candidate corpora can be obtained. Each candidate corpus is a child node of the child node. In this way, a Monte Carlo tree containing the root node and n-level child nodes is constructed. Each even-level child node represents a candidate corpus. Since such candidate corpora are expanded based on their upstream nodes, they not only conform to the context, but also can capture the uncertainty and diversity in language, and contain rich generated content. Further, by calculating the similarity between the n-level child node and the root node, the first node is determined from the n-level child nodes, and the synthesized corpus corresponding to the target processing task is generated based on the first node, that is, evaluating whether the candidate corpus is consistent with the natural corpus in terms of cycle consistency, obtaining the candidate corpus that is consistent with the natural language in terms of cycle consistency, and generating the synthesized corpus based on such candidate corpora can ensure that the synthesized corpus is highly consistent with the natural corpus, and solve the problem that the synthesized corpus may overfit the distribution of natural corpus or be unable to fully simulate natural corpus.
[0057] In addition, abandoning the traditional idea of training large language models based on natural corpus and instead combining natural corpus and synthesized corpus to train large language models enables the synthesized corpus to effectively make up for the deficiencies existing in natural corpus, thereby improving the training effect of large language models.
[0058] The corpus synthesis method and the training method of the large language model proposed in the embodiments of this specification can both be executed by an electronic device. As an example, it can be executed by software in the electronic device. The so-called electronic device here can include terminal devices, such as smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, smart home appliances, smart watches, vehicle terminals, aircraft, etc.; or, the electronic device can also include a server, such as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0059] The following will combine with the accompanying drawings to detail the technical solutions provided by each embodiment of this specification.
[0060] Please refer to Figure 1 , which is a schematic flowchart of a corpus synthesis method provided by an embodiment of this specification. The method may include: S102, taking the original corpus of the large language model as the root node, and expanding the root node layer by layer to obtain n-level child nodes.
[0061] The original corpus refers to the natural corpus used to train the large language model. By taking the original corpus as the root node and based on MCTS, the root node is expanded layer by layer to obtain n-level child nodes, where n is an integer and n≥2. The number of child nodes at each level can be multiple.
[0062] The tree structure composed of the root node and the n-level child nodes is called a Monte Carlo tree. In this Monte Carlo tree, the (i - 1)-level child node represents a task processing result obtained by performing a target processing task on the parent node of the (i - 1)-level child node through the large language model, and the i-level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-level child node through the large language model, where i is an even number and 1<i≤n.
[0063] In applications, each node of the Monte Carlo tree has a corresponding pointer or reference for pointing to the node, so as to better maintain the structure of the Monte Carlo tree. In addition, the Monte Carlo tree can be stored in a storage engine, and the storage engine can be based on persistent storage such as memory or hard disk.
[0064] The target processing task can be determined according to the training purpose of the large language model. For example, if the training purpose is to optimize the large language model for translation tasks, the target processing task is a translation task, that is, translating the corpus of one language into the corpus of another language. Another example is that if the training purpose is to optimize the large language model for text generation tasks, the target processing task is a text generation task, that is, generating text based on the given context.
[0065] The reverse operation of the target processing task refers to the operation that is inverse to the target processing task and depends on the specific content of the target processing task. For example, if the target processing task is to translate Chinese into English, the reverse operation of the target processing task is to translate English into Chinese. Another example, if the target processing task is to convert text from one encoding format to another, the reverse operation of the target processing task is to restore the text from the converted encoding format to the original encoding format. Still another example, if the target processing task is to concatenate multiple texts or strings into a complete text, the reverse operation of the target processing task is to split the concatenated text into the original multiple files or strings according to the concatenation rules or delimiters.
[0066] Large language models can be used through Prompt Engineering (PE). Specifically, by designing appropriate prompts to guide the large language model to perform the target processing task or the reverse operation of the target processing task.
[0067] For example, if the target processing task is to translate Chinese into English, the corresponding prompt for the target processing task is "Please translate the following Chinese text into English:", to guide the large language model to translate Chinese corpus into English corpus; the prompt corresponding to the reverse operation of the target processing task is "Translate the following sentence from English to Chinese:", to guide the large language model to back-translate English corpus into Chinese corpus.
[0068] In the embodiments of this specification, the root node can be expanded layer by layer in various appropriate ways to obtain n-level child nodes.
[0069] In one implementation, as Figure 2 shown, the above S102 includes the following steps: S121, determine the root node as the second node; repeat the following expansion for the second node, that is, steps S122 to S125.
[0070] S122, perform the target processing task on the second node through the large language model to obtain the child nodes of the second node.
[0071] For example, assume that the target processing task is to translate Chinese into English, and the corpus represented by the second node is "The setting sun and the solitary wild duck soar up together; The autumn waters merge with the vast sky in one hue." Concatenating this corpus with the prompt of the target processing task "Please translate the following Chinese text into English:" and inputting it into the large language model, the translation result "The setting sun and the solitary wild duck soar up together; The autumn waters merge with the vast sky in one hue." can be obtained.
[0072] In the application, the number of child nodes of the second node can be multiple. In this case, multiple different prompts corresponding to the target processing task can be generated, such as "Translate the following text into English, try to use complex and difficult words as much as possible:", "Translate the following text into English, try to use simple words as much as possible", etc.; for each prompt, concatenate the prompt with the corpus represented by the second node and input it into the large language model to obtain the translation result corresponding to the prompt; further, take each translation result corresponding to the prompt as a child node, and multiple child nodes of the second node can be obtained.
[0073] Or, generate a prompt corresponding to the target processing task, input the prompt into the large language model to obtain a translation result; then, introduce randomness in the generation of the large language model, such as adjusting the temperature variable (temperature) of sampling to increase the possibility of search to obtain another translation result; repeat the above process multiple times to obtain multiple translation results; further, take each translation result as a child node, and multiple child nodes of the second node can be obtained.
[0074] S123, determine the third node from the child nodes of the second node based on the similarity between the child nodes of the second node and the root node.
[0075] For each child node of the second node, the higher the similarity between the child node and the root node, the more accurate the task processing result represented by the child node. Therefore, selecting the child node with a higher similarity as the third node for expansion based on the similarity between the child nodes of the second node and the root node helps to improve the quality of the synthesized corpus.
[0076] As an example, determine at least one child node with a relatively high similarity ranking between the child nodes of the second node and the root node as the third node. For example, sort each child node of the second node in descending order of similarity and determine the first 2 child nodes as the third node.
[0077] As another example, in the case where the number of child nodes of the second node is large, a number of child nodes are selected from these child nodes by means of random or uniform sampling; then, at least one child node with a relatively high similarity ranking between the selected child nodes and the root node is determined as the third node. For example, first, 5 child nodes are selected from the child nodes of the second node by means of random or uniform sampling, and then 2 child nodes that are most similar to the root node are selected from these 5 child nodes as the third node. By determining the third node in this way, not only can the diversity of exploration be increased, premature convergence to a local optimal solution be avoided, but also the number of child nodes that need to be evaluated can be reduced, the overall expansion efficiency can be improved, and thus the corpus synthesis efficiency can be improved.
[0078] As yet another example, based on the similarity between the child nodes of the second node and the root node, the scores of the child nodes of the second node are determined; at least one child node with a relatively high score ranking among the child nodes of the second node is determined as the first node.
[0079] Specifically, for each child node of the second node, the similarity between the child node and the root node can be determined as the score of the child node.
[0080] Alternatively, a value can be sampled from the standard Gaussian distribution, and the sum of this value and the similarity between the node and the root node can be determined as the score of the child node. For example, the similarity between a certain child node and the root node is 0.85, and a value of 0.005 is obtained by sampling from the standard Gaussian distribution, then this value is added to the similarity to obtain the score of the child node, which is 0.855.
[0081] The above method is equivalent to introducing a certain degree of randomness into the similarity between the child node and the root node, simulating the uncertainty or error in the real world, alleviating the limitations of similarity evaluation, and continuing to expand on the basis of the third node determined thereby, which helps to increase the diversity of exploration and avoid falling into a local optimal solution.
[0082] S124, perform the inverse operation of the target processing task on the third node through the large language model to obtain the child nodes of the third node.
[0083] For example, assume that the target processing task is to translate Chinese into English. Then the reverse operation of the target processing task is to translate English into Chinese. The translation result represented by the third node is “The setting sun and the solitary wildduck soar up together; The autumn waters merge with the vast sky in one hue.”. After concatenating this translation result with the prompt corresponding to the reverse operation, “Translate the following sentence from English to Chinese:”, and inputting it into the large language model, the candidate corpus “The setting sun and the solitary wild duck soar up together; The autumn waters merge with the vast sky in one hue.” can be obtained.
[0084] In an application, the number of child nodes of the third node can be multiple. In this case, multiple different prompts corresponding to the reverse operation can be generated. For each prompt, after concatenating the prompt with the task processing result represented by the third node and inputting it into the large language model, the candidate corpus corresponding to this prompt can be obtained. Further, taking each candidate sentence corresponding to a prompt as a child node of the third node, multiple child nodes of the third node can be obtained.
[0085] Alternatively, generate a prompt corresponding to the reverse operation, input the prompt into the large language model to obtain a candidate corpus. Then, introduce randomness in the generation of the large language model, such as adjusting the temperature variable of sampling to increase the possibility of search to obtain another candidate corpus. Repeat the above process multiple times to obtain multiple candidate corpora. Further, taking each candidate corpus as a child node, multiple child nodes of the third node can be obtained.
[0086] S125, if the child nodes of the third node do not meet the preset stop expansion condition, then determine a new second node from the child nodes of the third node based on the similarity between the child nodes of the third node and the root node.
[0087] The stop expansion condition can be set according to actual needs. For example, the depth of the Monte Carlo tree (i.e., the number of levels of child nodes) reaches the depth threshold, or there is a child node among the child nodes of the third node whose similarity with the root node is greater than or equal to the similarity threshold, etc. The embodiments of this specification do not limit this.
[0088] In the above S125, the specific implementation manner of determining a new second node from the child nodes of the third node based on the similarity between the child nodes of the third node and the root node is similar to the specific implementation manner of the above S123 and will not be elaborated here.
[0089] It can be seen that in the above-mentioned layer-by-layer expansion process, each level of child nodes is expanded on the basis of its upstream node, so each level of child nodes not only conforms to the context, but also can capture the uncertainty and diversity in the language, and contains rich generated content. In addition, in each level of child nodes, the starting point of expansion is selected according to the similarity between the child nodes of this level and the root node, which helps to obtain high-quality candidate corpus.
[0090] For ease of understanding, the following Figure 3 The above expansion process is described.
[0091] like Figure 3 As shown, the original corpus is taken as the root node, which is recorded as Root. The root node is selected as the starting point, and the target processing task is performed on the root node through the large language model to obtain three first-level child nodes. Since the similarity between the rightmost first-level child node and the root node is high, this child node is selected as the starting point, and the large language model performs the inverse operation of the target processing task on the child node to obtain two second-level child nodes. Since these second-level child nodes do not meet the preset expansion stop condition, and the similarity between the leftmost second-level child node and the root node is high, the leftmost second-level child node is selected as the starting point, and the target processing task is performed on the child node through the large language model to obtain three third-level child nodes. Since the similarity between the middle third-level child node and the root node is high, the middle third-level child node is selected as the starting point, and the large language model performs the inverse operation of the target processing task on the child node to obtain two fourth-level child nodes. Since these fourth-level child nodes meet the expansion stop condition, the expansion is stopped, and a Monte Carlo tree containing the root node and fourth-level child nodes is obtained.
[0092] The above text shows some implementations of the above S102. Of course, it should be understood that the above S102 can also be implemented in other ways, and the embodiments of this specification do not limit this. For example, when expanding each level of child nodes, first randomly select some child nodes from the child nodes of this level, and then select at least one child node with the highest similarity to the root node from these child nodes as the starting point, and perform the target processing task or inverse operation on the child node through the large language model to obtain the next level of child nodes.
[0093] In the embodiments of this specification, various appropriate methods may be used to evaluate the similarity between child nodes at each level and the root node.
[0094] In one implementation, for each level of child nodes, the child nodes and the root node are respectively mapped into vectors in a high-dimensional space, and then the similarity between the child nodes and the root node is determined based on the distance between the two vectors, such as Euclidean distance or cosine distance.
[0095] In another implementation, the original corpus has a label corresponding to the target processing task. In this case, the similarity between the i-1th level child node and the root node is determined by: Step A1, segment the tag to obtain a plurality of first character strings, and segment the task processing result represented by the i-1th level child node to obtain a plurality of second character strings.
[0096] Specifically, the n-gram segmentation method is used to segment the labels and the task processing results represented by the i-1th level child nodes, where n-gram refers to a sequence of n consecutive words (or characters) in the text. For example, when n=2, "I love" and "love the motherland" are 2-grams (also called bigrams); when n=3, "I love the motherland" is a 3-gram (also called trigram).
[0097] For example, assuming that the target processing task is to translate Chinese into English, n=3, the label of the original corpus is "The dogruns fast", and the translation result corresponding to the i-1th level child node is "The dog is running fast", then, the label is segmented to obtain 2 first strings: "The dog is", "dog runs fast", and the translation result is segmented to obtain 3 second strings: "The dog is", "dog is running", "is running fast".
[0098] Step A2: obtaining the recall rate of the i-1th level child node based on the proportion of the first character string in the tag that matches the second character string.
[0099] Recall is a performance indicator commonly used in the field of information retrieval and machine learning, especially in processing classification tasks and search engines. Recall refers to the proportion of samples that are actually positive that are correctly predicted as positive by the model. It measures the ability of the model to find all positive samples. Mathematically, recall can be expressed as: ,in, Indicates the number of samples that are incorrectly predicted as negative (i.e., actually positive), Indicates the number of samples correctly predicted as positive.
[0100] In the embodiments of this specification, the recall rate of the (i - 1)-th level sub-node = the number of first strings in the labels of the original corpus that match the second string / the total number of multiple first strings. For example, if the labels of the original corpus contain 8 first strings, and 4 of them match the second string in the candidate corpus represented by the (i - 1)-th level sub-node, then the recall rate of the (i - 1)-th level sub-node is 4 / 8 = 0.5.
[0101] Step A3: Based on the proportion of the second strings that match the first string in the task processing result represented by the (i - 1)-th level sub-node, obtain the precision rate of the (i - 1)-th level sub-node.
[0102] Precision is a performance metric commonly used in the fields of information retrieval and machine learning, especially in scenarios such as dealing with classification tasks and search engines. Precision refers to the proportion of samples predicted as positive classes that are actually positive classes. It measures the accuracy of the model's predictions. Mathematically, precision can be expressed as: , where represents the number of samples mispredicted as positive classes (i.e., actually negative classes).
[0103] In the embodiments of this specification, the precision rate of the (i - 1)-th level sub-node = the number of second strings that match the first string in the candidate corpus represented by the (i - 1)-th level sub-node / the total number of multiple second strings. For example, if the candidate corpus represented by the (i - 1)-th level sub-node contains 10 second strings, and 4 of them match the first string in the labels of the original corpus, then the precision rate of the (i - 1)-th level sub-node is 4 / 10 = 0.4.
[0104] Step A4: Based on the recall rate and precision rate of the (i - 1)-th level sub-node, determine the similarity between the (i - 1)-th level sub-node and the root node.
[0105] There is a certain trade-off relationship between precision and recall. Generally, increasing precision may decrease recall, and vice versa. For example, in the spam filtering task, if the filtering criteria are tightened, it may reduce misjudgments (i.e., reduce FP), thereby increasing precision, but at the same time, it may also miss some real spam emails (i.e., increase FN), thereby decreasing recall.
[0106] As an example, in order to consider both precision and recall at the same time, the embodiments of this specification use the F1 score as a comprehensive evaluation metric to evaluate the similarity between the (i - 1)-th level sub-node and the root node, that is, taking the F1 score as the similarity between the (i - 1)-th level sub-node and the root node. The F1 score is the harmonic mean of precision and recall, expressed as , where represents the precision rate of the (i - 1)-th level sub-node, Represents the recall rate of the (i - 1)-th level child node. The F1 score ranges from 0 to 1. The closer the F1 score is to 1, the more similar the (i - 1)-th level child node is to the root node, indicating that the candidate corpus represented by the (i - 1)-th level child node is more similar to the label of the original corpus at the n-gram level, and thus the accuracy and fluency of this candidate corpus are relatively better.
[0107] As another example, one of precision and recall can also be determined as the similarity between the i-th level child node and the root node according to the field to which the original corpus belongs. For example, in medical data, recall is usually more emphasized because the cost of missing a positive sample may be very high. Therefore, if the original corpus is medical data, the recall is determined as the similarity between the i-th level child node and the root node. Another example is that in Internet media data, precision is usually more emphasized to avoid recommending irrelevant media content to users. Therefore, if the original corpus is Internet media data, the precision is determined as the similarity between the i-th level child node and the root node.
[0108] The similarity between the i-th level child node and the root node is determined in the following way: Step B1, segment the original corpus to obtain a plurality of third strings, and segment the candidate corpus represented by the i-th level child node to obtain a plurality of fourth strings.
[0109] Among them, the specific implementation manner of step B1 is similar to that of step A1 above and will not be elaborated.
[0110] Step B2, based on the proportion of the third strings in the original corpus that match the fourth strings, obtain the recall rate of the i-th level child node.
[0111] Among them, the specific implementation manner of step B2 is similar to that of step A2 above and will not be elaborated.
[0112] Step B3, based on the proportion of the fourth strings in the candidate corpus represented by the i-th level child node that match the third strings, obtain the precision of the i-th level child node.
[0113] Among them, the specific implementation manner of step B3 is similar to that of step A3 above and will not be elaborated.
[0114] Step B4, based on the recall rate and precision of the i-th level child node, determine the similarity between the i-th level child node and the root node.
[0115] Among them, the specific implementation manner of step B4 is similar to that of step A4 above and will not be elaborated.
[0116] The similarity determined by the above method can more accurately reflect the cycle consistency in the process from the original corpus to the candidate corpus, providing reliable data support for generating high-quality synthetic corpora.
[0117] S104. Based on the similarity between the nth-level child node and the root node, determine the first node from the nth-level child nodes.
[0118] In the step-by-step expansion operation of the root node, first perform the target processing task on the node, and then perform the inverse operation of the target processing task on the obtained child nodes, which is equivalent to performing a round of cyclic processing on the corpus represented by the node. If the processed corpus is similar to the original corpus, it indicates that the processed corpus is in cycle consistency with the original corpus and meets the quality standards for training large language models.
[0119] In one implementation, the above S104 includes the following steps: Determine at least one child node with a relatively high similarity ranking between the nth-level child nodes and the root node as the first node. For example, sort each child node of the nth-level child nodes in descending order of similarity, and determine the first 2 child nodes as the first node.
[0120] In another implementation, the above S104 includes the following steps: Based on the similarity between the nth-level child node and the root node, determine the score of the nth-level child node; Determine the child nodes with scores greater than or equal to the score threshold among the nth-level child nodes as the first node.
[0121] Specifically, for each nth-level child node, the similarity between the nth-level child node and the root node can be determined as the score of the nth-level child node.
[0122] Alternatively, sample from the standard Gaussian distribution to obtain a target value, and determine the sum of the similarity between the nth-level child node and the root node and the target value as the score of the nth-level child node. For example, the similarity between a certain nth-level child node and the root node is 0.75, and a value of 0.005 is obtained by sampling from the standard Gaussian distribution. Then add this value to the similarity to obtain the score of this nth-level child node, which is 0.755.
[0123] The above method is equivalent to introducing a certain degree of randomness into the similarity between the nth-level child node and the root node, simulating the uncertainty or error in the real world, alleviating the limitations of similarity evaluation, helping to increase the diversity of exploration, and avoiding falling into local optimal solutions.
[0124] S106. Based on the first node, generate the synthetic corpus corresponding to the target processing task.
[0125] In one implementation, the candidate corpus represented by the first node is determined as the synthesized corpus corresponding to the target processing task, and the task processing result represented by the parent node of the first node is determined as the label corresponding to the synthesized corpus.
[0126] In another implementation, the path from the root node to the first node is determined; all the nodes representing candidate corpora on this path are selected. For each selected node, the candidate corpus represented by this node is determined as the synthesized corpus corresponding to the target processing task, and the task processing result represented by the parent node of this node is determined as the label corresponding to the synthesized corpus.
[0127] For example, Figure 3 taking the Monte Carlo tree shown as an example, assuming that the fourth-level child node on the far right is the first node, then the path represented by the thick implementation arrow is the path from the root node to the first node, and this path can represent the corpus synthesis path. Based on this, the candidate corpus represented by the second-level child node on the left is determined as a synthesized corpus, the task processing result represented by the first-level child node on the far right is determined as the label of this synthesized corpus, and the fourth-level child node on the right is determined as another synthesized corpus, and the third-level child node in the middle is determined as the label of this synthesized corpus.
[0128] Since the first node is similar to the root node, the path from the root node to the first node conforms to cyclic consistency. Therefore, all the candidate corpora on this path meet the quality standard. Furthermore, these candidate corpora are determined as synthesized corpora, and the task processing results for obtaining these candidate corpora are determined as the labels of the synthesized corpora. In this way, while ensuring the quality of the synthesized corpora, the quantity of the synthesized corpora is enriched.
[0129] In yet another implementation, the candidate corpus represented by the first node is determined as the synthesized corpus corresponding to the target processing task, and the task processing result represented by the parent node of the first node is determined as the label of the synthesized corpus pair; starting from the first node, the parent node is searched layer by layer upwards until the found parent node is a first-level child node, and based on the similarity between the found parent node and the root node, a fourth node is determined from the found parent nodes; the candidate corpus represented by the fourth node is determined as the synthesized corpus corresponding to the target processing task, and the task processing result represented by the parent node of the fourth node is determined as the label of the synthesized corpus.
[0130] For example, Figure 3Taking the Monte Carlo tree shown as an example, assuming that the fourth-level child node on the far right is the first node, the candidate corpus represented by the first node is determined as a synthetic corpus, and the candidate processing result represented by the middle third-level child node is determined as the label of the synthetic corpus; continuing to search upward, the parent node of the third-level child node is the candidate corpus represented by the second-level child node on the left, but the similarity between the second-level child node and the root node is relatively low, so this child node and its parent node are discarded.
[0131] Compared with the second implementation method, this implementation method can further improve the quality of the synthetic corpus and its label.
[0132] The above shows some implementation methods of the above S106. Of course, it should be understood that the above S106 can also be implemented in other ways, and the embodiments of this specification do not limit this.
[0133] The corpus synthesis method proposed in the embodiments of this specification adopts the technical concept of corpus synthesis based on MCTS and cycle consistency. The original corpus is used as the root node, and the root node is expanded layer by layer based on MCTS. That is, first, the root node is used as the starting point, and the target processing task is executed on the starting point through a large language model. The multiple task processing results obtained are respectively used as a child node, and then the inverse operation of the target processing task is executed on the child node, and multiple candidate corpora can be obtained. Each candidate corpus is the child node of the child node. In this way, a Monte Carlo tree including the root node and n-level child nodes is constructed. Each even-level child node represents a candidate corpus. Since such candidate corpora are expanded based on their upstream nodes, they not only conform to the context, but also can capture the uncertainty and diversity in the language and contain rich generated content; further, by calculating the similarity between the n-level child node and the root node, the first node is determined from the n-level child nodes, and the synthetic corpus corresponding to the target processing task is generated based on the first node, that is, it is evaluated whether the candidate corpus conforms to the cycle consistency with the original corpus, and the candidate corpus that conforms to the cycle consistency with the original corpus is obtained. Generating a synthetic corpus based on such candidate corpora can ensure that the synthetic corpus is highly consistent with the original corpus and solve the problem that the synthetic corpus may overfit the distribution of the original corpus or cannot fully simulate the original corpus.
[0134] Please refer to Figure 4 , the embodiments of this specification also provide a training method for a large language model, and this method includes the following steps: S402, obtain a training corpus set for the target processing task.
[0135] The training corpus set includes the original corpus and the synthetic corpus. Each corpus has a label corresponding to the target processing task. The synthetic corpus is obtained by processing the original corpus through the above corpus synthesis method.
[0136] S404: Train the large language model based on each corpus and its label in the training corpus.
[0137] Training the large language model is a process of obtaining the parameters of the large language model from the training corpus. This process is generally divided into two stages: pre-training and fine-tuning.
[0138] Pre-training refers to the process of training the large language model using a large amount of text data without specific task labels. In this process, the large language model learns the general features and patterns of language. Pre-training usually adopts unsupervised learning or self-supervised learning methods, such as letting the large language model predict the next word in the corpus or letting the large language model predict the masked word. The purpose of pre-training is to enable the large language model to capture rich language knowledge, so that in subsequent tasks, only a small amount of data is needed for fine-tuning to achieve good results.
[0139] The large language model is an autoregressive sequence prediction model, where the output of the model depends not only on the current input but also on the previous outputs generated by the model. In natural language processing, autoregressive models are usually used for language modeling, that is, predicting the next word or character in a sentence. To this end, in the embodiments of this specification, the pre-training of the large language model is achieved by letting the large language model predict the next word in the corpus.
[0140] Specifically, for each corpus in the training corpus, assuming that the corpus contains T words, given the first t - 1 words in the corpus, the large language model predicts the probability of the t-th word. Based on the probability of each word, with the goal of maximizing the likelihood function, the parameters of the large language model are adjusted.
[0141] Among them, the likelihood function is shown in the following formula (1): (1) Among them, represents the parameters of the large language model, represents the probability, represents the probability that the large language model predicts the t-th word given the first t - 1 words.
[0142] Fine-tuning is a process of further training the model using labeled data for specific tasks on the basis of pre-training. Through fine-tuning, the large language model can be optimized for specific tasks. In the fine-tuning stage, most of the parameters of the pre-trained large language model are usually fixed, and only some layers (such as the output layer) or a small number of parameters of the large language model are trained. The purpose of fine-tuning is to make the pre-trained large language model adapt to new tasks and improve the performance of the large language model on specific tasks.
[0143] In the embodiments of this specification, in the fine-tuning stage, for each corpus in the training corpus, the large language model performs target task processing on the corpus to obtain the task processing result corresponding to the corpus; then, based on the task processing result and label corresponding to each corpus, the parameters of the large language model are adjusted.
[0144] Exemplarily, based on the task processing result and label corresponding to each corpus, the loss of the large language model is determined; then, with the goal of minimizing the loss, algorithms such as the backpropagation algorithm and the stochastic gradient descent algorithm are used to adjust the parameters of the large language model.
[0145] Among them, the loss function of the large language model usually depends on the specific content of the target processing task. For example, for a classification task, the objective function can be the cross-entropy loss, as shown in the following formula (2): (2) Among them, represents the loss of the large language model, represents the number of corpora in the training corpus, represents the i-th corpus, represents the label of the i-th corpus, represents the task processing result corresponding to the i-th corpus.
[0146] Backpropagation updates these parameters by calculating the gradient of the loss with respect to the model parameters. The following are the basic steps of the backpropagation algorithm: Forward propagation: The input data is passed through the network layer by layer, and each layer calculates its output until the prediction result is obtained at the last layer.
[0147] Calculate the loss: Use the prediction result and the true label to calculate the value of the loss function, and the loss function measures the error of the model prediction.
[0148] Backpropagation gradient: Starting from the output layer, calculate the gradient of the loss function with respect to the parameters of each layer in reverse. This process involves the chain rule, and the gradient of each layer is passed to the previous layer.
[0149] Update the parameters: Use the calculated gradient to update the weights and biases of the network, with the aim of reducing the value of the loss function.
[0150] Stochastic gradient descent is a mathematical optimization algorithm used to minimize the loss function when training a neural network. The following is the basic principle of stochastic gradient descent: Initialize the parameters: Randomly initialize the weights and biases of the network.
[0151] Select the batch size: Randomly select a batch from the training data. For example, train with 1 piece of text or 20 pieces of text, or it can be other values.
[0152] Calculate the gradient: For the selected batch of data, use the backpropagation algorithm to calculate the gradient of the loss function with respect to the network parameters.
[0153] Update the parameters: Update the weights and biases of the network according to the calculated gradient. The update rule is , where \(\alpha\) represents the learning rate, represents the gradient of the objective function with respect to the parameters.
[0154] Repeat the process: Repeat the above steps, each time selecting different batches of data until a certain stopping condition is reached, such as reaching a predetermined number of iterations, the value of the loss function is lower than a certain threshold, or the performance on the validation set no longer improves.
[0155] In summary, large language models learn extensive language knowledge through pre-training and then adapt to specific application scenarios through fine-tuning. This method enables large language models to perform well in a variety of different language tasks, greatly promoting the development of the natural language processing field.
[0156] In the embodiments of this specification, the above S404 can be implemented in various appropriate ways.
[0157] In one implementation, the synthetic corpus in the training corpus can be obtained by processing the original corpus with a large language model before training, and the training corpus is not updated during each subsequent round of training.
[0158] In another implementation, before training the large language model, first process the original corpus with the large language model before training to obtain a synthetic corpus, and the initial training corpus is composed of the original corpus and the synthetic corpus; further, after each round of training of the large language model, use the large language model after each round of training to update the training corpus for the next round of training.
[0159] Specifically, the above S404 includes the following steps: Perform multiple rounds of training on the large language model. Each round of training includes: performing the target processing task on each corpus in the training corpus through the large language model to obtain the task processing result corresponding to each corpus; adjusting the parameters of the large language model based on the task processing result and label corresponding to each corpus to obtain the large language model after this round of training; obtaining a new synthetic corpus based on the large language model after this round of training and the original corpus, and using the new synthetic corpus for the next round of training.
[0160] For ease of understanding, the following combinesFigure 5 The above training process will be described.
[0161] As Figure 5 shown, first, the large language model is preliminarily trained using the original corpus and its labels, and the parameters of the large language model are stored in the computer memory.
[0162] Then, the preliminarily trained large language model is called to execute the above corpus synthesis method to generate a synthetic corpus based on the original corpus. Specifically, it includes: using the MCTS planner, taking the original corpus as the root node, calling the preliminarily trained large language model to expand the root node layer by layer to obtain n-level child nodes; using the similarity score calculator, based on the similarity between the n-level child nodes and the root node, determining the first node from the n-level child nodes, and generating the synthetic corpus corresponding to the target processing task based on the first node, and sending the synthetic corpus to the storage engine.
[0163] Furthermore, the storage engine provides the synthetic data to the language model training algorithm. Through the language model training algorithm, based on the original corpus and the synthetic corpus, the large language model is trained for multiple rounds until the training stop condition is met. For each round of training, for each corpus in the training corpus set of this round, the target processing task is executed on the corpus based on the large language model to obtain the task processing result corresponding to the corpus; then, based on the task processing result and label corresponding to each corpus, the parameters of the large language model are adjusted to obtain the large language model after this round of training; then, the large language model after this round of training is called to execute the above corpus synthesis method to generate a new synthetic corpus on the basis of the original corpus, and the new synthetic corpus is added to the training corpus set as the training corpus set for the next round.
[0164] The training method of the large language model provided in the embodiments of this specification abandons the traditional idea of training the large language model based on the original corpus, and instead combines the original corpus and the synthetic corpus to train the large language model, so that the synthetic corpus can effectively make up for the deficiencies of the original corpus, thereby improving the training effect of the large language model.
[0165] Corresponding to the above Figure 1 shown corpus synthesis method, the embodiments of this specification also provide a corpus synthesis device. Figure 6 It is a schematic structural diagram of a corpus synthesis device 600 provided in the embodiments of this specification, including: an expansion module 610, a first determination module 620, and a generation module 630.
[0166] An expansion module 610 is used to take the original corpus of the large language model as the root node, and expand the root node layer by layer to obtain n-level child nodes; the (i-1)-th level child node represents a task processing result obtained by performing a target processing task on the parent node of the (i-1)-th level child node through the large language model, and the i-th level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level child node. i is an even number, and 1 < i ≤ n.
[0167] A first determination module 620 is used to determine a first node from the n-level child nodes based on the similarity between the n-level child nodes and the root node.
[0168] A generation module 630 is used to generate a synthetic corpus corresponding to the target processing task based on the first node.
[0169] The corpus synthesis device provided by the embodiments of this specification adopts the technical concept of corpus synthesis based on Monte Carlo tree search and cycle consistency. It takes the original corpus as the root node and expands the root node layer by layer based on Monte Carlo tree search. That is, first take the root node as the starting point, perform a target processing task on the starting point through the large language model, and take the obtained multiple task processing results as a child node respectively. Then perform the inverse operation of the target processing task on the child node, and multiple candidate corpora can be obtained. Each candidate corpus is the child node of the child node. In this way, a Monte Carlo tree containing the root node and n-level child nodes is constructed. Each even-level child node represents a candidate corpus. Since such candidate corpora are expanded based on their upstream nodes, they not only conform to the context but also can capture the uncertainty and diversity in the language, containing rich generated content. Further, by calculating the similarity between the n-level child nodes and the root node, a first node is determined from the n-level child nodes, and a synthetic corpus corresponding to the target processing task is generated based on the first node, that is, to evaluate whether the candidate corpus conforms to cycle consistency with the original corpus, obtain a candidate corpus that conforms to cycle consistency with the original corpus, and generate a synthetic corpus based on such candidate corpora, which can ensure that the synthetic corpus is highly consistent with the original corpus and solve the problem that the synthetic corpus may overfit the distribution of the original corpus or cannot fully simulate the original corpus.
[0170] In another embodiment, the expansion module is used for: Determine the root node as the second node, and repeat the following expansion for the second node: Perform the target processing task on the second node through the large language model to obtain the child nodes of the second node; Determine a third node from the child nodes of the second node based on the similarity between the child nodes of the second node and the root node; Performing the inverse operation of the target processing task on the third node through the large language model to obtain the child nodes of the third node; If the child nodes of the third node do not meet the preset stop expansion condition, then based on the similarity between the child nodes of the third node and the root node, new second nodes are determined from the child nodes of the third node.
[0171] In another embodiment, the first determination module is configured to: Determine the score of the nth-level child node based on the similarity between the nth-level child node and the root node; Determine the child nodes in the nth-level child nodes whose scores are greater than or equal to the score threshold as the first nodes.
[0172] In another embodiment, when the first determination module determines the score of the nth-level child node based on the similarity between the nth-level child node and the root node, the following steps are performed: Sampling from the standard Gaussian distribution to obtain a target value; Determine the sum of the similarity between the nth-level child node and the root node and the target value as the score of the nth-level child node.
[0173] In another embodiment, the generation module is configured to: Determine the candidate corpus represented by the first node as the synthetic corpus corresponding to the target processing task, and determine the task processing result represented by the parent node of the first node as the label of the synthetic corpus pair; Starting from the first node, search for the parent nodes layer by layer upward until the found parent node is the first-level child node, and based on the similarity between the found parent node and the root node, determine the fourth node from the found parent nodes; Determine the candidate corpus represented by the fourth node as the synthetic corpus corresponding to the target processing task, and determine the task processing result represented by the parent node of the fourth node as the label of the synthetic corpus.
[0174] In another embodiment, the original corpus has a label corresponding to the target processing task; The similarity between the (i - 1)th-level child node and the root node is determined by the following method: Performing word segmentation on the label to obtain a plurality of first strings, and performing word segmentation on the task processing result represented by the (i - 1)th-level child node to obtain a plurality of second strings; Based on the proportion of the first strings in the label that match the second strings, obtain the recall rate of the (i - 1)th-level child node; Based on the proportion of the second strings that match the first string in the task processing results represented by the (i - 1)-th level child nodes, obtain the precision rate of the (i - 1)-th level child nodes. Based on the recall rate and precision rate of the (i - 1)-th level child nodes, determine the similarity between the (i - 1)-th level child nodes and the root node.
[0175] In another embodiment, the similarity between the i-th level child nodes and the root node is determined by the following method: Perform word segmentation on the original corpus to obtain a plurality of third strings, and perform word segmentation on the candidate corpus represented by the i-th level child nodes to obtain a plurality of fourth strings; Based on the proportion of the third strings that match the fourth strings in the original corpus, obtain the recall rate of the i-th level child nodes; Based on the proportion of the fourth strings that match the third strings in the candidate corpus represented by the i-th level child nodes, obtain the precision rate of the i-th level child nodes; Based on the recall rate and precision rate of the i-th level child nodes, determine the similarity between the i-th level child nodes and the root node.
[0176] Obviously, the corpus synthesis device in the embodiments of this specification can be used as the execution subject of the Figure 1 corpus synthesis method shown above, and thus can implement the functions achieved by the Figure 1 corpus synthesis method. Since the principle is the same, it will not be elaborated here.
[0177] Corresponding to the Figure 4 training method of the large language model shown above, the embodiments of this specification also provide a training device for a large language model. Figure 7 FIG. is a schematic structural diagram of a training device 700 for a large language model provided by the embodiments of this specification, including: an acquisition module 710 and a training module 720.
[0178] The acquisition module 710 is configured to acquire a training corpus set for a target processing task, where the training corpus set includes an original corpus and a synthesized corpus, and the synthesized corpus is obtained from the original corpus by using the corpus synthesis method according to any one of claims 1 to 7.
[0179] The training module 720 is configured to train the large language model based on each corpus in the training corpus set and the label of each corpus.
[0180] The training device for a large language model provided by the embodiments of this specification abandons the traditional idea of training a large language model based on original corpus, and instead combines the original corpus and synthetic corpus to train the large language model, enabling the synthetic corpus to effectively make up for the deficiencies of the original corpus, thereby improving the training effect of the large language model.
[0181] In another embodiment, the training module is configured to: Perform multiple rounds of training on the large language model, and each round of training includes: Execute the target processing task on each corpus in the training corpus set through the large language model to obtain the task processing result corresponding to each corpus; Adjust the parameters of the large language model based on the task processing result and label corresponding to each corpus to obtain the large language model after this round of training; Obtain a new synthetic corpus based on the large language model after this round of training and the original corpus, and use the new synthetic corpus for the next round of training. The new synthetic corpus is obtained by the corpus synthesis method according to any one of claims 1 to 7.
[0182] Obviously, the training device for a large language model in the embodiments of this specification can be used as the execution subject of the above Figure 4 shown large language model training method, and thus can implement the functions achieved by the large language model training method in Figure 4 Since the principle is the same, it will not be elaborated here.
[0183] Figure 8 is a schematic structural diagram of an electronic device according to an embodiment of this specification. Please refer to Figure 8 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0184] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a bidirectional arrow is used in
[0185] Memory, used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0186] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a corpus synthesis device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Taking the original corpus of the large language model as the root node, expanding the root node layer by layer to obtain n-level child nodes; the (i-1)-th level child node represents a task processing result obtained by performing a target processing task on the parent node of the (i-1)-th level child node through the large language model, and the i-th level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level child node through the large language model, where i is an even number and 1 < i ≤ n; Determining a first node from the n-level child nodes based on the similarity between the n-level child nodes and the root node; Generating a synthesized corpus corresponding to the target processing task based on the first node.
[0187] Alternatively, the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a training device for the large language model at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Obtaining a training corpus set for the target processing task, where the training corpus set includes an original corpus and a synthesized corpus, and each corpus has a label corresponding to the target processing task, and the synthesized corpus is obtained by processing the original corpus through the corpus synthesis method provided in the embodiments of this specification; Training the large language model based on each corpus and the label of each corpus in the training corpus set.
[0188] As described above in this specification Figure 1 The method executed by the corpus synthesis device disclosed in the embodiments shown, or Figure 4 The method executed by the training device of the large language model disclosed in the embodiments shown can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this specification can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0189] It should be understood that the electronic device in the embodiments of this specification can implement the functions of the corpus synthesis device in Figure 1 the embodiments shown, or the electronic device in the embodiments of this specification can implement the functions of the training device of the large language model in Figure 4 the embodiments shown. Since the principles are the same, they will not be elaborated herein in the embodiments of this specification.
[0190] Of course, in addition to the software implementation method, the electronic device of this specification does not exclude other implementation methods, such as a logical device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and may also be hardware or a logical device.
[0191] The embodiments of this specification also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can cause the electronic device to execute Figure 1 the method of the illustrated embodiment, and specifically used to perform the following operations: Taking the original corpus of the large language model as the root node, expanding the root node layer by layer to obtain n-level child nodes; the (i-1)-th level child node represents a task processing result obtained by performing a target processing task on the parent node of the (i-1)-th level child node through the large language model, and the i-th level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level child node through the large language model, where i is an even number and 1 < i ≤ n; Based on the similarity between the n-level child nodes and the root node, determining a first node from the n-level child nodes; Based on the first node, generating a synthetic corpus corresponding to the target processing task.
[0192] Alternatively, when the instruction is executed by an electronic device including multiple application programs, it can cause the electronic device to execute Figure 4 the method of the illustrated embodiment, and specifically used to perform the following operations: Obtaining a training corpus set for the target processing task, the training corpus set including an original corpus and a synthetic corpus, each corpus having a label corresponding to the target processing task, and the synthetic corpus being obtained by processing the original corpus through the corpus synthesis method provided by the embodiments of this specification; Based on each corpus and the label of each corpus in the training corpus set, training the large language model.
[0193] The embodiments of this specification also provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute some or all of the steps in the corpus synthesis method or the training method of the large language model provided by the embodiments of this specification.
[0194] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0195] In summary, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.
[0196] The systems, devices, modules or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0197] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0198] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the said element.
[0199] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.
Claims
1. A corpus synthesis method, characterized in that, Including: Taking the original corpus of the large language model as the root node, and expanding the root node layer by layer to obtain n-level child nodes; The (i-1)-th level child node represents a task processing result obtained by performing a target processing task on the parent node of the (i-1)-th level child node through the large language model, and the i-th level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level child node. i is an even number, and 1 < i ≤ n; Based on the similarity between the n-level child nodes and the root node, determining the first node from the n-level child nodes; Generating a synthetic corpus corresponding to the target processing task based on the first node.
2. The method according to claim 1, characterized in that, The expanding the root node layer by layer to obtain n-level child nodes includes: Determining the root node as the second node, and repeatedly expanding the second node as follows: Performing the target processing task on the second node through the large language model to obtain the child nodes of the second node; Based on the similarity between the child nodes of the second node and the root node, determining the third node from the child nodes of the second node; Performing the inverse operation of the target processing task on the third node through the large language model to obtain the child nodes of the third node; If the child nodes of the third node do not meet the preset stop expansion condition, then based on the similarity between the child nodes of the third node and the root node, determining a new second node from the child nodes of the third node.
3. The method according to claim 1, wherein The determining the first node from the n-level child nodes based on the similarity between the n-level child nodes and the root node includes: Determining the score of the n-level child node based on the similarity between the n-level child node and the root node; Determining the child nodes in the n-level child nodes with scores greater than or equal to the score threshold as the first node.
4. The method according to claim 3, wherein The determining the score of the n-level child node based on the similarity between the n-level child node and the root node includes: Sampling from the standard Gaussian distribution to obtain a target value; Determining the sum of the similarity between the n-level child node and the root node and the target value as the score of the n-level child node.
5. The method according to claim 1, characterized in that The generating a synthetic corpus corresponding to the target processing task based on the first node includes: Determining the candidate corpus represented by the first node as the synthetic corpus corresponding to the target processing task, and determining the task processing result represented by the parent node of the first node as the label of the synthetic corpus pair; Starting from the first node, searching for the parent node layer by layer upward until the parent node found is the 1st level child node, and based on the similarity between the found parent node and the root node, determining the fourth node from the found parent nodes; Determining the candidate corpus represented by the fourth node as the synthetic corpus corresponding to the target processing task, and determining the task processing result represented by the parent node of the fourth node as the label of the synthetic corpus.
6. The method according to claim 1, wherein The original corpus has a label corresponding to the target processing task; The similarity between the (i - 1)-th level child node and the root node is determined as follows: Tokenize the label to obtain a plurality of first strings, and tokenize the task processing result represented by the (i - 1)-th level child node to obtain a plurality of second strings; Based on the proportion of the first strings in the label that match the second strings, obtain the recall rate of the (i - 1)-th level child node; Based on the proportion of the second strings in the task processing result represented by the (i - 1)-th level child node that match the first strings, obtain the precision rate of the (i - 1)-th level child node; Based on the recall rate and precision rate of the (i - 1)-th level child node, determine the similarity between the (i - 1)-th level child node and the root node.
7. The method according to claim 1, wherein The similarity between the i-th level child node and the root node is determined as follows: Tokenize the original corpus to obtain a plurality of third strings, and tokenize the candidate corpus represented by the i-th level child node to obtain a plurality of fourth strings; Based on the proportion of the third strings in the original corpus that match the fourth strings, obtain the recall rate of the i-th level child node; Based on the proportion of the fourth strings in the candidate corpus represented by the i-th level child node that match the third strings, obtain the precision rate of the i-th level child node; Based on the recall rate and precision rate of the i-th level child node, determine the similarity between the i-th level child node and the root node.
8. A training method for a large language model, characterized in that, It includes: Obtain a training corpus set for the target processing task, where the training corpus set includes an original corpus and a synthetic corpus, each corpus has a label corresponding to the target processing task, and the synthetic corpus is obtained by processing the original corpus through the corpus synthesis method according to any one of claims 1 to 7; Train the large language model based on each corpus and its label in the training corpus set.
9. The method according to claim 8, wherein The training of the large language model based on each corpus and its label in the training corpus set includes: Perform multiple rounds of training on the large language model, and each round of training includes: Execute the target processing task on each corpus in the training corpus set through the large language model to obtain a task processing result corresponding to each corpus; Based on the task processing result and label corresponding to each corpus, adjust the parameters of the large language model to obtain the large language model after this round of training; Based on the large language model after this round of training and the original corpus, obtain a new synthetic corpus, and use the new synthetic corpus for the next round of training. The new synthetic corpus is obtained through the corpus synthesis method according to any one of claims 1 to 7.
10. A corpus synthesis device, characterized in that, It includes: An expansion module, which is used to use the original corpus of the large language model as the root node, and expand the root node layer by layer to obtain n-level child nodes; The (i - 1)-th level child node represents a task processing result obtained by performing a target processing task on the parent node of the (i - 1)-th level child node through the large language model. The i-th level child node represents a candidate corpus obtained by performing the inverse operation of the target processing task on the parent node of the i-th level child node through the large language model, where i is an even number and 1 < i ≤ n; A first determination module, configured to determine a first node from the n-th level child nodes based on the similarity between the n-th level child nodes and the root node; A generation module, configured to generate a synthetic corpus corresponding to the target processing task based on the first node.
11. A training device for a large language model, characterized in that, Comprising: An acquisition module, configured to acquire a training corpus set for a target processing task, where the training corpus set includes an original corpus and a synthetic corpus, and the synthetic corpus is obtained from the original corpus by using the corpus synthesis method according to any one of claims 1 to 7; A training module, configured to train the large language model based on each corpus in the training corpus set and the label of each corpus.
12. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the corpus synthesis method according to any one of claims 1 to 7 or the training method of the large language model according to any one of claims 8 to 9.
13. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the corpus synthesis method according to any one of claims 1 to 7 or the training method of the large language model according to any one of claims 8 to 9.
14. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps in the corpus synthesis method according to any one of claims 1 to 7 or the training method of the large language model according to any one of claims 8 to 9.