Text sequence generation method, processing method, model training method, and device for enhancing context learning ability
By performing N iterative training on large language models and synthesizing texts, text sequences that enhance context learning ability are generated, solving the problem of high cost of training large models, and achieving rapid acquisition of better responses and improving model capabilities.
Patent Information
- Application Number
- CN202410323574.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-03-20
AI Technical Summary
The parameter scale of large language models is huge, the model training cost is high, and it is difficult to obtain better responses through continuous training.
By performing N iterative training on the initial model, N batches of synthetic text are generated and spliced into text sequences that enhance context learning ability. The evaluation text and reply text are used to optimize the model iteratively, reducing manual intervention.
It reduces the cost of model training, quickly obtains better reply text, improves the context learning ability of the model, and reduces resource consumption during the training process.
Smart Images

Figure CN118193996B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to the fields of large models, large language models, and generative model technologies. More specifically, the present disclosure provides a text sequence generation method for enhancing context learning ability, a processing method for enhancing context learning ability, a model training method for enhancing context learning ability, an apparatus, an electronic device, a storage medium, and a computer program product. Background Art
[0002] With the development of artificial intelligence technology, the application of large language models has received extensive attention. In the scenario of using large language models for intelligent responses, to obtain better responses, it is necessary to continuously train the large model so that the large model continuously generates better responses until a response that meets the requirements is generated.
[0003] However, the parameter scale of large language models is huge, and the cost of model training is high. Continuously training the large model to obtain better responses will bring a large training cost. Summary of the Invention
[0004] The present disclosure provides a text sequence generation method for enhancing context learning ability, a processing method for enhancing context learning ability, a model training method for enhancing context learning ability, an apparatus, an electronic device, a storage medium, and a computer program product.
[0005] According to a first aspect, there is provided a text sequence generation method for enhancing context learning ability, the method comprising: performing N iterative trainings on an initial model using a query text to obtain N batches of synthetic texts, where N is an integer greater than 1; and concatenating the N batches of synthetic texts to obtain a text sequence with enhanced context learning ability; wherein performing N iterative trainings on the initial model using the query text to obtain N batches of synthetic texts includes: for the nth training, performing the nth training using the evaluation text of the model obtained after the (n - 1)th training, the evaluation text being obtained by evaluating the response text of the model obtained after the (n - 1)th training, the response text being a response to the query text, n being an integer greater than 1 and less than or equal to N; and generating the nth batch of synthetic texts according to the query text, response text, and evaluation text corresponding to the nth training.
[0006] According to a second aspect, a processing method for enhancing context learning ability is provided. The method includes: inputting the text to be queried into a trained model to obtain an initial response text, where the trained model is trained using a text sequence for enhancing context learning ability, and the text sequence for enhancing context learning ability is generated according to the method of the first aspect above; evaluating the initial response text to obtain an initial evaluation text; and using the trained model to generate a new response text based on the initial evaluation text and the context information of the text sequence.
[0007] According to a third aspect, a model training method for enhancing context learning ability is provided. The method includes: obtaining a text sequence for enhancing context learning ability, where the text sequence for enhancing context learning ability is generated according to the method of the first aspect above; inputting the text sequence into the model to be trained to obtain a probability sequence corresponding to the text sequence; determining a loss according to the probability sequence; and adjusting the parameters of the model to be trained according to the loss to obtain a trained model.
[0008] According to a fourth aspect, a text sequence generation device for enhancing context learning ability is provided. The device includes: an iteration module for obtaining N batches of synthetic texts by performing N - time iterative training on an initial model using a query text, where N is an integer greater than 1; and a splicing module for splicing the N batches of synthetic texts to obtain a text sequence for enhancing context learning ability; where the iteration module includes: a training unit for, for the n - th training, using the evaluation text of the model obtained after the (n - 1)-th training to perform the n - th training, the evaluation text being obtained by evaluating the response text of the model obtained after the (n - 1)-th training, the response text being a response to the query text, and n being an integer greater than 1 and less than or equal to N; and a synthetic text generation unit for generating the n - th batch of synthetic texts according to the query text, response text, and evaluation text corresponding to the n - th training.
[0009] According to a fifth aspect, a processing device for enhancing context learning ability is provided. The device includes: a first response module for inputting the text to be queried into a trained model to obtain an initial response text, where the trained model is trained using a text sequence for enhancing context learning ability, and the text sequence for enhancing context learning ability is generated according to the device of the fourth aspect above; an evaluation module for evaluating the initial response text to obtain an initial evaluation text; and a second response module for using the trained model to generate a new response text based on the initial evaluation text and the context information of the text sequence.
[0010] According to a sixth aspect, there is provided a model training apparatus for enhancing context learning ability, the apparatus including: an acquisition module configured to acquire a text sequence for enhancing context learning ability, where the text sequence for enhancing context learning ability is generated according to the apparatus of the fourth aspect above; a processing module configured to input the text sequence into a model to be trained to obtain a probability sequence corresponding to the text sequence; a loss determination module configured to determine a loss according to the probability sequence; and an adjustment module configured to adjust parameters of the model to be trained according to the loss to obtain a trained model.
[0011] According to a seventh aspect, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0012] According to an eighth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided by the present disclosure.
[0013] According to a ninth aspect, there is provided a computer program product, including a computer program, where the computer program is stored on at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method provided by the present disclosure.
[0014] It should be understood that the content described in this part is not intended to identify key or important features of embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0016] Figure 1 is a schematic diagram of an exemplary system architecture to which at least one of a method for generating a text sequence for enhancing context learning ability, a processing method, and a model training method according to an embodiment of the present disclosure can be applied;
[0017] Figure 2 is a flowchart of a method for generating a text sequence for enhancing context learning ability according to an embodiment of the present disclosure;
[0018] Figure 3 is a schematic diagram of a method for performing N - iteration training on an initial model according to an embodiment of the present disclosure;
[0019] Figure 4AIt is a schematic diagram of splicing synthetic texts of N batches according to an embodiment of the present disclosure;
[0020] Figure 4B It is a schematic diagram of splicing synthetic texts of N batches according to another embodiment of the present disclosure;
[0021] Figure 5 It is a flowchart of a processing method for enhancing context learning ability according to an embodiment of the present disclosure;
[0022] Figure 6 It is a flowchart of a model training method for enhancing context learning ability according to an embodiment of the present disclosure;
[0023] Figure 7 It is a block diagram of a text sequence generation device for enhancing context learning ability according to an embodiment of the present disclosure;
[0024] Figure 8 It is a block diagram of a processing device for enhancing context learning ability according to an embodiment of the present disclosure;
[0025] Figure 9 It is a block diagram of a model training device for enhancing context learning ability according to an embodiment of the present disclosure;
[0026] Figure 10 It is a block diagram of an electronic device for at least one of a text sequence generation method, a processing method, and a model training method for enhancing context learning ability according to an embodiment of the present disclosure. Detailed implementation manners
[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0028] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0029] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0030] Figure 1It is a schematic diagram of an exemplary system architecture that can apply at least one of a text sequence generation method, a processing method, and a model training method with enhanced context learning ability according to an embodiment of the present disclosure. It should be noted that Figure 1 What is shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0031] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices, including but not limited to smartphones, tablets, laptop computers, etc.
[0033] The processing method with enhanced context learning ability provided by the embodiments of the present disclosure can generally be executed by the terminal devices 101, 102, 103. Correspondingly, the text processing device provided by the embodiments of the present disclosure can generally be set in the terminal devices 101, 102, 103.
[0034] At least one of the text sequence generation method with enhanced context learning ability and the model training method with enhanced context learning ability provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, at least one of the sample text sequence generation device and the model training device provided by the embodiments of the present disclosure can generally be set in the server 105.
[0035] Figure 2 It is a flowchart of a text sequence generation method with enhanced context learning ability according to an embodiment of the present disclosure.
[0036] As Figure 2 shown, the text sequence generation method 200 with enhanced context learning ability includes operation S210 to operation S220.
[0037] In operation S210, the initial model is iteratively trained N times using the query text to obtain N batches of synthetic text.
[0038] N, as the number of iterations, is an integer greater than 1. The query text can be a batch of query texts, such as including query text A, query text B, query text C, etc. The initial model can be a large language model. The initial model is trained iteratively N times using the query text, where each of the N trainings can be fine-tuning based on reinforcement learning. The fine-tuning process based on reinforcement learning includes using the "question text - response text - evaluation text" triple for reinforcement learning.
[0039] The query text can be various statements used for querying in different application scenarios according to the application scenario. The query statement can be an interrogative sentence, such as "What is the highest mountain in the world?" It can also be other forms of statements such as declarative sentences and imperative sentences, such as "Translate the following content".
[0040] The question text is input into the model, and the model generates the corresponding response text. For example, the response text can include the answer to the question, the translation result, etc.
[0041] The response text can be evaluated to obtain the evaluation text. Another large model can be used to automatically evaluate the response text generated by the current large model, or it can be evaluated based on humans. The evaluation text can include the satisfaction level with the response text, such as very good, moderate, poor. The evaluation text can also be an evaluation value, such as 60 points, 80 points, etc. In addition, the evaluation text can also include statement sentences such as opinions, such as "Answer the question in a humorous tone", "Adjust the format of the translation text", etc.
[0042] Operation S210 specifically includes operations S211 to S212.
[0043] In operation S211, for the nth training, the evaluation text of the model obtained after the (n - 1)th training is used for the nth training.
[0044] For example, for each training, the evaluation text of the model obtained from the previous training can be used to adjust the parameters of the model, so that the response text output by the model obtained from this training is better than the response text output by the model obtained from the previous training.
[0045] According to an embodiment of the present disclosure, for the model obtained after the nth training, the query text is input into the model obtained after the nth training to obtain the response text corresponding to the nth training; and the response text corresponding to the nth training is evaluated to obtain the evaluation text corresponding to the nth training.
[0046] For example, for each trained model obtained after each training, the query text is input into the trained model of that time to obtain a response text, and the response text is evaluated to obtain an evaluation text. The evaluation text is used to adjust the model parameters during the next model training to increase or decrease the probability of the model outputting the previous response text. The model after parameter adjustment is the new trained model. The query text is input into the new trained model to obtain a new response text, and then the new response text is evaluated to obtain a new evaluation text. The new evaluation text is used for the next training, and so on, until the number of iterations reaches a threshold or the evaluation text meets the preset requirements (for example, the evaluation value reaches 90).
[0047] For example, before model training, the query text can be input into the untrained model to obtain a response text, and the response text is evaluated to obtain an evaluation text. The evaluation text can be used for the first model training.
[0048] In operation S212, according to the query text, response text, and evaluation text corresponding to the nth training, the nth batch of synthetic text is generated.
[0049] In the above iterative process, for each trained model obtained after each training, the model is used to process the query text to obtain a response text, and the response text is evaluated to obtain an evaluation text. The query text, response text, and evaluation text in this process can be combined to obtain synthetic text, which is used as a batch of synthetic text. After N iterative trainings, N batches of synthetic text can be obtained.
[0050] In operation S220, the N batches of synthetic text are concatenated to obtain a text sequence that enhances the context learning ability.
[0051] For example, the N batches of synthetic text can be concatenated in the iterative order to obtain a text sequence that enhances the context learning ability. In addition, the same query texts in the N batches of synthetic text can be extracted, and then the query text, as well as the response texts and evaluation texts of different batches corresponding to the query text, are concatenated in sequence to obtain a text sequence that enhances the context learning ability.
[0052] Embodiments of the present disclosure concatenate multiple batches of query texts, response texts, and evaluation texts generated by iterative training of the model to obtain a text sequence that enhances the context learning ability. The text sequence that enhances the context learning ability includes response texts optimized through multiple evaluations, so that the text sequence includes context information for continuously generating better responses. Using this text sequence as a sample for training a new model can enable the new model to learn the context information of the text sequence, thereby improving the response effect of the model.
[0053] Figure 3 Schematic diagram of a method for training an initial model for N iterations according to an embodiment of the present disclosure.
[0054] As Figure 3 shown, the initial model can be Model 1 of version 0, and Model 1 can be a large language model. Queries A, B, C, etc. can be a batch of query texts. Inputting this batch of query texts into Model 1 of version 0 respectively obtains reply texts such as Reply A1, Reply B1, Reply C1, etc. Evaluating the reply texts such as Reply A1, Reply B1, Reply C1 respectively obtains evaluation texts such as Evaluation A1, Evaluation B1, Evaluation C1, etc. The query texts such as Queries A, B, C, etc., the reply texts such as Reply A1, Reply B1, Reply C1, etc., and the evaluation texts such as Evaluation A1, Evaluation B1, Evaluation C1, etc. are combined into the synthetic texts of the first batch.
[0055] Using the synthetic texts of the first batch to train the version 0 model, the first version of the model is obtained. The training process can be fine-tuning based on reinforcement learning. Taking Query A, Reply A1, and Evaluation A1 as an example, training the version 0 model includes: Inputting Query A, Reply A1, and Evaluation A1 into the version 0 model. The version 0 model processes Query A based on Evaluation A1, causing the version 0 model to adjust its parameters, thereby increasing or decreasing the probability of outputting Reply A1. The version 0 model after parameter adjustment is used as the version 1 model.
[0056] Next, inputting the query texts such as Queries A, B, C, etc. into the version 1 model, the version 1 model outputs reply texts such as Reply A2, Reply B2, Reply C2, etc. Evaluating the reply texts such as Reply A2, Reply B2, Reply C2 respectively obtains evaluation texts such as Evaluation A2, Evaluation B2, Evaluation C2, etc. The query texts such as Queries A, B, C, etc., the reply texts such as Reply A2, Reply B2, Reply C2, etc., and the evaluation texts such as Evaluation A2, Evaluation B2, Evaluation C2, etc. are combined into the synthetic texts of the second batch.
[0057] Using the synthetic texts of the second batch to train the version 1 model, the training method is similar to that of training the version 0 model, which will not be elaborated here. The version 1 model is obtained as the version 2 model after training.
[0058] And so on, after multiple iterations, the final version of the model can be obtained. Each version of the model generates a batch of synthetic data. N iterations can obtain N batches of synthetic data, that is, N synthetic texts.
[0059] In the above iterative process, each iteration includes fine-tuning based on reinforcement learning and generation of synthetic data. The query text used for generating synthetic data each time can remain unchanged, and only automatic fine-tuning is performed through the evaluation text without manual rewriting of the query, achieving a fully automated iterative process and reducing the manual workload.
[0060] According to an embodiment of the present disclosure, operation S220 includes, for the same query text, concatenating the query text, and the response text and evaluation text in different batches of synthetic text corresponding to the query text to obtain a text sequence that enhances the context learning ability.
[0061] Figure 4A It is a schematic diagram of concatenating N batches of synthetic text according to an embodiment of the present disclosure.
[0062] As Figure 4A shown, the concatenation method in this embodiment is to concatenate the response text and evaluation text in different batches generated by the same query text.
[0063] For example, for query text A (Query A), the response text in the synthetic text of the first batch for this query text is Response A1, and the evaluation text is Evaluation A1. The response text in the synthetic text of the second batch for this query text is Response A2, and the evaluation text is Evaluation A2. Query A, Response A1, Evaluation A1, Response A2, Evaluation A2, etc. are concatenated.
[0064] Similarly, for query text B (Query B), Query B, Response B1, Evaluation B1, Response B2, Evaluation B2, etc. are concatenated. For query text C (Query C), Query C, Response C1, Evaluation C1, Response C2, Evaluation C2, etc. are concatenated.
[0065] According to an embodiment of the present disclosure, operation S220 includes, for each batch of synthetic text, concatenating the query text, response text, and evaluation text within the batch of synthetic text to obtain the concatenated text of this batch; and concatenating the concatenated texts of N batches to obtain a text sequence that enhances the context learning ability.
[0066] Figure 4B It is a schematic diagram of concatenating N batches of synthetic text according to another embodiment of the present disclosure.
[0067] As Figure 4B shown, the concatenation method in this embodiment is to concatenate all the query text, response text, and evaluation text within the same batch first, and then concatenate the concatenated texts of different batches.
[0068] For example, the query texts of one batch include Query A and Query B. The Query A, Reply A1, Evaluation A1, Query B, Reply B1, and Evaluation B1 within the first batch are concatenated to obtain the concatenated text of the first batch.
[0069] Similarly, the Query A, Reply A2, Evaluation A2, Query B, Reply B2, and Evaluation B2 within the second batch are concatenated to obtain the concatenated text of the second batch.
[0070] And so on, to obtain the concatenated texts of each of the N batches.
[0071] The concatenated texts of the N batches are further concatenated to obtain a text sequence with enhanced context learning ability.
[0072] According to an embodiment of the present disclosure, by concatenating the query texts, reply texts, and evaluation texts in the N batches, a text sequence with enhanced context learning ability is obtained. This text sequence with enhanced context learning ability contains the reply texts obtained after multiple evaluations, enabling the text sequence to contain the context information for continuously generating better replies.
[0073] According to an embodiment of the present disclosure, category hint information and batch information can be added to the query texts, reply texts, and evaluation texts in the text sequence with enhanced context learning ability, respectively.
[0074] For example, for the query text, a category identifier similar to "User Query" can be added. For the reply text, a category identifier similar to "Possible Reply Generated by the Model" can be added. For the evaluation text, a category identifier similar to "Evaluation of the Reply" can be added.
[0075] For the query texts, reply texts, and evaluation texts generated in each iteration, batch information can also be added. Since the query text remains unchanged during the iteration process, batch information can be added only to the reply text and the evaluation text. The batch information corresponds to the iteration number. Therefore, for the reply text, something like "Possible Reply Generated by the Model after the 1st Iteration", "Possible Reply Generated by the Model after the 2nd Iteration", etc. can be added. For the evaluation text, something like "Evaluation of the Reply after the 1st Iteration", "Evaluation of the Reply after the 2nd Iteration", etc. can be added.
[0076] When using the above sample text sequence to train a new model (such as Model 2), the above category hint information and batch information can help Model 2 learn the context information of the text sequence with enhanced context learning ability.
[0077] According to an embodiment of the present disclosure, the present disclosure also provides a processing method for enhancing context learning ability.
[0078] Figure 5 It is a flowchart of a processing method for enhancing context learning ability according to an embodiment of the present disclosure.
[0079] As Figure 5 shown, the processing method 500 for enhancing context learning ability includes operations S510 to S530.
[0080] In operation S510, the text to be queried is input into the trained model to obtain an initial response text.
[0081] In operation S520, the initial response text is evaluated to obtain an initial evaluation text.
[0082] In operation S530, the trained model generates a new response text based on the initial evaluation text and the context information of the text sequence.
[0083] The trained model can be a model trained using a text sequence with enhanced context learning ability, and this model can be a large language model. The text sequence with enhanced context learning ability can be generated by the above-mentioned text sequence generation method for enhancing context learning ability.
[0084] When the text to be queried is input into the trained model, the model can generate an initial response message, and this initial response message may not be the optimal response. Other large models can be used to evaluate the initial response text to obtain an initial evaluation text. Inputting this initial evaluation text into the trained model, the trained model can generate a new response text based on the context information of the sample text sequence. This new response text is better than the initial response text.
[0085] Since the text sequence with enhanced context learning ability contains multiple batches of response texts and evaluation texts, these multiple batches of response texts are tuned after being evaluated based on multiple batches of evaluation texts. Therefore, the text sequence with enhanced context learning ability contains responses tuned based on multiple evaluations, that is, it contains context information for continuously generating better responses. Using this text sequence with enhanced context learning ability to train a new model can enable the new model to learn the context information of the tuned responses. Thus, in the application stage, for the text to be queried, an initial response text is obtained, for the initial response text, an initial evaluation text is obtained, and based on this initial evaluation text, the model can directly generate a better response based on the context information of the tuned response.
[0086] Compared with the related art where continuous training of a model is required to generate better responses, embodiments of the present disclosure generate better responses based on the context information of a text sequence that enhances context learning ability and includes response texts and evaluation texts in multiple batches. This can obtain better response texts more quickly without the need for model training, reducing costs.
[0087] According to embodiments of the present disclosure, the text sequence that enhances context learning ability includes a query text, a response text, and an evaluation text with category hint information and batch information. The above operation S530 includes using a trained model to determine the context information of the text sequence corresponding to the initial evaluation text based on the category hint information and batch information, and generating a new response text according to the context information.
[0088] The category hint information in the text sequence that enhances context learning ability indicates the query text, response text, and evaluation text in the text, and the batch information indicates the context information of the text sequence. After inputting the initial evaluation text into the model, the model can determine the context information corresponding to the initial evaluation text from the text sequence based on the category hint information and batch information, and can quickly generate a new response text with better effects based on the context information corresponding to the initial evaluation text.
[0089] According to embodiments of the present disclosure, the new response text can be evaluated to obtain a new evaluation text; use the trained model to generate a new round of response text based on the new evaluation text and the context information of the text sequence that enhances context learning ability, as the new response text, and return to the step of evaluating the new response text until the target response text is obtained.
[0090] For example, the new response text can be evaluated to obtain a new evaluation text. If the new evaluation text indicates that the new response text does not meet the requirements, the evaluation text of the new response can continue to be input into the trained model, so that the model continues to generate a better response based on the context information of the text sequence that enhances context learning ability. This better response is then used as the new response text, and return to the above step of evaluating the new response text, repeating this process until a response text that meets the requirements is obtained.
[0091] Embodiments of the present disclosure input the evaluation text into the model, enabling the model to generate a new response text based on the context of the text sequence that enhances context learning ability. By continuously evaluating the new response text and inputting the evaluation text into the model, the model continuously generates new response texts based on the context information, and can quickly obtain better response texts without training.
[0092] Embodiments of the present disclosure continuously generate better responses through continuous self - feedback. In this process, the model's capabilities are continuously improved without the need to retrain the parameters, which is fully achieved through in - context learning.
[0093] After obtaining data in a certain domain, embodiments of the present disclosure can directly customize a specialized model through in - context learning.
[0094] According to an embodiment of the present disclosure, the present disclosure also provides a model training method for enhancing in - context learning ability.
[0095] Figure 6 It is a flowchart of a model training method for enhancing in - context learning ability according to an embodiment of the present disclosure.
[0096] As Figure 6 shown, the model training method 600 for enhancing in - context learning ability includes operations S610 - S640.
[0097] In operation S610, obtain a text sequence for enhancing in - context learning ability.
[0098] In operation S620, input the text sequence into the model to be trained to obtain a probability sequence corresponding to the text sequence.
[0099] In operation S630, determine the loss according to the probability sequence.
[0100] In operation S640, adjust the parameters of the model to be trained according to the loss to obtain a trained model.
[0101] The text sequence for enhancing in - context learning ability can be generated according to the above - mentioned text generation method for enhancing in - context learning ability. Input the text sequence into the model to be trained to obtain a probability sequence corresponding to the text sequence. The probabilities in the probability sequence represent the probabilities of generating the text at the corresponding positions in the text sequence. The loss can be calculated based on this probability. For example, determine the loss based on the difference between this probability and the probability (probability is 1) of actually generating the corresponding text. The parameters of the model can be adjusted according to the loss.
[0102] According to an embodiment of the present disclosure, the text sequence for enhancing in - context learning ability includes synthetic texts in multiple batches. Each batch of synthetic texts includes a query text, a response text, and an evaluation text. Operation S630 includes determining the loss according to the probabilities at the positions corresponding to the response text and the evaluation text in the probability sequence.
[0103] Multiple batches of query texts in a text sequence for enhancing context learning ability can be invariant. Therefore, the loss of the query texts can be not calculated, and only the losses of the response texts and the evaluation texts are calculated. For example, the loss is calculated based on the probabilities at the corresponding positions of the response texts and the evaluation texts, and the parameters of the model are adjusted according to this loss.
[0104] According to an embodiment of the present disclosure, a text sequence for enhancing context learning ability includes responses tuned based on multiple evaluations, that is, it includes context information for continuously generating better responses. Using this text sequence for enhancing context learning ability to train a model, a trained model is obtained. This trained model can learn the context information of the text sequence, and thus can quickly generate better responses based on the context information.
[0105] According to an embodiment of the present disclosure, the present disclosure also provides a device for generating a text sequence for enhancing context learning ability, a processing device for enhancing context learning ability, and a model training device for enhancing context learning ability.
[0106] Figure 7 is a block diagram of a device for generating a text sequence for enhancing context learning ability according to an embodiment of the present disclosure.
[0107] As Figure 7 shown, the device 700 for generating a text sequence for enhancing context learning ability includes an iterative module 710 and a splicing module 720. The iterative module 710 includes a training unit and a synthetic text generation unit.
[0108] The iterative module 710 is configured to perform N - time iterative training on an initial model by using query texts to obtain N batches of synthetic texts, where N is an integer greater than 1.
[0109] The splicing module 720 is configured to splice the N batches of synthetic texts to obtain a text sequence for enhancing context learning ability.
[0110] The training unit is configured to, for the n - th training, use the evaluation text of the model obtained after the (n - 1)-th training to perform the n - th training. The evaluation text is obtained by evaluating the response text of the model obtained after the (n - 1)-th training. The response text is a response to the query text, and n is an integer greater than 1 and less than or equal to N.
[0111] The synthetic text generation unit is configured to generate the n - th batch of synthetic texts according to the query text, response text, and evaluation text corresponding to the n - th training.
[0112] The splicing module 720 is used to splice the query text, as well as the response text and evaluation text in different batches of synthesized texts corresponding to the query text, for the same query text, to obtain a text sequence that enhances the context learning ability.
[0113] The splicing module 720 includes a first splicing unit and a second splicing unit.
[0114] The first splicing unit is used to splice the query text, response text, and evaluation text within the synthesized text of each batch to obtain the spliced text of that batch.
[0115] The second splicing unit is used to splice the spliced texts of N batches to obtain a text sequence that enhances the context learning ability.
[0116] The text sequence generation device 700 for enhancing the context learning ability further includes an adding module. The adding module is used to add category hint information and batch information to the query text, response text, and evaluation text in the text sequence for enhancing the context learning ability, respectively.
[0117] The iteration module 710 further includes a response text generation unit and an evaluation text generation unit.
[0118] The response text generation unit is used to input the query text into the model obtained after the nth training for the model obtained after the nth training, to obtain the response text corresponding to the nth training.
[0119] The evaluation text generation unit is used to evaluate the response text corresponding to the nth training to obtain the evaluation text corresponding to the nth training.
[0120] Figure 8 It is a block diagram of a processing device for enhancing the context learning ability according to an embodiment of the present disclosure.
[0121] As Figure 8 shown, the processing device 800 for enhancing the context learning ability includes a first response module 810, an evaluation module 820, and a second response module 830.
[0122] The first response module 810 is used to input the text to be queried into the trained model to obtain an initial response text, where the trained model is trained using the text sequence for enhancing the context learning ability, and the text sequence for enhancing the context learning ability is generated according to the above-mentioned text sequence generation device for enhancing the context learning ability.
[0123] The evaluation module 820 is used to evaluate the initial response text to obtain an initial evaluation text.
[0124] The second response module 830 is used to generate a new response text based on the initial evaluation text and the context information of the text sequence using the trained model.
[0125] The text sequence enhancing context learning ability includes query text, response text, and evaluation text with category hint information and batch information.
[0126] The second response module 830 is used to use the trained model to determine the context information of the text sequence corresponding to the initial evaluation text based on the category hint information and the batch information, and generate a new response text according to the context information.
[0127] The evaluation module 820 is further used to evaluate the new response text to obtain a new evaluation text.
[0128] The second response module 830 is further used to generate a new round of response text based on the new evaluation text and the context information of the text sequence using the trained model.
[0129] The text processing device 800 further includes a return module. The return module is used to return the new round of response text generated by the second response module as the new response text to the evaluation module for evaluation until the second response module generates the target response text.
[0130] Figure 9 It is a block diagram of a model training device for enhancing context learning ability according to an embodiment of the present disclosure.
[0131] As Figure 9 As shown, the model training device 900 includes an acquisition module 910, a processing module 920, a loss determination module 930, and an adjustment module 940.
[0132] The acquisition module 910 is used to acquire a text sequence for enhancing context learning ability, where the text sequence for enhancing context learning ability is generated according to the above-mentioned text sequence generation device for enhancing context learning ability.
[0133] The processing module 920 is used to input the text sequence into the model to be trained to obtain a probability sequence corresponding to the text sequence.
[0134] The loss determination module 930 is used to determine the loss according to the probability sequence.
[0135] The adjustment module 940 is used to adjust the parameters of the model to be trained according to the loss to obtain a trained model.
[0136] The text sequence enhancing context learning ability includes synthetic texts of multiple batches, and each batch of synthetic texts includes query text, response text, and evaluation text.
[0137] The loss determination module 930 is used to determine the loss according to the probabilities at the corresponding positions of the response text and the evaluation text in the probability sequence.
[0138] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0139] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] As Figure 10 shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0141] Multiple components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0142] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as at least one of the text sequence generation method, the processing method, and the model training method for enhancing context learning ability. For example, in some embodiments, at least one of the text sequence generation method, the processing method, and the model training method for enhancing context learning ability can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of at least one of the text sequence generation method, the processing method, and the model training method for enhancing context learning ability described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute at least one of the text sequence generation method, the processing method, and the model training method for enhancing context learning ability by any other suitable means (e.g., by means of firmware).
[0143] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0145] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0147] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0148] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs that run on the respective computers and have a client-server relationship with each other.
[0149] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0150] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a text sequence with enhanced context learning ability, comprising: Performing N - times iterative training on an initial model using query texts to obtain N batches of synthetic texts, where N is an integer greater than 1; and Concatenating the N batches of synthetic texts to obtain a text sequence with enhanced context learning ability; Wherein, the step of performing N - times iterative training on the initial model using query texts to obtain N batches of synthetic texts includes: for the n - th training, Using the evaluation text of the model obtained from the (n - 1)-th training to perform the n - th training, the evaluation text is obtained by evaluating the response text of the model obtained from the (n - 1)-th training, and the response text is a response to the query text, where n is an integer greater than 1 and less than or equal to N; and Generating the n - th batch of synthetic texts according to the query text, response text, and evaluation text corresponding to the n - th training; The method further includes: Training a target model using the text sequence, so that the target model learns the context information of the text sequence, the text sequence contains response texts that have been optimized through multiple evaluations, and the context information is information used to optimize the response text; and For an initial response text to a query text, using the target model to optimize the initial response text according to the context information corresponding to the evaluation text of the initial response text to obtain a new response text for the query text.
2. The method according to claim 1, wherein The step of concatenating the N batches of synthetic texts to obtain a text sequence with enhanced context learning ability includes: For the same query text, concatenating the query text, and the response texts and evaluation texts in different batches of synthetic texts corresponding to the query text to obtain the text sequence with enhanced context learning ability.
3. The method according to claim 1, wherein, The step of concatenating the N batches of synthetic texts to obtain a text sequence with enhanced context learning ability includes: For each batch of synthetic texts, concatenating the query text, response text, and evaluation text within the batch of synthetic texts to obtain a concatenated text for the batch; and Concatenating the concatenated texts of N batches to obtain the text sequence with enhanced context learning ability.
4. The method according to any one of claims 1 to 3, further comprising: Adding category hint information and batch information to the query text, response text, and evaluation text in the text sequence with enhanced context learning ability respectively.
5. The method according to claim 1 further comprises: For the model obtained from the n - th training, Inputting the query text into the model obtained from the n - th training to obtain a response text corresponding to the n - th training; And Evaluating the response text corresponding to the n - th training to obtain an evaluation text corresponding to the n - th training.
6. A processing method for enhancing context learning ability, comprising: Inputting a query text into a trained model to obtain an initial response text, where the trained model is trained using a text sequence with enhanced context learning ability, and the text sequence with enhanced context learning ability is generated according to the method of any one of claims 1 to 5. Evaluate the initial response text to obtain an initial evaluation text; and Use the trained model to generate a new response text based on the initial evaluation text and the context information of the text sequence; wherein the text sequence includes response texts that have been evaluated and optimized multiple times, and the context information is information used to optimize the response text; the generating of the new response text includes: Optimize the initial response text according to the context information corresponding to the initial evaluation text to obtain the new response text.
7. The method according to claim 6, wherein, The text sequence for enhancing context learning ability includes query texts, response texts, and evaluation texts with category hint information and batch information; The using the trained model to generate a new response text based on the initial evaluation text and the context information of the text sequence includes: Use the trained model to determine the context information corresponding to the initial evaluation text based on the category hint information and the batch information, and generate the new response text according to the context information.
8. The method according to claim 6, further comprising: Evaluate the new response text to obtain a new evaluation text; Use the trained model to generate a new round of response text based on the new evaluation text and the context information of the text sequence, as the new response text, and return to the step of evaluating the new response text until a target response text is obtained.
9. A method for training a model with enhanced context learning ability, comprising: Obtain a text sequence with enhanced context learning ability, wherein the text sequence with enhanced context learning ability is generated according to the method described in any one of claims 1 to 5; Input the text sequence into the model to be trained to obtain a probability sequence corresponding to the text sequence; Determine a loss according to the probability sequence; and Adjust the parameters of the model to be trained according to the loss to obtain a trained model, so that the trained model learns the context information of the text sequence, the text sequence includes response texts that have been evaluated and optimized multiple times, and the context information is information used to optimize the response text; The method further comprises: For the initial response text of the text to be queried, use the trained model to optimize the initial response text according to the context information corresponding to the evaluation text of the initial response text to obtain a new response text for the text to be queried.
10. The method according to claim 9, wherein, The text sequence for enhancing context learning ability includes synthetic texts in multiple batches, and each batch of synthetic texts includes query texts, response texts, and evaluation texts; the determining of the loss according to the probability sequence includes: Determine the loss according to the probabilities at the positions corresponding to the response text and the evaluation text in the probability sequence.
11. A device for generating a text sequence with enhanced context learning ability, comprising: An iterative module for iteratively training an initial model N times using query texts to obtain N batches of synthetic texts, where N is an integer greater than 1; and A splicing module for splicing the synthetic texts of the N batches to obtain a text sequence with enhanced context learning ability; Wherein, the iteration module includes: A training unit for, for the nth training, using the evaluation text of the model obtained after the (n - 1)th training to perform the nth training, the evaluation text being obtained by evaluating the response text of the model obtained after the (n - 1)th training, the response text being a response to the query text, and n being an integer greater than 1 and less than or equal to N; and A synthetic text generation unit for generating the synthetic text of the nth batch according to the query text, response text, and evaluation text corresponding to the nth training; The apparatus further includes: A target model training module for training a target model using the text sequence, so that the target model learns the context information of the text sequence, the text sequence includes response texts that have been evaluated and optimized multiple times, and the context information is information for optimizing the response text; and A first optimization module for, for the initial response text of the text to be queried, using the target model to optimize the initial response text according to the context information corresponding to the evaluation text of the initial response text to obtain a new response text for the text to be queried.
12. The apparatus according to claim 11, wherein, The splicing module for, for the same query text, splicing the query text, the response texts and evaluation texts in the synthetic texts of different batches corresponding to the query text to obtain the text sequence with enhanced context learning ability.
13. The device according to claim 11, wherein, The splicing module includes: A first splicing unit for, for the synthetic text of each batch, splicing the query text, response text, and evaluation text in the synthetic text of this batch to obtain the spliced text of this batch; and A second splicing unit for splicing the spliced texts of N batches to obtain the text sequence with enhanced context learning ability.
14. The apparatus according to any one of claims 11 to 13, further includes: An adding module for adding category hint information and batch information to the query text, response text, and evaluation text in the text sequence with enhanced context learning ability respectively.
15. The apparatus according to claim 11, the iteration module further includes: A response text generation unit for, for the model obtained after the nth training, inputting the query text into the model obtained after the nth training to obtain the response text corresponding to the nth training; And An evaluation text generation unit for evaluating the response text corresponding to the nth training to obtain the evaluation text corresponding to the nth training.
16. A processing apparatus with enhanced context learning ability, including: A first response module for inputting the text to be queried into a trained model to obtain an initial response text, wherein the trained model is trained using a text sequence with enhanced context learning ability, and the text sequence with enhanced context learning ability is generated according to the apparatus according to any one of claims 11 to 15; An evaluation module for evaluating the initial response text to obtain an initial evaluation text; and A second response module for using the trained model to generate a new response text based on the initial evaluation text and the context information of the text sequence; Wherein, the text sequence includes response texts that have been evaluated and optimized multiple times, and the context information is information for optimizing the response text; the second response module is configured to optimize the initial response text according to the context information corresponding to the initial evaluation text to obtain the new response text.
17. The apparatus according to claim 16, wherein, The text sequence enhancing context learning ability includes query texts, response texts, and evaluation texts with category hint information and batch information; the second response module is configured to use the trained model to determine the context information corresponding to the initial evaluation text based on the category hint information and batch information, and generate the new response text according to the context information.
18. The apparatus according to claim 16, wherein The evaluation module is further configured to evaluate the new response text to obtain a new evaluation text; The second response module is further configured to use the trained model to generate a new round of response text based on the new evaluation text and the context information of the text sequence; The apparatus further includes: A return module for returning the new round of response text generated by the second response module as the new response text to the evaluation module for evaluation until the second response module generates a target response text.
19. A model training apparatus for enhancing context learning ability, comprising: An acquisition module for acquiring a text sequence for enhancing context learning ability, wherein the text sequence for enhancing context learning ability is generated according to the apparatus according to any one of claims 11 to 15; A processing module for inputting the text sequence into a model to be trained to obtain a probability sequence corresponding to the text sequence; A loss determination module for determining a loss according to the probability sequence; and An adjustment module for adjusting the parameters of the model to be trained according to the loss to obtain a trained model, so that the trained model learns the context information of the text sequence, the text sequence includes response texts that have been evaluated and optimized multiple times, and the context information is information for optimizing the response text; The apparatus further includes: A second optimization module for, for the initial response text of the text to be queried, using the trained model to optimize the initial response text according to the context information corresponding to the evaluation text of the initial response text to obtain a new response text of the text to be queried.
20. The apparatus according to claim 19, wherein, The text sequence enhancing context learning ability includes synthetic texts in multiple batches, and each batch of synthetic texts includes query texts, response texts, and evaluation texts; the loss determination module is configured to determine the loss according to the probabilities at the positions corresponding to the response texts and evaluation texts in the probability sequence.
21. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1 to 10.
23. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1 to 10 when executed by a processor.
Citation Information
Patent Citations
Quality evaluation model training method, multi-round dialogue quality evaluation method and multi-round dialogue quality evaluation device
CN117556005A