Text processing model training method, text generation method and related device
By training the target domain model without labeling and training with the general domain model in combination, the inefficiency and text quality problems of general large language models when migrating to small domains are solved, and more efficient migration and more standardized text generation are achieved.
Patent Information
- Application Number
- CN202410015506.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, when migrating the general large language model to a specific small field, it requires a lot of manpower to annotate text, resulting in inefficient migration, and the generated text may have semantic ambiguity, automatic speech recognition errors and inconvenient sentence breaking.
The domain text processing model of the target domain is trained through the unlabeled domain data, and the general domain generated text is used as a reference, and the general text processing model to be trained in the general domain is jointly trained, and the model parameters are adjusted to improve the standardization and accuracy of text generation.
It reduces the workload of manual labeling, improves model migration efficiency, and standardizes and accuracy of the generated text speech, which can better answer detailed questions in the target field.
Smart Images

Figure CN120258068A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a training method for a text processing model, a text generation method, and related devices. Background Art
[0002] With the development of electronic technology, the use of LLM (Large Language Model) is becoming more and more popular. An LLM refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. The LLM model can be used to process various natural language tasks.
[0003] In practical applications, in order to apply a general LLM model to a specific small domain and generate text that meets expectations in the scenarios of this small domain, it is necessary to first spend a lot of manpower to annotate the original text of the small domain to generate text samples suitable for this small domain, and then use these text samples for model training. Therefore, the process of migrating the capabilities of a general large model to a specific small domain is time-consuming and laborious, and the model migration efficiency is low. Summary of the Invention
[0004] Embodiments of this application provide a training method for a text processing model, a text generation method, and related devices to improve model migration efficiency.
[0005] In a first aspect, embodiments of this application provide a training method for a text processing model, including:
[0006] A domain text processing model to be trained in a target domain performs text generation processing based on the instruction text in the target domain to obtain a target domain generated text, and a general text processing model to be trained in a general domain performs text generation processing based on the instruction text to obtain a general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data; the general text processing model to be trained in the general domain is trained with general data;
[0007] The domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain a first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain a second generated text;
[0008] Determine a model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
[0009] In a second aspect, an embodiment of the present application provides a text generation method, including:
[0010] Obtain a to-be-processed instruction text;
[0011] A general text processing model in the general domain performs text generation processing based on the to-be-processed instruction text to obtain a first response text, and a domain text processing model in the target domain performs text generation processing based on the to-be-processed instruction to obtain a second response text; the general text processing model and the domain text processing model are trained by the training method of the text processing model as described in the first aspect;
[0012] Output a target text in response to the to-be-processed instruction text based on the first response text and the second response text.
[0013] In a third aspect, an embodiment of the present application provides a training device for a text processing model, including:
[0014] A generation unit is configured to enable a domain text processing model to be trained in the target domain to perform text generation processing based on the instruction text in the target domain to obtain a target domain generated text, and enable a general text processing model to be trained in the general domain to perform text generation processing based on the instruction text to obtain a general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain using unlabeled domain data; the general text processing model to be trained in the general domain is trained using general data;
[0015] The generation unit is further configured to enable the domain text processing model to be trained in the target domain to perform text generation processing based on the general domain generated text and the instruction text to obtain a first generated text; and enable the general text processing model to be trained in the general domain to perform text generation processing based on the target domain generated text and the instruction text to obtain a second generated text;
[0016] A training unit is configured to determine a model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
[0017] In a fourth aspect, an embodiment of the present application provides a text generation device, including:
[0018] An obtaining unit is configured to obtain a to-be-processed instruction text;
[0019] A generation unit is configured to perform text generation processing on the to-be-processed instruction text based on a general text processing model in a general domain to obtain a first response text, and perform text generation processing on the to-be-processed instruction based on a domain text processing model in a target domain to obtain a second response text; the general text processing model and the domain text processing model are trained by the training method of the text processing model as described in the first aspect;
[0020] An output unit is configured to output a target text in response to the to-be-processed instruction text based on the first response text and the second response text.
[0021] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a processor; and a memory configured to store computer-executable instructions, where the computer-executable instructions, when executed, cause the processor to execute the training method of the text processing model as described in the first aspect, or the text generation method as described in the second aspect.
[0022] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium for storing computer-executable instructions, where the computer-executable instructions, when executed by a processor, implement the training method of the text processing model as described in the first aspect, or the text generation method as described in the second aspect.
[0023] It can be seen that in the embodiments of the present application, first, the domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain using unlabeled domain data; the general text processing model to be trained in the general domain is trained using general data; then, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text; finally, the model training loss is determined according to the first generated text and the second generated text, and the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss. Thus, considering that the text generated by the domain text processing model to be trained in the target domain obtained by training with unlabeled domain data may have negative impacts brought by noise data such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation, etc., in the embodiments of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the general domain generated text generated by the general text processing model to be trained in the general domain as a reference during the process of performing text generation processing to obtain the first generated text, and utilize the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, which can enable the domain text processing model to be trained in the target domain to improve the speech normativity of the generated text and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation after model training.
[0024] Considering that the general text processing model to be trained in the general domain can perform question answering and assistance on general social knowledge, but in the case of being refined to the target domain, it is difficult for this general text processing model to answer detailed questions in the target domain. In the embodiments of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the target domain generated text generated by the domain text processing model to be trained in the target domain as a reference during the process of performing text generation processing to obtain the second generated text, and utilize the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, which can enable the general text processing model to be trained in the general domain to be able to answer detailed questions in this target domain after model training.
[0025] By determining the model training loss based on the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be achieved, enabling the texts generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to tend to be consistent after model training, which is beneficial for the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to learn from each other's advantages.
[0026] In addition, in the process of transferring the capabilities of the general text processing model to be trained in the general domain to the target domain in the embodiments of the present application, the training samples input to the domain text processing model to be trained in the target domain include instruction texts and general domain generated texts, and the general domain generated texts are obtained by the general text processing model to be trained in the general domain through text generation processing based on the instruction texts; in the process of transferring the capabilities of the domain text processing model to be trained in the target domain to the general domain in the embodiments of the present application, the training samples input to the general text processing model to be trained include instruction texts and target domain generated texts, and the target domain generated texts are obtained by the domain text processing model in the target domain through text generation processing based on the instruction texts. Therefore, in the process of model transfer, the training samples of the domain text processing model to be trained in the target domain and the general text processing model to be trained in the general domain do not rely on manual annotation, reducing the manual workload and improving the training sample generation efficiency, thereby improving the overall model transfer efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;
[0028] Figure 1 It is a processing flow chart of a method for training a text processing model provided by an embodiment of the present application;
[0029] Figure 2 It is a processing flow chart of another method for training a text processing model provided by an embodiment of the present application;
[0030] Figure 3 It is a data flow diagram of a method for training a text processing model provided by an embodiment of the present application;
[0031] Figure 4 It is a processing flowchart of a text generation method provided by an embodiment of this application;
[0032] Figure 5 It is a schematic diagram of a training device for a text processing model provided by an embodiment of this application;
[0033] Figure 6 It is a schematic diagram of a text generation device provided by an embodiment of this application;
[0034] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of this application. Detailed implementation manners
[0035] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0036] An embodiment of a training method for a text processing model provided by this specification:
[0037] Text generation refers to using natural language processing technology to automatically generate text content that meets grammatical and semantic requirements through learning and understanding a large amount of text data and mastering language rules. Text generation is an important application field of natural language processing technology.
[0038] Some users have the need for text generation in a specific small domain. To meet this need, it is usually necessary to train a general LLM model using domain data to transfer the capabilities of the general large model to the specific small domain. The training of the general large model is mainly divided into three steps: SFT (Supervised Fine-Tuning) fine-tuning training, RM (Reward Model) model training, and reinforcement learning training based on the RM model. The preliminary fine-tuning enables the model to have basic answering knowledge, and the training samples used in the SFT fine-tuning training are unlabeled samples. After the SFT fine-tuning training, the obtained model is used to generate a variety of different results, and the generated results are manually labeled and optimized in large quantities, and then the RM model is trained, and then the trained RM model is used for automated reinforcement training.
[0039] In the process of migrating the capabilities of a general large model to a specific small domain, if the training samples are manually annotated, it is time-consuming and laborious, and the manually annotated samples are only applicable to a specific small domain. When migrating the general LLM model to other domains, re-annotation is required, so the model migration efficiency is low; if only SFT fine-tuning training is performed without any manual annotation or optimization of the training samples, the summary ability of the trained model is not strong, and there is a lot of noise in the training samples, which may lead to poor text generation effects of the trained model and it is difficult to meet the user's quality requirements for the text.
[0040] Therefore, to solve the above problems, an embodiment of the present application provides a training method for a text processing model. This training method can be executed by an electronic device, which can be a terminal device or a server. Among them, the terminal device can include a laptop computer and an intelligent interaction device, and the server can include an independent physical server, a server cluster, or a cloud server capable of cloud computing.
[0041] The training method for the text processing model provided by the embodiment of the present application is as follows. First, the domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data; the general text processing model to be trained in the general domain is obtained by training with general data; then, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text; finally, the model training loss is determined according to the first generated text and the second generated text, and the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss.
[0042] Therefore, considering that the text generated by the domain text processing model to be trained in the target domain, which is trained with data from unlabeled domains, may have negative impacts caused by noisy data, such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation. In the embodiments of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the general domain generated text generated by the general text processing model to be trained in the general domain as a reference during the process of text generation to obtain the first generated text. By using the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, the domain text processing model to be trained in the target domain can improve the speech normativity of the generated text after model training, and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation. Considering that the general text processing model to be trained in the general domain can answer questions and provide assistance on general social knowledge, but in the case of being refined to the target domain, it is difficult for the general text processing model to answer detailed questions in the target domain. In the embodiments of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the target domain generated text generated by the domain text processing model to be trained in the target domain as a reference during the process of text generation to obtain the second generated text. By using the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, the general text processing model to be trained in the general domain can answer detailed questions in the target domain after model training.
[0043] By determining the model training loss based on the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be achieved. This enables the text generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to tend to be consistent after model training, which is beneficial for the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to learn from each other's advantages.
[0044] In addition, during the process of transferring the capabilities of the general text processing model to be trained in the general domain to the target domain in the embodiments of the present application, the training samples input to the domain text processing model to be trained in the target domain include instruction texts and general domain generated texts, where the general domain generated texts are obtained by the general text processing model to be trained in the general domain through text generation processing based on the instruction texts; during the process of transferring the capabilities of the domain text processing model to be trained in the target domain to the general domain in the embodiments of the present application, the training samples input to the general text processing model to be trained include instruction texts and target domain generated texts, where the target domain generated texts are obtained by the domain text processing model in the target domain through text generation processing based on the instruction texts. Therefore, during the model transfer process, the training samples of the domain text processing model to be trained in the target domain and the general text processing model to be trained in the general domain do not rely on manual annotation, reducing the manual workload and improving the training sample generation efficiency, thereby improving the overall model transfer efficiency.
[0045] Figure 1 FIG. is a processing flow chart of a method for training a text processing model provided by an embodiment of the present application. Referring to Figure 1 , the method for training the text processing model provided in this embodiment specifically includes steps S102 to S106.
[0046] Step S102, the domain text processing model to be trained in the target domain performs text generation processing based on the instruction texts in the target domain to obtain target domain generated texts, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction texts to obtain general domain generated texts; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain using unlabeled domain data; the general text processing model to be trained in the general domain is obtained by training using general data.
[0047] The general domain is a general domain that involves all walks of life and has a relatively large scope. The target domain can be understood as a specific domain, different from the general domain, or the target domain can also be interpreted as a private domain. The private domain and the general domain are relative concepts. The private domain can focus on a specified industry and has a relatively small scope compared to the general domain.
[0048] The instruction text can be used to describe the requirements for the text to be generated or to ask questions. The instruction text in the target domain can be used to describe the requirements for the text to be generated in the target domain or to ask questions regarding the target domain.
[0049] For example, the instruction text can require generating text in a specific scenario in the target domain, or the instruction text can pose a professional question within the target domain.
[0050] The domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the target domain generated text, which can be the text to be generated in response to the instruction text generated by the domain text processing model to be trained in the target domain.
[0051] The general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the general domain generated text, which can be the text to be generated in response to the instruction text generated by the general text processing model to be trained in the general domain.
[0052] When the instruction text is used to describe the requirements for the text to be generated, the text to be generated in response to the instruction text can be the text that meets the requirements.
[0053] For example, Instruction Text 1 is: Please generate a conversation between an agent and a customer in an information recommendation scenario.
[0054] The target domain generated text is Conversation 1 that meets the requirement of "a conversation between an agent and a customer in an information recommendation scenario", and the general domain generated text is Conversation 2 that meets the requirement of "a conversation between an agent and a customer in an information recommendation scenario".
[0055] When the instruction text is used for questioning, the text to be generated in response to the instruction text can be the text that answers the question.
[0056] For example, Instruction Text 2 is: When does the discount on Product A end?
[0057] The target domain generated text is Reply Text 1 that answers the question "When does the discount on Product A end", and the general domain generated text is Reply Text 2 that answers the question "When does the discount on Product A end".
[0058] The LLM model includes a general large model and a domain large model. The general large model is a large model pre-trained using a large amount of general data and has cross-task generality and cross-domain generality. The domain large model is a model obtained by training and optimizing the general large model in a specific domain or industry.
[0059] Exemplarily, the general large model can be open-source models such as chatglm and bloom.
[0060] Compared with the general large model, the domain large model has the following advantages:
[0061] (a1) Domain specialization: The domain large model is specially trained and can better understand and process knowledge, terms, and contexts in a specific domain.
[0062] (a2) High-quality output: Due to optimization in a specific domain, the output quality of domain large models in that specific domain is usually higher than that of general large models.
[0063] (a3) Better performance in specific tasks: For tasks in a specific domain, domain large models usually perform better than general large models.
[0064] Compared with general large models, domain large models have the following disadvantages;
[0065] (b1) Data requirements and training costs: Domain large models require a large amount of domain-specific data for training.
[0066] (b2) Domain large models have strong adaptability in a specific domain and relatively weak performance in other domains outside that specific domain.
[0067] (b3) Due to the frequent changes in domain-specific knowledge and requirements, domain large models need to be updated and maintained regularly to keep up with new developments.
[0068] In the embodiments of this specification, the domain text processing model to be trained in the target domain can be a domain large model, which can be used for text generation. The general text processing model to be trained in the general domain can be a general large model, which can be used for text generation.
[0069] The domain text processing model to be trained in the target domain can be obtained by training the general text processing model to be trained in the general domain with unlabeled domain data in the target domain.
[0070] Unlabeled domain data refers to domain data that has not been manually labeled and cleaned. Domain data can be dialogue text, Q&A text, etc.
[0071] Exemplarily, the general text processing model to be trained in the general domain can be a Bloom 7B model. The domain text processing model to be trained in the target domain can be a model obtained by training the Bloom 7B model with unlabeled and uncleaned domain data in the target domain. Since there are many noises in the natural domain data that has not been manually labeled and cleaned, such as off-topic answers, obvious filler words, interrupted conversations, etc. Therefore, although the trained domain text processing model to be trained in the target domain has certain domain knowledge, it will carry the characteristics of noisy data during the Q&A process.
[0072] The general text processing model to be trained in the general domain is obtained by training with general data.
[0073] General data may include large-scale and diverse labeled data sets that cover knowledge in various fields.
[0074] The general text processing model to be trained in the general domain has a vast amount of general knowledge and good summarization ability. The text generated by the general text processing model to be trained in the general domain has no semantic ambiguity, is fluent and natural, but cannot answer detailed questions in the target domain.
[0075] The domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the target domain generated text. Specifically, in implementation, the instruction text can be input into the domain text processing model to be trained in the target domain for text generation processing to obtain the target domain generated text that responds to the instruction text.
[0076] In the case where the instruction text is used to describe the requirements for the text to be generated, the instruction text is input into the domain text processing model to be trained in the target domain for text generation processing to obtain the target domain generated text that meets the requirements.
[0077] In the case where the instruction text is used for questioning, the instruction text is input into the domain text processing model to be trained in the target domain for text generation processing to obtain the target domain generated text that answers the question.
[0078] The general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the general domain generated text. Specifically, in implementation, the instruction text can be input into the general text processing model to be trained in the general domain for text generation processing to obtain the general domain generated text that responds to the instruction text.
[0079] In the case where the instruction text is used to describe the requirements for the text to be generated, the instruction text is input into the general text processing model to be trained in the general domain for text generation processing to obtain the general domain generated text that meets the requirements.
[0080] In the case where the instruction text is used for questioning, the instruction text is input into the general text processing model to be trained in the general domain for text generation processing to obtain the general domain generated text that answers the question.
[0081] Serial numbers such as "first" and "second" in this specification are only used to distinguish similar features and have no actual meaning, and will not be elaborated further below.
[0082] In a specific implementation method, the instruction text includes background information of the target domain and text generation instructions; the domain text processing model to be trained in the target domain includes a keyword extraction module and a text generation module connected in series; the domain text processing model to be trained in the target domain performs text generation processing based on the instruction text of the target domain to obtain the target domain generated text, including: the keyword extraction module performs keyword extraction processing on the background information to obtain keyword information in the background information; the text generation module performs text prediction processing according to the text generation instructions under the prompt of the keyword information to obtain the target domain generated text.
[0083] Text generation instructions can be used to describe requirements for the text to be generated, or to ask questions.
[0084] The background information of the target domain can be a supplementary information for the text generation instruction, and is used for the model to refer to how to generate text in response to the text generation instruction.
[0085] For example, instruction text 1 is: Please refer to the following information to generate a conversation between an agent and a customer in an information recommendation scenario: "background information 1". The text generation instruction is "Please refer to the following information to generate a conversation between an agent and a customer in an information recommendation scenario", and the background information is "background information 1".
[0086] The domain text processing model to be trained in the target domain includes a keyword extraction module and a text generation module connected in series, and the output of the keyword extraction model is the input of the text generation module.
[0087] The keyword extraction module performs keyword extraction processing on the background information to obtain keyword information in the background information. In specific implementation, the background information included in the instruction text can be input into the keyword extraction module in the domain text processing model to be trained in the target domain to perform keyword extraction processing to obtain keyword information in the background information.
[0088] The text generation module performs text prediction processing according to the text generation instruction under the prompt of the keyword information to obtain the target domain generated text. In specific implementation, the keyword information and the text generation instruction included in the instruction text can be input into the text generation module in the domain text processing model to be trained in the target domain, and the text prediction processing is performed under the prompt of the keyword information to obtain the target domain generated text that responds to the text generation instruction.
[0089] The target domain generated text may or may not contain keyword information.
[0090] In the case that the target domain generated text does not include keyword information, the target domain generated text may include mapping words, and the mapping words may be determined based on the keyword information and the pre-configured correspondence between keywords and words.
[0091] Considering that the text generated in the embodiments of this specification for the target domain can be a domain large model obtained by training a general large model using unannotated and uncleaned domain data in the target domain, the text generated in the target domain using the text generated in the target domain may have problems such as excessive filler words, unsmooth speech, and ASR (Automatic Speech Recognition) translation errors.
[0092] For example, the instruction text 1 includes:
[0093] Please imitate a telesales agent to make a call in combination with the following information. Background information: #Ms. A# Initial credit limit a1# Remaining credit limit a2# Linked institution A. Please start making calls round by round.
[0094] Among them, the background information includes: "Background information: #Ms. A# Initial credit limit a1# Remaining credit limit a2# Linked institution A", and the text generation instructions include: "Please imitate a telesales agent to make a call in combination with the following information" and "Please start making calls round by round".
[0095] Input the background information included in the instruction text into the keyword extraction module for keyword extraction processing, and the keyword information obtained from the background information includes: Initial credit limit a1, Remaining credit limit a2.
[0096] Input the keyword information and the text generation instructions included in the instruction text 1 into the text generation module, and perform text prediction processing under the prompt of the keyword information to obtain the target domain generated text 1 in response to the text generation instructions. The target domain generated text 1 is as follows:
[0097] Agent: Hello, may I ask if you are Ms. Wu?
[0098] Customer: Yes, this is me.
[0099] Agent: Hello, Ms. Wu. This is the account manager of XX institution.
[0100] Customer: Hmm. What's the matter?
[0101] Agent: We have given you a preliminary credit limit increase here. Then we also invite you to withdraw the remaining credit limit in your account according to your needs. After you withdraw successfully and maintain a good credit record, the company will give you an additional opportunity for an activity to increase the credit limit and reduce the interest rate. You can check it later.
[0102] Due to a similar concept, the general text processing model to be trained in the general domain can include a keyword extraction module and a text generation module connected in series in sequence, and the output of the keyword extraction model is the input of the text generation module.
[0103] The keyword extraction module in the general text processing model performs keyword extraction processing on the background information to obtain keyword information in the background information. In specific implementation, the background information included in the instruction text can be input into the keyword extraction module in the general text processing model for keyword extraction processing to obtain keyword information in the background information.
[0104] The text generation module in the general text processing model performs text prediction processing according to the text generation instruction under the prompt of the keyword information to obtain the general domain generated text. In specific implementation, the keyword information and the text generation instruction included in the instruction text can be input into the text generation module in the general text processing model, and the text prediction processing is performed under the prompt of the keyword information to obtain the general domain generated text that responds to the text generation instruction.
[0105] The general domain generated text may contain keyword information or may not contain keyword information.
[0106] In the case that the general domain generated text does not include keyword information, the general domain generated text may include mapping words, and the mapping words may be determined based on the keyword information and the pre-configured correspondence between keywords and words.
[0107] Considering that the general text processing model to be trained in the general field in the embodiments of this specification can be a general large model, and the general large model has insufficient understanding of the details of specific fields, the general domain generated text generated by the general text processing model to be trained in the general field may not be suitable for answering questions in the target field.
[0108] Exemplarily, the target domain is the telemarketing domain, and the general domain generated text is the question and answer text obtained by the general text processing model to be trained in the general domain performing text prediction processing under the prompt of keyword information. Since the general text processing model to be trained in the general domain does not have detailed knowledge of the target domain, the question and answer text is far from the actual marketing process in the telemarketing field.
[0109] For example, instruction text 1 includes:
[0110] Please imitate a telemarketing agent and make a call based on the following information. Background information: #Ms. A# Initial credit limit a1# Remaining credit limit a2# Binding institution A. Please start the call round by round.
[0111] Among them, the background information includes: "Background information: #Ms. A#Initial credit limit a1#Remaining credit limit a2#Bound institution A", and the text generation instructions include: "Please imitate a telemarketing agent to make a call based on the following information" and "Please start a round of calls."
[0112] Input the background information included in the instruction text into the keyword extraction module in the second language model for keyword extraction processing, and the obtained keyword information in the background information includes: initial amount a1, remaining amount a2.
[0113] Input the keyword information and the text generation instruction included in the instruction text 1 into the text generation module of the second language model, and perform text prediction processing under the prompt of the keyword information to obtain the general domain generated text 1 in response to the text generation instruction. The general domain generated text 1 is as follows:
[0114] Agent: Hello, I am a telemarketing agent from XX institution. Thank you for taking the precious time to contact us. First of all, I would like to confirm some information about your resource lending. Your initial amount is a1 and the remaining amount is a2, right? (Waiting)
[0115] Customer: Yes.
[0116] Agent: Okay, thank you very much. Now I would like to know if you have any questions or need any help with our service recently?
[0117] Step S104, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text.
[0118] The domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text. Specifically, in implementation, a first training sample can be constructed according to the general domain generated text and the instruction text, and the first training sample is input into the domain text processing model to be trained in the target domain for text generation processing to obtain the first generated text.
[0119] The general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text. Specifically, in implementation, a second training sample can be constructed according to the target domain generated text and the instruction text, and the second training sample is input into the general text processing model to be trained in the general domain for text generation processing to obtain the second generated text.
[0120] It should be noted that although during the execution of step S102, the instruction text in the target domain is input into the domain text processing model to be trained in the target domain for text generation processing to obtain the generated text in the target domain, this instruction text is not used as an independent training sample, and after the domain text processing model to be trained in the target domain generates the generated text in the target domain, the model parameters of the domain text processing model to be trained in the target domain are not adjusted. Similarly, although during the execution of step S102, the instruction text is input into the general text processing model to be trained in the general domain for text generation processing to obtain the generated text in the general domain, this instruction text is not used as an independent training sample, and after the general text processing model to be trained in the general domain generates the generated text in the general domain, the model parameters of the general text processing model to be trained in the general domain are not adjusted.
[0121] The purpose of obtaining the generated text in the general domain and the generated text in the target domain by executing step S102 is to construct the first training sample required for model training of the domain text processing model to be trained in the target domain, and to construct the second training sample required for model training of the general text processing model to be trained in the general domain.
[0122] To construct the first training sample based on the generated text in the general domain and the instruction text, it can be that when the number of instruction texts is multiple, for each instruction text, a first training sample of the domain text processing model to be trained in the target domain is constructed according to the instruction text and the generated text in the general domain that responds to the instruction text. Among them, the generated text in the general domain that responds to the instruction text can be the text obtained by the general text processing model to be trained in the general domain through text generation processing based on the instruction text.
[0123] To construct the second training sample based on the generated text in the target domain and the instruction text, it can be that when the number of instruction texts is multiple, for each instruction text, a second training sample of the general text processing model to be trained in the general domain is constructed according to the instruction text and the generated text in the target domain that responds to the instruction text. Among them, the generated text in the target domain that responds to the instruction text can be the text obtained by the domain text processing model to be trained in the target domain through text generation processing based on the instruction text.
[0124] By using the generated text in the general domain generated by the general text processing model to be trained in the general domain to construct the first training sample required for model training of the domain text processing model to be trained in the target domain, the knowledge and summarization ability of the general large model trained with a large amount of high-quality labeled data can be used to guide and optimize the training of the domain large model, so that while the domain large model has domain knowledge, it can learn the regular and semantically unambiguous expression methods of the general large model, and obtain a domain large model with better-quality speech.
[0125] By constructing a second training sample required for model training of a general text processing model to be trained in the general domain by using the target domain generated text generated by a domain text processing model to be trained in the target domain, the domain Q&A ability of a domain large model trained with domain data can be used to guide the optimization of the training of the general large model, so that while the general large model has general knowledge, it can learn the expression ways of domain jargons and obtain a general large model with domain knowledge.
[0126] During the process of the domain text processing model to be trained in the target domain performing text generation processing on the instruction text, the general domain generated text can be used as a reference.
[0127] The domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain a first generated text, which can be to input the instruction text and the general domain generated text into the domain text processing model to be trained in the target domain, and perform text prediction processing under the prompt of the general domain generated text to obtain the first generated text in response to the instruction text.
[0128] The instruction text may include background information of the target domain and a text generation instruction. The domain text processing model to be trained in the target domain may include a keyword extraction module and a text generation module connected in series in sequence, and the output of the keyword extraction model is the input of the text generation module.
[0129] The keyword extraction module performs keyword extraction processing on the background information to obtain keyword information in the background information; the text generation module performs text prediction processing according to the text generation instruction under the prompt of the keyword information and the general domain generated text to obtain the first generated text.
[0130] In specific implementation, the background information included in the instruction text can be input into the keyword extraction module in the domain text processing model to be trained in the target domain for keyword extraction processing to obtain keyword information in the background information; the keyword information, the text generation instruction included in the instruction text, and the general domain generated text are input into the text generation module in the domain text processing model to be trained in the target domain, and text prediction processing is performed under the prompt of the keyword information and the general domain generated text to obtain the first generated text in response to the text generation instruction.
[0131] Considering that the domain text processing model to be trained in the target domain can be a domain large model, the general domain generated text generated by the general text processing model to be trained in the general domain can be used as reference information for the domain large model to learn the process words and expressions of the general large model. The domain large model can summarize and answer by combining the background information and the smooth words and expressions organized by the general large model. Furthermore, using the general large model to guide the training of the domain large model can improve the normativity of the words and expressions of the domain large model and reduce the influence of noise data such as semantic ambiguity, ASR errors, and unsmooth sentence breaks.
[0132] For example, instruction text 1 includes:
[0133] Please imitate a telemarketing agent to make a call in combination with the following information. Background information: #Ms. A#Initial credit limit a1#Remaining credit limit a2#Linked institution A. Please start a round of calls.
[0134] Among them, the background information includes: "Background information: #Ms. A#Initial credit limit a1#Remaining credit limit a2#Linked institution A", and the text generation instructions include: "Please imitate a telemarketing agent to make a call in combination with the following information" and "Please start a round of calls".
[0135] Input the background information included in the instruction text into the keyword extraction module in the domain text processing model to be trained in the target domain for keyword extraction processing, and the keyword information obtained from the background information includes: initial credit limit a1, remaining credit limit a2.
[0136] Input the keyword information, the text generation instructions included in instruction text 1, and the general domain generated text into the text generation module of the domain text processing model to be trained in the target domain, and perform text prediction processing under the prompt of the keyword information and the general domain generated text to obtain the first generated text 1 in response to the text generation instructions. The first generated text 1 is as follows:
[0137] Agent: Hello, I'm a telemarketing agent from XX institution. Thank you for taking the time to answer our call. Are you Ms. Wu?
[0138] Customer: Yes.
[0139] Agent: Okay, thank you very much. We have given you a preliminary credit limit increase. If you need it, you are welcome to withdraw cash as needed.
[0140] The first generated text and the target-domain generated text are both texts in response to an instruction text, and both the first generated text and the target-domain generated text are generated by a domain text processing model to be trained in the target domain. The difference between the first generated text and the target-domain generated text is that the domain text processing model to be trained in the target domain does not refer to the general-domain generated text when generating the target-domain generated text, while the domain text processing model to be trained in the target domain refers to the general-domain generated text when generating the first generated text.
[0141] During the process of text generation processing of an instruction text by a general text processing model to be trained in the general domain, the target-domain generated text can be used as a reference.
[0142] The general text processing model to be trained in the general domain performs text generation processing based on the target-domain generated text and the instruction text to obtain a second generated text. It can be to input the instruction text and the target-domain generated text into the general text processing model to be trained in the general domain, and perform text prediction processing under the prompt of the target-domain generated text to obtain the second generated text in response to the instruction text.
[0143] The instruction text can include background information of the target domain and a text generation instruction. The general text processing model to be trained in the general domain can include a keyword extraction module and a text generation module connected in series in sequence. The output of the keyword extraction model is the input of the text generation module.
[0144] The keyword extraction module performs keyword extraction processing on the background information to obtain keyword information in the background information; the text generation module performs text prediction processing according to the text generation instruction under the prompt of the keyword information and the target-domain generated text to obtain the second generated text.
[0145] In specific implementation, the background information included in the instruction text can be input into the keyword extraction module in the general text processing model to be trained in the general domain for keyword extraction processing to obtain keyword information in the background information; the keyword information, the text generation instruction included in the instruction text, and the target-domain generated text are input into the text generation module in the general text processing model to be trained in the general domain, and text prediction processing is performed under the prompt of the keyword information and the target-domain generated text to obtain the second generated text in response to the text generation instruction.
[0146] Considering that the general text processing model to be trained in the general domain can be a general large model, the target domain generated text generated by the domain text processing model to be trained in the target domain can be used as reference information for the general large model to learn the professional knowledge of the target domain, and the general large model can reorganize and output in combination with background knowledge and the speech of the domain large model. Furthermore, using the domain large model to guide the training of the general large model is conducive to the general large model with standard dialogue speech to quickly learn the ability of domain answering and the ability of the proprietary process of the target domain.
[0147] For example, Instruction Text 1 includes:
[0148] Please imitate a telemarketing agent to make a call in combination with the following information. Background information: #Ms. A#Initial credit limit a1#Remaining credit limit a2#Linked institution A. Please start a round of calls.
[0149] Among them, the background information includes: "Background information: #Ms. A#Initial credit limit a1#Remaining credit limit a2#Linked institution A", and the text generation instructions include: "Please imitate a telemarketing agent to make a call in combination with the following information" and "Please start a round of calls".
[0150] Input the background information included in the instruction text into the keyword extraction module in the general text processing model to be trained for keyword extraction processing, and the keyword information obtained from the background information includes: Initial credit limit a1, Remaining credit limit a2.
[0151] Input the keyword information, the text generation instructions included in Instruction Text 1, and the target domain generated text into the text generation module of the general text processing model to be trained, and perform text prediction processing under the prompt of the keyword information and the target domain generated text to obtain the second generated text 1 in response to the text generation instructions. The second generated text 1 is as follows:
[0152] Agent: Hello, may I ask if you are Ms. Wu?
[0153] Customer: Yes, this is me.
[0154] Agent: Hello, Ms. Wu. I'm the account manager of XX institution.
[0155] Customer: Hmm. What's the matter?
[0156] Agent: We have given you a preliminary credit limit increase here. Then we also invite you to withdraw the remaining credit limit in your account according to your needs. After you withdraw successfully and maintain a good credit record, the company will give you an additional opportunity for a credit limit increase and interest rate reduction activity. You can check it later.
[0157] The second generated text and the general domain generated text are both texts in response to instruction texts, and both the second generated text and the general domain generated text are generated by a general text processing model to be trained in the general domain. The difference between the second generated text and the general domain generated text is that the general text processing model to be trained in the general domain does not refer to the target domain generated text when generating the general domain generated text, and the general text processing model to be trained in the general domain refers to the target domain generated text when generating the second generated text.
[0158] Step S106, determine the model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
[0159] Adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss may be to adjust the model parameters of the general text processing model to be trained in the general domain according to the model training loss, and, adjust the model parameters of the domain text processing model to be trained in the target domain according to the model training loss until the training end condition is met. The training end condition may be that the value of the model training loss no longer decreases, or the value of the model training loss is less than or equal to a preset numerical threshold, or the number of training rounds of the model reaches a preset quantity threshold, and so on.
[0160] By generating the model training loss according to the first generated text and the second generated text and adjusting the model parameters of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain during the process of jointly training the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain, the first generated text and the second generated text can complement each other and simultaneously possess the advantages of the general large model and the domain large model. Furthermore, two models with different focus directions are obtained through joint training. The trained general text processing model can process both the text instructions in the target domain and the text instructions in the general domain, but the processing ability of the text instructions in the general domain will be relatively stronger; the trained domain text processing model can process both the text instructions in the target domain and the text instructions in the general domain, but the processing ability of the text instructions in the target domain will be relatively stronger.
[0161] In addition, in the training method of the text processing model provided in the embodiments of this specification, using the knowledge and summarization ability of the general large model to guide the training of the domain large model can reduce the cost of manually cleaning data and manually annotating samples, improve the iterative optimization speed, reduce the utilization of the pre-trained model resources again, and can also quickly transfer the large model capabilities to various small domain scenarios, enabling each small domain to quickly possess the large model assistance ability.
[0162] In specific implementation, in the first joint training, first, the domain text processing model to be trained in the target domain performs text generation processing based on the instruction text 1 in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text 1 to obtain the general domain generated text; then, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text 1 to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text 1 to obtain the second generated text; finally, the model training loss is determined according to the first generated text and the second generated text. If it is determined that the training end condition is satisfied, the model training is ended, and the trained domain text processing model in the target domain and the trained general text processing model in the general domain are obtained. If it is determined that the training end condition is not satisfied, the model parameters of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss, and the second joint training is continued.
[0163] In the second joint training, first, the domain text processing model to be trained in the target domain before the first model parameter adjustment performs text generation processing based on the instruction text 2 in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain before the first model parameter adjustment performs text generation processing based on the instruction text 2 to obtain the general domain generated text; then, the domain text processing model to be trained in the target domain after 1 model parameter adjustment performs text generation processing based on the general domain generated text and the instruction text 2 to obtain the first generated text; and the general text processing model to be trained in the general domain after 1 model parameter adjustment performs text generation processing based on the target domain generated text and the instruction text 2 to obtain the second generated text; finally, the model training loss is determined according to the first generated text and the second generated text. If it is determined that the training end condition is satisfied, the model training is ended, and the trained domain text processing model in the target domain and the trained general text processing model in the general domain are obtained. If it is determined that the training end condition is not satisfied, the model parameters of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss, and the third joint training is continued.
[0164] In the Nth joint training, first, the domain text processing model to be trained in the target domain before the first model parameter adjustment performs text generation processing based on the instruction text N in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain before the first model parameter adjustment performs text generation processing based on the instruction text N to obtain the general domain generated text; then, the domain text processing model to be trained in the target domain after (N - 1) model parameter adjustments performs text generation processing based on the general domain generated text and the instruction text N to obtain the first generated text; and the general text processing model to be trained in the general domain after (N - 1) model parameter adjustments performs text generation processing based on the target domain generated text and the instruction text N to obtain the second generated text; finally, the model training loss is determined according to the first generated text and the second generated text. If it is determined that the training end condition is satisfied, the model training is ended, and the domain text processing model that has been trained in the target domain and the general text processing model that has been trained in the general domain are obtained. If it is determined that the training end condition is not satisfied, the model parameters of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss, and the (N + 1)th joint training is continued. N is a natural number greater than 1.
[0165] In the example of the above Nth joint training, in the first joint training, neither the domain text processing model to be trained in the target domain nor the general text processing model to be trained in the general domain has undergone model parameter adjustment. Therefore, only two models are required to implement this joint training. Thus, the domain text processing model to be trained in the target domain can be used to generate both the target domain generated text and the first generated text; the general text processing model to be trained in the general domain can be used to generate both the general domain generated text and the second generated text. If the training end condition is satisfied after the first joint training and the model training is ended, only two models are required throughout the joint training process.
[0166] In the example of the above Nth joint training, in the joint training in other rounds after the first joint training, the domain text processing model to be trained in the target domain before the first model parameter adjustment is used to generate the target domain generated text, and the domain text processing model to be trained in the target domain after one or more model parameter adjustments is used to generate the first generated text; the general text processing model to be trained in the general domain before the first model parameter adjustment is used to generate the general domain generated text, and the general text processing model to be trained in the general domain after one or more model parameter adjustments is used to generate the second generated text. In this case, four models are required throughout the joint training process.
[0167] In order to implement the above example of N - time joint training, where in the same joint training, both the domain text processing model to be trained in the target domain before the first model parameter adjustment and the domain text processing model to be trained in the target domain after one or more model parameter adjustments are used, and also in the same joint training, both the general text processing model to be trained in the general domain before the first model parameter adjustment and the general text processing model to be trained in the general domain after one or more model parameter adjustments are used, the following method can be adopted: Before the first joint training, perform a replication process on the domain text processing model to be trained in the target domain, and perform a replication process on the general text processing model to be trained in the general domain in advance.
[0168] For example, perform a replication process on the domain text processing model to be trained in the target domain to obtain two domain text processing models to be trained in the target domain. For ease of distinction, one can be called the domain large model_alter, and the other can be called the domain large model_frozen. Perform a parameter freezing process on the domain large model_frozen. Then, use the domain large model_frozen when executing step S102, and this domain large model_frozen is the domain text processing model to be trained in the target domain before the first model parameter adjustment; use the domain large model_alter when executing step S104, and this domain large model_alter is the domain text processing model to be trained in the target domain after one or more model parameter adjustments.
[0169] Perform a replication process on the general text processing model to be trained in the general domain to obtain two general text processing models to be trained in the general domain. For ease of distinction, one can be called the general large model_alter, and the other can be called the general large model_frozen. Perform a parameter freezing process on the general large model_frozen. Then, use the general large model_frozen when executing step S102, and this general large model_frozen is the general text processing model to be trained in the general domain before the first model parameter adjustment; use the general large model_alter when executing step S104, and this general large model_alter is the general text processing model to be trained in the general domain after one or more model parameter adjustments.
[0170] The output of the general large model_frozen can be the input of the domain large model_alter. The output of the domain large model_frozen can be the input of the general large model_alter.
[0171] In specific implementation, the domain large model _frozen performs text generation processing based on the instruction text of the target domain to obtain the generated text of the target domain, and the general large model _frozen performs text generation processing based on the instruction text to obtain the generated text of the general domain; the domain large model _alter performs text generation processing based on the generated text of the general domain and the instruction text to obtain the first generated text; and the general large model _alter performs text generation processing based on the generated text of the target domain and the instruction text to obtain the second generated text; determine the model training loss according to the first generated text and the second generated text, and adjust the model parameters of the domain large model _alter and the general large model _alter according to the model training loss.
[0172] It should be noted that the models that need to adjust the model parameters at the end of each joint training include the domain large model _alter and the general large model _alter, and do not include the domain large model _frozen and the general large model _frozen. The domain large model _alter used in the last joint training before meeting the model training end condition is the domain text processing model that has been trained in the target domain, and the general large model _alter used in the last joint training before meeting the model training end condition is the general text processing model that has been trained in the general domain.
[0173] In a specific implementation manner, the training method of the text processing model further includes: obtaining an initial text processing model; inputting general data into the initial text processing model for iterative training to obtain a general text processing domain to be trained in the general domain; inputting the unlabeled domain data of the target domain into the general text processing domain to be trained in the general domain for iterative training to obtain a domain text processing model to be trained in the target domain.
[0174] The initial text processing model can be an LLM that has not undergone any model training.
[0175] Input general data into the initial text processing model for iterative training to obtain a general text processing domain to be trained in the general domain.
[0176] The general data can include large-scale and diverse labeled data sets that cover knowledge in various fields.
[0177] In specific implementation, it is also possible to perform a replication process on the general text processing domain to be trained in the general domain, so as to obtain two general text processing domains to be trained in the general domain.
[0178] Input the unlabeled domain data of the target domain into the general text processing domain to be trained in the general domain for iterative training to obtain a domain text processing model to be trained in the target domain.
[0179] The unlabeled domain data in the field of vision can be domain data in the target field that has not been manually labeled and cleaned.
[0180] During specific implementation, the domain text processing model to be trained in the target field can also be copied to obtain two domain text processing models to be trained in the target field.
[0181] Among the two general domain general text processing fields to be trained obtained, one of the parameters can be frozen, and the other remains in the state where the parameters are not frozen. Then, the general domain general text processing field with frozen parameters can be used to execute step S102, and the general domain general text processing field with unfrozen parameters can be used to execute step S104.
[0182] Similarly, among the two domain text processing fields to be trained in the target field obtained, one of the parameters can be frozen, and the other remains in the state where the parameters are not frozen. Then, the domain text processing field with frozen parameters in the target field can be used to execute step S102, and the domain text processing field with unfrozen parameters in the target field can be used to execute step S104.
[0183] In a specific implementation manner, determining the model training loss according to the first generated text and the second generated text includes: calculating the contrast loss between the first generated text and the second generated text, and using the contrast loss as the model training loss; or, determining the first text vector representing the first generated text and determining the second text vector representing the second generated text; performing vector distance calculation processing according to the first text vector and the second text vector, and using the calculated vector distance as the model training loss.
[0184] Determining the model training loss according to the first generated text and the second generated text can be calculating the contrast loss between the first generated text and the second generated text, and using the contrast loss as the model training loss.
[0185] The contrast loss can be InfoNCE loss. InfoNCE loss is a loss function based on contrast, which can reflect the similarity degree between the first generated text and the second generated text.
[0186] The larger the value of the InfoNCE loss, the less similar the first generated text and the second generated text are. The smaller the value of the InfoNCE loss, the more the first generated text and the second generated text tend to be consistent.
[0187] By taking the contrastive loss as the model training loss, the second generated text generated by the general text processing model to be trained in the general domain after multiple rounds of iterative training and the first generated text generated by the domain text processing model to be trained in the target domain after multiple rounds of iterative training can tend to be consistent, so that the general text processing model to be trained in the general domain can learn the advantages of the domain large model, and the domain text processing model to be trained in the target domain can learn the advantages of the general large model.
[0188] In the case of taking the contrastive loss as the model training loss, when the model training loss satisfies the training end condition, it can be that the value of the model training loss no longer decreases.
[0189] Determining the model training loss according to the first generated text and the second generated text can also include: determining the first text vector representing the first generated text and determining the second text vector representing the second generated text; performing vector distance calculation processing based on the first text vector and the second text vector, and taking the calculated vector distance as the model training loss.
[0190] Determining the first text vector representing the first generated text can be to perform encoding processing on the first generated text to obtain the first text vector representing the first generated text.
[0191] Determining the second text vector representing the second generated text can be to perform encoding processing on the second generated text to obtain the second text vector representing the second generated text.
[0192] Performing vector distance calculation processing based on the first text vector and the second text vector can be to perform a difference processing on the first text vector and the second text vector to obtain the target vector distance.
[0193] The target vector distance can reflect the similarity degree between the first generated text and the second generated text. The larger the value of the target vector distance, the less similar the first generated text and the second generated text are. The smaller the value of the target vector distance, the more the first generated text and the second generated text tend to be consistent.
[0194] By taking the target vector distance as the model training loss, the second generated text generated by the general text processing model to be trained in the general domain after multiple rounds of iterative training and the first generated text generated by the domain text processing model to be trained in the target domain after multiple rounds of iterative training can tend to be consistent, so that the general text processing model to be trained in the general domain can learn the advantages of the domain large model, and the domain text processing model to be trained in the target domain can learn the advantages of the general large model.
[0195] In a specific implementation, the first generated text and the second generated text are obtained during the i-th joint training process corresponding to the instruction text, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1. After obtaining the first generated text and the second generated text, the training method of the text processing model further includes: if i is less than N, then increment i by 1 and assign the result to i, and trigger the step of the domain text processing model to be trained in the target domain to perform text generation processing based on the general domain generated text and the instruction text, until i is equal to N.
[0196] The step of the domain text processing model to be trained in the target domain to perform text generation processing based on the general domain generated text and the instruction text refers to the following steps: the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text, that is, the aforementioned step S104.
[0197] The number of instruction texts can be multiple.
[0198] For each instruction text, during the i-th joint training process corresponding to the instruction text, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text. i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1.
[0199] For example, N = 3.
[0200] If i = 1, during the 1st joint training process corresponding to instruction text 1, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and instruction text 1 to obtain the first generated text 1; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and instruction text 1 to obtain the second generated text 1.
[0201] Since 1 is less than 3, increment i by 1 and assign the result to i, so that i = 2.
[0202] During the 2nd joint training process corresponding to instruction text 1, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and instruction text 1 to obtain the first generated text 2; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and instruction text 1 to obtain the second generated text 2.
[0203] Since 2 is less than 3, after incrementing i by 1 and assigning the result to i, i becomes 3.
[0204] During the 3rd joint training process corresponding to instruction text 1, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and instruction text 1 to obtain the first generated text 3; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and instruction text 1 to obtain the second generated text 3.
[0205] Since 3 = 3, the loop ends.
[0206] In the conventional model training process, a training sample set including multiple samples can be pre-generated, and the training sample set is input into the model to be trained for iterative training. Furthermore, the training samples used in different training rounds are different. For example, the training sample input into the model in the 1st training is sample 1, and sample 1 can be used to generate a text pair, and this text pair can be used to generate the model training loss for the 1st training; the training sample input into the model in the 2nd training is sample 2, and sample 2 can be used to generate a text pair, and this text pair can be used to generate the model training loss for the 2nd training... The training sample input into the model in the kth training is sample k, and sample k can be used to generate a text pair, and this text pair can be used to generate the model training loss for the kth training, and so on. k is an integer greater than 1.
[0207] This implementation method takes into account that the text generation results of the LLM have certain randomness and generalization. For each instruction text, multiple first generated texts and multiple second generated texts in response to this instruction text can be generated. Thus, during the process of calculating the model training loss, multiple model training losses are generated using the multiple second generated texts and multiple first generated texts and averaged, which is beneficial to reducing the fluctuations caused by model prediction randomness and improving the optimization efficiency and performance.
[0208] For example, input sample 1 into the model 3 times to obtain three first generated texts and three second generated texts. Then, use the 9 text pairs obtained by permutation and combination to generate the average loss, and this average loss is the model training loss corresponding to sample 1, which can be used to replace the model training loss for 1 time of training in the aforementioned conventional model training.
[0209] In a specific implementation manner, determining the model training loss based on the first generated text and the second generated text includes: adding the first generated text to the first generated text set, and adding the second generated text to the second generated text set, where the first generated text set includes N first generated texts, and the second generated text set includes N second generated texts; performing permutation and combination on the N first generated texts and the N second generated texts to obtain M text pair sets, where a text pair set includes N text pairs, and a text pair is composed of a first generated text and a second generated text; for each text pair set, determining the minimum contrast loss of the text pair set based on the N contrast losses of the N text pairs in the text pair set; the contrast loss of a text pair is calculated based on the comparison of the first generated text and the second generated text in the text pair; performing a mean operation on the M minimum contrast losses of the M text pair sets to obtain the model training loss.
[0210] N can be a natural number greater than 1, and M can be a natural number greater than 1. M and N can be the same or different.
[0211] For the same instruction text, add the first generated text in response to the instruction text to the first generated text set, and add the second generated text in response to the instruction text to the second generated text set, where the first generated text set includes N first generated texts, and the second generated text set includes N second generated texts.
[0212] For example, N = 3. In the first joint training process corresponding to instruction text 1, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and instruction text 1 to obtain the first generated text a1; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and instruction text 1 to obtain the second generated text b1. Add the first generated text a1 to the first generated text set, and add the second generated text b1 to the second generated text set.
[0213] Since 1 is less than 3, perform an increment operation on i and assign the result to i, so that i = 2.
[0214] In the second joint training process corresponding to instruction text 1, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and instruction text 1 to obtain the first generated text a2; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and instruction text 1 to obtain the second generated text b2. Add the first generated text a2 to the first generated text set, and add the second generated text b2 to the second generated text set.
[0215] Since 2 is less than 3, after incrementing i by 1 and assigning the result to i, i becomes 3.
[0216] During the 3rd joint training process corresponding to instruction text 1, the domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and instruction text 1 to obtain the first generated text a3; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and instruction text 1 to obtain the second generated text b3. The first generated text a3 is added to the first generated text set, and the second generated text b3 is added to the second generated text set.
[0217] Since 3 = 3, the loop ends. The first generated text set includes 3 first generated texts: a1, a2, a3. The second generated text set includes 3 second generated texts: b1, b2, b3.
[0218] Arrange and combine the N first generated texts and N second generated texts to obtain M text pair sets. One text pair set includes N text pairs, and one text pair consists of one first generated text and one second generated text.
[0219] For example, N = 3, M = 3. The 3 first generated texts include: a1, a2, a3. The 3 second generated texts include: b1, b2, b3. Arrange and combine the 3 first generated texts and 3 second generated texts to obtain 3 text pair sets. Among them, text pair set 1 includes 3 text pairs: (a1, b1), (a1, b2), (a1, b3); text pair set 2 includes 3 text pairs: (a2, b1), (a2, b2), (a2, b3); text pair set 3 includes 3 text pairs: (a3, b1), (a3, b2), (a3, b3).
[0220] For each text pair set, based on the N contrast losses of the N text pairs in the text pair set, determine the minimum contrast loss of the text pair set. It can be to calculate the contrast loss of each of the N text pairs in the N text pairs to obtain N contrast losses, determine the minimum value among the N contrast losses, and determine this minimum value as the minimum contrast loss of the text pair set.
[0221] The contrast loss of a text pair is calculated based on the first generated text and the second generated text in the text pair. The contrast loss can be InfoNCE loss (contrast learning loss). InfoNCE loss is a loss function based on contrast, which can reflect the similarity degree between the first generated text and the second generated text.
[0222] Perform a mean operation on the M minimum contrast losses of the M text pair sets to obtain the model training loss.
[0223] For example, N = 3. The set 1 of text pairs includes 3 text pairs: (a1, b1), (a1, b2), (a1, b3); the set 2 of text pairs includes 3 text pairs: (a2, b1), (a2, b2), (a2, b3); the set 3 of text pairs includes 3 text pairs: (a3, b1), (a3, b2), (a3, b3).
[0224] For the set 1 of text pairs, calculate the contrastive loss of the text pair (a1, b1) to obtain l1, calculate the contrastive loss of the text pair (a1, b2) to obtain l2, and calculate the contrastive loss of the text pair (a1, b3) to obtain l3. Determine the minimum value among the three contrastive losses l1, l2, and l3 of the set 1 of text pairs to obtain the minimum contrastive loss of the set 1 of text pairs, which can be represented by v1.
[0225] For the set 2 of text pairs, calculate the contrastive loss of the text pair (a2, b1) to obtain l4, calculate the contrastive loss of the text pair (a2, b2) to obtain l5, and calculate the contrastive loss of the text pair (a2, b3) to obtain l6. Determine the minimum value among the three contrastive losses l4, l5, and l6 of the set 2 of text pairs to obtain the minimum contrastive loss of the set 2 of text pairs, which can be represented by v2.
[0226] For the set 3 of text pairs, calculate the contrastive loss of the text pair (a3, b1) to obtain l7, calculate the contrastive loss of the text pair (a3, b2) to obtain l8, and calculate the contrastive loss of the text pair (a3, b3) to obtain l9. Determine the minimum value among the three contrastive losses l1, l2, and l3 of the set 3 of text pairs to obtain the minimum contrastive loss of the set 3 of text pairs, which can be represented by v3.
[0227] Perform a mean calculation on the minimum contrastive loss v1 of the set 1 of text pairs, the minimum contrastive loss v2 of the set 2 of text pairs, and the minimum contrastive loss v3 of the set 3 of text pairs to obtain the model training loss V, where V = (v1 + v2 + v3) / 3.
[0228] By generating multiple training losses using multiple second generated texts and multiple first generated texts during the calculation of the model training loss, and performing operations such as finding the minimum value and finding the mean, the fluctuations caused by the randomness of model prediction can be reduced, and the optimization efficiency and performance can be improved.
[0229] In Figure 2In the illustrated embodiment, first, the domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the generated text in the target domain, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the generated text in the general domain; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data; the general text processing model to be trained in the general domain is trained with general data; then, the domain text processing model to be trained in the target domain performs text generation processing based on the generated text in the general domain and the instruction text to obtain the first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the generated text in the target domain and the instruction text to obtain the second generated text; finally, the model training loss is determined according to the first generated text and the second generated text, and the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss. Thus, considering that the text generated by the domain text processing model to be trained in the target domain obtained by training with unlabeled domain data may have negative impacts brought by noise data such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence breaks, etc., in the embodiment of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the generated text in the general domain generated by the general text processing model to be trained in the general domain as a reference during the process of performing text generation processing to obtain the first generated text, and use the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, which can enable the domain text processing model to be trained in the target domain to improve the speech normativity of the generated text after model training and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence breaks.
[0230] Considering that the general text processing model to be trained in the general domain can perform question answering and assistance for general social knowledge, but in the case of refining to the target domain, it is difficult for the general text processing model to be trained in the general domain to answer the detailed questions in the target domain. In the embodiment of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the generated text in the target domain generated by the domain text processing model to be trained in the target domain as a reference during the process of performing text generation processing to obtain the second generated text, and use the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, which can enable the general text processing model to be trained in the general domain to answer the detailed questions in the target domain after model training.
[0231] By determining the model training loss based on the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be achieved, so that the texts generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain tend to be consistent after model training, which is beneficial for the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to learn from each other's advantages.
[0232] In addition, in the process of transferring the capabilities of the general text processing model to be trained in the general domain to the target domain in the embodiments of the present application, the training samples input to the domain text processing model to be trained in the target domain include instruction texts and general domain generated texts, and the general domain generated texts are obtained by the general text processing model to be trained in the general domain performing text generation processing based on the instruction texts; in the process of transferring the capabilities of the domain text processing model to be trained in the target domain to the general domain in the embodiments of the present application, the training samples input to the general text processing model to be trained include instruction texts and target domain generated texts, and the target domain generated texts are obtained by the domain text processing model in the target domain performing text generation processing based on the instruction texts. Therefore, in the process of model transfer, the training samples of the domain text processing model to be trained in the target domain and the general text processing model to be trained in the general domain do not rely on manual annotation, reducing the manual workload and improving the training sample generation efficiency, thereby improving the overall model transfer efficiency.
[0233] Figure 2 It is a processing flow chart of another text processing model training method provided by the embodiments of the present application.
[0234] Figure 2 It shows another one or more acquisition methods of the domain text processing model in the target domain, the general text processing model in the general domain, the domain text processing model to be trained, and the general text processing model to be trained.
[0235] As Figure 2 shown, in step S202, the general large model is fine-tuned using the in-vertical-domain dialogue data to obtain the domain large model.
[0236] The in-vertical-domain dialogue data can be domain data in the target domain that has not been manually annotated and cleaned.
[0237] Step S204: Make a copy of the general large model, one being the general large model 1 and the other being the general large model 2; make a copy of the domain large model obtained in the previous step, one being the domain large model 1 and the other being the domain large model 2.
[0238] Making a copy of the general large model, one being the general large model 1 and the other being the general large model 2, can be taking the existing general large model before copying as the general large model 1 and the copied general large model as the general large model 2. Additionally, parameter thawing processing can be performed on the general large model 2.
[0239] Making a copy of the general large model, one being the general large model 1 and the other being the general large model 2, can be taking the existing general large model before copying as the general large model 2 and the copied general large model as the general large model 1. Additionally, parameter freezing processing can be performed on the general large model 1.
[0240] The general large model 1 can be the general large model with frozen parameters, referring to Figure 2 the general large model_frozen in the embodiment.
[0241] The general large model 2 can be the general large model with model parameters to be trained, referring to Figure 2 the general large model_alter in the embodiment.
[0242] Making a copy of the domain large model, one being the domain large model 1 and the other being the domain large model 2, can be taking the existing domain large model before copying as the domain large model 1 and the copied domain large model as the domain large model 2. Additionally, parameter thawing processing can be performed on the domain large model 2.
[0243] Making a copy of the domain large model, one being the domain large model 1 and the other being the domain large model 2, can also be taking the existing domain large model before copying as the domain large model 2 and the copied domain large model as the domain large model 1. Additionally, parameter freezing processing can be performed on the domain large model 1.
[0244] The domain large model 1 can be the domain large model with frozen parameters, referring to Figure 2 the domain large model_frozen in the embodiment.
[0245] The domain large model 2 can be the domain large model with model parameters to be trained, referring to Figure 2 the domain large model_alter in the embodiment.
[0246] Step S206: Use domain data to train the general large model 2 and the domain large model 2. The general large model 1 and the domain large model 1 are used as aids during the training process. During the training process, improve the answering ability and speech summarization ability of the domain large model 2, and improve the domain knowledge ability of the general large model 2.
[0247] The general large model 1 can be used as an aid for the domain large model 2 during the training process, and the domain large model 1 can be used as an aid for the general large model 2 during the training process.
[0248] This step can refer to Figure 2 Steps S202 - S206 in the embodiment.
[0249] Step S208: Stop training when the capabilities of the general large model 2 and the domain large model 2 are aligned with each other, and obtain a domain large model that has both domain knowledge and the ability to generalize summary speech.
[0250] The domain large model that has both domain knowledge and the ability to generalize summary speech includes the trained general large model 2 and the trained domain large model 2.
[0251] Due to the same technical concept, the description in this embodiment is relatively simple. For the relevant parts, please refer to the corresponding descriptions in the above - provided method embodiments.
[0252] Figure 3 This is a data flow diagram of a training method for a text processing model provided by an embodiment of the present application.
[0253] As Figure 3 shown, input the training data 302 into the domain large model 304 for text generation processing to obtain the target domain generated text. The training data 302 can refer to Figure 2 the corresponding description part of the "instruction text of the target domain" in the embodiment.
[0254] Input the training data 302 into the general large model 306 for text generation processing to obtain the general domain generated text.
[0255] Input the training data 302 and the target domain generated text output by the domain large model 304 into the general large model 308 for text generation processing to obtain the second generated text.
[0256] Input the training data 302 and the general domain generated text output by the general large model 306 into the domain large model 310 for text generation processing to obtain the first generated text.
[0257] The model training loss 312 is generated based on the second generated text output by the general large model 308 and the first generated text output by the domain large model 310, and this model training loss 312 is used to adjust the model parameters of the general large model 308 and the domain large model 310.
[0258] Due to the same technical concept, the description in this embodiment is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above.
[0259] An embodiment of a text generation method provided in this specification:
[0260] Due to the same technical concept as the training method embodiment of the foregoing text processing model, this specification also provides an embodiment of a text generation method. Figure 4 It is a processing flow chart of a text generation method provided for the embodiments of this application.
[0261] As Figure 4 shown, in step S402, the instruction text to be processed is obtained.
[0262] In step S404, the general text processing model in the general domain performs text generation processing based on the instruction text to be processed to obtain a first response text, and the domain text processing model in the target domain performs text generation processing based on the instruction to be processed to obtain a second response text; the general text processing model and the domain text processing model are trained through the training method of the text processing model.
[0263] The training method of the text processing model in this step can refer to the training method embodiment of the foregoing text processing model.
[0264] In step S406, a target text in response to the instruction text to be processed is output based on the first response text and the second response text.
[0265] The target text can be either one of the first response text and the second response text, or can include both the first response text and the second response text.
[0266] After obtaining the first response text and the second response text, both the first response text and the second response text can be determined as the target text in response to the instruction text to be processed, and this target text is displayed for the user to select.
[0267] After obtaining the first response text and the second response text, it is also possible to score the first response text and the second response text according to a preset evaluation method to obtain the evaluation score of the first response text and the evaluation score of the second response text, and determine the one with the higher evaluation score as the target text in response to the instruction text to be processed.
[0268] In as Figure 4In the illustrated embodiment, first, an instruction text to be processed is obtained; then, a general text processing model in the general domain performs text generation processing based on the instruction text to be processed to obtain a first response text, and a domain text processing model in the target domain performs text generation processing based on the instruction to be processed to obtain a second response text; the general text processing model and the domain text processing model are obtained through a training method of the text processing model; finally, a target text in response to the instruction text to be processed is output based on the first response text and the second response text. Thus, the general text processing model and the domain text processing model are obtained by jointly training the general text processing model to be trained and the domain text processing model to be trained through the training method of the text processing model. During the joint training process, considering that the text generated by the domain text processing model to be trained in the target domain obtained by training with unlabeled domain data may have negative impacts brought by noise data such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation, in the embodiment of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the general domain generated text generated by the general text processing model to be trained in the general domain as a reference during the process of performing text generation processing to obtain a first generated text, and use the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, which can enable the domain text processing model to be trained in the target domain to improve the speech normativity of the generated text and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation after model training. Considering that the general text processing model to be trained in the general domain can perform question answering and assistance for general social knowledge, but in the case of refining to the target domain, it is difficult for the general text processing model to be trained in the general domain to answer detailed questions in the target domain. In the embodiment of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the target domain generated text generated by the domain text processing model to be trained in the target domain as a reference during the process of performing text generation processing to obtain a second generated text, and use the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, which can enable the general text processing model to be trained in the general domain to answer detailed questions in the target domain after model training.By determining the model training loss based on the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, the joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be achieved, so that the texts generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain tend to be consistent after model training, which is beneficial for the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to learn from each other's advantages. Therefore, on the basis that the general text processing model in the general domain and the domain text processing model in the target domain are trained by the text processing model training method provided in the foregoing method embodiments, there are differences in the focus of the first response text generated by the general text processing model in the general domain based on the instruction text to be processed and the second response text generated by the domain text processing model in the target domain based on the instruction text to be processed, but both have the advantage of being able to answer the detailed questions in the target domain, and have the advantages of improving the normativity of the conversation, reducing semantic ambiguity, automatic speech recognition errors, unsmooth sentence segmentation and other problems. Furthermore, the target text output based on the first response text and the second response text can better respond to the instruction text to be processed.
[0269] Due to the same technical concept, the description in this embodiment is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above.
[0270] In the above embodiment, a training method for a text processing model is provided. Correspondingly, based on the same technical concept, the embodiment of the present application also provides a training device for a text processing model, which will be described below with reference to the drawings.
[0271] Figure 5 It is a schematic diagram of a training device for a text processing model provided by an embodiment of the present application.
[0272] This embodiment provides a training device 500 for a text processing model, including:
[0273] A generation unit 502, configured to perform text generation processing on the domain text processing model to be trained in the target domain based on the instruction text in the target domain to obtain a target domain generated text, and perform text generation processing on the general text processing model to be trained in the general domain based on the instruction text to obtain a general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data; the general text processing model to be trained in the general domain is obtained by training with general data;
[0274] The generating unit 502 is further configured to perform text generation processing on the domain text processing model to be trained in the target domain based on the general domain generated text and the instruction text to obtain a first generated text; and perform text generation processing on the general text processing model to be trained in the general domain based on the target domain generated text and the instruction text to obtain a second generated text;
[0275] The training unit 504 is configured to determine a model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
[0276] Optionally, when determining the model training loss according to the first generated text and the second generated text, the training unit 504 performs the following steps:
[0277] Calculate a contrast loss between the first generated text and the second generated text, and use the contrast loss as the model training loss;
[0278] Or,
[0279] Determine a first text vector representing the first generated text, and determine a second text vector representing the second generated text;
[0280] Perform vector distance calculation processing according to the first text vector and the second text vector, and use the calculated vector distance as the model training loss.
[0281] Optionally, the first generated text and the second generated text are obtained in the i-th joint training process corresponding to the instruction text, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1; the training device 500 of the text processing model further includes:
[0282] A loop unit, configured to, if i is less than N, perform an increment operation on i and then assign the incremented value to i, and trigger the step of performing text generation processing on the domain text processing model to be trained in the target domain based on the general domain generated text and the instruction text until i is equal to N.
[0283] Optionally, when determining the model training loss according to the first generated text and the second generated text, the training unit 504 performs the following steps:
[0284] Add the first generated text to a first generated text set, and add the second generated text to a second generated text set, where the first generated text set includes N first generated texts, and the second generated text set includes N second generated texts;
[0285] Arrange and combine the N first generated texts and the N second generated texts to obtain M sets of text pairs. One set of text pairs includes N text pairs, and one text pair consists of one first generated text and one second generated text;
[0286] For each set of text pairs, determine the minimum contrast loss of the set of text pairs based on the N contrast losses of the N text pairs in the set of text pairs; the contrast loss of one text pair is calculated based on the contrast loss between the first generated text and the second generated text in the text pair;
[0287] Perform a mean operation on the M minimum contrast losses of the M sets of text pairs to obtain the model training loss.
[0288] Optionally, the training device 500 of the text processing model further includes:
[0289] An acquisition unit for acquiring an initial text processing model;
[0290] The training unit 504 is further configured to iteratively train the general data by inputting it into the initial text processing model to obtain a general text processing field to be trained in the general field;
[0291] The training unit 504 is further configured to iteratively train the unlabeled domain data in the target domain by inputting it into the general text processing field to be trained in the general domain to obtain a domain text processing model to be trained in the target domain.
[0292] Optionally, the instruction text includes background information and text generation instructions in the target domain; the domain text processing model to be trained in the target domain includes a keyword extraction module and a text generation module connected in series in sequence; when the generation unit 502 performs text generation processing on the instruction text in the target domain based on the domain text processing model to be trained in the target domain to obtain a generated text in the target domain, the following steps are performed:
[0293] The keyword extraction module performs keyword extraction processing on the background information to obtain keyword information in the background information;
[0294] The text generation module performs text prediction processing according to the text generation instructions under the prompt of the keyword information to obtain the generated text in the target domain.
[0295] The training device for the text processing model provided by the embodiments of the present application includes: a generation unit, configured to perform text generation processing on the domain text processing model to be trained in the target domain based on the instruction text in the target domain to obtain the generated text in the target domain, and perform text generation processing on the general text processing model to be trained in the general domain based on the instruction text to obtain the generated text in the general domain; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data; the general text processing model to be trained in the general domain is obtained by training with general data; the generation unit is further configured to perform text generation processing on the domain text processing model to be trained in the target domain based on the generated text in the general domain and the instruction text to obtain the first generated text; and perform text generation processing on the general text processing model to be trained in the general domain based on the generated text in the target domain and the instruction text to obtain the second generated text; a training unit, configured to determine the model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss. Thus, considering that the text generated by the domain text processing model to be trained in the target domain obtained by training with unlabeled domain data may have negative impacts brought by noise data such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence breaks, etc., in the embodiments of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the generated text in the general domain generated by the general text processing model to be trained in the general domain as a reference during the process of performing text generation processing to obtain the first generated text, and use the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, which can make the domain text processing model to be trained in the target domain improve the speech norm of the generated text after model training, and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence breaks.
[0296] Considering that the general text processing model to be trained in the general domain can perform question answering and assistance on general social knowledge, but in the case of being refined to the target domain, it is difficult for the general text processing model to answer the detailed questions in the target domain. In the embodiments of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the generated text in the target domain generated by the domain text processing model to be trained in the target domain as a reference during the process of performing text generation processing to obtain the second generated text, and use the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, which can make the general text processing model to be trained in the general domain be able to answer the detailed questions in the target domain after model training.
[0297] By determining the model training loss according to the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, the joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be realized, so that the texts generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain tend to be consistent after model training, which is beneficial for the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to learn the advantages of each other.
[0298] In addition, in the process of transferring the capabilities of the general text processing model to be trained in the general domain to the target domain in the embodiments of the present application, the training samples input to the domain text processing model to be trained in the target domain include instruction texts and general domain generated texts, and the general domain generated texts are obtained by the general text processing model to be trained in the general domain performing text generation processing based on the instruction texts; in the process of transferring the capabilities of the domain text processing model to be trained in the target domain to the general domain in the embodiments of the present application, the training samples input to the general text processing model to be trained include instruction texts and target domain generated texts, and the target domain generated texts are obtained by the domain text processing model in the target domain performing text generation processing based on the instruction texts. Therefore, in the process of model transfer, the training samples of the domain text processing model to be trained in the target domain and the general text processing model to be trained in the general domain do not rely on manual annotation, reducing the manual workload and improving the training sample generation efficiency, thereby improving the overall model transfer efficiency.
[0299] In the above embodiments, a text generation method is provided. Correspondingly, based on the same technical concept, the embodiments of the present application also provide a text generation device, which will be described below with reference to the accompanying drawings.
[0300] Figure 6 It is a schematic diagram of a text generation device provided by the embodiments of the present application.
[0301] This embodiment provides a text generation device 600, including:
[0302] An obtaining unit 602, configured to obtain an instruction text to be processed;
[0303] A generating unit 604, configured to perform text generation processing on the instruction text to be processed by a general text processing model in the general domain to obtain a first response text, and perform text generation processing on the instruction text to be processed by a domain text processing model in the target domain to obtain a second response text; the general text processing model and the domain text processing model are obtained by training through a training method of a text processing model;
[0304] An output unit 606, configured to output a target text for responding to the to-be-processed instruction text based on the first response text and the second response text.
[0305] The text generation device provided by the embodiments of the present application includes: an acquisition unit, configured to acquire an instruction text to be processed; a generation unit, configured to perform text generation processing on the instruction text to be processed based on a general text processing model in the general domain to obtain a first response text, and perform text generation processing on the instruction text to be processed based on a domain text processing model in the target domain to obtain a second response text; the general text processing model and the domain text processing model are obtained by training through a training method of the text processing model; an output unit, configured to output a target text in response to the instruction text to be processed based on the first response text and the second response text. Thus, the general text processing model and the domain text processing model are obtained by jointly training the general text processing model to be trained and the domain text processing model to be trained in the target domain through the training method of the text processing model. During the joint training process, considering that the text generated by the domain text processing model to be trained in the target domain obtained by training with unlabeled domain data may have negative impacts caused by noise data such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation, in the embodiments of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the general domain generated text generated by the general text processing model to be trained in the general domain as a reference during the process of performing text generation processing to obtain a first generated text, and use the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, so that the domain text processing model to be trained in the target domain can improve the speech normativity of the generated text after model training and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence segmentation. Considering that the general text processing model to be trained in the general domain can perform question answering and assistance for general social knowledge, but in the case of being refined to the target domain, it is difficult for the general text processing model to be trained in the general domain to answer detailed questions in the target domain. In the embodiments of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the target domain generated text generated by the domain text processing model to be trained in the target domain as a reference during the process of performing text generation processing to obtain a second generated text, and use the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, so that the general text processing model to be trained in the general domain can answer detailed questions in the target domain after model training.By determining the model training loss based on the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be achieved, enabling the texts generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to tend to be consistent after model training, which is conducive to the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain learning from each other's advantages. Therefore, on the basis that the general text processing model in the general domain and the domain text processing model in the target domain are trained by the text processing model training method provided in the foregoing method embodiments, there are differences in the focus of the first response text generated by the general text processing model in the general domain based on the to-be-processed instruction text and the second response text generated by the domain text processing model in the target domain based on the to-be-processed instruction text, but both have the advantage of being able to answer detailed questions in the target domain, and have the advantages of improving the normativity of the conversation, reducing semantic ambiguity, automatic speech recognition errors, unsmooth sentence segmentation and other problems. Furthermore, the target text output based on the first response text and the second response text can better respond to the to-be-processed instruction text.
[0306] Corresponding to the training method of a text processing model or a text generation method described above, based on the same technical concept, an embodiment of the present application further provides an electronic device, which is used to execute the training method of the text processing model or the text generation method provided above. Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0307] As Figure 7 shown, the electronic device may vary greatly due to configuration or performance, and may include one or more processors 701 and a memory 702. One or more application programs or data may be stored in the memory 702. Among them, the memory 702 may be a transient storage or a persistent storage. The application programs stored in the memory 702 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the electronic device. Further, the processor 701 may be set to communicate with the memory 702 and execute a series of computer executable instructions in the memory 702 on the electronic device. The electronic device may further include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, one or more keyboards 706, etc.
[0308] In a specific embodiment, an electronic device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the electronic device and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:
[0309] The domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain using unlabeled domain data; the general text processing model to be trained in the general domain is obtained by training using general data;
[0310] The domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain a first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain a second generated text;
[0311] Determine the model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
[0312] In another specific embodiment, an electronic device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the electronic device and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:
[0313] Obtain the instruction text to be processed;
[0314] The general text processing model in the general domain performs text generation processing based on the instruction text to be processed to obtain a first response text, and the domain text processing model in the target domain performs text generation processing based on the instruction to be processed to obtain a second response text; the general text processing model and the domain text processing model are obtained by training through the training method of the text processing model as described in the first aspect;
[0315] Output a target text in response to the to-be-processed instruction text based on the first response text and the second response text.
[0316] Embodiments of the computer-readable storage medium provided in this specification are as follows:
[0317] Corresponding to the training method of a text processing model described above, based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium.
[0318] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following processes can be implemented:
[0319] The domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain a target domain generated text, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain a general domain generated text; the domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain using unlabeled domain data; the general text processing model to be trained in the general domain is trained using general data;
[0320] The domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain a first generated text; and the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain a second generated text;
[0321] Determine a model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
[0322] It should be noted that the embodiment of the computer-readable storage medium in this specification and the embodiment of the training method of the text processing model in this specification are based on the same inventive concept. Therefore, for the specific implementation of this embodiment, reference can be made to the implementation of the corresponding method described above, and repeated parts will not be elaborated.
[0323] Corresponding to the text generation method described above, based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium.
[0324] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following processes can be implemented:
[0325] Obtain the instruction text to be processed;
[0326] The general text processing model in the general field performs text generation processing based on the instruction text to be processed to obtain a first response text, and the domain text processing model in the target field performs text generation processing based on the instruction to obtain a second response text; the general text processing model and the domain text processing model are trained by the training method of the text processing model as described in the first aspect;
[0327] Output a target text in response to the instruction text to be processed based on the first response text and the second response text.
[0328] It should be noted that the embodiments of the computer-readable storage medium in this specification and the embodiments of the text generation method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.
[0329] In the embodiments of the present application, first, the domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the generated text in the target domain, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the generated text in the general domain. The domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data. The general text processing model to be trained in the general domain is obtained by training with general data. Then, the domain text processing model to be trained in the target domain performs text generation processing based on the generated text in the general domain and the instruction text to obtain the first generated text. And the general text processing model to be trained in the general domain performs text generation processing based on the generated text in the target domain and the instruction text to obtain the second generated text. Finally, the model training loss is determined according to the first generated text and the second generated text, and the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain are adjusted according to the model training loss. In this way, considering that the text generated by the domain text processing model to be trained in the target domain obtained by training with unlabeled domain data may have negative impacts brought by noise data such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence breaks, etc., in the embodiments of the present application, for the same instruction text, the domain text processing model to be trained in the target domain can use the generated text in the general domain generated by the general text processing model to be trained in the general domain as a reference during the process of performing text generation processing to obtain the first generated text, and use the general text processing model to be trained in the general domain to guide the model training of the domain text processing model to be trained in the target domain, which can make the domain text processing model to be trained in the target domain improve the speech normativity of the generated text after model training and reduce problems such as semantic ambiguity, automatic speech recognition errors, and unsmooth sentence breaks. Considering that the general text processing model to be trained in the general domain can perform question answering and assistance for general social knowledge, but in the case of being refined to the target domain, it is difficult for the general text processing model to be trained in the general domain to answer the detailed questions in the target domain. In the embodiments of the present application, for the same instruction text, the general text processing model to be trained in the general domain can use the generated text in the target domain generated by the domain text processing model to be trained in the target domain as a reference during the process of performing text generation processing to obtain the second generated text, and use the domain text processing model to be trained in the target domain to guide the model training of the general text processing model to be trained in the general domain, which can make the general text processing model to be trained in the general domain be able to answer the detailed questions in the target domain after model training.By determining the model training loss based on the first generated text and the second generated text, and adjusting the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss, joint training of the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain can be achieved, enabling the texts generated by the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain to tend to be consistent after model training, which is conducive to the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain learning from each other's advantages. In addition, in the process of transferring the capabilities of the general text processing model to be trained in the general domain to the target domain in the embodiments of the present application, the training samples input to the domain text processing model to be trained in the target domain include instruction texts and general domain generated texts, and the general domain generated texts are obtained by the general text processing model to be trained in the general domain performing text generation processing based on the instruction texts; in the process of transferring the capabilities of the domain text processing model to be trained in the target domain to the general domain in the embodiments of the present application, the training samples input to the general text processing model to be trained include instruction texts and target domain generated texts, and the target domain generated texts are obtained by the domain text processing model in the target domain performing text generation processing based on the instruction texts. Therefore, in the process of model transfer, the training samples of the domain text processing model to be trained in the target domain and the general text processing model to be trained in the general domain do not rely on manual annotation, reducing the manual workload and improving the training sample generation efficiency, thereby improving the overall model transfer efficiency.
[0330] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0331] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0332] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable devices produce means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or in one or more blocks.
[0333] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or in one or more blocks.
[0334] These computer program instructions can also be loaded onto a computer or other programmable device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or in one or more blocks.
[0335] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0336] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0337] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0338] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0339] Embodiments of the present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0340] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0341] The above are only examples of this document and are not intended to limit this document. For those skilled in the art, various changes and modifications can be made to this document. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this document shall be included within the scope of the claims of this document.
Claims
1. A training method for a text processing model, characterized in that Including: The domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the target domain generated text, and the general text processing model to be trained in the general domain performs text generation processing based on the instruction text to obtain the general domain generated text; The domain text processing model to be trained in the target domain performs text generation processing based on the general domain generated text and the instruction text to obtain the first generated text; And the general text processing model to be trained in the general domain performs text generation processing based on the target domain generated text and the instruction text to obtain the second generated text; Determine the model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
2. The method according to claim 1, wherein The determining the model training loss according to the first generated text and the second generated text includes: Calculating the contrast loss between the first generated text and the second generated text, and using the contrast loss as the model training loss; Or, Determining a first text vector representing the first generated text, and determining a second text vector representing the second generated text; Performing vector distance calculation processing according to the first text vector and the second text vector, and using the calculated vector distance as the model training loss.
3. The method according to claim 1, characterized in that, The first generated text and the second generated text are obtained in the i-th joint training process corresponding to the instruction text, where i is an integer greater than or equal to 1 and less than or equal to N, and N is an integer greater than 1; After obtaining the first generated text and the second generated text, the method further includes: If i is less than N, then perform an increment operation on i and assign the result to i, triggering the step of the domain text processing model to be trained in the target domain to perform text generation processing based on the general domain generated text and the instruction text, until i is equal to N.
4. The method according to claim 3, characterized in that The determining the model training loss according to the first generated text and the second generated text includes: Adding the first generated text to the first generated text set, and adding the second generated text to the second generated text set, where the first generated text set includes N first generated texts, and the second generated text set includes N second generated texts; Performing permutation and combination on the N first generated texts and the N second generated texts to obtain M text pair sets, where a text pair set includes N text pairs, and a text pair consists of a first generated text and a second generated text; For each text pair set, determining the minimum contrast loss of the text pair set based on the N contrast losses of the N text pairs in the text pair set; the contrast loss of a text pair is calculated based on the first generated text and the second generated text in the text pair; Performing a mean operation on the M minimum contrast losses of the M text pair sets to obtain the model training loss.
5. The method according to claim 1, wherein The method further includes: Obtaining an initial text processing model; Input the general data into the initial text processing model for iterative training to obtain the general text processing field to be trained in the general domain; Input the unlabeled domain data in the target domain into the general text processing field to be trained in the general domain for iterative training to obtain the domain text processing model to be trained in the target domain.
6. The method according to claim 1, characterized in that, The instruction text includes the background information and text generation instructions in the target domain; the domain text processing model to be trained in the target domain includes a keyword extraction module and a text generation module connected in series in sequence; The domain text processing model to be trained in the target domain performs text generation processing based on the instruction text in the target domain to obtain the generated text in the target domain, including: The keyword extraction module performs keyword extraction processing on the background information to obtain keyword information in the background information; The text generation module performs text prediction processing according to the text generation instructions under the prompt of the keyword information to obtain the generated text in the target domain.
7. A text generation method, characterized in that, Including: Obtain the instruction text to be processed; The general text processing model in the general domain performs text generation processing based on the instruction text to be processed to obtain the first response text, and the domain text processing model in the target domain performs text generation processing based on the instruction to be processed to obtain the second response text; the general text processing model and the domain text processing model are trained by the training method of the text processing model according to any one of claims 1-6; Output the target text in response to the instruction text to be processed based on the first response text and the second response text.
8. A training device for a text processing model, characterized in that, Including: A generation unit, configured to perform text generation processing on the domain text processing model to be trained in the target domain based on the instruction text in the target domain to obtain the generated text in the target domain, and perform text generation processing on the general text processing model to be trained in the general domain based on the instruction text to obtain the generated text in the general domain; The domain text processing model to be trained in the target domain is obtained by training the general text processing model to be trained in the general domain with unlabeled domain data; the general text processing model to be trained in the general domain is obtained by training with general data; The generation unit is further configured to perform text generation processing on the domain text processing model to be trained in the target domain based on the generated text in the general domain and the instruction text to obtain the first generated text; And perform text generation processing on the general text processing model to be trained in the general domain based on the generated text in the target domain and the instruction text to obtain the second generated text; A training unit, configured to determine the model training loss according to the first generated text and the second generated text, and adjust the general text processing model to be trained in the general domain and the domain text processing model to be trained in the target domain according to the model training loss.
9. A text generation device, characterized in that, Including: An acquisition unit, configured to acquire the instruction text to be processed; A generation unit is configured to perform text generation processing on the to-be-processed instruction text by a general text processing model in a general domain to obtain a first response text, and perform text generation processing on the to-be-processed instruction by a domain text processing model in a target domain to obtain a second response text; the general text processing model and the domain text processing model are trained by the training method of the text processing model according to any one of claims 1-6; An output unit is configured to output a target text in response to the to-be-processed instruction text based on the first response text and the second response text.
10. An electronic device, characterized in that, The device includes: A processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to execute the training method of the text processing model according to any one of claims 1-6, or the text generation method according to claim 7.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store computer-executable instructions, which, when executed by a processor, implement the training method of the text processing model according to any one of claims 1-6, or the text generation method according to claim 7.