Training data generation method, training method, text processing method and related products

By generating training data of natural language processing models, using similarity and tasks associated with tasks to generate training data, the problem of high labor costs caused by manual annotation is solved, and the accuracy and efficiency of training data are improved.

CN120296415APending Publication Date: 2025-07-11SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343429.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing natural language processing model training data leads to high labor costs through manual annotation.

Method used

By obtaining the first text, performing multiple tasks to generate the second text, and generating training data based on similarity and task-association-related tasks, the training data is generated using similarity and style evaluation to reduce labor costs.

Benefits of technology

It reduces the labor cost of training data generation and improves the accuracy and efficiency of training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296415A_ABST
    Figure CN120296415A_ABST
Patent Text Reader

Abstract

The invention discloses a training data generation method, a training method, a text processing method and a related product. The training data generation method comprises the steps of obtaining a first text; executing the first task based on the first text to obtain a second text; a second task associated with the first task is executed based on the second text to obtain a third text, the second task comprises an execution result based on the first task, an original text of the first task is determined, and the execution result of the first task is obtained by executing the first task based on the original text; and determining a first similarity between the first text and the third text. Based on the first similarity, the first text and the second text, first training data is generated, the first training data comprises the first text and the second text, and in the first training data, the second text is used for supervising an execution result obtained by executing the first task based on the first text by the first model. Therefore, the labor cost for generating the training data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular, to a method for generating training data, a training method, a text processing method, and related products. Background Art

[0002] With the development of natural language technology, the application of natural language processing models is becoming more and more extensive. Before using a natural language processing model to perform natural language processing tasks, it is necessary to train the natural language processing model with training data. Currently, the training data is obtained through manual annotation, but this method has a high labor cost. Summary of the Invention

[0003] This application provides a method for generating training data, a training method, a text processing method, and related products to reduce the labor cost of generating training data. Among them, the related products include a training data generation device, a training device, a text processing device, an electronic device, a computer-readable storage medium, and a computer program product.

[0004] In a first aspect, a method for generating training data is provided. The method includes:

[0005] Obtain a first text;

[0006] Execute a first task based on the first text to obtain a second text;

[0007] Execute a second task associated with the first task based on the second text to obtain a third text. The second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by executing the first task based on the original text;

[0008] Determine a first similarity between the first text and the third text; generate first training data based on the first similarity, the first text, and the second text. The first training data includes the first text and the second text. In the first training data, the second text is used to supervise the execution result obtained by a first model executing the first task based on the first text.

[0009] Combined with any implementation manner of this application, the executing a first task based on the first text to obtain a second text includes:

[0010] Execute the first task based on a first processing method and the first text to obtain the second text;

[0011] The method further includes:

[0012] Performing the first task based on a second processing method different from the first processing method and the first text to obtain a fourth text;

[0013] Performing the second task based on the fourth text to obtain a fifth text;

[0014] Determining a second similarity between the first text and the fifth text;

[0015] The generating the first training data based on the first similarity, the first text, and the second text includes:

[0016] Generating the first training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text.

[0017] Combined with any embodiment of the present application, the performing the second task associated with the first task based on the second text to obtain a third text includes:

[0018] Performing the second task based on the first processing method and the second text to obtain the third text;

[0019] The performing the second task based on the fourth text to obtain a fifth text includes:

[0020] Performing the second task based on the second processing method and the fourth text to obtain the fifth text.

[0021] Combined with any embodiment of the present application, the generating the first training data based on the first similarity, the first text, and the second text includes:

[0022] Generating the first training data based on the magnitude relationship between the first similarity and the second similarity, the first text, the second text, and the fourth text;

[0023] In the case where the first training data includes the first text and the second text, the second text is used to supervise the execution result obtained by the first model performing the first task based on the first text; in the case where the first training data includes the first text and the fourth text, the fourth text is used to supervise the execution result obtained by the first model performing the first task based on the first text.

[0024] Combined with any embodiment of the present application, the method further includes:

[0025] Generate second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text, where the second training data is used to optimize the training of the first model, and the optimization training is different from the training of the first model based on the first training data.

[0026] Combined with any implementation manner of the present application, the generating of the second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text includes:

[0027] Determine the quality of the second text based on whether the style of the first similarity and the second text meets the first preset requirement;

[0028] Determine the quality of the fourth text based on whether the style of the second similarity and the fourth text meets the first preset requirement;

[0029] When the quality of the second text is higher than the quality of the fourth text, the generating of the second training data includes the first text, the second text, and the fourth text. In the second training data, the second text is the execution result obtained by performing the first task based on the first text, and the fourth text is not the execution result obtained by performing the first task based on the first text.

[0030] Combined with any implementation manner of the present application, the method further includes:

[0031] When the quality of the second text and the quality of the fourth text do not meet the second preset requirement, obtain a sixth text, where the sixth text includes the execution result obtained by performing the first task based on the first text, and the quality of the sixth text meets the second preset requirement;

[0032] The generating of the second training data includes the first text, the second text, and the sixth text. In the second training data, the sixth text is the execution result obtained by performing the first task based on the first text, and the second text is not the execution result obtained by performing the first task based on the first text;

[0033] and / or, the generating of the second training data includes the first text, the fourth text, and the sixth text. In the second training data, the sixth text is the execution result obtained by performing the first task based on the first text, and the fourth text is not the execution result obtained by performing the first task based on the first text.

[0034] In combination with any embodiment of the present application, before generating the first training data based on the first similarity, the first text, and the second text, the method further includes:

[0035] Based on the content quality rule and / or content verification instruction, determine whether the content of the first text and the content of the second text are reasonable, and obtain a judgment result;

[0036] The generating the first training data based on the first similarity, the first text, and the second text includes:

[0037] In the case where the judgment result includes that the content of the first text and the content of the second text are both reasonable, generate the first training data based on the first similarity, the first text, and the second text.

[0038] In a second aspect, a method for generating training data is provided, and the method includes:

[0039] Obtain a first text, where the first text includes characters in a first language;

[0040] Translate the characters in the first language in the first text into characters in a second language to obtain a second text;

[0041] Translate the characters in the second language in the second text into the characters in the first language to obtain a third text;

[0042] Determine a first similarity between the first text and the third text;

[0043] Generate first training data based on the first similarity, the first text, and the second text, where the first training data includes the first text and the second text, and in the first training data, the second text is used to supervise the text obtained by the first model translating the characters in the first language in the first text into the characters in the second language.

[0044] In a third aspect, a training method is provided, and the method includes:

[0045] Obtain the first training data generated based on the first aspect and any of its embodiments;

[0046] Train a first model based on the first training data, where the first model is used to perform a first task.

[0047] In a fourth aspect, a training method is provided, and the method includes:

[0048] Obtain a first model and a second model based on a third party, where the second model is used to evaluate the quality of the execution result of a first task, and the quality of the execution result of the first task includes the accuracy of the execution result of the first task and whether the style of the execution result of the first task meets a first preset requirement. The second model is trained based on second training data, and the second training data is generated based on an implementation manner of a first aspect;

[0049] Based on the second model, train the first model to obtain a third model, where the third model is used to execute the first task, and the style of the execution result obtained by executing the first task meets the first preset requirement.

[0050] Combined with any implementation manner of this application, the training the first model based on the second model to obtain a third model includes:

[0051] Obtain a first text and at least one candidate text, where the candidate text is obtained by the first model executing the first task based on the first text;

[0052] Based on the second model, determine a compliant text that meets the first preset requirement and a non-compliant text that does not meet the first preset requirement from the at least one candidate text;

[0053] Based on the first model, determine a first probability that the compliant text is the result of executing the first task based on the first text, and a second probability that the non-compliant text is the result of executing the first task based on the first text;

[0054] Based on a first difference between the first probability and the second probability, determine a first loss of the first model, where the first difference is negatively correlated with the first loss;

[0055] Based on the first loss, update the parameters of the first model to obtain the third model.

[0056] In a fifth aspect, a text processing method is provided, and the method includes:

[0057] Obtain a text to be processed and a third model trained based on a fourth aspect and any of its implementation manners;

[0058] Based on the third model and the text to be processed, execute the first task to obtain a target text, and the style of the target text meets the first preset requirement.

[0059] In a sixth aspect, a training data generation device is provided, and the training data generation device includes:

[0060] An acquisition unit, configured to acquire a first text;

[0061] An execution unit, configured to execute a first task based on the first text to obtain a second text;

[0062] The execution unit is further configured to execute a second task associated with the first task based on the second text to obtain a third text, where the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by executing the first task based on the original text;

[0063] A determination unit, configured to determine a first similarity between the first text and the third text;

[0064] A generation unit, configured to generate first training data based on the first similarity, the first text, and the second text, where the first training data includes the first text and the second text, and in the first training data, the second text is used to supervise the execution result obtained by the first model executing the first task based on the first text.

[0065] Combined with any embodiment of the present application, the execution unit is further configured to:

[0066] Execute the first task based on a first processing method and the first text to obtain the second text;

[0067] Execute the first task based on a second processing method different from the first processing method and the first text to obtain a fourth text;

[0068] Execute the second task based on the fourth text to obtain a fifth text;

[0069] The determination unit is further configured to determine a second similarity between the first text and the fifth text;

[0070] The generation unit is further configured to generate first training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text.

[0071] Combined with any embodiment of the present application, the execution unit is further configured to execute the second task based on the first processing method and the second text to obtain the third text;

[0072] Execute the second task based on the second processing method and the fourth text to obtain the fifth text.

[0073] Combined with any embodiment of the present application, the generation unit is further configured to:

[0074] Generate the first training data based on the magnitude relationship between the first similarity and the second similarity, the first text, the second text, and the fourth text;

[0075] When the first training data includes the first text and the second text, the second text is used to supervise the execution result obtained by the first model performing the first task based on the first text; when the first training data includes the first text and the fourth text, the fourth text is used to supervise the execution result obtained by the first model performing the first task based on the first text.

[0076] Combined with any implementation manner of the present application, the generating unit is further configured to generate second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text, and the second training data is used to perform optimization training on the first model, and the optimization training is different from the training performed on the first model based on the first training data.

[0077] Combined with any implementation manner of the present application, the generating unit is further configured to:

[0078] Determine the quality of the second text based on whether the style of the first similarity and the second text meets the first preset requirement;

[0079] Determine the quality of the fourth text based on whether the style of the second similarity and the fourth text meets the first preset requirement;

[0080] When the quality of the second text is higher than the quality of the fourth text, generate the second training data including the first text, the second text, and the fourth text. In the second training data, the second text is the execution result obtained by the first model performing the first task based on the first text, and the fourth text is not the execution result obtained by the first model performing the first task based on the first text.

[0081] Combined with any implementation manner of the present application, the generating unit is further configured to:

[0082] When the quality of the second text and the quality of the fourth text do not meet the second preset requirement, obtain a sixth text, where the sixth text includes the execution result obtained by the first model performing the first task based on the first text, and the quality of the sixth text meets the second preset requirement;

[0083] Generating the second training data includes the first text, the second text, and the sixth text. In the second training data, the sixth text is an execution result obtained by performing the first task based on the first text, and the second text is not an execution result obtained by performing the first task based on the first text;

[0084] And / or, generating the second training data includes the first text, the fourth text, and the sixth text. In the second training data, the sixth text is an execution result obtained by performing the first task based on the first text, and the fourth text is not an execution result obtained by performing the first task based on the first text.

[0085] Combined with any implementation manner of the present application, the execution unit is further configured to:

[0086] Based on the content quality rule and / or the content verification instruction, determine whether the content of the first text and the content of the second text are reasonable, and obtain a determination result;

[0087] The generating unit is further configured to, when the determination result includes that the content of the first text and the content of the second text are both reasonable, generate first training data based on the first similarity, the first text, and the second text.

[0088] In a seventh aspect, a training data generation device is provided. The training data generation device includes:

[0089] An obtaining unit, configured to obtain a first text, where the first text includes characters in a first language;

[0090] A translation unit, configured to translate the characters in the first language in the first text into characters in a second language to obtain a second text;

[0091] The translation unit is further configured to translate the characters in the second language in the second text into the characters in the first language to obtain a third text;

[0092] A determination unit, configured to determine a first similarity between the first text and the third text;

[0093] A generating unit, configured to generate first training data based on the first similarity, the first text, and the second text. The first training data includes the first text and the second text. In the first training data, the second text is used to supervise the text obtained by a first model translating the characters in the first language in the first text into the characters in the second language.

[0094] In an eighth aspect, a training device is provided. The training device includes:

[0095] An acquisition unit, configured to acquire first training data generated based on the first aspect and any of its embodiments;

[0096] A training unit, configured to train a first model based on the first training data, where the first model is used to perform a first task.

[0097] In a ninth aspect, a training device is provided, and the training device includes:

[0098] An acquisition unit, configured to acquire a first model and a second model based on the third aspect, where the second model is used to evaluate the quality of the execution result of the first task, and the quality of the execution result of the first task includes the accuracy of the execution result of the first task and whether the style of the execution result of the first task meets a first preset requirement. The second model is trained based on second training data, and the second training data is generated based on an embodiment of the first aspect;

[0099] A training unit, configured to train the first model based on the second model to obtain a third model, where the third model is used to perform the first task, and the style of the execution result obtained by performing the first task meets the first preset requirement.

[0100] In combination with any embodiment of the present application, the training unit is further configured to:

[0101] Acquire a first text and at least one candidate text, where the candidate text is the result obtained by the first model performing the first task based on the first text;

[0102] Determine a compliant text that meets the first preset requirement and a non-compliant text that does not meet the first preset requirement from the at least one candidate text based on the second model;

[0103] Determine a first probability that the compliant text is the result of the first model performing the first task based on the first text, and a second probability that the non-compliant text is the result of the first model performing the first task based on the first text;

[0104] Determine a first loss of the first model based on a first difference between the first probability and the second probability, where the first difference is negatively correlated with the first loss;

[0105] Update the parameters of the first model based on the first loss to obtain the third model.

[0106] In a tenth aspect, a text processing device is provided, and the text processing device includes:

[0107] An acquisition unit, configured to acquire the text to be processed and a third model trained based on the training method and its implementation manners of the fourth aspect;

[0108] An execution unit, configured to execute a first task based on the third model and the text to be processed, so as to obtain a target text, and the style of the target text meets a first preset requirement.

[0109] In an eleventh aspect, there is provided an electronic device, including: a processor and a memory, where the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any of its implementation manners as described above, or the electronic device executes the technical solution of the second aspect as described above, or the electronic device executes the technical solution of the third aspect as described above, or the electronic device executes the technical solution of the fourth aspect as described above, or the electronic device executes the fifth aspect and any of its implementation manners as described above.

[0110] In a twelfth aspect, there is provided another electronic device, including: a processor, a sending device, an input device, an output device, and a memory, where the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any of its implementation manners as described above, or the electronic device executes the technical solution of the second aspect as described above, or the electronic device executes the technical solution of the third aspect as described above, or the electronic device executes the technical solution of the fourth aspect as described above, or the electronic device executes the fifth aspect and any of its implementation manners as described above.

[0111] In a thirteenth aspect, there is provided a computer-readable storage medium, in which a computer program is stored, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the first aspect and any of its implementation manners as described above, or the processor is caused to execute the technical solution of the second aspect as described above, or the processor is caused to execute the technical solution of the third aspect as described above, or the processor is caused to execute the technical solution of the fourth aspect as described above, or the processor is caused to execute the fifth aspect and any of its implementation manners as described above.

[0112] In a fourteenth aspect, a computer program product is provided. The computer program product includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer is caused to execute the above first aspect and any of its embodiments, or the computer is caused to execute the technical solution of the above second aspect, or the computer is caused to execute the technical solution of the above third aspect, or the computer is caused to execute the technical solution of the above fourth aspect, or the computer is caused to execute the above fifth aspect and any of its embodiments.

[0113] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit this application.

[0114] In this application, after the generation device obtains the first text, it executes a first task based on the first text to obtain a second text. Then it executes a second task based on the second text to obtain a third text. Among them, the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by executing the first task based on the original text. Then the first similarity between the first text and the third text is determined. At this time, the greater the first similarity, the more similar the content of the first text is to the content of the third text. Also, since the third text is determined based on the second text, the similarity between the content of the third text and the content of the second text is high. Therefore, the similarity between the content of the second text and the content of the first text is high. That is to say, the accuracy of the second text as the execution result obtained by executing the first task based on the first text is high. That is, the first similarity is positively correlated with the accuracy of the second text as the execution result of executing the first task based on the first text, and the first similarity can measure the accuracy of the second text. Since the first similarity can measure the accuracy of the second text, it can be determined whether the second text can be used to supervise the execution result obtained by the first model executing the first task based on the first text based on the first similarity. Therefore, the first training data can be generated based on the first similarity, the first text, and the second text, thereby reducing the labor cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] In order to more clearly illustrate the technical solutions in the embodiments of this application or the background art, the following will describe the drawings required to be used in the embodiments of this application or the background art.

[0116] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show the embodiments that conform to this application and are used together with the specification to illustrate the technical solutions of this application.

[0117] Figure 1 It is a schematic flowchart of a method for generating training data provided by an embodiment of this application;

[0118] Figure 2Schematic flowchart of another training data generation method provided by an embodiment of the present application;

[0119] Figure 3 Schematic flowchart of yet another training data generation method provided by an embodiment of the present application;

[0120] Figure 4 Schematic flowchart of a training method provided by an embodiment of the present application;

[0121] Figure 5 Schematic flowchart of another training method provided by an embodiment of the present application;

[0122] Figure 6 Schematic flowchart of a text processing method provided by an embodiment of the present application;

[0123] Figure 7 Schematic structural diagram of a training data generation device provided by an embodiment of the present application;

[0124] Figure 8 Schematic structural diagram of another training data generation device provided by an embodiment of the present application;

[0125] Figure 9 Schematic structural diagram of a training device provided by an embodiment of the present application;

[0126] Figure 10 Schematic structural diagram of another training device provided by an embodiment of the present application;

[0127] Figure 11 Schematic structural diagram of a text processing device provided by an embodiment of the present application;

[0128] Figure 12 Schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0129] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0130] In the description, claims and the above-mentioned drawings of this application, terms such as "first", "second", etc. are used to distinguish different objects rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0131] Reference to "embodiment" herein means that a particular feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the description and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0132] An embodiment of this application provides a method for generating training data. Among them, the execution subject of the training data generation method is a training data generation device (hereinafter simply referred to as the generation device). Among them, the generation device can be any electronic device that can execute the technical solutions disclosed in the method embodiments of this application. Optionally, the generation device can be one of the following: a computer, a server.

[0133] It should be understood that the method embodiments of this application can also be implemented by a processor executing computer program code. The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for generating training data provided by an embodiment of this application.

[0134] 101. Obtain a first text.

[0135] In the embodiments of this application, the first text can be any text. In a possible implementation, the first text is a news report containing key elements such as characters, events, time, and location, such as "Yesterday, in the square in the city center, a grand public welfare concert was held, with many well-known singers gathering together, attracting thousands of citizens to come and watch." Another example is that the first text is an abstract of an academic paper, such as "This paper conducts research on the application of artificial intelligence in medical image diagnosis. By comparing multiple algorithm models, an optimized deep learning architecture is proposed. Experimental results show that this architecture can effectively improve the diagnostic accuracy." Another example is that the first text is a product description, "This smart watch has functions such as heart rate monitoring, sleep tracking, and exercise recording. It adopts a thin and light design, has a battery life of up to one week, and supports the switching of multiple personalized watch faces."

[0136] Optionally, the text includes characters and data that are not text in modality. For example, the text includes characters and images. For another example, the text includes characters and videos. For still another example, the text includes characters and audio.

[0137] In an implementation of obtaining the first text, the generating device receives the first text input by the user through the input component, where the input component includes: a mouse, a keyboard, a touch screen, a touch pad, and an audio input device.

[0138] In another implementation of obtaining the first text, the generating device receives the first text sent by the user through the terminal, where the terminal includes: a mobile phone, a computer, a tablet computer, and a smart wearable device.

[0139] 102. Execute a first task based on the first text to obtain a second text.

[0140] In the embodiments of the present application, the first task includes a task of generating text based on text. In a possible implementation manner, the first task includes translation. Optionally, the first text includes characters in a first language, and the first task includes translating the characters in the first language into characters in a second language. Executing the first task based on the first text includes translating the characters in the first language in the first text into characters in the second language. Correspondingly, the second text obtained by executing the first task based on the first text includes characters in the second language. For example, if the first language is Chinese, correspondingly, the first text includes Chinese characters, that is, the first text includes Chinese. If the second language is English, then executing the first task based on the first text includes translating the Chinese in the first text into English. Correspondingly, the translated second text includes English. In another possible implementation manner, the first task includes a dialogue task. Specifically, executing the first task based on the first text includes generating a reply to the first text. Correspondingly, the second text includes the answer to the question in the first text. Optionally, the first text includes a question, and the first task includes generating an answer to the question. In another possible implementation manner, the first task includes generating antithetical couplets. Specifically, the first text includes requirements for the antithetical couplets to be generated. Executing the first task based on the first text includes generating antithetical couplets that match the requirements in the first text. Correspondingly, the second text includes the antithetical couplets generated by executing the first task. In another possible implementation manner, the first task includes composing poems. Specifically, the first text includes requirements for the poems to be composed. Executing the first task based on the first text includes generating poems that match the requirements in the first text. Correspondingly, the second text includes the poems generated by executing the first task.

[0141] Optionally, the generating device implements step 102 based on a first model, where the first model has the ability to perform a first task based on text. For example, the first task includes translation, and the first model has translation ability. For another example, the first task includes a dialogue task, and the first model has the ability to conduct a dialogue. For yet another example, the first task includes generating couplets, and the first model has the ability to generate couplets. For still another example, the first task includes poem writing, and the first model has the ability to generate poems. Optionally, the first model is a large language model (LLM).

[0142] 103. Perform a second task associated with the first task based on the second text to obtain a third text.

[0143] In the embodiments of the present application, the second task is associated with the first task. Specifically, the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by performing the first task based on the original text. That is to say, the original text is the text required to obtain the execution result of the first task, and the second task includes determining the text required to obtain the execution result of the first task based on the execution result of the first task. Specifically, the first task includes a task of generating text based on text, and the execution result of the first task includes the text generated based on the text. In a possible implementation manner, the first task includes translation, and the second task includes determining the text to be translated based on the translation result. Optionally, the first task includes translating the text in the first language into the text in the second language, then the second task includes translating the text in the second language into the text in the first language. For example, the text in the first language is Chinese, and the text in the second language is English, then the second task includes translating the English in the text into Chinese. In another possible implementation manner, the first task includes a dialogue task, and the second task includes determining the speech corresponding to the reply based on the reply in the dialogue. Optionally, the first task includes generating an answer to a question, and the second task includes determining the question based on the answer. In yet another possible implementation manner, the first task includes generating couplets, and the second task includes determining the requirements for generating the couplets based on the couplets. In yet another possible implementation manner, the first task includes poem writing, and the second task includes determining the requirements for generating the poem based on the poem.

[0144] Optionally, the generating device implements step 103 based on a second model, where the second model has the ability to perform a second task based on text. For example, the second task includes determining the text to be translated based on the translation result, and the second model has the ability to determine the text to be translated based on the translation result. Another example is that the second task includes determining the speech corresponding to the reply based on the reply in the conversation, and the second model has the ability to determine the speech corresponding to the reply based on the reply in the conversation. Still another example is that the second task includes determining the requirements for generating the couplet based on the couplet, and the second model has the ability to determine the requirements for generating the couplet based on the couplet. Another example is that the second task includes determining the requirements for generating the poem based on the poem, and the second model has the ability to determine the requirements for generating the poem based on the poem. Optionally, the second model is an LLM.

[0145] 104. Determine the first similarity between the first text and the third text.

[0146] In the embodiments of the present application, the first similarity is the similarity between the first text and the third text. The greater the first similarity, the more similar the content of the first text is to the content of the third text. Also, since the third text is determined based on the second text, the similarity between the content of the third text and the content of the second text is high. Therefore, the similarity between the content of the second text and the content of the first text is high. That is to say, the second text has a high accuracy as the execution result obtained by performing the first task based on the first text, that is, the first similarity can be used to measure the accuracy of the second text. Specifically, the first similarity is positively correlated with the accuracy of the second text.

[0147] In a possible implementation manner, the generating device determines the similarity between texts based on bilingual evaluation understudy (BLEU) between texts. Correspondingly, the generating device determines the BLEU between the first text and the third text to determine the first similarity. In another possible implementation manner, the generating device encodes the first text to obtain a vector of the first text, and encodes the third text to obtain a vector of the third text. Based on the cosine similarity between the vector of the first text and the vector of the third text, the first similarity is determined. In still another possible implementation manner, the generating device encodes the first text to obtain a vector of the first text, and encodes the third text to obtain a vector of the third text. Based on the distance between the vector of the first text and the vector of the third text, the first similarity is determined.

[0148] In a possible scenario, the first task includes translating the text in the first language into the text in the second language, and the second task includes translating the text in the second language into the text in the first language. The first text includes the text in the first language, the second text includes the text in the second language, and the third text includes the text in the first language. Correspondingly, the first similarity between the first text and the third text can be used to measure the accuracy of the second text as the translation result of the first text, and the greater the first similarity, the higher the accuracy.

[0149] 105. Generate first training data based on the first similarity, the first text, and the second text.

[0150] In the embodiments of the present application, the first training data is used to train a first model capable of performing the first task based on the text. The first training data may include the first text and the second text. When the first training data includes the first text and the second text, the second text is used to supervise the execution result obtained by the first model performing the first task based on the first text.

[0151] Optionally, in the process of training the first model using the first training data, after inputting the first text into the first model to be trained, the first model to be trained performs the first task based on the first text to obtain a first training result. Then, use the second text to supervise the execution result obtained by the first model performing the first task based on the first text. Specifically, based on the second difference between the first training result and the second text, determine the second loss of the first model to be trained, where the second difference is positively correlated with the second loss. Based on the second loss, update the parameters of the first model to be trained until the second loss converges, stop updating the parameters of the first model to be trained, and use the first model to be trained as the first model.

[0152] Optionally, the first training data can be used for SFT, and the first model includes a model obtained by performing supervised fine-tuning (SFT) on a pre-trained model. Optionally, the pre-trained model is an LLM. Specifically, using the first training data to perform SFT on the pre-trained model can obtain the first model. For example, when the first task includes translating the text in the first language into the text in the second language, the first text in the first training data includes the text in the first language, the second text in the first training data includes the text in the second language, and the first model to be trained is the pre-trained model. After inputting the first text into the pre-trained model, the pre-trained model translates the text in the first language in the first text into the text in the second language to obtain a first training result. Based on the second difference between the first training result and the second text, determine the second loss of the pre-trained model. Based on the second loss, update the parameters of the pre-trained model until the second loss converges, stop updating the parameters of the pre-trained model, and use the pre-trained model as the first model.

[0153] Since the first similarity can be used to measure the accuracy of the second text, when the generating device determines that the accuracy of the second text is high based on the first similarity, generating the first training data based on the first text and the second text can improve the accuracy of the first training data.

[0154] In a possible implementation manner, because the first similarity is positively correlated with the accuracy of the second text, the first similarity is large when the accuracy of the second text is high. In the embodiments of the present application, the generating device determines whether the first similarity is large or small based on the first threshold, and further determines whether the accuracy of the second text is high or low. Specifically, when the first similarity is greater than or equal to the first threshold, it indicates that the first similarity is large, which also means that the accuracy of the second text is high. On the contrary, when the first similarity is less than the first threshold, it indicates that the first similarity is small, which also means that the accuracy of the second text is low. Therefore, when the first similarity is greater than or equal to the first threshold, the generating device generates the first training data based on the first text and the second text. Optionally, when the first similarity is less than the first threshold, the first training data is not generated based on the first text and the second text.

[0155] As an optional implementation manner, before executing step 105, the generating device determines whether the content of the first text and the content of the second text are reasonable based on the content quality rule and / or the content verification instruction, and obtains a judgment result.

[0156] In the embodiments of the present application, the content of the text includes the grammar of the text, the spelling of the words, and the semantics expressed by the words. Whether the content of the text is reasonable includes at least one of the following: the grammar of the text is correct, the spelling of the words is correct, the words are fluent, the semantics expressed by the words conforms to logic, and the logic of the text is clear.

[0157] The reasonableness of the content of the text indicates that the accuracy of the content of the text is high. Therefore, when the judgment result includes that the content of the first text and the content of the second text are both reasonable, generating the first training data based on the first similarity, the first text, and the second text can improve the accuracy of the first training data.

[0158] Optionally, the first task includes translating the text in the first language into the text in the second language. Since in the process of translating the text in the first language in the first text into the text in the second language to obtain the second text, it is possible to retain the text in the first language in the first text, which may lead to the second text including the text in the first language. Then, translating the text in the second language in the second text into the text in the first language to obtain the third text can make the third text retain the text in the first language in the first text, thereby making the first similarity between the first text and the third text large. However, the second text obviously has errors, that is, the accuracy of the second text is not high. For example, the first text includes: Hello, keyboard. The second text includes: XXX, keyboard. The third text includes: Hello, keyboard. Among them, the first language is Chinese, and XXX is the text in the second language.

[0159] Therefore, the content of the text in the embodiments of the present application may reasonably include that the grammar of the text is correct and the spelling of the words is correct. In this way, by detecting the content of the second text, it can be detected whether the second text includes the text in the first language, thereby improving the accuracy of the first similarity.

[0160] The content quality rule includes a rule for judging whether the content of the first text is reasonable. For example, the content quality rule includes that if the text includes the text in two or more languages, it is determined that the content of the text is unreasonable. Another example is that the content quality rule includes that if the format of the text is not the preset format, it is determined that the content of the text is unreasonable. Based on the content quality rule, the generating device can judge whether the content of the first text and the content of the second text are reasonable.

[0161] The content verification instruction includes an instruction for instructing the LLM to judge whether the content of the text is reasonable. Based on the content verification instruction, the generating device can instruct the LLM to judge whether the content of the text is reasonable. Optionally, the first language includes Chinese and the second language includes English. The generating device generates a first prompt word based on the first text and the second text, where the first prompt word includes a content verification instruction, and the content verification instruction is used to instruct the LLM to evaluate whether the first text and the second text are grammatically correct and spelled correctly.

[0162] For example, the first prompt word includes the following content:

[0163] You are an expert in evaluating text content. I will show you a Chinese sentence and its corresponding English translation. Please judge whether this pair of sentences has grammar and word spelling problems.

[0164] If there are no grammar and word spelling problems in the Chinese and English sentences, please output "No problem", otherwise output "There is a problem". Only output your judgment without any explanation.

[0165] Chinese sentence: {zh_sent}

[0166] English sentence: {en_sent}

[0167] Please give your judgment. Again, only output "no problem" or "problem", without giving any explanation.

[0168] In the first prompt, zh_sent represents the Chinese sentence in the first text, and en_sent represents the English translation in the second text.

[0169] In the embodiment of the present application, after obtaining the first text, the generating device performs a first task based on the first text to obtain a second text. Then, a second task is performed based on the second text to obtain a third text. Among them, the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by performing the first task based on the original text. Then, the first similarity between the first text and the third text is determined. At this time, the greater the first similarity, the more similar the content of the first text is to the content of the third text. Also, because the third text is determined based on the second text, the similarity between the content of the third text and the content of the second text is high. Therefore, the similarity between the content of the second text and the content of the first text is high. This also means that the second text has a high accuracy as the execution result obtained by performing the first task based on the first text. That is, the first similarity is positively correlated with the accuracy of the second text as the execution result of performing the first task based on the first text, and the first similarity can measure the accuracy of the second text. Since the first similarity can measure the accuracy of the second text, it can be determined whether the second text can be used to supervise the execution result obtained by the first model performing the first task based on the first text based on the first similarity. Therefore, the first training data can be generated based on the first similarity, the first text, and the second text, thereby reducing the labor cost.

[0170] As an alternative implementation, the generating device realizes "performing a first task based on the first text to obtain a second text" by performing the following steps: performing the first task based on the first processing method and the first text to obtain a second text.

[0171] In the embodiments of the present application, the first processing method is a method for executing a first task. For example, the first task includes translating the text in the first language into the text in the second language. The first processing method may be rule-based machine translation (RBMT), or may be statistical machine translation (SMT), or may also be translation using a recurrent neural network (RNN), or may also be a long-short term memory network (LSTM). In a possible implementation, the first processing method includes a first task model for executing the first task, where the first task model is a deep learning model. For example, the first task includes translating the text in the first language into the text in the second language, and the first task model may be an RNN or an LSTM. Another example is that the first task includes a dialogue task, and the first task model may be an LLM capable of executing the dialogue task.

[0172] The generating device executes the first task based on the first processing method and the first text, and may execute the first task based on the first text according to the first processing method to obtain a second text. For example, the first task includes translating the text in the first language into the text in the second language, and the first processing method includes a first task model. The generating device may input the first text into the first task model to obtain the second text obtained by the first task model executing the first task based on the first text.

[0173] Optionally, during the execution of step 103, the generating device performs the following steps: executes a second task based on the first processing method and the second text to obtain a third text. That is, the generating device executes the first task and the second task based on the same processing method. Optionally, the first processing method includes a first model, where the first model has the ability to execute the first task and the second task. For example, the first task includes translating the text in the first language into the text in the second language, and the second task includes translating the text in the second language into the text in the first language. Accordingly, the first model has the ability to translate the text in the first language into the text in the second language and the ability to translate the text in the second language into the text in the first language.

[0174] In this implementation, the generating device further performs the following steps: performing a first task based on a second processing method different from the first processing method and a first text to obtain a fourth text; performing a second task based on the fourth text to obtain a fifth text; determining a second similarity between the first text and the fifth text, where the second similarity is positively correlated with the accuracy of the fourth text.

[0175] In the embodiments of the present application, the second processing method is also a method for performing the first task, and the first processing method is different from the second processing method. For example, the first task includes translating the text in the first language into the text in the second language. The first processing method may be RBMT, and the second processing method may also be SMT. In a possible implementation, the first processing method includes a first task model for performing the first task, and the second processing method includes a second task model for performing the first task. Among them, both the first task model and the second task model are deep learning models, and the first task model is different from the second task model. For example, the first task includes translating the text in the first language into the text in the second language. The first task model may be an RNN, and the second task model may be an LSTM. Another example is that the first task includes translating the text in the first language into the text in the second language. The model structures of the first task model and the second task model are the same, but the parameters of the first task model and the parameters of the second task model are different.

[0176] After performing the first task based on the second processing method and the first text to obtain the fourth text, perform the second task based on the fourth text to obtain the fifth text. Then determine the second similarity between the first text and the fifth text. The greater the second similarity, the more similar the content of the first text is to the content of the fifth text. Also, because the fifth text is determined based on the fourth text, the similarity between the content of the fifth text and the content of the fourth text is high. Therefore, the similarity between the content of the fourth text and the content of the first text is high. That is to say, the accuracy of the fourth text as the execution result obtained by performing the first task based on the first text is high. That is, the second similarity can be used to measure the accuracy of the fourth text. Specifically, the second similarity is positively correlated with the accuracy of the fourth text.

[0177] Optionally, the generating device performs the second task based on the second processing method and the fourth text to obtain the fifth text. That is, the generating device performs the first task and the second task based on the same processing method.

[0178] After obtaining the second similarity, the generating device further performs the following steps: generating second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text.

[0179] In the embodiments of the present application, the second training data is used to optimize the training of the first model, where the optimization training is different from the training of the first model based on the first training data. Optionally, the second training data is used to train the first model so that the style of the text generated by the first model meets the first preset requirement. In other words, through the optimization training, the style of the text generated by the first model can meet the first preset requirement.

[0180] In a possible implementation manner, a second model is trained based on the second training data, where the second model is used to evaluate the quality of the execution result of the first task, and the quality of the execution result of the first task includes the accuracy of the execution result of the first task and whether the style of the execution result of the first task meets the first preset requirement. Optionally, the second model includes a reward model (RM). Then, based on the second model, the first model is trained, and the style of the text generated by the first model can meet the first preset requirement.

[0181] Since the execution result of the first task is text, whether the style of the execution result of the first task meets the first preset requirement is whether the style of the text obtained by executing the first task meets the first preset requirement. For example, the style of the text includes a lively style, a humorous style, an elegant style, a rigorous style, a sharp style, and an Internet style.

[0182] In a possible implementation manner, when the style of the text obtained by executing the first task is the first preset style, the style of the text obtained by executing the first task meets the first preset requirement. On the contrary, when the style of the text obtained by executing the first task is not a lively style, the style of the text obtained by executing the first task does not meet the first preset requirement. For example, the first task includes translation, and the first preset style is a lively style. The style of the translated text is a lively style, indicating that the style of the text obtained by executing the first task meets the first preset requirement. On the contrary, the style of the translated text is not a lively style, indicating that the style of the text obtained by executing the first task does not meet the first preset requirement. For example, for the two sentences "You are not my type" and "You are not my cup of tea", the style of "You are not my type" is not a lively style, and the style of "You are not my cup of tea" is a lively style.

[0183] In another possible implementation manner, when the style of the text obtained by executing the first task is the second preset style, the style of the text obtained by executing the first task meets the first preset requirement. For example, the second preset style is a lively style, and the style of the text obtained by executing the first task is a lively style, indicating that the style of the text obtained by executing the first task meets the first preset requirement. On the contrary, the style of the text obtained by executing the first task is not a lively style, indicating that the style of the text obtained by executing the first task does not meet the first preset requirement.

[0184] Optionally, the second training data includes an original text and execution results of two different first tasks, where the execution results of the two different first tasks are both obtained by performing the first task based on the original text. In this way, when the quality of the execution results of the two different first tasks is different, training the second model based on the second training data enables the second model to learn the ability to evaluate the quality of the execution results of the first task by distinguishing the quality of the execution results of different first tasks. When the quality of the execution results of the two different first tasks is the same, training the second model based on the second training data enables the second model to learn the ability to evaluate the quality of the execution results of the first task by identifying the execution results of the first tasks with the same quality. Since the second text and the fourth text are both obtained by performing the first task based on the first text, the generating device can generate the second training data based on the first text, the second text, and / or the fourth text. In a possible implementation manner, the generating device can generate the second training data based on the first text and the second text. In another possible implementation manner, the generating device can generate the second training data based on the first text and the fourth text. In yet another possible implementation manner, the generating device can generate the second training data based on the first text, the second text, and the fourth text.

[0185] As a possible implementation manner, when the generating device executes the process of "generating the second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text", the following steps are performed: The generating device determines the quality of the second text based on the first similarity and whether the style of the second text meets the first preset requirement. The generating device determines the quality of the fourth text based on the second similarity and whether the style of the fourth text meets the first preset requirement. When the quality of the second text is higher than the quality of the fourth text, the generated second training data includes the first text, the second text, and the fourth text. In the second training data, the second text is the execution result obtained by performing the first task based on the first text, and the fourth text is not the execution result obtained by performing the first task based on the first text.

[0186] Optionally, the generating device determines whether the style of the text meets the first preset requirement through a style evaluation model. Optionally, after a relevant person determines whether the style of the text meets the first preset requirement through manual annotation, the annotation result is input into the generating device so that the generating device determines whether the style of the text meets the first preset requirement.

[0187] Optionally, when the style of the text meets the first preset requirement, the generation device determines that the style of the text is divided into a first preset value. When the style of the text does not meet the first preset requirement, the generation device determines that the style of the text is divided into a second preset value, where the first preset value is greater than the second preset value and less than 1. The style score and the similarity used to measure the accuracy of the text are weighted and summed to determine the quality of the text. For example, the first preset value is 1 and the second preset value is 0. If the style of the second text meets the first preset requirement, the style score of the second text is 1. If the first similarity is 0.8, then by weighted summing 1 and 0.8, the quality of the second text can be determined.

[0188] Optionally, the generation device determines the quality of the text based on a preset mapping relationship, the similarity used to measure the accuracy of the text, and whether the style of the text meets the first preset requirement, where the preset mapping includes: when the style of the text meets the first preset requirement, the mapping relationship between the similarity used to measure the accuracy of the text and the quality of the text; when the style of the text does not meet the first preset requirement, the mapping relationship between the similarity used to measure the accuracy of the text and the quality of the text.

[0189] In this implementation, the second training data includes the original text of a first task (i.e., the first text), the execution result of a first task with high quality (i.e., the second text), and the execution result of a first task with low quality (i.e., the fourth text). Since in the second training data, the second text is the execution result obtained by executing the first task based on the first text, and the fourth text is not the execution result obtained by executing the first task based on the first text, training the second model based on the second training data enables the second model to learn the ability to distinguish between the execution results of high-quality first tasks and low-quality first tasks, and further enables the second model to have the ability to evaluate the quality of the execution results of the first task.

[0190] As an alternative implementation, during the process of the generating device executing "generating second training data based on the first text, the second text, and the fourth text", the following steps are performed: When the quality of the second text and the quality of the fourth text both do not meet the second preset requirement, obtain a sixth text, where the sixth text includes the execution result obtained by performing a first task based on the first text, and the quality of the sixth text meets the second preset requirement. The generated second training data includes the first text, the second text, and the sixth text. In the second training data, the sixth text is the execution result obtained by performing the first task based on the first text, and the second text is not the execution result obtained by performing the first task based on the first text. And / or, the generated second training data includes the first text, the fourth text, and the sixth text. In the second training data, the sixth text is the execution result obtained by performing the first task based on the first text, and the fourth text is not the execution result obtained by performing the first task based on the first text.

[0191] In the embodiments of the present application, that the quality of the text does not meet the second preset requirement indicates that the quality of the text is low. Therefore, when the quality of the second text and the quality of the fourth text both do not meet the second preset requirement, the generating device obtains the sixth text. Then, based on the first text, the second text, and the sixth text, the second training data is generated. Or, based on the first text, the fourth text, and the sixth text, the second training data is generated. Or, both based on the first text, the second text, and the sixth text, the second training data is generated, and based on the first text, the fourth text, and the sixth text, the second training data is generated. At this time, the second training data includes two sets of data, one set includes the first text, the second text, and the sixth text, and the other set includes the first text, the fourth text, and the sixth text.

[0192] In this implementation, the second training data includes the original text of the first task (i.e., the first text), the execution result of the first task with high quality (i.e., the sixth text), and the execution result of the first task with low quality (i.e., the second text or the fourth text). Training the second model based on the second training data can enable the second model to learn the ability to distinguish between the execution results of the first task with high quality and the execution results of the first task with low quality, and further enable the second model to have the ability to evaluate the quality of the execution results of the first task.

[0193] As an alternative implementation, during the process of the generating device executing step 105, the following steps are performed: Generate first training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text.

[0194] In the embodiments of the present application, the first training data includes the first text and the second text, or the first training data includes the first text and the fourth text. When the first training data includes the first text and the second text, the second text is used to supervise the execution result obtained by the first model executing the first task based on the first text. When the first training data includes the first text and the fourth text, the fourth text is used to supervise the execution result obtained by the first model executing the first task based on the first text.

[0195] In a possible implementation manner, the first training data generated by the generating device based on the first text and the second text includes the first text and the second text. In another possible implementation manner, the first training data generated by the generating device based on the first text and the fourth text includes the first text and the fourth text. In yet another possible implementation manner, the first training data generated by the generating device based on the first text, the second text, and the fourth text may include the first text and the second text, or may include the first text and the fourth text.

[0196] Optionally, the first training data is generated based on the magnitude relationship between the first similarity and the second similarity, the first text, the second text, and the fourth text.

[0197] Since the first similarity can be used to measure the accuracy of the second text, and the second similarity can be used to measure the accuracy of the fourth text, the generating device can determine the magnitude relationship between the accuracy of the second text and the accuracy of the fourth text based on the magnitude relationship between the first similarity and the second similarity, and then can generate the first training data based on the first text, the second text, and the fourth text.

[0198] In a possible implementation manner, the magnitude relationship between the first similarity and the second similarity includes that the first similarity and the second similarity differ greatly. This indicates that for executing the first task based on the first text, the difference in the accuracy of the execution results obtained by different processing methods is relatively large, which also means that for executing the first task, the first text is a hard sample. Since among the second text and the fourth text, the text corresponding to the larger value of the first similarity and the second similarity has a higher accuracy, the generating device can generate the first training data based on the first text and the text corresponding to the larger value of the first similarity and the second similarity. In this way, the first training data can include the hard sample and the ground truth (GT) of the hard sample, where the GT of the hard sample includes the text corresponding to the larger value of the first similarity and the second similarity. In this case, when the first model is trained based on the first training data, the first model can learn how to execute the first task based on the hard sample based on the hard sample and the GT of the hard sample, and thus can improve the accuracy of the processing result obtained by the first model executing the first task based on the hard sample.

[0199] In the embodiment of the present application, the generating device determines whether the difference between the first similarity and the second similarity is large or small based on the second threshold. Specifically, if the difference between the first similarity and the second similarity is greater than the second threshold, it indicates that the difference between the first similarity and the second similarity is large. Conversely, if the difference between the first similarity and the second similarity is less than or equal to the second threshold, it indicates that the difference between the first similarity and the second similarity is small.

[0200] Optionally, the size relationship between the first similarity and the second similarity includes that the first similarity is greater than the second similarity. Then, among the second text and the fourth text, the text corresponding to the larger value of the first similarity and the second similarity is the second text. Correspondingly, the generating device generates the first training data based on the first text and the second text. Optionally, the size relationship between the first similarity and the second similarity includes that the second similarity is greater than the first similarity. Then, among the second text and the fourth text, the text corresponding to the larger value of the first similarity and the second similarity is the fourth text. Correspondingly, the generating device generates the first training data based on the first text and the fourth text.

[0201] In another possible implementation manner, the size relationship between the first similarity and the second similarity includes that the difference between the first similarity and the second similarity is small. This indicates that for performing the first task based on the first text, the difference in the accuracy of the execution results obtained based on different processing methods is small, which also means that for performing the first task, the first text is an easy sample. Since among the second text and the fourth text, the text corresponding to the larger value of the first similarity and the second similarity is the text with higher accuracy, the generating device can generate the first training data based on the first text and the text corresponding to the larger value of the first similarity and the second similarity. In this way, the first training data can include easy samples and the GT of the easy samples, where the GT of the easy samples includes the text corresponding to the larger value of the first similarity and the second similarity. In this case, when the first model is trained based on the first training data, the first model can learn how to perform the first task based on the easy samples and the GT of the easy samples, thereby improving the accuracy of the processing results obtained by the first model when performing the first task based on the easy samples.

[0202] Optionally, the magnitude relationship between the first similarity and the second similarity includes that the first similarity is greater than the second similarity. Then, among the second text and the fourth text, the text corresponding to the larger value of the first similarity and the second similarity is the second text. Correspondingly, the generating device generates first training data based on the first text and the second text. Optionally, the magnitude relationship between the first similarity and the second similarity includes that the second similarity is greater than the first similarity. Then, among the second text and the fourth text, the text corresponding to the larger value of the first similarity and the second similarity is the fourth text. Correspondingly, the generating device generates first training data based on the first text and the fourth text.

[0203] The embodiments of the present application also provide another method for generating training data. The first task in this method for generating training data includes translating the text in the first language into the text in the second language. The execution subject of this method for generating training data can be the generating device described above. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another method for generating training data provided by the embodiments of the present application.

[0204] 201. Obtain a first text, where the first text includes the text in the first language.

[0205] 202. Translate the text in the first language in the first text into the text in the second language to obtain a second text.

[0206] Optionally, the generating device executes step 202 based on a first model, where the first model includes an LLM.

[0207] 203. Translate the text in the second language in the second text into the text in the first language to obtain a third text.

[0208] Optionally, the first model has both the ability to translate the text in the first language into the text in the second language and the ability to translate the text in the second language into the text in the first language. The generating device can execute step 203 based on the first model.

[0209] 204. Determine the first similarity between the first text and the third text.

[0210] In the embodiments of the present application, the first similarity is positively correlated with the accuracy of the second text.

[0211] 205. Generate first training data based on the first similarity, the first text, and the second text.

[0212] In the embodiments of the present application, the first training data includes a first text and a second text. In the first training data, the second text is used to supervise the text obtained by the first model translating the text in the first language in the first text into the text in the second language. The implementation of this step can refer to the implementation of step 105 and will not be elaborated here.

[0213] In the embodiments of the present application, after obtaining the first text, the generating device translates the text in the first language in the first text into the text in the second language to obtain a second text. Then, the text in the second language in the second text is translated into the text in the first language to obtain a third text. Then, the first similarity between the first text and the third text is determined. At this time, the greater the first similarity, the more similar the content of the first text is to the content of the third text. Also, since the third text is obtained by translating the second text, the similarity between the content of the third text and the content of the second text is high. Therefore, the similarity between the content of the second text and the content of the first text is high. That is to say, the accuracy of the second text as the translation result of the first text is high, that is, the first similarity is positively correlated with the accuracy of the second text as the translation result of the first text, and the first similarity can measure the accuracy of the second text. Since the first similarity can measure the accuracy of the second text, it can be determined whether the second text can be used to supervise the translation result obtained by the first model translating the text in the first language in the first text into the text in the second language based on the first similarity. Therefore, the first training data can be generated based on the first similarity, the first text, and the second text, thereby reducing the labor cost.

[0214] As an optional implementation manner, the generating device translates the text in the first language in the first text into the text in the second language based on a first processing manner to obtain a second text. Optionally, the first processing manner includes a first model, that is, the generating device translates the text in the first language in the first text into the text in the second language based on the first model to obtain a second text.

[0215] In this implementation manner, the generating device further performs the following steps: Translate the text in the first language in the first text into the text in the second language based on a second processing manner different from the first processing manner to obtain a fourth text. Optionally, the second processing manner includes a second model. Translate the text in the second language in the fourth text into the text in the first language to obtain a fifth text. Determine the second similarity between the first text and the fifth text. The second similarity is positively correlated with the accuracy of the fourth text. Generate second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text, where the second training data is used to train the first model so that the translation result generated by the first model conforms to the first preset requirement.

[0216] Please refer to Figure 3 , Figure 3The flowchart shows another method for generating training data provided by an embodiment of this application. As Figure 3 shown, after obtaining the first text, the generating device performs translation based on the first processing method to obtain the second text. Specifically, based on the first processing method, the text in the first language in the first text is translated into the text in the second language to obtain the second text. Optionally, the first processing method includes a first model, and the generating device uses the first model to translate the text in the first language in the first text into the text in the second language to obtain the second text. Then, by back-translating the second text, the third text is obtained. Specifically, the text in the second language in the second text is translated into the text in the first language to obtain the third text. Optionally, the generating device uses the first model to translate the text in the second language in the second text into the text in the first language to obtain the third text. After obtaining the third text, the first similarity between the first text and the third text is determined. Optionally, the generating device determines the BLEU between the first text and the third text to obtain the first similarity.

[0217] In addition, after obtaining the first text, the generating device performs translation based on the second processing method to obtain the fourth text. Specifically, based on the second processing method, the text in the first language in the first text is translated into the text in the second language to obtain the fourth text. Optionally, the second processing method includes a second model, and the generating device uses the second model to translate the text in the first language in the first text into the text in the second language to obtain the fourth text. Then, by back-translating the fourth text, the fifth text is obtained. Specifically, the text in the second language in the fourth text is translated into the text in the first language to obtain the fifth text. Optionally, the generating device uses the second model to translate the text in the second language in the fourth text into the text in the first language to obtain the fifth text. After obtaining the fifth text, the second similarity between the first text and the fifth text is determined. Optionally, the generating device determines the BLEU between the first text and the fifth text to obtain the second similarity.

[0218] If the first similarity - the second similarity is greater than the second threshold, then the first training data is generated based on the first text and the second text, and the second training data is generated based on the first text, the second text, and the fourth text. In the first training data, the first text is the input, and the second text is the label, that is, the second text is used to supervise the text obtained by the first model translating the text in the first language in the first text into the second language. In the second training data, the first text is the input, the second text is the chosen option, and the fourth text is the rejected option, that is, the second text is the text obtained by translating the text in the first language in the first text into the second language, and the fourth text is not the text obtained by translating the text in the first language in the first text into the second language.

[0219] Optionally, the second training data can be generated based on the text quality of the second text and the text quality of the fourth text. Specifically, based on the text quality of the second text and the text quality of the fourth text, the second text and the fourth text can be divided into the following four cases:

[0220] 1. If the text quality of the second text is higher than that of the fourth text, then mark the first text, the second text, and the fourth text as good (G).

[0221] 2. If the text quality of the second text is lower than that of the fourth text, then mark the first text, the second text, and the fourth text as bad (B).

[0222] 3. If the text quality of the second text and the text quality of the fourth text are not much different, then mark the first text, the second text, and the fourth text as the same (S).

[0223] 4. If the text quality of the second text and the text quality of the fourth text both do not meet the second preset requirement, then mark the first text, the second text, and the fourth text as all bad (A).

[0224] For the cases marked as G or B, a set of second training data can be generated based on the first text, the second text, and the fourth text. Specifically, in the case marked as G, the first text can be determined as the input in the second training data, the second text can be determined as the option in the second training data, and the fourth text can be determined as the rejection in the second training data. In the case marked as B, the first text can be determined as the input in the second training data, the fourth text can be determined as the option in the second training data, and the second text can be determined as the rejection in the second training data.

[0225] For the case marked as S, the second training data is not generated based on the first text, the second text, and the fourth text.

[0226] For the case marked as A, after obtaining the sixth text whose text quality meets the second preset requirement, two sets of second training data can be generated based on the first text, the second text, the fourth text, and the sixth text. Specifically, after determining the sixth text based on the first text, the first text can be determined as the input in the second training data, the sixth text can be determined as the option in the second training data, and the second text can be determined as the rejection in the second training data. And the first text can be determined as the input in the second training data, the sixth text can be determined as the option in the second training data, and the fourth text can be determined as the rejection in the second training data.

[0227] Optionally, generating the second training data based on the text quality of the second text and the text quality of the fourth text can be achieved through manual annotation.

[0228] Optionally, determine whether the style of the text meets the first preset requirement by manual annotation. For example, when the generating device determines the first similarity for measuring the accuracy of the second text, determine whether the style of the second text meets the first preset requirement by manual annotation, and then determine the quality of the second text.

[0229] Optionally, when the quality of the second text and the quality of the fourth text do not meet the second preset requirement, determine the sixth text based on the first text by manual annotation. For example, the first task includes translating the text in the first language into the second language. Correspondingly, both the second text and the fourth text are translation results of the first text. If the quality of the second text and the quality of the fourth text do not meet the second preset requirement, then the translation result of the first text can be determined by manual annotation as the sixth text.

[0230] By Figure 3 The first training data and the second training data can be obtained through the training generation method shown, and subsequently, the first model can be trained based on the first training data, and the second model can be trained based on the second training data.

[0231] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a training method provided by an embodiment of the present application. Through this training method, the first model can be trained using the first training data. The execution subject of this training method is a training device, where the training device can be any electronic device that can execute the technical solutions disclosed in the method embodiments of the present application. Optionally, the training device can be one of the following: a computer, a server.

[0232] 401. Obtain the first training data generated by the training data generation method based on the foregoing.

[0233] The first training data generated by the training data generation method described above includes an input and a label, where the input is the data input to the model during the training process, and the label is used to supervise the data output by the model. For example, in Figure 3 the training data generation method shown, the input of the first training data includes the first text, and the label of the first training data includes the second text.

[0234] 402. Train the first model based on the first training data, where the first model is used to execute the first task.

[0235] In a possible implementation manner, during the training process, the first model to be trained determines to perform a first task based on the input in the first training data, and obtains a third probability of the label. Based on the third probability, a second loss of the first model to be trained is determined, where the third probability is positively correlated with the second loss. Based on the second loss, the parameters of the first model to be trained are updated until the second loss converges, and then the update of the parameters of the first model to be trained is stopped, and the first model to be trained is used as the first model.

[0236] Optionally, the first model to be trained is a pre-trained model pre-trained based on the third training data, where the third training data also includes an input and a label, and the third training data may not be generated based on the training data generation method described above.

[0237] Optionally, the input in the first training data includes text of a target type. For example, the target type includes comments on documents on the Internet platform, the body of the document, and the title of the document.

[0238] Optionally, the generation device determines that the second loss of the first model to be trained is represented by the following formula:

[0239] Loss sft =-P θ (y|x)…Formula (1)

[0240] where x is the input in the first training data, y is the label in the first training data, P θ (y|x) is the third probability that the first model to be trained determines to perform the first task based on the input in the first training data and obtains the label, and θ is the parameter of the first model to be trained.

[0241] In the embodiments of the present application, after the training device obtains the first training data, based on the first training data, the first model is trained, so that the first model can have the ability to perform the first task.

[0242] Benefiting from its powerful performance, the application of LLM is becoming more and more extensive. For example, intelligent question answering is realized through LLM, translation is performed through LLM, and abstracts of text are extracted through LLM. Before applying LLM, it is necessary to train LLM. Among them, the process of training LLM can be realized through reinforcement learning from human feedback (RLHF). RLHF includes: obtaining a pre-trained model by pre-training a basic language model, obtaining an SFT model by performing SFT on the pre-trained model, obtaining an RM based on a supervised fine-tuning model, and performing reinforcement learning on the SFT model based on the RM to obtain an LLM model.

[0243] And through Figure 4The first model trained by the training method shown above can be used as an SFT model, and the second model trained based on the second training data obtained from the previous text can be used as an RM model. Therefore, the embodiments of the present application also provide another training method. Through this training method, the first model can be strengthened and learned based on the second model to obtain an LLM model (hereinafter referred to as the third model). Among them, the third model has the ability to execute the first task, and the style of the text obtained by executing the first task meets the first preset requirement.

[0244] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of another training method provided by the embodiments of the present application. Through this training method, the third model can be trained. The execution subject of this training method is a training device. Among them, the training device can be any electronic device that can execute the technical solutions disclosed in the method embodiments of the present application. Optionally, the training device can be one of the following: a computer, a server. Optionally, the training device that executes Figure 4 the training method shown above and the training device that executes Figure 5 the training method shown above can be different devices.

[0245] 501. Obtain the first model and the second model trained by the training method described above.

[0246] In the embodiments of the present application, the first model is trained by the training method shown in Figure 4 . The second model is used to evaluate the quality of the execution result of the first task. The quality of the execution result of the first task includes the accuracy of the execution result of the first task and whether the style of the execution result of the first task meets the first preset requirement. The second model is trained based on the second training data, and the second training data is generated based on the method of any one of claims 2 to 6.

[0247] In one implementation manner of obtaining the second model, obtain the second training data generated based on the training data generation method described above. Based on the second training data, train the second model.

[0248] The second training data generated by the training data generation method described above includes an input, an option, and a rejection item. Among them, the input is the data input to the model during the training process, the option is the execution result obtained by the model executing the first task based on the input, and the rejection item is not the execution result obtained by the model executing the first task based on the input. For example, in Figure 3 the training data generation method shown, the input of the second training data includes the first text, the option of the second training data includes the second text, and the rejection item of the second training data includes the fourth text.

[0249] In a possible implementation manner, during the training process, the second model to be trained determines the quality of the option as the execution result of the first task, and obtains a first evaluation result, where the second model to be trained is different from the first model to be trained. The second model to be trained determines the quality of the rejection item as the execution result of the first task, and obtains a second evaluation result. Based on the magnitude relationship between the first evaluation result and the second evaluation result, the third loss of the second model to be trained is determined. Based on the third loss, the parameters of the second model to be trained are updated until the third loss converges, and the update of the parameters of the second model to be trained is stopped, and the second model to be trained is used as the second model.

[0250] Optionally, the generating device first determines the fourth probability that the quality of the option as the execution result of the first task is higher than the quality of the rejection item as the execution result of the first task through the following formula:

[0251]

[0252] where P(i beats j) represents the fourth probability, i represents the option, and j represents the option. represents taking the first evaluation result as the exponent of the natural number e, represents taking the second evaluation result as the exponent of the natural number e.

[0253] Then, based on the fourth probability, the Bradley-Terry (BT) loss is determined as the third loss.

[0254] Optionally, the determination of the Bradley-Terry (BT) loss is represented by the following formula:

[0255] L = -∑ (i,j) [y ij log(P(i beats j)) + (1 - y ij )log(1 - P(i beats j))]… Formula (3)

[0256] where ∑ represents summation, i represents the option, and j represents the rejection item. y ij represents the label of the fourth probability, where, when i represents the option and j represents the rejection item, y ij = 1. P(i beats j) represents the fourth probability. log(·) represents the logarithmic function with base 10.

[0257] After determining the third loss based on Formula (3), updating the parameters of the second model to be trained based on the third loss can make the fourth probability approach the label y of the fourth probability ij。In the case where the third loss converges, using the second model to be trained as the second model enables the second model to evaluate the accuracy of distinguishing the quality of the execution result of the first task.

[0258] 502. Based on the second model, train the first model to obtain a third model, where the third model is used to execute the first task, and the style of the execution result obtained by executing the first task conforms to the first preset requirement.

[0259] In a possible implementation, obtain a first text and at least one candidate text, where the candidate text is obtained by executing the first task based on the first text. Based on the second model, determine a conforming text whose style conforms to the first preset requirement and a non-conforming text whose style does not conform to the first preset requirement from the at least one candidate text. Based on the first model, determine a first probability that the conforming text is the result of executing the first task based on the first text, and a second probability that the non-conforming text is the result of executing the first task based on the first text. Based on the first difference between the first probability and the second probability, determine the first loss of the first model, where the first difference is negatively correlated with the first loss. Based on the first loss, update the parameters of the first model to obtain the third model.

[0260] Optionally, during the training process, the training device can set different hyperparameters for the first model so that the first model obtains at least one candidate text based on the first text, where when the number of candidate texts is greater than 1, any two candidate texts are different.

[0261] Optionally, the second model can evaluate the quality of each candidate text to obtain at least one candidate evaluation result. Then, based on the at least one evaluation result, a conforming text whose style conforms to the first preset requirement and a non-conforming text whose style does not conform to the first preset requirement can be determined from the at least one candidate text. It should be understood that the number of conforming texts can be greater than 1, and the number of non-conforming texts can also be greater than 1.

[0262] Optionally, after determining the conforming text and the non-conforming text, generate fourth training data based on the conforming text, the non-conforming text, and a second prompt word, where the second prompt word is used to instruct the first model to execute the first task based on the first text.

[0263] For example, the second prompt word includes the following content:

[0264] {

[0265] Second prompt word: Translate the text in the first language in the first text into the text in the second language.

[0266] Option: Conforming text.

[0267] Rejection item: Non-conforming text.

[0268] }。

[0269] Training the first model based on the fourth training data enables the first model to determine a first probability that the text conforms to the result of performing the first task based on the first text, and a second probability that the text does not conform to the result of performing the first task based on the first text. Furthermore, the first loss can be determined based on the first probability and the second probability, and the parameters of the first model can be updated based on the first loss to obtain the third model.

[0270] Optionally, training the first model based on the fourth training data can be achieved by the direct preference optimization (DPO) method.

[0271] Specifically, let π θ be the first model, and let the probability that the y determined by the first model is the result of performing the first task based on x be: r θ (x,y) = logπ θ (y|x) … Formula (4).

[0272] The first probability that the first model determines that the text conforms to the result of performing the first task based on the first text is represented by the following formula:

[0273]

[0274] The second probability that the first model determines that the text does not conform to the result of performing the first task based on the first text is represented by the following formula:

[0275]

[0276] Then, the first loss determined based on DPO is represented by the following formula:

[0277]

[0278] where L(θ) represents the first loss, represents the expectation, x represents the first text, y + represents the text that conforms, y - represents the text that does not conform. log(·) represents the logarithmic function with base 10. β is the temperature hyperparameter. represents the first probability. represents the second probability.

[0279] In Formula (7), if is greater than or equal to the fourth threshold, it indicates that This indicates that the first model can accurately determine that the compliant text is suitable as the execution result and the non-compliant text is not suitable as the execution result. At this time, the first loss is small. If is less than the fourth threshold, it indicates that this means that the first model cannot accurately determine which of the compliant text and the non-compliant text is suitable as the execution result. At this time, the first loss is large.

[0280] Updating the parameters of the first model based on the first loss can make greater than or equal to the fourth threshold, so that the first model can accurately determine that the compliant text is suitable as the execution result and the non-compliant text is not suitable as the execution result.

[0281] In the embodiments of the present application, after obtaining the first model and the second model, the training device trains the first model based on the second model to obtain a third model. Thus, the third model can be made to have the ability to execute the first task, and the style of the execution result obtained by the third model executing the first task conforms to the first preset requirement.

[0282] After training the third model through the training method shown in Figure 5 executing the first task based on the third model to obtain the execution result of the first task can improve the accuracy of the execution result of the first task, and the style of the execution result of the first task conforms to the first preset requirement. In this way, the style of the execution result obtained by executing the first task can be made more in line with human preferences. For example, if human preferences include texts in a lively style, then by executing the first task through the third model, the obtained execution result is a text in a lively style.

[0283] Based on this, the embodiments of the present application further provide a text processing method. The execution subject of the text processing method is a text processing device, where the text processing device can be any electronic device that can execute the technical solutions disclosed in the method embodiments of the present application. Optionally, the text processing device can be one of the following: a computer, a server. Please refer to Figure 6 , Figure 6 which is a schematic flowchart of a text processing method provided by the embodiments of the present application.

[0284] 601. Obtain the text to be processed and the third model trained by the training method described above.

[0285] Specifically, the third model is trained by the training method shown in Figure 5 .

[0286] 602. Execute the first task based on the third model and the text to be processed to obtain a target text, where the style of the target text conforms to the first preset requirement.

[0287] In the embodiments of the present application, after obtaining the third model, the text processing device performs a first task based on the third model and the text to be processed, and obtains the target text, which can improve the accuracy of the target text and make the style of the target text meet the first preset requirement.

[0288] In a possible scenario, the text to be processed includes documents on an Internet platform, and the first task includes translating the documents on the Internet platform. Specifically, the text in the first language in the document is translated into the text in the second language, and it is required that the style of the translation result is lively. Then, in the case where the style of the text meets the first preset requirement, that is, the style of the text is a lively style, the text in the first language in the text to be processed is translated into the text in the second language based on the third model to obtain the target text, which can improve the accuracy of the target text, that is, improve the accuracy of the translation result, and moreover, the style of the target text can be matched with the lively style.

[0289] In another possible scenario, the text to be processed includes the speech in a conversation, and the first task includes a conversation task, and it is required that the style of the reply in the conversation is lively. Then, in the case where the style of the text meets the first preset requirement, that is, the style of the text is a lively style, the reply to the text to be processed is determined based on the third model to obtain the target text, which can improve the accuracy of the target text, that is, improve the accuracy of the reply, and moreover, the style of the target text can be matched with the lively style.

[0290] In still another possible scenario, the first task includes poem writing, and the text to be processed includes the theme of the poem and the requirements of the poem. For example, the theme of the poem is spring, and the requirements of the poem include a bold style. Then, in the case where the style of the text meets the first preset requirement, that is, the style of the text is a bold style, the text to be processed is input into the third model, and the third model can generate a poem related to spring based on the text to be processed, and the generated poem is in a bold style.

[0291] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0292] The above details the method of the embodiments of the present application. Below, the device of the embodiments of the present application is provided.

[0293] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a training data generation device provided by the embodiments of the present application. The training data generation device 1 includes: an acquisition unit 11, an execution unit 12, a determination unit 13, and a generation unit 14, where:

[0294] The acquisition unit 11 is configured to acquire the first text;

[0295] Execution unit 12, configured to perform a first task based on the first text to obtain a second text;

[0296] The execution unit 12 is further configured to perform a second task associated with the first task based on the second text to obtain a third text, where the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by performing the first task based on the original text;

[0297] Determination unit 13, configured to determine a first similarity between the first text and the third text;

[0298] Generation unit 14, configured to generate first training data based on the first similarity, the first text, and the second text, where the first training data includes the first text and the second text, and in the first training data, the second text is used to supervise the execution result obtained by the first model performing the first task based on the first text.

[0299] Combined with any embodiment of the present application, the execution unit 12 is further configured to:

[0300] Perform the first task based on a first processing method and the first text to obtain the second text;

[0301] Perform the first task based on a second processing method different from the first processing method and the first text to obtain a fourth text;

[0302] Perform the second task based on the fourth text to obtain a fifth text;

[0303] The determination unit 13 is further configured to determine a second similarity between the first text and the fifth text;

[0304] The generation unit 14 is further configured to generate first training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text.

[0305] Combined with any embodiment of the present application, the execution unit 12 is further configured to perform the second task based on the first processing method and the second text to obtain the third text;

[0306] Perform the second task based on the second processing method and the fourth text to obtain the fifth text.

[0307] Combined with any embodiment of the present application, the generation unit 14 is further configured to:

[0308] Generate the first training data based on the magnitude relationship between the first similarity and the second similarity, the first text, the second text, and the fourth text;

[0309] When the first training data includes the first text and the second text, the second text is used to supervise the execution result obtained by the first model executing the first task based on the first text; when the first training data includes the first text and the fourth text, the fourth text is used to supervise the execution result obtained by the first model executing the first task based on the first text.

[0310] Combined with any implementation manner of the present application, the method further includes:

[0311] Generate second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text, where the second training data is used to perform optimization training on the first model, and the optimization training is different from the training performed on the first model based on the first training data.

[0312] Combined with any implementation manner of the present application, the generating unit 14 is further configured to:

[0313] Determine the quality of the second text based on the first similarity and whether the style of the second text meets the first preset requirement;

[0314] Determine the quality of the fourth text based on the second similarity and whether the style of the fourth text meets the first preset requirement;

[0315] When the quality of the second text is higher than the quality of the fourth text, generate the second training data including the first text, the second text, and the fourth text. In the second training data, the second text is the execution result obtained by the first model executing the first task based on the first text, and the fourth text is not the execution result obtained by the first model executing the first task based on the first text.

[0316] Combined with any implementation manner of the present application, the generating unit 14 is further configured to:

[0317] When the quality of both the second text and the fourth text does not meet the second preset requirement, obtain a sixth text, where the sixth text includes the execution result obtained by the first model executing the first task based on the first text, and the quality of the sixth text meets the second preset requirement;

[0318] Generating the second training data includes the first text, the second text, and the sixth text. In the second training data, the sixth text is an execution result obtained by performing the first task based on the first text, and the second text is not an execution result obtained by performing the first task based on the first text;

[0319] And / or, generating the second training data includes the first text, the fourth text, and the sixth text. In the second training data, the sixth text is an execution result obtained by performing the first task based on the first text, and the fourth text is not an execution result obtained by performing the first task based on the first text.

[0320] Combined with any embodiment of the present application, the execution unit 12 is further configured to:

[0321] Based on the content quality rule and / or the content verification instruction, determine whether the content of the first text and the content of the second text are reasonable, and obtain a judgment result;

[0322] The generating unit 14 is further configured to, when the judgment result includes that the content of the first text and the content of the second text are both reasonable, generate first training data based on the first similarity, the first text, and the second text.

[0323] In an embodiment of the present application, after obtaining the first text, the generating device performs a first task based on the first text to obtain a second text. Then, a second task is performed based on the second text to obtain a third text. Here, the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by performing the first task based on the original text. Then, the first similarity between the first text and the third text is determined. At this time, the greater the first similarity, the more similar the content of the first text is to the content of the third text. Also, since the third text is determined based on the second text, the similarity between the content of the third text and the content of the second text is high. Therefore, the similarity between the content of the second text and the content of the first text is high. That is to say, the accuracy of the second text as an execution result obtained by performing the first task based on the first text is high. That is, the first similarity is positively correlated with the accuracy of the second text as an execution result of performing the first task based on the first text, and the first similarity can measure the accuracy of the second text. Since the first similarity can measure the accuracy of the second text, it can be determined whether the second text can be used to supervise the execution result obtained by the first model performing the first task based on the first similarity. Therefore, first training data can be generated based on the first similarity, the first text, and the second text, thereby reducing the labor cost.

[0324] Please refer to Figure 8 , Figure 8Schematic structural diagram of another training data generation device provided by an embodiment of this application. The training data generation device 2 includes: an acquisition unit 21, a translation unit 22, a determination unit 23, and a generation unit 24, where:

[0325] The acquisition unit 21 is configured to acquire a first text, where the first text includes characters in a first language;

[0326] The translation unit 22 is configured to translate the characters in the first language in the first text into characters in a second language to obtain a second text;

[0327] The translation unit 22 is further configured to translate the characters in the second language in the second text into the characters in the first language to obtain a third text;

[0328] The determination unit 23 is configured to determine a first similarity between the first text and the third text;

[0329] The generation unit 24 is configured to generate first training data based on the first similarity, the first text, and the second text. The first training data includes the first text and the second text. In the first training data, the second text is used to supervise the text obtained by the first model translating the characters in the first language in the first text into the characters in the second language.

[0330] In the embodiment of this application, after the generation device acquires the first text, it obtains the second text based on translating the characters in the first language in the first text into the characters in the second language. Then it translates the characters in the second language in the second text into the characters in the first language to obtain the third text. Then it determines the first similarity between the first text and the third text. At this time, the greater the first similarity, the more similar the content of the first text is to the content of the third text. Also, because the third text is obtained by translating the second text, the similarity between the content of the third text and the content of the second text is high. Therefore, the similarity between the content of the second text and the content of the first text is high. That is to say, the accuracy of the second text as the translation result of the first text is high. That is, the first similarity is positively correlated with the accuracy of the second text as the translation result of the first text, and the first similarity can measure the accuracy of the second text. Since the first similarity can measure the accuracy of the second text, it can be determined whether the second text can be used to supervise the translation result obtained by the first model translating the characters in the first language in the first text into the characters in the second language based on the first similarity. Therefore, the first training data can be generated based on the first similarity, the first text, and the second text, thereby reducing the labor cost.

[0331] Please refer to Figure 9 , Figure 9Schematic structural diagram of a training device provided by an embodiment of the present application. The training device 3 includes: an acquisition unit 31 and a training unit 32, where:

[0332] The acquisition unit 31 is configured to acquire first training data generated based on the first aspect and any of its embodiments;

[0333] The training unit 32 is configured to train a first model based on the first training data, and the first model is used to perform a first task.

[0334] In an embodiment of the present application, after the training device acquires the first training data and trains a first model based on the first training data, the first model can be enabled to have the ability to perform the first task.

[0335] Please refer to Figure 10 , Figure 10 Schematic structural diagram of another training device provided by an embodiment of the present application. The training device 4 includes: an acquisition unit 41 and a training unit 42, where:

[0336] The acquisition unit 41 is configured to acquire a first model and a second model based on the third aspect, where the second model is used to evaluate the quality of the execution result of the first task. The quality of the execution result of the first task includes the accuracy of the execution result of the first task and whether the style of the execution result of the first task meets a first preset requirement. The second model is trained based on second training data, and the second training data is generated based on an embodiment of the first aspect;

[0337] The training unit 42 is configured to train the first model based on the second model to obtain a third model, and the third model is used to perform the first task, and the style of the execution result obtained by performing the first task meets the first preset requirement.

[0338] Combined with any embodiment of the present application, the training unit 42 is further configured to:

[0339] Acquire a first text and at least one candidate text, where the candidate text is the result obtained by the first model performing the first task based on the first text;

[0340] Determine, based on the second model, a compliant text that meets the first preset requirement and a non-compliant text that does not meet the first preset requirement from the at least one candidate text;

[0341] Determine, based on the first model, a first probability that the compliant text is the result of performing the first task based on the first text, and a second probability that the non-compliant text is the result of performing the first task based on the first text;

[0342] Determine a first loss of the first model based on a first difference between the first probability and the second probability, where the first difference is negatively correlated with the first loss;

[0343] Update parameters of the first model based on the first loss to obtain the third model.

[0344] In an embodiment of the present application, after obtaining a first model and a second model, a training device trains the first model based on the second model to obtain a third model, whereby the third model can be enabled to have the ability to execute a first task, and the style of the execution result obtained by the third model executing the first task meets a first preset requirement.

[0345] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a text processing device provided in an embodiment of the present application. The text processing device 5 includes: an acquisition unit 51 and an execution unit 52, where: the acquisition unit 51 is configured to acquire a text to be processed and a third model trained based on the training method and its implementation manner in the fourth aspect; the execution unit 52 is configured to execute a first task based on the third model and the text to be processed to obtain a target text, and the style of the target text meets a first preset requirement.

[0346] In an embodiment of the present application, after obtaining the third model, the text processing device executes a first task based on the third model and the text to be processed to obtain a target text, which can improve the accuracy of the target text and make the style of the target text meet a first preset requirement.

[0347] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.

[0348] Figure 12 which is a schematic hardware structure diagram of an electronic device provided in an embodiment of the present application. The electronic device 6 includes a processor 61 and a memory 62. Optionally, the electronic device 6 further includes an input device 63 and an output device 64. The processor 61, the memory 62, the input device 63, and the output device 64 are coupled through a connector, and the connector includes various interfaces, transmission lines, or buses, etc., which are not limited in the embodiments of the present application. It should be understood that in various embodiments of the present application, coupling means being interconnected in a specific manner, including being directly connected or indirectly connected through other devices, for example, being connected through various interfaces, transmission lines, buses, etc.

[0349] The processor 61 may include one or more processors, such as including one or more central processing units (CPUs). When the processor is a CPU, the CPU may be a single-core CPU or a multi-core CPU. Optionally, the processor 61 may be a processor group composed of multiple CPUs, and the multiple processors are coupled to each other through one or more buses. Optionally, the processor may also be other types of processors, etc., which are not limited in the embodiments of the present application.

[0350] The memory 62 can be used to store computer program instructions and various computer program codes including the program codes for executing the solutions of the present application. Optionally, the memory includes but is not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), and this memory is used for relevant instructions and data.

[0351] The input device 63 is used to input data and / or signals, and the output device 64 is used to output data and / or signals. The input device 63 and the output device 64 may be independent devices or an integrated device.

[0352] It can be understood that in the embodiments of the present application, the memory 62 can not only be used to store relevant instructions, but also be used to store relevant data, and the embodiments of the present application do not limit the specific data stored in this memory.

[0353] It can be understood that Figure 12 Only a simplified design of an electronic device is shown. In practical applications, the electronic device may also separately include necessary other elements, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of the present application are within the protection scope of the present application.

[0354] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0355] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art can also clearly understand that each embodiment of the present application has its own emphasis in description. For the convenience and brevity of description, the same or similar parts may not be repeated in different embodiments. Therefore, the parts not described or not described in detail in a certain embodiment can be referred to the descriptions of other embodiments.

[0356] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces. The indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.

[0357] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0358] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0359] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital versatile disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0360] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by computer programs instructing relevant hardware. The programs can be stored in a computer-readable storage medium. When the programs are executed, they can include the processes of the above method embodiments. The aforementioned storage media include various media that can store program codes, such as read-only memory (ROM) or random access memory (RAM), magnetic disks, or optical discs.

Claims

1. A method for generating training data, characterized in that, The method includes: Obtain a first text; Execute a first task based on the first text to obtain a second text; Execute a second task associated with the first task based on the second text to obtain a third text, where the second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by executing the first task based on the original text; Determine a first similarity between the first text and the third text; Generate first training data based on the first similarity, the first text, and the second text, where the first training data includes the first text and the second text, and in the first training data, the second text is used to supervise the execution result obtained by the first model executing the first task based on the first text.

2. The method according to claim 1, wherein The executing the first task based on the first text to obtain a second text includes: Execute the first task based on a first processing method and the first text to obtain the second text; The method further includes: Execute the first task based on a second processing method different from the first processing method and the first text to obtain a fourth text; Execute the second task based on the fourth text to obtain a fifth text; Determine a second similarity between the first text and the fifth text; The generating the first training data based on the first similarity, the first text, and the second text includes: Generate first training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text.

3. The method according to claim 2, wherein The executing the second task associated with the first task based on the second text to obtain a third text includes: Execute the second task based on the first processing method and the second text to obtain the third text; The executing the second task based on the fourth text to obtain a fifth text includes: Execute the second task based on the second processing method and the fourth text to obtain the fifth text.

4. The method according to claim 2 or 3, characterized in that, The generating the first training data based on the first similarity, the first text, and the second text includes: Generate the first training data based on the magnitude relationship between the first similarity and the second similarity, the first text, the second text, and the fourth text; In the case where the first training data includes the first text and the second text, the second text is used to supervise the execution result obtained by the first model executing the first task based on the first text; in the case where the first training data includes the first text and the fourth text, the fourth text is used to supervise the execution result obtained by the first model executing the first task based on the first text.

5. The method according to claim 2 or 3, characterized in that, The method further includes: Generate second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text, where the second training data is used to perform optimization training on the first model, and the optimization training is different from the training performed on the first model based on the first training data.

6. The method according to claim 5, wherein Generating second training data based on the first similarity, the second similarity, the first text, the second text, and / or the fourth text includes: Determining the quality of the second text based on the first similarity and whether the style of the second text meets a first preset requirement; Determining the quality of the fourth text based on the second similarity and whether the style of the fourth text meets the first preset requirement; When the quality of the second text is higher than the quality of the fourth text, generating the second training data includes the first text, the second text, and the fourth text. In the second training data, the second text is an execution result obtained by performing the first task based on the first text, and the fourth text is not an execution result obtained by performing the first task based on the first text.

7. The method according to claim 6, wherein The method further includes: When the quality of both the second text and the fourth text does not meet a second preset requirement, obtaining a sixth text, where the sixth text includes an execution result obtained by performing the first task based on the first text, and the quality of the sixth text meets the second preset requirement; Generating the second training data includes the first text, the second text, and the sixth text. In the second training data, the sixth text is an execution result obtained by performing the first task based on the first text, and the second text is not an execution result obtained by performing the first task based on the first text; and / or, generating the second training data includes the first text, the fourth text, and the sixth text. In the second training data, the sixth text is an execution result obtained by performing the first task based on the first text, and the fourth text is not an execution result obtained by performing the first task based on the first text.

8. The method according to any one of claims 1 to 3, characterized in that, Before generating the first training data based on the first similarity, the first text, and the second text, the method further includes: Judging whether the content of the first text and the content of the second text are reasonable based on a content quality rule and / or a content verification instruction to obtain a judgment result; Generating the first training data based on the first similarity, the first text, and the second text includes: When the judgment result includes that the content of both the first text and the second text is reasonable, generating the first training data based on the first similarity, the first text, and the second text.

9. A method for generating training data, characterized in that, The method includes: Obtaining a first text, where the first text includes characters in a first language; Translating the characters in the first language in the first text into characters in a second language to obtain a second text; Translating the characters in the second language in the second text into the characters in the first language to obtain a third text; Determining a first similarity between the first text and the third text; Generate first training data based on the first similarity, the first text, and the second text. The first training data includes the first text and the second text. In the first training data, the second text is used to supervise the text obtained by the first model translating the text in the first language in the first text into the text in the second language.

10. A training method, characterized in that, The method includes: Obtain first training data generated based on the method described in any one of claims 1 to 7; Based on the first training data, train to obtain a first model, where the first model is used to perform a first task.

11. A training method, characterized in that, The method includes: Obtain a first model and a second model trained based on the method described in claim 10. The second model is used to evaluate the quality of the execution result of the first task. The quality of the execution result of the first task includes the accuracy of the execution result of the first task and whether the style of the execution result of the first task meets a first preset requirement. The second model is trained based on second training data, and the second training data is generated based on the method described in any one of claims 5 to 7; Based on the second model, train the first model to obtain a third model, where the third model is used to perform the first task and the style of the execution result obtained by performing the first task meets the first preset requirement.

12. The method according to claim 11, wherein The training the first model based on the second model to obtain a third model includes: Obtain a first text and at least one candidate text, where the candidate text is obtained by the first model performing the first task based on the first text; Based on the second model, determine a compliant text that meets the first preset requirement and a non-compliant text that does not meet the first preset requirement from the at least one candidate text; Based on the first model, determine a first probability that the compliant text is the result of performing the first task based on the first text, and a second probability that the non-compliant text is the result of performing the first task based on the first text; Based on a first difference between the first probability and the second probability, determine a first loss of the first model, where the first difference is negatively correlated with the first loss; Based on the first loss, update the parameters of the first model to obtain the third model.

13. A text processing method, characterized in that, The method includes: Obtain a text to be processed and a third model trained based on the method described in claim 11 or 12; Based on the third model and the text to be processed, perform the first task to obtain a target text, where the style of the target text meets the first preset requirement.

14. A training data generation device, characterized in that, The training data generation device includes: An acquisition unit, configured to acquire a first text; An execution unit, configured to perform a first task based on the first text to obtain a second text; The execution unit is further configured to perform a second task associated with the first task based on the second text to obtain a third text. The second task includes determining the original text of the first task based on the execution result of the first task, and the execution result of the first task is obtained by performing the first task based on the original text; A determining unit, configured to determine a first similarity between the first text and the third text; A generating unit, configured to generate first training data based on the first similarity, the first text, and the second text, where the first training data includes the first text and the second text, and in the first training data, the second text is used to supervise an execution result obtained by a first model when executing the first task based on the first text.

15. A training data generation device, characterized in that, The training data generating device includes: An obtaining unit, configured to obtain a first text, where the first text includes characters in a first language; A translating unit, configured to translate the characters in the first language in the first text into characters in a second language to obtain a second text; The translating unit is further configured to translate the characters in the second language in the second text into characters in the first language to obtain a third text; A determining unit, configured to determine a first similarity between the first text and the third text; A generating unit, configured to generate first training data based on the first similarity, the first text, and the second text, where the first training data includes the first text and the second text, and in the first training data, the second text is used to supervise a text obtained by a first model when translating the characters in the first language in the first text into characters in the second language.

16. A text processing device, characterized in that, The text processing device includes: An obtaining unit, configured to obtain a text to be processed and a third model trained based on the method described in claim 11 or 12; An executing unit, configured to execute a first task based on the third model and the text to be processed to obtain a target text, where a style of the target text meets a first preset requirement.

17. An electronic device, characterized in that, including: A processor and a memory, where the memory is configured to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method described in any one of claims 1 to 8, or the electronic device executes the method described in claim 9, or the electronic device executes the method described in claim 10, or the electronic device executes the method described in claim 11 or 12, or the electronic device executes the method described in claim 13.

18. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method described in any one of claims 1 to 8, or the processor is caused to execute the method described in claim 9, or the processor is caused to execute the method described in claim 10, or the processor is caused to execute the method described in claim 11 or 12, or the processor is caused to execute the method described in claim 13.

19. A computer program product, characterized in that, The computer program product includes a computer program or instructions; when the computer program or instructions are run on a computer, the computer is caused to perform the method described in any one of claims 1 to 8, or the computer is caused to perform the method described in claim 9, or the computer is caused to perform the method described in claim 10, or the computer is caused to perform the method described in claim 11 or 12, or the computer is caused to perform the method described in claim 13.