Data generation method, model training method, and data processing method

Through the collaborative work and feedback optimization of multiple expert large models, the diversity and quality issues in the generation of large model data were resolved, achieving efficient and accurate data generation and model training results.

CN120561595BActive Publication Date: 2026-02-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510788012.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2026-02-06
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing large models suffer from problems such as insufficient diversity of generated data, poor quality of response results, weak generalization ability, and uneven distribution of generated data during the data generation process.

Method used

By working collaboratively with multiple expert models, the system uses data synthesis expert units to generate response results to be evaluated. The response results are then optimized through evaluation and reflection correction processes until the generation requirements are met, resulting in high-quality and diverse training data.

Benefits of technology

It achieves efficient, accurate, and diverse generation of large model data, provides high-quality training data for model optimization, and improves the quality of model response results and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561595B_ABST
    Figure CN120561595B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data generation method, relates to the technical field of artificial intelligence, and particularly relates to the technical field of large models and agents. A specific implementation scheme is as follows: at least one to-be-evaluated response result is generated by using at least one of a plurality of data synthesis expert units; at least one evaluation result for the at least one to-be-evaluated response result is determined by using at least one of the plurality of data synthesis expert units; in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, at least one correction data for at least one data synthesis expert unit is determined according to the at least one evaluation result for the at least one to-be-corrected response result; and the operation of generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units is returned to. The present disclosure also provides a model training method, a data processing method, a device, an electronic device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of large models and agents. More specifically, the present disclosure provides a data generation method, a model training method, a data processing method, an apparatus, an electronic device and a storage medium. BACKGROUND

[0002] Sample data is the basis for training large language models. With the rapid development of large language models (LLM), the demand for large-scale and high-quality sample data is increasing when training large models to continuously improve the capabilities of large models. SUMMARY

[0003] The present disclosure provides a data generation method, a model training method, a data processing method, an apparatus, an electronic device and a storage medium.

[0004] According to an aspect of the present disclosure, a data generation method is provided, which includes: generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units; determining at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determining at least one correction data for the at least one data synthesis expert unit according to the at least one evaluation result for the at least one to-be-corrected response result; returning to the operation of generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, wherein the data synthesis expert unit is configured to generate the to-be-evaluated response result according to at least one of target data and received correction data.

[0005] According to another aspect of the present disclosure, a model training method is provided, which includes: inputting target data in a sample data pair into a to-be-trained model to obtain a to-be-optimized response result; training the to-be-trained model according to the to-be-optimized response result and a target response result in the sample data pair, wherein the sample data pair is determined by: generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units, the data synthesis expert unit being configured to generate the to-be-evaluated response result according to target data; determining at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; in response to determining that the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, determining the sample data pair according to the target response result in the at least one to-be-evaluated response result and the target data.

[0006] According to another aspect of the present disclosure, a data processing method is provided, which comprises: inputting to-be-processed data into a target model to obtain a response result corresponding to the to-be-processed data, wherein the target model is obtained by training a training model according to the method described in the embodiments of the present disclosure.

[0007] According to another aspect of the present disclosure, a data generation apparatus is provided, which comprises: a result generation module configured to generate at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units; a result determination module configured to determine at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; a data determination module configured to, in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determine at least one correction data for at least one data synthesis expert unit according to the at least one evaluation result for the at least one to-be-corrected response result; and return to the operation of generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, wherein the data synthesis expert unit is configured to generate the to-be-evaluated response result according to at least one of target data and received correction data.

[0008] According to another aspect of the present disclosure, a model training apparatus is provided, which comprises: a first input module configured to input target data in a sample data pair into a to-be-trained model to obtain a to-be-optimized response result; and a training module configured to train the to-be-trained model according to the to-be-optimized response result and a target response result in the sample data pair, wherein the sample data pair is determined by: a result generation module configured to generate at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units, the data synthesis expert unit being configured to generate the to-be-evaluated response result according to target data; a result determination module configured to determine at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; and a data determination module configured to, in response to determining that the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, determine the sample data pair according to the target response result in the at least one to-be-evaluated response result and the target data.

[0009] According to another aspect of the present disclosure, a data processing apparatus is provided, which comprises: a second input module configured to input to-be-processed data into a target model to obtain a response result corresponding to the to-be-processed data, wherein the target model is obtained by training a training model according to the apparatus described in the embodiments of the present disclosure.

[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the present disclosure.

[0011] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions for causing a computer to perform the method provided by the present disclosure is provided.

[0012] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method provided by the present disclosure.

[0013] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0015] Figure 1 is an exemplary system architecture schematic diagram of the data generation method, model training method, data processing method and corresponding device according to one embodiment of the present disclosure;

[0016] Figure 2 is a flowchart of the data generation method according to one embodiment of the present disclosure;

[0017] Figure 3A is a schematic diagram of the data generation method according to one embodiment of the present disclosure;

[0018] Figure 3B is a schematic diagram of data synthesis according to one embodiment of the present disclosure;

[0019] Figure 4 is a flowchart of the data generation method according to another embodiment of the present disclosure;

[0020] Figure 5 is a flowchart of the model training method according to one embodiment of the present disclosure;

[0021] Figure 6 is a flowchart of the data processing method according to one embodiment of the present disclosure;

[0022] Figure 7 is a block diagram of the data generation device according to one embodiment of the present disclosure;

[0023] Figure 8 is a block diagram of a model training apparatus according to an embodiment of the disclosure;

[0024] Figure 9 is a block diagram of a data processing apparatus according to an embodiment of the disclosure; and

[0025] Figure 10 is a block diagram of an electronic device to which a data generation method, a model training method, and a data processing method can be applied according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0026] Exemplary embodiments of the disclosure are described herein with reference to the accompanying drawings, which are presented for the purpose of illustration and description. They are not intended to limit the scope of the disclosure, and they are considered as merely exemplary. Therefore, those of ordinary skill in the art will recognize various modifications and changes that can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. Also, descriptions of known functions and constructions are omitted in the following description for clarity and conciseness.

[0027] At present, the large model data synthesis technology mainly includes the following four methods:

[0028] I. Using a large model to generate a sample pair of instructions and response results, training the large model through the sample pair, which can reduce the cost of manual annotation, but is easy to cause the large model to produce false cognition.

[0029] II. Replacing manual annotation with feedback of an artificial intelligence large model to realize automatic alignment of instructions and response results. However, this method depends on a strong teacher model, and the feedback reliability of the response result is low.

[0030] III. Training the model by artificially manufacturing adversarial sample pairs during training to improve the stability of the model and avoid the model producing false responses due to interference factors. However, this method has insufficient diversity of generated data, and the artificially generated adversarial samples may deviate from the real scene.

[0031] IV. Through knowledge distillation, the ability of a large model is transferred to a small model, which can improve the robustness of the model and the generation ability of the small model. However, this method depends on the generation quality of the teacher model, and may receive false knowledge under the influence of the teacher model. In addition, the diversity of the data generated by the teacher model is insufficient, and the generalization ability of the student model is weak.

[0032] The above four methods all have problems such as insufficient diversity of generated data, poor quality of response results, weak generalization ability, and uneven distribution of generated data.

[0033] Based on this, the present disclosure proposes a data generation method, a model training method and a data processing method, aiming to realize efficient, accurate and diversified large model data generation through the collaborative work of multiple expert large models.

[0034] Figure 1 is an exemplary system architecture schematic diagram to which the data generation method, the model training method, the data processing method and the corresponding device according to one embodiment of the present disclosure can be applied. It should be noted that, Figure 1 The system architecture shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0035] As Figure 1 shown, the system architecture 100 according to the embodiment can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.

[0036] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers and desktop computers, etc.

[0037] The server 105 can be a server providing various services, such as a background management server supporting the website browsed by the user using the terminal devices 101, 102, 103 (only as an example). The background management server can analyze and process the received user requests, to-be-processed data, etc., and feed back the processing results (such as processing results generated according to the input information of the user, etc.) to the terminal device. The server 105 can be deployed with a predetermined model trained, and by inputting the input information of the user into the predetermined model, the final response result is obtained, and the response result is returned to the user. The predetermined model can be a large language model (Large Language Model, LLM).

[0038] It should be noted that the data generation method, model training method and data processing method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the data generation apparatus, model training apparatus and data processing apparatus provided by the embodiments of the present disclosure can generally be arranged in the server 105. The data generation method, model training method and data processing method provided by the embodiments of the present disclosure can also be executed by the terminal device 101, 102 or 103. Correspondingly, the data generation apparatus, model training apparatus and data processing apparatus provided by the embodiments of the present disclosure can also be arranged in the terminal device 101, 102 or 103. The data generation method, model training method and data processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102 or 103 and / or the server 105. Correspondingly, the data generation apparatus, model training apparatus and data processing apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102 or 103 and / or the server 105.

[0039] Figure 2 FIG. 2 is a flowchart of a data generation method according to an embodiment of the present disclosure.

[0040] As shown in FIG. 2, the data generation method 200 can include operations S210-S230. Figure 2

[0041] At operation S210, at least one to-be-evaluated response result is generated by using at least one of a plurality of data synthesis expert units.

[0042] The data synthesis expert unit can refer to a large model learned for data features in a specific field. The plurality of data synthesis expert units can refer to large models focusing on different fields or tasks and determining response results through different solution strategies. The to-be-evaluated response result can refer to a response result generated by the data synthesis expert unit after receiving input text or the like.

[0043] For example, the input text can be input into three data synthesis expert units, and three response results related to the input text but different in content can be obtained as the to-be-evaluated response results.

[0044] At operation S220, at least one evaluation result for the at least one to-be-evaluated response result is determined by using at least one of the plurality of data synthesis expert units.

[0045] One data synthesis expert unit can be used to evaluate the generated to-be-evaluated response result, or multiple data synthesis expert units can be selected for evaluation to obtain evaluation results from multiple dimensions.

[0046] ​For example, the generated three to-be-evaluated response results can be evaluated using one or more of the three data synthesis expert units described above. The data synthesis expert units can evaluate the to-be-evaluated response results from different evaluation perspectives such as data diversity, rationality, and the like, to obtain corresponding evaluation results.

[0047] The evaluation results can represent whether the to-be-evaluated response results meet the generated requirements.

[0048] For example, the three data synthesis expert units described above include a first data synthesis expert unit, a second data synthesis expert unit, and a third data synthesis expert unit. The second data synthesis expert unit can evaluate the to-be-evaluated response results from the perspective of rationality. If the evaluation result given by the second data synthesis expert unit is that it is in line with common sense, it can be determined that the to-be-evaluated response result is correct. If the evaluation result given by the second data synthesis expert unit is that it is not in line with common sense, it can be determined that the to-be-evaluated response result has a problem.

[0049] In operation S230, in response to determining that at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, at least one correction data for at least one data synthesis expert unit is determined according to at least one evaluation result for at least one to-be-corrected response result, and the operation of generating at least one to-be-evaluated response result using at least one of the plurality of data synthesis expert units is returned to until at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result.

[0050] The to-be-corrected response result can be an incorrect response result that does not meet the generated requirements. The correction data can be optimized text for the to-be-corrected response result, used to correct errors of the data synthesis expert unit in the understanding and answering process.

[0051] In an embodiment of the present disclosure, whether the at least one to-be-evaluated response result meets the generated requirements can be determined according to the evaluation results given by the data synthesis expert units, and a response result that does not meet the generated requirements is determined as a to-be-corrected response result. Correction data for optimization is generated according to the to-be-corrected response result and input into the corresponding data synthesis expert unit, so that the data synthesis expert unit adjusts and optimizes based on the correction data to generate data that meets the requirements.

[0052] For example, the second data synthesis expert unit described above evaluates the plurality of to-be-evaluated response results from the perspective of rationality. Taking a first to-be-evaluated response result and a second to-be-evaluated response result included in the plurality of to-be-evaluated response results as an example, if the evaluation result for the first to-be-evaluated response result indicates that the first to-be-evaluated response result is not in line with common sense, correction data for the first to-be-evaluated response result can be generated so that the corresponding correction data is input into the first data synthesis expert unit that generates the first to-be-evaluated response result.

[0053] In the embodiments of the present disclosure, the data synthesis expert unit is configured to generate the response result to be evaluated based on at least one of the target data and the received modified data.

[0054] The target data can refer to text information input into the data synthesis expert unit for generating the response result, and the target data can be a material text for generating the response result. Based on the modified data generated in operation S230 and the target data, the response result can be regenerated in operation S210. Next, the processing procedures of operations S210 to S230 can be repeated, and the modified data generated based on the evaluation result and the target data are input into the data synthesis expert unit together, and the data synthesis expert unit is iterated for multiple times until the response result generated by the data synthesis expert unit based on the target data in the later iteration process meets the requirements, that is, the evaluation result given by the data synthesis expert unit indicates that there is no response result to be modified in the response result to be evaluated.

[0055] Through the embodiments of the present disclosure, through the collaborative work of multiple expert units, diverse response results can be generated, and the response results can be adjusted and optimized through the evaluation and reflection modification process, so that the response results can be quickly iterated. The generated response results can be used as training data for large models, and high-quality and accurate training data can be provided for subsequent large model optimization.

[0056] In the embodiments of the present disclosure, the data synthesis expert unit can be a data synthesis large model or a data synthesis agent.

[0057] The large model or agent focused on different fields can be obtained by pre-training, supervised fine-tuning (SFT) or direct preference optimization (DPO) to train and optimize the model, and can be used for generating response results for input text.

[0058] In the embodiments of the present disclosure, the target data can be material data for generating the response result, including multiple data types. The target data can include at least one of target text data, target audio data, target image data and target video data.

[0059] For example, the target data can be target text data such as news, novels, article materials, etc. The target data can also be target audio data such as songs, pure music, sound effects, etc. The target data can also be target image data such as photos, emoticons, dynamic pictures, etc. The target data can also be target video data such as animations, videos, TV series, etc.

[0060] Next, in combination with the above description of the data synthesis expert unit, the following embodiments of the present disclosure are described. Figure 3AThe data generation method disclosed herein will be further explained.

[0061] Figure 3A This is a schematic diagram of a data generation method according to an embodiment of the present disclosure.

[0062] like Figure 3A As shown, in the embodiments of this disclosure, at least one target data d320 can be obtained by preprocessing multiple initial data d310. Preprocessing methods may include filtering and annotation, which will be described below.

[0063] In embodiments of this disclosure, the preprocessing process may include: filtering multiple initial data d310 according to preset filtering rules to obtain at least one intermediate data; and labeling the at least one intermediate data to obtain at least one target data d320.

[0064] Preset filtering rules can refer to filtering data that is irrelevant to the generation task, as well as filtering data that lacks practical meaning or is of low quality. Low-quality data can be semantically contradictory.

[0065] The preprocessing step can be used to filter low-quality data according to preset filtering rules, preventing data that does not meet the preset filtering rules from entering subsequent processing procedures. At the same time, data that meets the preset filtering rules is labeled to obtain target data including at least one label, so that the target data can be provided to the large model or agent corresponding to the label.

[0066] Tags can indicate the source, content, format, and expression of target data.

[0067] For example, the initial data includes a text describing the moon. By preprocessing the initial data, target data containing tags can be obtained. The tags included in the target data can include: a domain tag of "literature", a language type tag of "Chinese", and a source tag of "a certain magazine".

[0068] According to embodiments of this disclosure, by preprocessing the input corpus, questions, and other target data through filtering and tagging, data lacking practical meaning or irrelevant to the subsequent process can be prevented from entering. In the subsequent synthesis and processing stages, explicit representations corresponding to the target data can be obtained through tags, improving the rationality, accuracy, and diversity of data synthesis.

[0069] Next, the instruction i330 to be processed can be determined based on the target data.

[0070] For example, the target data is a piece of text describing the moon, and the to-be-processed instruction can be "generate a picture about the moon" or "generate a poem about the moon" based on the target data. Based on the target data, to-be-processed instructions of different generation difficulties, different expression modes, different languages, and different data quantities can be determined, thereby improving the diversity of data synthesis.

[0071] In an embodiment of the present disclosure, the data synthesis expert unit can generate a to-be-evaluated response result according to at least one of the at least one label, the target data, and the received correction data.

[0072] When the data synthesis expert unit generates a to-be-evaluated response result based on the target data for the first time, the generation can be based on the label or the target data alone.

[0073] In the first synthesis process in the data synthesis process that can be cycled multiple times, after the data synthesis expert unit receives the target data d320, the at least one label, and the to-be-processed instruction i330, the data synthesis expert unit can generate a to-be-evaluated response result r340 corresponding to the to-be-processed instruction i330 according to the target data d320 and the at least one label.

[0074] After the to-be-evaluated response result is evaluated and the corresponding correction data is generated, after the data synthesis expert unit receives the correction data, the data synthesis expert unit can generate a corresponding to-be-evaluated response result r340 according to at least one of the target data d320, the at least one label, the to-be-processed instruction i330, and the correction data. Then, the to-be-evaluated response result r340 is audited in terms of quality, format, compliance, etc., and the sample data pair p350 is determined if the audit is passed. In the following, the process of data synthesis will be described in combination with Figure 3B The process of data synthesis is described.

[0075] Figure 3B is a schematic diagram of data synthesis according to an embodiment of the present disclosure.

[0076] As shown in Figure 3B , a plurality of expert large models can be used as a question understanding expert unit 311, a step decomposition expert unit 312, a plurality of data synthesis expert units 313, an evaluation expert unit 313', and a reflection and correction expert unit 314, respectively. The evaluation expert unit 313' can evaluate the to-be-evaluated response result. One or more of the plurality of expert large models serving as the plurality of data synthesis expert units 313 can serve as one or more evaluation expert units 313'.

[0077] In embodiments of the present disclosure, in some implementations of operation S210, generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units comprises: based on the at least one to-be-processed sub-problem determined by the data understanding result, generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units.

[0078] The data understanding result is obtained according to at least one of the target data and the to-be-processed instruction, and the to-be-processed instruction is obtained according to the target data. The to-be-processed instruction can be text, including one or more characters. Taking an example in which the data understanding result is obtained according to the target data and the to-be-processed instruction, the to-be-processed instruction can be understood by using the problem understanding expert unit to determine the data understanding result. The data understanding result can be a result of understanding information such as background, generation condition, and requirement of the input to-be-processed instruction. It can be understood that, for different to-be-processed instructions, different problem understanding results can be determined by using different solving paths to achieve generation of the to-be-evaluated response result from a suitable path.

[0079] For example, the target data is a paragraph of text describing the moon. The to-be-processed instruction determined according to the target data is "generate a poem about the moon". The data understanding result generated by using the problem understanding expert unit 311 according to the target data and the to-be-processed instruction can be "generate a seven-character quatrains about the target data, and the content is a poem praising the moon".

[0080] It can be understood that the data understanding of the present disclosure is described above, and the step decomposition based on the data understanding result will be described below.

[0081] According to embodiments of the present disclosure, based on the at least one to-be-processed sub-problem determined by the data understanding result, generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units comprises: determining the at least one to-be-processed sub-problem by using the step decomposition expert unit 312 according to the data understanding result. Based on the at least one to-be-processed sub-problem, generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units.

[0082] The to-be-processed sub-problem can be a plurality of sub-problems obtained by decomposing the to-be-processed instruction which is originally difficult to understand into sub-problems that are easy for the data synthesis expert unit to understand by understanding and analyzing the to-be-processed instruction and the data understanding result. In this way, the data synthesis expert unit can solve the to-be-processed instruction by decomposing a plurality of steps or sub-problems to obtain the to-be-evaluated response result.

[0083] For example, as described above, the target data is a piece of text describing the moon. According to the target data, the to-be-processed instruction can be generated as "generate a poem describing the moon". According to the data understanding result determined based on the to-be-processed instruction and the target data, the data understanding result can be "generate a seven-character quatrains in combination with the target data, and the content is a poem praising the moon". Next, according to the data understanding result, the to-be-processed sub-problems obtained by disassembling the data understanding result can include "extract the words related to the moon in the target data", "get the format of the poem", "determine the words according to the rhythm regulation of the seven-character quatrains", "generate a poem about the moon", and the like. One or more data synthesis expert units can obtain one or more to-be-evaluated response results by solving the to-be-processed sub-problems one by one.

[0084] It can be understood that the above describes the step disassembly and the manner of data synthesis of the present disclosure. The evaluation expert unit will be further described below.

[0085] According to an embodiment of the present disclosure, in some implementations of the above operation S220, determining at least one evaluation result for at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units includes: determining N intermediate evaluation data for the to-be-evaluated intermediate result by using N data synthesis expert units 313. The to-be-evaluated intermediate result is obtained by masking the identification information of the to-be-evaluated response result related to the data synthesis expert unit generating the to-be-evaluated response result. N can be an integer greater than or equal to 1. According to the N intermediate evaluation data for the to-be-evaluated intermediate result, the evaluation result for the to-be-evaluated response result is determined.

[0086] In an embodiment of the present disclosure, the data synthesis expert unit can serve as an evaluation expert unit to blindly evaluate the to-be-evaluated response result. Blind evaluation means that when evaluating the to-be-evaluated response result, the data synthesis expert unit as the evaluation expert unit cannot obtain the relevant information of the data synthesis expert unit generating the to-be-evaluated response result. By masking the identification information of at least the to-be-evaluated response result, a plurality of to-be-evaluated intermediate results with unknown sources are obtained, so as to realize data synthesis and evaluation by using fewer expert large models, and further reduce the computational resource overhead for obtaining sample data.

[0087] For example, a first to-be-evaluated response result, a second to-be-evaluated response result and a third to-be-evaluated response result can be respectively generated by using the first data synthesis expert unit, the second data synthesis expert unit and the third data synthesis expert unit. By masking the identification information of the to-be-evaluated response result, three to-be-evaluated intermediate results are obtained. Two of the three data synthesis expert units are randomly determined as evaluation expert units, and the three to-be-evaluated intermediate results are respectively evaluated, so that two intermediate evaluation data for each to-be-evaluated intermediate result are obtained. According to the two intermediate evaluation data for each to-be-evaluated intermediate result, the evaluation result of the to-be-evaluated response result can be determined.

[0088] According to an embodiment of the present disclosure, the evaluation result can include at least one evaluation index value, and the at least one evaluation index value includes at least one of a first evaluation index value and a second evaluation index value.

[0089] The first evaluation index value may, for example, be an index value of positive evaluation, such as an index value of sentence fluency, an index value of rationality, etc. The second evaluation index value may, for example, be an index value of negative evaluation, such as an index value of contradiction degree, an index value of redundancy and repetition degree, etc. For example, in the two intermediate evaluation data for each to-be-evaluated intermediate result, each intermediate evaluation data can include one first intermediate evaluation value and one second intermediate evaluation value. The average of the two first intermediate evaluation values respectively from the two intermediate evaluation data is determined as one first evaluation index value. The average of the two second intermediate evaluation values respectively from the two intermediate evaluation data is determined as one second evaluation index value.

[0090] It can be understood that the above describes the evaluation expert unit of the present disclosure, and some ways of determining the correction data will be described below.

[0091] In some embodiments, the above method can further include: in response to determining that the plurality of evaluation index values for the to-be-evaluated response result do not satisfy at least one of the at least one preset evaluation condition, determining that the to-be-evaluated response result is a to-be-corrected response result. The at least one preset evaluation condition includes at least one of: the first evaluation index value is greater than or equal to a first evaluation threshold; and the second evaluation index value is less than or equal to a second evaluation threshold.

[0092] For example, as Figure 3BAs shown, the first evaluation threshold can be preset as 6, and the second evaluation threshold can be preset as 5. The first evaluation result of the first to-be-evaluated response result includes the first evaluation index value Result 11 and the second evaluation index value Result 21. The first evaluation index value Result 11 can be 5, and the second evaluation index value Result 21 can be 3. It can be determined that the first to-be-evaluated response result does not meet the preset evaluation condition, and it can be determined that the first to-be-evaluated response result is the first to-be-corrected response result. The second evaluation result of the second to-be-evaluated response result includes the first evaluation index value Result 12 and the second evaluation index value Result 22. The first evaluation index value Result 12 can be 7, and the second evaluation index value Result 22 can be 6. It can be determined that the second to-be-evaluated response result does not meet the preset evaluation condition, and it can be determined that the second to-be-evaluated response result is the second to-be-corrected response result. Thus, at least one evaluation result indicates that there is a to-be-corrected response result in the plurality of to-be-evaluated response results. The above operation S230 can be performed.

[0093] In some embodiments, in some implementations of the above operation S230, in response to determining that at least one of the evaluation results indicates that there is at least one to-be-corrected response result in at least one of the to-be-evaluated response results, at least one correction data for at least one of the data synthesis experts is determined by the reflection correction expert unit 314 according to at least one of the evaluation results for at least one of the to-be-corrected response results. For example, the first correction data for the first data synthesis expert can be determined by the reflection correction expert unit 314 according to the first evaluation result for the above first to-be-corrected response result. Thus, by evaluating the response results of the model from a multi-dimensional perspective, the response results are quantitatively analyzed using multiple evaluation indexes to determine whether the response results meet the requirements. In the case where the response results do not meet the requirements, correction data is generated in a timely manner for feedback, which can ensure the quality of the finally generated data and realize precise and diversified data synthesis.

[0094] Next, the operation S210 can be returned to until at least one of the evaluation results determined later indicates that there is no to-be-corrected response result in at least one of the to-be-evaluated response results. In the case where there is correction data, the data synthesis expert unit can regenerate the response results in combination with the correction data and the target data, one or more labels.

[0095] For example, the correction data is a correction text. The correction text for the above first data synthesis expert can be “the moon of the poem should appear at night”, and the correction text is provided to the data synthesis expert unit so that the data synthesis expert unit regenerates the to-be-evaluated response result later.

[0096] According to an embodiment of the present disclosure, by reflecting and analyzing the response results generated by the diversity of the expert model, errors existing in the generated results are found in time, and correction data is fed back to the expert model for the errors, so that the synthesis strategy of the expert model is adjusted and optimized, helping the expert model to realize iterative upgrading.

[0097] It can be understood that the above describes the present disclosure by taking the existence of a to-be-corrected response result in one or more to-be-evaluated response results as an example. After a plurality of data generations and corrections, one or more to-be-evaluated response results can all meet the above-mentioned preset evaluation condition, and there is no to-be-corrected response result, which will be described below.

[0098] In some embodiments, the above-mentioned method further comprises: in response to determining that the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, determining a sample data pair according to a target response result in the at least one to-be-evaluated response result and target data.

[0099] The target response result can be determined from the at least one to-be-evaluated response result. The target response result can be the response result with the highest quality among the response results that meet the preset evaluation condition. For the selection of the target response result, comparison can be made in multiple dimensions such as diversity, data length, language style, and generation professionalism, so as to determine the target response result from one or more to-be-evaluated response results. The target response result and the target data of the input data synthesized by the expert unit can be used as a sample data pair for model training.

[0100] For example, as described above, according to the target data (text describing the moon), three to-be-evaluated response results generated by three data synthesis expert units can be obtained. In the case where the evaluation results of the three to-be-evaluated response results all meet the preset evaluation condition, it can be determined that the three evaluation results indicate that there is no to-be-corrected response result in the three to-be-evaluated response results. Thus, the three to-be-evaluated response results can be ranked in multiple dimensions, and finally a to-be-evaluated response result is determined as the target response result. The target response result can form a sample data pair with the target data. In one example, the to-be-evaluated response results can be ranked from large to small based on the first evaluation index value, so as to take the to-be-evaluated response result with the largest first evaluation index value as the target response result. It can be understood that ranking can also be performed in other ways.

[0101] It can be understood that the above describes the present disclosure by taking the target data and the target response result as a sample data pair as an example. However, the present disclosure is not limited thereto, which will be described below.

[0102] In some embodiments, the to-be-processed instruction and the target response result can be used as a sample data pair. For example, the to-be-processed instruction "generate a poem about the moon" and the above-mentioned target response result can be used as a sample data pair.

[0103] It can be understood that the above describes the present disclosure by taking the example of determining the sample data pair according to at least one of the target data and the to-be-processed instruction and the target response result. However, the present disclosure is not limited thereto, and the target sample result can be further evaluated, which will be further described below.

[0104] According to an embodiment of the present disclosure, the determining the sample data pair according to the target response result and the target data in the at least one to-be-evaluated response result comprises: in response to determining that the target response result meets at least one preset data generation condition, determining the sample data pair, the sample data pair comprising at least one of the target data and the to-be-processed instruction derived from the target data and the target response result. The at least one preset data generation condition comprises: a data format of the target response result being a preset data format.

[0105] In an embodiment of the present disclosure, the quality audit expert unit can perform multi-dimensional quality audit and evaluation on the target response result generated by the data synthesis expert unit. For example, the data format of the target response result can be audited from the data format dimension.

[0106] In an embodiment of the present disclosure, the at least one preset data generation condition can further comprise: a content quality of the target response result meeting a preset data quality condition; and the target response result meeting a preset data rule. The preset data rule can be a relevant law, a regulation or a public order and good custom.

[0107] For example, if the target response result is a piece of code, the content quality of the target response result can refer to the processing speed, time complexity and other attribute data of the code. The preset data quality condition can refer to the time complexity of the target response result being less than a preset complexity threshold. In a case where the time complexity of the target response result meets the preset data generation condition, the target response result, and the corresponding target data and / or to-be-processed instruction can be taken as the sample data pair.

[0108] According to an embodiment of the present disclosure, by performing multi-dimensional audit on the response result generated by the expert large model through the quality audit expert unit, the quality of the finally generated data can be ensured, so as to obtain high-quality sample data.

[0109] It can be understood that the above describes the present disclosure by taking the example of generating the to-be-processed instruction according to the target data. However, the present disclosure is not limited thereto, and the following describes the present disclosure by taking the example of understanding the question according to the target data.

[0110] In other embodiments, the data understanding result is obtained by the question understanding expert unit 311 according to the target data.

[0111] For example, the target data is a piece of text describing the moon. According to the target data, the question understanding expert unit 311 can determine that the data understanding result is "generate a poem with a style of quatrains, a language of Chinese, and a content related to the target data". For another example, in the case of question understanding according to the target data, the question understanding result can also be "generate a picture of the moon in a cartoon style referring to the target data".

[0112] It can be understood that the above describes the method of the present disclosure, and the method of the present disclosure will be further described in combination with the code generation scene.

[0113] Figure 4 is a flowchart of a data generation method according to another embodiment of the present disclosure.

[0114] As shown in Figure 4 , the data generation method includes operations S401-S402 and operations S410-S440.

[0115] In operation S401, a plurality of initial data is preprocessed to obtain at least one target data.

[0116] For example, a large amount of open source code can be obtained as a plurality of initial data d410. The initial data can include a piece of open source code. By preprocessing the initial data, target data containing labels can be obtained. The labels contained in the target data can include: a text type label "code", a domain label "finance", a function label "SQL query", a language type label "python", and a source label "open source code repository". For another example, the instructions input by the user to the large model in the code generation scene can also be obtained as initial instructions d411. Next, operation S402 can be performed.

[0117] In operation S402, at least one to-be-processed instruction is generated according to at least one of the target data and the initial instruction. For example, the target data is a piece of code, and different to-be-processed instructions can include "explain the meaning of the code", "modify the code", "add comments to the code", and the like.

[0118] In operation S410, at least one to-be-evaluated response result is generated by using at least one data synthesis expert unit based on at least one of the target data, the to-be-processed instruction, and the correction data.

[0119] In an embodiment of the present disclosure, in the case of initial data synthesis, there is no correction data for the data synthesis expert unit. The target data and the to-be-processed instruction can be input into a plurality of data synthesis expert units respectively to obtain a plurality of to-be-evaluated response results.

[0120] For example, according to different to-be-processed instructions, the plurality of data synthesis expert units can give diversified to-be-evaluated response results. Next, taking that the to-be-processed instruction is "adding comments to the code" as an example for description.

[0121] In embodiments of the present disclosure, the target data and some other labels can also be taken as model inputs, and the other labels are taken as supplementary descriptions of the target data, to obtain corresponding to-be-evaluated response results.

[0122] For example, the target data is a piece of open source code containing "finance", "code", and "python" labels, and the target data and the label "improvement" and "adding comments" are input into the data synthesis expert unit to obtain a plurality of to-be-evaluated response results.

[0123] In some other embodiments of the present disclosure, after operation S401 is completed, operation S402 can also be selected not to be performed, and operation S410 is performed on the basis of the target data obtained in operation S401, and the same number of to-be-evaluated response results are generated according to the target data by using at least one data synthesis expert unit. For example, the target data is a piece of code and the labels "finance" and "python". The target data is input into a plurality of data synthesis expert units, and each data synthesis expert unit can process the target data based on different angles to obtain to-be-evaluated response results after code modification, code correction, and code imitation. The to-be-evaluated response results are code comment results. Next, still taking that operation S402 is performed as an example, description is made.

[0124] In operation S420, the evaluation expert evaluates the to-be-evaluated response results to determine whether there is a to-be-corrected response result. If at least one evaluation result indicates that at least one to-be-evaluated response result includes a to-be-corrected response result, operation S430 is performed. If at least one evaluation result indicates that at least one to-be-evaluated response result does not include a to-be-corrected response result, operation S440 is performed.

[0125] In operation S430, at least one correction data for at least one data synthesis expert unit is determined according to at least one evaluation result for at least one to-be-corrected response result.

[0126] In embodiments of the present disclosure, the reflection correction expert unit can be used to give correction data corresponding to the to-be-corrected response result based on the evaluation result. The data synthesis expert unit can generate to-be-evaluated response results more in line with the preset evaluation condition on the basis of the correction data.

[0127] For example, the target data is a piece of code. The data synthesis expert unit determines the to-be-evaluated response result based on the target data to be the annotation result of the code. After obtaining the evaluation result for the to-be-evaluated response result, the reflection correction expert can determine the correction data to be "the target data is code instead of the to-be-translated sentence". Based on the correction data and the target data, the data synthesis expert unit can regenerate the to-be-evaluated response result that meets the preset evaluation condition.

[0128] After obtaining the at least one correction data for the at least one data synthesis expert unit, operation S410 can be performed again. After one or more data generations based on the correction data are performed, if the at least one evaluation result indicates that there is no to-be-corrected response result in the to-be-evaluated response result, operation S440 can be performed.

[0129] In operation S440, the target response result is determined from the at least one to-be-evaluated response result.

[0130] According to an embodiment of the present disclosure, through the plurality of to-be-processed instructions, the data synthesis expert unit can answer from different angles and different requirements, ensuring the diversity of the response. Moreover, the generated response result is reflected and analyzed, and the data is regenerated. Through such a synthesis method, high-quality and diverse training samples can be efficiently and accurately obtained, helping the large model to achieve iterative improvement.

[0131] Figure 5 is a flowchart of a model training method according to an embodiment of the present disclosure.

[0132] As shown in Figure 5 , the model training method 500 includes operation S510- operation S520.

[0133] In operation S510, the target data in the sample data pair is input into the to-be-trained model to obtain a to-be-optimized response result.

[0134] The to-be-optimized response result can be a response result generated by the to-be-trained model based on the target data.

[0135] In operation S520, the to-be-trained model is trained according to the to-be-optimized response result and the target response result in the sample data pair.

[0136] The loss between the target response result and the to-be-optimized response result can be calculated, and the parameters of the to-be-trained model are adjusted based on the loss until the loss between the target response result and the to-be-optimized response result is less than a preset loss threshold, and a trained target model is obtained.

[0137] In an embodiment of the present disclosure, the sample data pair is determined by: generating at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units, the data synthesis expert unit being configured to generate the to-be-evaluated response result according to the target data; determining at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; and in response to determining that the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, determining the sample data pair according to the target response result in the at least one to-be-evaluated response result and the target data. For example, the sample data pair in operation S510 can be determined based on the target data and the to-be-evaluated response result in the data processing method 200, and will not be described again for the sake of simplicity.

[0138] In an embodiment of the present disclosure, the sample data pair is also determined by: in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determining at least one correction data for at least one data synthesis expert unit according to the at least one evaluation result for the at least one to-be-corrected response result; and returning to the operation of generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, wherein the data synthesis expert unit is configured to generate the to-be-evaluated response result according to at least one of the target data and the received correction data.

[0139] According to an embodiment of the present disclosure, by performing a series of processes such as screening, synthesis, and correction on the target data by using the plurality of expert large models, a sample data pair with diversity and precision is obtained, and a high-quality sample data pair with diversity and balanced distribution is obtained by training the model according to the sample data, which helps the model to realize rapid iteration and improves the generalization ability of the model.

[0140] Figure 6 is a flowchart of a data processing method according to an embodiment of the present disclosure.

[0141] As shown in Figure 6 , the data processing method 600 includes operation S610.

[0142] In operation S610, the to-be-processed data is input into the target model to obtain a response result corresponding to the to-be-processed data.

[0143] The target model can be a model obtained by training a to-be-trained model based on the method 500 described above using the sample data pair. By inputting the to-be-processed data into the target model, a precise, diversified, and high-quality response result can be obtained.

[0144] Figure 7is a block diagram of a data generation apparatus according to an embodiment of the present disclosure.

[0145] As shown in Figure 7 The data generation apparatus 700 can include a result generation module 710, a result determination module 720, and a data determination module 730.

[0146] The result generation module 710 is configured to generate at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units.

[0147] The result determination module 720 is configured to determine at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units.

[0148] The data determination module 730 is configured to, in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determine at least one correction data for at least one data synthesis expert unit according to the at least one evaluation result for the at least one to-be-corrected response result. Return to the operation of generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units, wherein the data synthesis expert unit is configured to generate the to-be-evaluated response result according to at least one of the target data and the received correction data.

[0149] According to an embodiment of the present disclosure, the data generation apparatus 700 further includes a preprocessing module. The preprocessing module is configured to preprocess a plurality of initial data to obtain at least one target data, wherein the target data has at least one label.

[0150] According to an embodiment of the present disclosure, the preprocessing module includes a filtering submodule and a labeling submodule. The filtering submodule is configured to filter the plurality of initial data according to a preset filtering rule to obtain at least one intermediate data. The labeling submodule is configured to label the at least one intermediate data to obtain the at least one target data.

[0151] According to an embodiment of the present disclosure, the result generation module 710 includes a sub-problem determination submodule. The sub-problem determination submodule is configured to generate the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units based on at least one to-be-handled sub-problem determined by a data understanding result. The data understanding result is obtained according to at least one of the target data and a to-be-handled instruction, and the to-be-handled instruction is obtained according to the target data.

[0152] According to an embodiment of the present disclosure, the data understanding result is obtained by a question understanding expert unit according to the target data.

[0153] According to an embodiment of the present disclosure, the data understanding result is obtained by the problem understanding expert unit according to the to-be-processed instruction, and the to-be-processed instruction is obtained by the instruction synthesis expert unit according to the target data.

[0154] According to an embodiment of the present disclosure, the sub-problem determining sub-module comprises a disassembling unit and a generating unit. The disassembling unit is configured to determine at least one to-be-processed sub-problem by using the step disassembling expert unit according to the data understanding result. The generating unit is configured to generate at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units based on the at least one to-be-processed sub-problem.

[0155] According to an embodiment of the present disclosure, the result determining module 720 comprises an intermediate evaluation sub-module and an evaluation result determining sub-module. The intermediate evaluation sub-module is configured to determine N intermediate evaluation data for a to-be-evaluated intermediate result by using the N data synthesis expert units. The to-be-evaluated intermediate result is obtained by masking the identification information related to the data synthesis expert unit that generates the to-be-evaluated response result in the to-be-evaluated response result, and N is an integer greater than or equal to 1. The evaluation result determining sub-module is configured to determine an evaluation result for the to-be-evaluated response result according to the N intermediate evaluation data for the to-be-evaluated intermediate result.

[0156] According to an embodiment of the present disclosure, the evaluation result comprises at least one evaluation index value, the at least one evaluation index value comprises at least one of a first evaluation index value and a second evaluation index value, and the data generation apparatus 700 further comprises a first correction determining sub-module and a second correction determining sub-module. The first correction determining sub-module is configured to determine that the to-be-evaluated response result is not a to-be-corrected response result in response to determining that the plurality of evaluation index values for the to-be-evaluated response result satisfy at least one preset evaluation condition. The second correction determining sub-module is configured to determine that the to-be-evaluated response result is a to-be-corrected response result in response to determining that the plurality of evaluation index values for the to-be-evaluated response result do not satisfy at least one of the at least one preset evaluation condition. The at least one preset evaluation condition comprises at least one of the following: the first evaluation index value is greater than or equal to a first evaluation threshold value; and the second evaluation index value is less than or equal to a second evaluation threshold value.

[0157] According to an embodiment of the present disclosure, the data generation apparatus 700 further comprises a sample determining module configured to determine a sample data pair according to a target response result and the target data in the at least one to-be-evaluated response result in response to determining that the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result.

[0158] According to an embodiment of the present disclosure, the sample determination module comprises a response result determination submodule and a sample determination submodule. The response result determination submodule is configured to determine a target response result from the at least one to-be-evaluated response result. The sample determination submodule is configured to, in response to determining that the target response result meets at least one preset data generation condition, determine a sample data pair, the sample data pair comprising at least one of the target data and a to-be-processed instruction derived from the target data and the target response result. The at least one preset data generation condition comprises that a data format of the target response result is a preset data format.

[0159] According to an embodiment of the present disclosure, the data synthesis expert unit is configured to generate the to-be-evaluated response result according to at least one of the at least one label, the target data, and the received correction data.

[0160] According to an embodiment of the present disclosure, the data synthesis expert unit is a data synthesis large model or a data synthesis intelligent agent.

[0161] According to an embodiment of the present disclosure, the target data comprises at least one of target text data, target audio data, target image data, and target video data.

[0162] Figure 8 is a block diagram of a model training apparatus according to an embodiment of the present disclosure.

[0163] As shown in Figure 8 , the model training apparatus 800 can comprise a first input module 810 and a training module 820.

[0164] The first input module 810 is configured to input the target data in the sample data pair into the to-be-trained model to obtain a to-be-optimized response result.

[0165] The training module 820 is configured to train the to-be-trained model according to the to-be-optimized response result and the target response result in the sample data pair.

[0166] According to an embodiment of the present disclosure, in combination with Figure 7 and Figure 8 , the sample data pair is determined by the result generation module 710, the result determination module 720, and the data determination module 730 in the data generation apparatus 700: at least one to-be-evaluated response result is generated by using at least one of a plurality of data synthesis expert units, the data synthesis expert unit being configured to generate the to-be-evaluated response result according to the target data; at least one evaluation result for the at least one to-be-evaluated response result is determined by using at least one of the plurality of data synthesis expert units; in response to determining that the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, the sample data pair is determined according to the target response result in the at least one to-be-evaluated response result and the target data.

[0167] Figure 9 is a block diagram of a data processing apparatus according to one embodiment of the present disclosure.

[0168] As shown in Figure 9 , the data processing apparatus 900 includes a second input module 910.

[0169] The second input module 910 is configured to input the to-be-processed data into the target model to obtain a response result corresponding to the to-be-processed data, where the target model is obtained by training the training model according to the apparatus described in the embodiments of the present disclosure.

[0170] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.

[0171] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0172] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0173] As shown in Figure 10 , the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes according to a computer program stored in a Read-Only Memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a Random Access Memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0174] A number of the components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through computer networks, such as the Internet, and / or various telecommunication networks.

[0175] The computing unit 1001 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a Central Processing Unit (CPU), a Graph Processing Unit (GPU), various special-purpose Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the data generation method, the model training method, and the data processing method. For example, in some embodiments, the data generation method, the model training method, and the data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded to the RAM 1003 and executed by the computing unit 1001, one or more steps of the data generation method, the model training method, and the data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the data generation method, the model training method, and the data processing method by any other appropriate means, such as by means of firmware.

[0176] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), an Application Specific Standard Parts (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0177] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as part of a separate software package, and partially on a remote machine or entirely on a remote machine or server.

[0178] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0179] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a Cathode Ray Tube (CRT) monitor or a Liquid Crystal Display (LCD)) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0180] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0181] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0182] It should be understood that the steps shown in the various forms of flow above can be reordered, added to, or deleted from. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which are not limited herein.

[0183] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data generation method, comprising: generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units; determining at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determining at least one correction data for at least one of the data synthesis expert units according to the at least one evaluation result for the at least one to-be-corrected response result; returning to the operation of generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, wherein the data synthesis expert unit is used to generate the to-be-evaluated response result according to at least one of target data and received correction data. 2.The method of claim 1, further comprising: preprocessing a plurality of initial data to obtain at least one target data, wherein the target data has at least one label.

3. The method of claim 2, wherein, The preprocessing a plurality of initial data to obtain at least one target data comprises: filtering a plurality of the initial data according to a preset filtering rule to obtain at least one intermediate data; annotating at least one intermediate data to obtain at least one target data.

4. The method of claim 1, wherein, The generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units comprises: generating at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units based on at least one to-be-processed sub-problem determined by a data understanding result, wherein the data understanding result is obtained according to at least one of the target data and a to-be-processed instruction, and the to-be-processed instruction is obtained according to the target data.

5. The method of claim 4, wherein, The data understanding result is obtained by a question understanding expert unit according to the target data.

6. The method of claim 4, wherein, The data understanding result is obtained by a question understanding expert unit according to the to-be-processed instruction, and the to-be-processed instruction is obtained by an instruction synthesis expert unit according to the target data.

7. The method of claim 4, wherein, The generating at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units based on at least one to-be-processed sub-problem determined by a data understanding result comprises: determining at least one to-be-processed sub-problem by using a step decomposition expert unit according to the data understanding result; generating at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units based on at least one to-be-processed sub-problem.

8. The method of claim 1, wherein, The determining at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units comprises: N data synthesis expert units are used to determine N intermediate evaluation data for an intermediate result to be evaluated, wherein the intermediate result to be evaluated is obtained by masking identification information related to a data synthesis expert unit generating the response result to be evaluated in the response result to be evaluated, and N is an integer greater than or equal to 1; According to the N intermediate evaluation data for the intermediate result to be evaluated, the evaluation result for the response result to be evaluated is determined.

9. The method of claim 1, wherein, The evaluation result includes at least one evaluation index value, and the at least one evaluation index value includes at least one of a first evaluation index value and a second evaluation index value, Further comprising: In response to determining that the plurality of evaluation index values for the response result to be evaluated satisfy at least one preset evaluation condition, determining that the response result to be evaluated is not a response result to be corrected; In response to determining that the plurality of evaluation index values for the response result to be evaluated do not satisfy at least one of the at least one preset evaluation condition, determining that the response result to be evaluated is a response result to be corrected, The at least one preset evaluation condition includes at least one of: The first evaluation index value is greater than or equal to a first evaluation threshold value; The second evaluation index value is less than or equal to a second evaluation threshold value.

10. The method of claim 1, further comprising: In response to determining that at least one of the evaluation results indicates that there is no response result to be corrected in at least one of the response results to be evaluated, determining a sample data pair according to a target response result in at least one of the response results to be evaluated and the target data.

11. The method of claim 10, wherein the determining a sample data pair according to a target response result in at least one of the response results to be evaluated and the target data comprises: Determining a target response result from at least one of the response results to be evaluated; In response to determining that the target response result satisfies at least one preset data generation condition, determining a sample data pair, the sample data pair including at least one of the target data and a to-be-processed instruction derived from the target data and the target response result, Wherein, the at least one preset data generation condition includes that the data format of the target response result is a preset data format.

12. The method of claim 2, wherein, The data synthesis expert unit is used to generate the response result to be evaluated according to at least one of the at least one label, the target data, and the received correction data.

13. The method of claim 1, wherein, The data synthesis expert unit is a data synthesis large model or a data synthesis intelligent agent, The operation of returning to generating at least one response result to be evaluated using at least one of the plurality of data synthesis expert units until at least one of the evaluation results indicates that there is no response result to be corrected in at least one of the response results to be evaluated comprises: Returning to the operation of generating at least one response result to be evaluated using at least one of the plurality of data synthesis expert units until at least one of the evaluation results indicates that there is no response result to be corrected in at least one of the response results to be evaluated.

14. The method of claim 1, wherein, The target data includes at least one of target text data, target audio data, target image data, and target video data.

15. A model training method comprising: inputting target data in a sample data pair into a to-be-trained model to obtain a to-be-optimized response result; training the to-be-trained model according to the to-be-optimized response result and a target response result in the sample data pair, wherein the sample data pair is determined by: generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units; determining at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determining at least one correction data for at least one of the data synthesis expert units according to the at least one evaluation result for the at least one to-be-corrected response result; returning to the operation of generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, and determining a sample data pair according to a target response result in the at least one to-be-evaluated response result and the target data, the data synthesis expert unit being used to generate the to-be-evaluated response result according to at least one of the target data and received correction data.

16. A data processing method comprising: inputting to-be-processed data into a target model to obtain a response result corresponding to the to-be-processed data, wherein the target model is obtained by training a training model according to the method of claim 15.

17. A data generation apparatus comprising: a result generation module configured to generate at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units; a result determination module configured to determine at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; a data determination module configured to, in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determine at least one correction data for at least one of the data synthesis expert units according to the at least one evaluation result for the at least one to-be-corrected response result; returning to the operation of generating at least one to-be-evaluated response result by using at least one of a plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, wherein the data synthesis expert unit is used to generate the to-be-evaluated response result according to at least one of the target data and received correction data.

18. A model training apparatus comprising: a first input module configured to input target data in a sample data pair into a to-be-trained model to obtain a to-be-optimized response result; a training module configured to train the to-be-trained model according to the to-be-optimized response result and a target response result in the sample data pair, wherein the sample data pair is determined by: a result generation module configured to generate at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; a result determination module configured to determine at least one evaluation result for the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units; a data determination module configured to, in response to determining that the at least one evaluation result indicates that there is at least one to-be-corrected response result in the at least one to-be-evaluated response result, determine at least one correction data for at least one of the plurality of data synthesis expert units according to the at least one evaluation result for the at least one to-be-corrected response result; return to the operation of generating the at least one to-be-evaluated response result by using at least one of the plurality of data synthesis expert units until the at least one evaluation result indicates that there is no to-be-corrected response result in the at least one to-be-evaluated response result, the data synthesis expert unit being configured to generate the to-be-evaluated response result according to at least one of target data and received correction data; a sample determination module configured to determine sample data pairs according to target response results in the at least one to-be-evaluated response result and the target data.

19. A data processing apparatus, comprising: a second input module configured to input to-be-processed data into a target model to obtain a response result corresponding to the to-be-processed data, wherein the target model is obtained by training a training model by the apparatus of claim 18.

20. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 14.

21. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1 to 14.

22. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 14.

22. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 14.

Citation Information

Patent Citations

  • Dynamic analysis method and device based on simulation model and simulation equipment test result

    CN118261049A

  • Training sample generation method, training method and information evaluation method and device

    CN118964998A