Data generation method and device, electronic equipment and storage medium

By using a large model to determine the Q&A mode and generate the Q&A data pair, the problem of difficult data quality in the prior art is solved, and higher quality and flexible data generation is achieved.

CN120146196APending Publication Date: 2025-06-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308127.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When obtaining training data of large models, the prior art relies on manual annotation or single-stage automatic annotation, making it difficult to ensure data quality, especially when facing complex target text or data generation needs.

Method used

By obtaining the target text, the first large model determines the optional Q&A pattern related to the target text, and uses it as the target Q&A pattern, and uses the second large model to generate Q&A data pairs that match the target Q&A pattern, and performs pattern determination and data generation in stages.

Benefits of technology

This method does not require manual participation, reduces human resource consumption, reduces subjective bias, improves the quality of Q&A data pairs, and shows stronger task-orientedness when facing complex needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146196A_ABST
    Figure CN120146196A_ABST
Patent Text Reader

Abstract

The invention provides a data generation method and device, electronic equipment and a storage medium, and relates to the technical field of computers, in particular to the application fields of machine learning, large models, generative artificial intelligence, question and answer robots and the like. According to the specific implementation scheme, a target text is obtained; determining at least one selectable question and answer mode related to the target text by utilizing the first large model; and taking the selectable question-answering mode as a target question-answering mode to generate a question-answering data pair matched with the target question-answering mode based on the target text by utilizing the second large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to application fields such as machine learning, large models, generative artificial intelligence, and question-and-answer robots. Specifically, it relates to a data generation method, apparatus, electronic device, and storage medium. Background Art

[0002] A large model refers to a "large parameter" model trained using large-scale data and powerful computing capabilities, which usually has high generality and generalization capabilities. In order to further improve the specific capabilities of large models, especially those with a scale of billions or tens of billions of parameters (e.g., inference computing capabilities and / or discrimination capabilities), it is often necessary to provide some relevant training data for targeted training of large models. Summary of the Invention

[0003] The present disclosure provides a data generation method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of the present disclosure, there is provided a data generation method, including:

[0005] Obtaining a target text;

[0006] Using a first large model to determine at least one optional question-and-answer pattern related to the target text;

[0007] Taking the optional question-and-answer pattern as a target question-and-answer pattern, and using a second large model to generate a question-and-answer data pair that matches the target question-and-answer pattern based on the target text.

[0008] According to a second aspect of the present disclosure, there is provided a data generation apparatus, including:

[0009] A text acquisition unit for obtaining a target text;

[0010] A pattern determination unit for using a first large model to determine at least one optional question-and-answer pattern related to the target text;

[0011] A data generation unit for taking the optional question-and-answer pattern as a target question-and-answer pattern, and using a second large model to generate a question-and-answer data pair that matches the target question-and-answer pattern based on the target text.

[0012] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0013] At least one processor;

[0014] A memory communicatively connected to the at least one processor;

[0015] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to execute the method provided in the first aspect of the present disclosure.

[0016] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are for causing the computer to execute the method provided in the first aspect of the present disclosure.

[0017] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided in the first aspect of the present disclosure.

[0018] Adopting the present disclosure can improve the quality of question-and-answer data pairs.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0021] Figure 1 is a schematic flowchart of a data generation method provided by an embodiment of the present disclosure;

[0022] Figure 2 is a complete flowchart illustration of a data generation method provided by an embodiment of the present disclosure;

[0023] Figure 3 is a schematic diagram of an application scenario of a data generation method provided by an embodiment of the present disclosure;

[0024] Figure 4 is a schematic structural block diagram of a data generation device provided by an embodiment of the present disclosure;

[0025] Figure 5 is a schematic structural block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiments

[0026] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] As described above, in order to further improve the capabilities of large models, especially those with billions or tens of billions of parameters (e.g., inference computing capabilities and / or discrimination capabilities), it is often necessary to provide some relevant training data for targeted training of large models. Currently, these training data are usually obtained in the following ways:

[0028] Obtain the target text;

[0029] By means of manual annotation or automatic annotation of the target text, obtain question-and-answer data pairs as training data.

[0030] Among them, manual annotation is to analyze the target text by annotators and manually design the corresponding question-and-answer data pairs. This process not only consumes a large amount of human resources, but also strongly depends on the professional knowledge and discrimination capabilities of annotators. Therefore, there is a large degree of subjectivity and the quality of the question-and-answer data pairs cannot be ensured. Automatic annotation is to use a large model to generate question-and-answer data pairs at one time based on the target text. Since the question-and-answer data pairs are generated at one time without dividing the generation task of the question-and-answer data pairs into stages, the quality of the question-and-answer data pairs often cannot be ensured when facing complex target texts or complex data generation requirements.

[0031] In view of the above problems, the embodiments of the present disclosure provide a data generation method, which can be applied to an electronic device. Among them, the electronic device can be a server or a terminal device. Here, the terminal device can be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a tablet computer, etc.), a personal digital processor or other similar computing devices. Hereinafter, in combination with Figure 1 the following flow schematic diagram, a data generation method provided by the embodiments of the present disclosure will be described. It should be noted that although the logical order is shown in the flow schematic diagram, in some cases, the steps shown or described in the flowchart may also be executed in other orders.

[0032] Step S101, obtain the target text.

[0033] Among them, the target text can be sourced from the document to be annotated. For example, the target text can be a text segment intercepted from the document to be annotated; or, it can include the text segment intercepted from the document to be annotated and the title content of the text segment. Here, the document to be annotated can be a knowledge document in a specific field, such as the field of computer technology, the medical field, the chemical engineering field, the cultural and sports field, etc.

[0034] Step S102, use the first large model to determine at least one optional question-and-answer pattern (Schema) related to the target text.

[0035] Among them, the first large model can be a "large parameter" model trained using large-scale data and powerful computing capabilities, also known as a large language model (LLM). Specifically, it can be a pre-trained neural network model (e.g., an autoregressive generation model with a Transformer architecture), which has general language knowledge, world knowledge, professional knowledge in various fields, etc.

[0036] In addition, it should be noted that in the embodiments of the present disclosure, the optional Q&A mode can represent the Q&A data type, that is, it can represent the type of Q&A data pairs that can be generated based on the target text. Specifically, it can include an inference calculation type and a discrimination type. Among them, the inference calculation type can further include an inference calculation type related to a numerical range and an inference calculation type related to a threshold; the discrimination type can further include a discrimination type related to a numerical range and a discrimination type involving multi-condition judgment.

[0037] Step S103: Use each optional Q&A mode as the target Q&A mode to generate, by means of the second large model, a Q&A data pair that matches the target Q&A mode based on the target text.

[0038] Specifically, each optional Q&A mode in at least one optional Q&A mode can be used as the target Q&A mode to generate, by means of the second large model, a Q&A data pair that matches the target Q&A mode based on the target text.

[0039] Among them, the second large model can be a "large parameter" model trained using large-scale data and powerful computing capabilities, also known as an LLM. Specifically, it can be a pre-trained neural network model (e.g., an autoregressive generation model with a Transformer architecture), which has general language knowledge, world knowledge, professional knowledge in various fields, etc. Here, the second large model and the first large model can be the same large model or two independent large models. The embodiments of the present disclosure do not limit this.

[0040] In addition, it should be noted that in the embodiments of the present disclosure, the Q&A data pair can include a question sample and an answer sample for answering the question sample. It should also be noted that in the embodiments of the present disclosure, the matching of the Q&A data pair with the target Q&A mode can be understood as: the type of the Q&A data pair is the Q&A data type represented by the Q&A mode.

[0041] By using the data generation method provided in the embodiments of the present disclosure, the target text can be obtained, and at least one optional Q&A mode related to the target text can be determined by using the first large model. Then, the optional Q&A mode is used as the target Q&A mode to generate Q&A data pairs that match the target Q&A mode based on the target text by using the second large model. On the one hand, this process only relies on the invocation of large models and does not require manual participation. It can not only greatly reduce the consumption of human resources, but also effectively reduce the subjective biases that may be brought by manual participation, thereby improving the quality of Q&A data pairs. On the other hand, in this process, the generation task of Q&A data pairs is divided into two stages: mode determination and data generation. Compared with the traditional single-stage generation method (that is, the method of "using a large model to generate Q&A data pairs at one time based on the target text"), it will show stronger task pertinence when facing complex target texts or complex data generation requirements, thereby improving the quality of Q&A data pairs.

[0042] In some alternative embodiments, step S102, that is, "using the first large model to determine at least one optional Q&A mode related to the target text" may include:

[0043] Determine a plurality of candidate Q&A modes;

[0044] Generate a first prompt message based on the target text and the plurality of candidate Q&A modes;

[0045] Use the first large model to determine at least one optional Q&A mode related to the target text from the plurality of candidate Q&A modes based on the first prompt message.

[0046] Among them, the plurality of candidate Q&A modes can be a plurality of preset original Q&A modes, or can be obtained by performing a correction operation on the plurality of original Q&A modes in response to a user correction operation. Here, the plurality of original Q&A modes can include at least one of an inference calculation mode and a discrimination mode.

[0047] In one example, the inference calculation mode can include a first inference calculation mode and a second inference calculation mode. Among them, the first inference calculation mode can be an inference calculation mode related to a numerical range, which can represent an inference calculation type related to a numerical range; the second inference calculation mode can be an inference calculation mode related to a threshold, which can represent an inference calculation type related to a threshold.

[0048] In another example, the discrimination mode can include a first discrimination mode and a second discrimination mode. Among them, the first discrimination mode can be a discrimination mode related to a numerical range, which can represent a discrimination type related to a numerical range; the second discrimination mode can be a discrimination mode involving multi-condition judgment, which can represent a discrimination type involving multi-condition judgment.

[0049] Further, in the embodiments of the present disclosure, the following method can be used to determine multiple candidate Q&A patterns:

[0050] Determine multiple original Q&A patterns;

[0051] In the case where no pattern correction indication is detected, use the original Q&A patterns as candidate Q&A patterns;

[0052] Alternatively, in the case where a pattern correction indication is detected, based on the pattern correction indication, perform a correction operation on the multiple original Q&A patterns to obtain multiple candidate Q&A patterns.

[0053] In one example, in the case where no pattern correction indication is detected, each of the multiple original Q&A patterns can be used as a candidate Q&A pattern. In addition, in the embodiments of the present disclosure, the pattern correction indication can be a pattern correction instruction generated in response to a user correction operation; the correction operation can include at least one of an amplification operation, a replacement operation, and a deletion operation.

[0054] Further, in the embodiments of the present disclosure, in the case where a pattern correction indication is detected, the pattern correction indication can be parsed to obtain an indication parsing result. Thereafter, in the case where the indication parsing result indicates that an amplification operation needs to be performed on the multiple original Q&A patterns, perform an amplification operation on the multiple original Q&A patterns according to the indication parsing result to obtain multiple candidate Q&A patterns; in the case where the indication parsing result indicates that a replacement operation needs to be performed on the multiple original Q&A patterns, perform a replacement operation on the multiple original Q&A patterns according to the indication parsing result to obtain multiple candidate Q&A patterns; in the case where the indication parsing result indicates that a deletion operation needs to be performed on the multiple original Q&A patterns, perform a deletion operation on the multiple original Q&A patterns according to the indication parsing result to obtain multiple candidate Q&A patterns. Among them, performing a deletion operation on the multiple original Q&A patterns can be to delete some of the original Q&A patterns in the multiple original Q&A patterns, while retaining the other part of the original Q&A patterns, and using the remaining other part of the original Q&A patterns as candidate Q&A patterns.

[0055] In addition, in the embodiments of the present disclosure, after determining multiple candidate Q&A patterns, it is necessary to generate a first prompt message based on the target text and the multiple candidate Q&A patterns, so as to use a first large model to determine at least one optional Q&A pattern related to the target text from the multiple candidate Q&A patterns based on the first prompt message. Among them, the first indication information can include the following three parts:

[0056] (1) The target text;

[0057] (2) The multiple candidate Q&A patterns;

[0058] (3) First instruction statement: Please determine at least one optional Q&A mode related to the target text from multiple candidate Q&A modes.

[0059] In the above manner, in the embodiments of the present disclosure, multiple candidate Q&A modes can be determined, and based on the target text and multiple candidate Q&A modes, a first prompt message is generated. Then, using the first large model, based on the first prompt message, at least one optional Q&A mode related to the target text is determined from multiple candidate Q&A modes. During this process, multiple candidate Q&A modes are provided to the first large model in the first prompt message, so that the first large model can more accurately determine at least one optional Q&A mode related to the target text. In this way, when each optional Q&A mode in at least one optional Q&A mode is used as the target Q&A mode to utilize the second large model to generate a Q&A data pair that matches the target Q&A mode based on the target text, the quality of the Q&A data pair can be further improved.

[0060] Moreover, when determining multiple candidate Q&A modes, multiple original Q&A modes can be determined. In the case where no mode correction instruction is detected, the original Q&A mode is used as the candidate Q&A mode; or, in the case where a mode correction instruction is detected, based on the mode correction instruction, a correction operation is performed on multiple original Q&A modes to obtain multiple candidate Q&A modes. That is to say, in the embodiments of the present disclosure, multiple candidate Q&A modes can also be dynamically adjusted to improve the flexibility and controllability of the data generation system including the first large model and the second large model, so that the data generation system can more efficiently meet the generation requirements of different types of Q&A data pairs.

[0061] Further, in the embodiments of the present disclosure, "using the first large model, based on the first prompt message, determining at least one optional Q&A mode related to the target text from multiple candidate Q&A modes" can specifically be: using the first large model, based on the first prompt message, determining at least one optional Q&A mode related to the target text from multiple candidate Q&A modes, and the number of Q&A pairs corresponding to the optional Q&A mode, specifically, the number of Q&A pairs corresponding to each optional Q&A mode in at least one optional Q&A mode.

[0062] In one example, when the first large model determines at least one optional Q&A mode related to the target text from multiple candidate Q&A modes based on the first prompt message, it can obtain the number of Q&A pairs corresponding to each optional Q&A mode in at least one optional Q&A mode based on its general language knowledge, world knowledge, professional knowledge in various fields, etc.

[0063] In another example, hint supplementary information can also be used to supplement the first hint information to obtain the supplemented first hint information, so as to use the first large model to determine at least one optional Q&A pattern related to the target text and the number of Q&A pairs corresponding to each optional Q&A pattern in the at least one optional Q&A pattern based on the supplemented first hint information. Among them, the hint supplementary information can be: for each optional Q&A pattern in the at least one optional Q&A pattern, when the optional Q&A pattern is an inference calculation pattern, the number of knowledge points included in the target text is used as the number of Q&A pairs corresponding to the optional Q&A pattern; when the optional Q&A pattern is a discrimination pattern, twice the number of knowledge points included in the target text is used as the number of Q&A pairs corresponding to the optional Q&A pattern.

[0064] For example, the target text includes:

[0065] Title content: XXX part description;

[0066] Text fragment: The allowable working temperature of this part is 0 to 100 degrees Celsius. Among them, 0 to 10 degrees Celsius is the extreme working condition; 60 to 70 degrees Celsius is the suitable working condition.

[0067] Among them, the knowledge points included in the target text are: the allowable working temperature of the XXX part is 0 to 100 degrees Celsius, the extreme working condition of the XXX part is 0 to 10 degrees Celsius, and the suitable working condition of the XXX part is 60 to 70 degrees Celsius. That is, the number of knowledge points included in the target text is 3. Therefore, for each optional Q&A pattern in the at least one optional Q&A pattern, when the optional Q&A pattern is an inference calculation pattern, 3 is used as the number of Q&A pairs corresponding to the optional Q&A pattern; when the optional Q&A pattern is a discrimination pattern, 6 is used as the number of Q&A pairs corresponding to the optional Q&A pattern.

[0068] For another example, the target text includes:

[0069] Title content: XXX part description;

[0070] Text fragment: The lowest threshold of the allowable working temperature range of this part is 0 degrees Celsius, and the highest threshold is 100 degrees Celsius.

[0071] Among them, the knowledge points included in the target text are: the lowest threshold of the allowable working temperature range of the XXX part is 0 degrees Celsius, and the highest threshold of the allowable working temperature range of the XXX part is 100 degrees Celsius. That is, the number of knowledge points included in the target text is 2. Therefore, for each optional question-and-answer mode in at least one optional question-and-answer mode, when the optional question-and-answer mode is the reasoning calculation mode, 2 is used as the number of question-and-answer pairs corresponding to the optional question-and-answer mode; when the optional question-and-answer mode is the discrimination mode, 4 is used as the number of question-and-answer pairs corresponding to the optional question-and-answer mode.

[0072] Based on this, in the embodiments of the present disclosure, "using the second large model, based on the target text, to generate question-and-answer data pairs matching the target question-and-answer mode" in step S103 may include:

[0073] Generate a second prompt message based on the target text, the target question-and-answer mode, and the number of question-and-answer pairs corresponding to the target question-and-answer mode;

[0074] Using the second large model, based on the second prompt message, generate the target number of question-and-answer data pairs matching the target question-and-answer mode.

[0075] Among them, the target number may be the number of question-and-answer pairs corresponding to the target question-and-answer mode.

[0076] In one example, the second indication information may include the following four parts:

[0077] (1) The target text;

[0078] (2) The target question-and-answer mode;

[0079] (3) The number of question-and-answer pairs corresponding to the target question-and-answer mode;

[0080] (4) The second indication statement: Based on the target text, generate the target number of question-and-answer data pairs matching the target question-and-answer mode, and ensure that the target number is equal to the number of question-and-answer pairs corresponding to the target question-and-answer mode.

[0081] It should be noted that in the embodiments of the present disclosure, when the target question-and-answer mode is the reasoning calculation mode, the target number of question-and-answer data pairs may correspond one-to-one to the knowledge points included in the target text; when the target question-and-answer mode is the discrimination mode, among the target number of question-and-answer data pairs, every two question-and-answer data pairs may correspond to one knowledge point included in the target text, and moreover, these two question-and-answer data pairs may be question-and-answer data pairs with similar question samples and opposite reply samples.

[0082] For example, the target text includes:

[0083] Title content: XXX part description;

[0084] Text segment: The allowable operating temperature of this part is 0 to 100 degrees Celsius. Among them, 0 to 10 degrees Celsius is the extreme operating condition; 60 to 70 degrees Celsius is the suitable operating condition.

[0085] Among them, the knowledge points included in the target text are: the allowable operating temperature of the XXX part is 0 to 100 degrees Celsius, the extreme operating condition of the XXX part is 0 to 10 degrees Celsius, and the suitable operating condition of the XXX part is 60 to 70 degrees Celsius. That is, the number of knowledge points included in the target text is 3.

[0086] Then, in the case where the target Q&A mode is the reasoning calculation mode, 3 Q&A data pairs can be generated, and the 3 Q&A data pairs can include:

[0087] (1) The first Q&A data pair corresponding to the knowledge point "the allowable operating temperature of the XXX part is 0 to 100 degrees Celsius":

[0088] Question sample: What state is the XXX part in when the temperature environment is 90 degrees Celsius?

[0089] Answer sample: Operating state.

[0090] (2) The second Q&A data pair corresponding to the knowledge point "the extreme operating condition of the XXX part is 0 to 10 degrees Celsius":

[0091] Question sample: What operating condition is the XXX part in when the temperature environment is 5 degrees Celsius?

[0092] Answer sample: Extreme operating condition.

[0093] (3) The third Q&A data pair corresponding to the knowledge point "the suitable operating condition of the XXX part is 60 to 70 degrees Celsius":

[0094] Question sample: What operating condition is the XXX part in when the temperature environment is 65 degrees Celsius?

[0095] Answer sample: Suitable operating condition.

[0096] In the case where the target Q&A mode is the discrimination mode, 6 Q&A data pairs can be generated, and the 6 Q&A data pairs can include:

[0097] (1) The first Q&A data pair corresponding to the knowledge point "the allowable operating temperature of the XXX part is 0 to 100 degrees Celsius":

[0098] Question sample: Is the XXX part in the operating state when the temperature environment is 90 degrees Celsius?

[0099] Answer sample: Yes.

[0100] (2) The second Q&A data pair corresponding to the knowledge point "The allowable operating temperature of XXX parts is 0 to 100 degrees Celsius":

[0101] Question sample: Is XXX part in the operating state at a temperature of 200 degrees Celsius?

[0102] Answer sample: No.

[0103] (3) The third Q&A data pair corresponding to the knowledge point "The extreme operating conditions of XXX parts are 0 to 10 degrees Celsius":

[0104] Question sample: Is XXX part in the extreme operating conditions at a temperature of 5 degrees Celsius?

[0105] Answer sample: Yes.

[0106] (4) The fourth Q&A data pair corresponding to the knowledge point "The extreme operating conditions of XXX parts are 0 to 10 degrees Celsius":

[0107] Question sample: Is XXX part in the extreme operating conditions at a temperature of 20 degrees Celsius?

[0108] Answer sample: No.

[0109] (5) The fifth Q&A data pair corresponding to the knowledge point "The suitable operating conditions of XXX parts are 60 to 70 degrees Celsius":

[0110] Question sample: Is XXX part in the suitable operating conditions at a temperature of 65 degrees Celsius?

[0111] Answer sample: Yes.

[0112] (6) The sixth Q&A data pair corresponding to the knowledge point "The suitable operating conditions of XXX parts are 60 to 70 degrees Celsius":

[0113] Question sample: Is XXX part in the suitable operating conditions at a temperature of 80 degrees Celsius?

[0114] Answer sample: No.

[0115] In the above manner, in the embodiments of the present disclosure, the first large model can be used to determine, based on the first prompt information, at least one optional Q&A pattern related to the target text and the number of Q&A pairs corresponding to the optional Q&A pattern from multiple candidate Q&A patterns. When using the second large model to generate Q&A data pairs that match the target Q&A pattern based on the target text, the second prompt information is generated based on the target text, the target Q&A pattern, and the number of Q&A pairs corresponding to the target Q&A pattern, and the second large model is used to generate the target number of Q&A data pairs that match the target Q&A pattern. That is to say, in the embodiments of the present disclosure, for the optional Q&A pattern, the target number of Q&A data pairs that match it can be reasonably controlled, thereby further improving the flexibility and controllability of the data generation system, enabling the data generation system to more efficiently meet the generation requirements of different types of Q&A data pairs.

[0116] Furthermore, the data generation method provided in the embodiments of the present disclosure may further include:

[0117] Determine multiple candidate generation models;

[0118] Based on the target Q&A pattern, determine the second large model from multiple candidate generation models.

[0119] Among them, each candidate generation model among the multiple candidate generation models can be a "large parameter" model trained using large-scale data and powerful computing capabilities, also known as an LLM. Specifically, it can be a pre-trained neural network model (for example, an autoregressive generation model with a Transformer architecture), which has general language knowledge, world knowledge, professional knowledge in various fields, etc.

[0120] More specifically, the multiple candidate generation models may include a long inference large model and a general large model. Based on this, in the embodiments of the present disclosure, when the target Q&A pattern is an inference calculation pattern, the long inference large model can be selected from the multiple candidate generation models as the second large model; when the target Q&A pattern is a discrimination pattern, the general large model can be selected from the multiple candidate generation models as the second large model.

[0121] In the above manner, in the embodiments of the present disclosure, multiple candidate generation models can be determined, and the second large model can be determined from the multiple candidate generation models based on the target Q&A pattern. That is to say, in the embodiments of the present disclosure, in the data generation stage, after each optional Q&A pattern in at least one optional Q&A pattern is used as the target Q&A pattern, the candidate generation model that best matches the target Q&A model can be selected from the multiple candidate generation models as the second large model to further improve the quality of the Q&A data pairs.

[0122] Next, in combination with Figure 2, a complete process of a data generation method provided by an embodiment of the present disclosure will be described.

[0123] First, obtain the target text.

[0124] Among them, the target text can be sourced from the document to be annotated. For example, the target text can be a text snippet intercepted from the document to be annotated; or, it can include the text snippet intercepted from the document to be annotated and the title content of this text snippet. Here, the document to be annotated can be a knowledge document in a specific field, such as the field of computer technology, the medical field, the chemical engineering field, the cultural and sports field, etc.

[0125] Exemplarily, the target text includes:

[0126] Title content: XXX Part Description;

[0127] Text snippet: The allowable operating temperature of this part is 0 to 100 degrees Celsius. Among them, 0 to 10 degrees Celsius is the extreme working condition; 60 to 70 degrees Celsius is the suitable working condition.

[0128] After obtaining the target text, at least one optional Q&A mode related to the target text and the number of Q&A pairs corresponding to each optional Q&A mode in the at least one optional Q&A mode can be determined by using the first large model. Specifically, multiple candidate Q&A modes can be determined, and based on the target text and the multiple candidate Q&A modes, a first prompt message can be generated. Then, using the first large model, based on the first prompt message, at least one optional Q&A mode related to the target text and the number of Q&A pairs corresponding to each optional Q&A mode in the at least one optional Q&A mode can be determined from the multiple candidate Q&A modes. For example, the first prompt message can be supplemented with prompt supplementary information to obtain the supplemented first prompt message, so as to use the first large model to determine, based on the supplemented first prompt message, at least one optional Q&A mode related to the target text and the number of Q&A pairs corresponding to each optional Q&A mode in the at least one optional Q&A mode from the multiple candidate Q&A modes. Among them, the prompt supplementary information can be: for each optional Q&A mode in the at least one optional Q&A mode, when the optional Q&A mode is an inference calculation mode, the number of knowledge points included in the target text is used as the number of Q&A pairs corresponding to this optional Q&A mode; when the optional Q&A mode is a discrimination mode, twice the number of knowledge points included in the target text is used as the number of Q&A pairs corresponding to this optional Q&A mode.

[0129] Continuing with the foregoing example, the target text includes:

[0130] Title content: XXX Part Description;

[0131] Text fragment: The allowable operating temperature of this part is 0 to 100 degrees Celsius. Among them, 0 to 10 degrees Celsius is the extreme working condition; 60 to 70 degrees Celsius is the suitable working condition.

[0132] Among them, the knowledge points included in the target text are: the allowable operating temperature of the XXX part is 0 to 100 degrees Celsius, the extreme working condition of the XXX part is 0 to 10 degrees Celsius, and the suitable working condition of the XXX part is 60 to 70 degrees Celsius. That is, the number of knowledge points included in the target text is 3.

[0133] Suppose that multiple original Q&A modes include the first reasoning and calculation mode, the second reasoning and calculation mode, the first discrimination mode, and the second discrimination mode. Among them, the first reasoning and calculation mode is a reasoning and calculation mode related to a numerical range, which represents the type of reasoning and calculation related to a numerical range; the second reasoning and calculation mode is a reasoning and calculation mode related to a threshold value, which represents the type of reasoning and calculation related to a threshold value; the first discrimination mode is a discrimination mode related to a numerical range, which represents the type of discrimination related to a numerical range; the second discrimination mode is a discrimination mode involving multi-condition judgment, which represents the type of discrimination involving multi-condition judgment. Then, based on the target text and multiple candidate Q&A modes, the first prompt information is generated, and using the first large model, based on the first prompt information, at least one optional Q&A mode related to the target text determined from multiple candidate Q&A modes may include the first reasoning and calculation mode and the first discrimination mode. Moreover, the number of Q&A pairs corresponding to the first reasoning and calculation mode is 3, and the number of Q&A pairs corresponding to the first discrimination mode is 6.

[0134] After using the first large model to determine at least one optional Q&A mode related to the target text and the number of Q&A pairs corresponding to each optional Q&A mode in at least one optional Q&A mode, each optional Q&A mode in at least one optional Q&A mode can be used as the target Q&A mode, and based on the target text, the target Q&A mode, and the number of Q&A pairs corresponding to the target Q&A mode, the second prompt information is generated, and then using the second large model, based on the second prompt information, the target number of Q&A data pairs matching the target Q&A mode is generated. Among them, the second large model can be determined from multiple candidate generation models based on the target Q&A mode after re-determining multiple candidate generation models. More specifically, multiple candidate generation models may include a long reasoning large model and a general large model. Based on this, in the embodiments of the present disclosure, when the target Q&A mode is a reasoning and calculation mode, the long reasoning large model can be selected from multiple candidate generation models as the second large model; when the target Q&A mode is a discrimination mode, the general large model can be selected from multiple candidate generation models as the second large model.

[0135] Continuing with the previous example, when the first inference calculation mode in the first inference calculation mode and the first discrimination mode is used as the target Q&A mode, based on the target text, the target Q&A mode, and the number of Q&A pairs corresponding to the target Q&A mode (i.e., 3), the second prompt information is generated, and the long inference large model is selected from the long inference large model and the general large model as the second large model. Then, using the second large model, the 3 Q&A data pairs that match the target Q&A mode generated based on the second prompt information can include:

[0136] (1) The first Q&A data pair corresponding to the knowledge point "The allowable operating temperature of XXX parts is 0 to 100 degrees Celsius":

[0137] Question sample: What is the state of XXX parts in a temperature environment of 90 degrees Celsius?

[0138] Answer sample: Operating state.

[0139] (2) The second Q&A data pair corresponding to the knowledge point "The extreme operating condition of XXX parts is 0 to 10 degrees Celsius":

[0140] Question sample: What is the operating condition of XXX parts in a temperature environment of 5 degrees Celsius?

[0141] Answer sample: Extreme operating condition.

[0142] (3) The third Q&A data pair corresponding to the knowledge point "The suitable operating condition of XXX parts is 60 to 70 degrees Celsius":

[0143] Question sample: What is the operating condition of XXX parts in a temperature environment of 65 degrees Celsius?

[0144] Answer sample: Suitable operating condition.

[0145] When the first discrimination mode in the first inference calculation mode and the first discrimination mode is used as the target Q&A mode, based on the target text, the target Q&A mode, and the number of Q&A pairs corresponding to the target Q&A mode (i.e., 6), the second prompt information is generated, and the general large model is selected from the long inference large model and the general large model as the second large model. Then, using the second large model, the 6 Q&A data pairs that match the target Q&A mode generated based on the second prompt information can include:

[0146] (1) The first Q&A data pair corresponding to the knowledge point "The allowable operating temperature of XXX parts is 0 to 100 degrees Celsius":

[0147] Question sample: Is XXX parts in an operating state in a temperature environment of 90 degrees Celsius?

[0148] Answer sample: Yes.

[0149] (2) The second Q&A data pair corresponding to the knowledge point "The allowable operating temperature of XXX parts is 0 to 100 degrees Celsius":

[0150] Question sample: Is XXX part in the working state at a temperature of 200 degrees Celsius?

[0151] Answer sample: No.

[0152] (3) The third Q&A data pair corresponding to the knowledge point "The extreme operating conditions of XXX parts are 0 to 10 degrees Celsius":

[0153] Question sample: Is XXX part in the extreme operating conditions at a temperature of 5 degrees Celsius?

[0154] Answer sample: Yes.

[0155] (4) The fourth Q&A data pair corresponding to the knowledge point "The extreme operating conditions of XXX parts are 0 to 10 degrees Celsius":

[0156] Question sample: Is XXX part in the extreme operating conditions at a temperature of 20 degrees Celsius?

[0157] Answer sample: No.

[0158] (5) The fifth Q&A data pair corresponding to the knowledge point "The suitable operating conditions of XXX parts are 60 to 70 degrees Celsius":

[0159] Question sample: Is XXX part in the suitable operating conditions at a temperature of 65 degrees Celsius?

[0160] Answer sample: Yes.

[0161] (6) The sixth Q&A data pair corresponding to the knowledge point "The suitable operating conditions of XXX parts are 60 to 70 degrees Celsius":

[0162] Question sample: Is XXX part in the suitable operating conditions at a temperature of 80 degrees Celsius?

[0163] Answer sample: No.

[0164] The data generation method provided by the embodiments of the present disclosure has at least the following beneficial effects:

[0165] (1) Enhance cost-effectiveness: The data generation method provided by the embodiments of the present disclosure only relies on the invocation of large models and does not require manual participation. It can not only greatly reduce the consumption of human resources but also effectively reduce the subjective bias that may be brought by manual participation, thereby improving the quality of Q&A data pairs.

[0166] (2) Improve pertinence, flexibility, and controllability: The data generation method provided in the embodiments of the present disclosure divides the generation task of question-and-answer data pairs into two stages: mode determination and data generation. Compared with the traditional single-stage generation method, it will show stronger task pertinence when facing complex target texts or complex data generation requirements. Moreover, the staged data generation method endows the data generation system including the first large model and the second large model with higher flexibility and controllability, enabling the data generation system to more efficiently handle different types of data pair generation requirements.

[0167] (3) Improve accuracy: The data generation method provided in the embodiments of the present disclosure can provide at least one optional question-and-answer mode, and use each optional question-and-answer mode in the at least one optional question-and-answer mode as the target question-and-answer mode to generate question-and-answer data pairs that match the target question-and-answer mode based on the target text by using the second large model. Compared with the traditional single-stage generation method, this significantly reduces the complexity of the data generation task. This divide-and-conquer data generation method not only simplifies the data generation process but also significantly improves the quality (e.g., accuracy) of the question-and-answer data pairs.

[0168] Please refer to Figure 3 , which is a schematic diagram of an application scenario of a data generation method provided in the embodiments of the present disclosure.

[0169] The data generation method provided in the embodiments of the present disclosure is applied to an electronic device. Among them, the electronic device can be a server or a terminal device. Here, the terminal device can be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a tablet computer, etc.), a personal digital processor, or other similar computing devices.

[0170] Here, the electronic device is used to:

[0171] Obtain the target text;

[0172] Use the first large model to determine at least one optional question-and-answer mode related to the target text;

[0173] Use each optional question-and-answer mode in the at least one optional question-and-answer mode as the target question-and-answer mode to generate question-and-answer data pairs that match the target question-and-answer mode based on the target text by using the second large model.

[0174] It should be noted that in the embodiments of the present disclosure, Figure 4 The shown application scenario schematic diagram is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 4 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0175] To better implement the foregoing data generation method, an embodiment of the present disclosure further provides a data generation device 400, which can be integrated into an electronic device. Among them, the electronic device can be a server or a terminal device. Here, the terminal device can be a workbench, a mainframe computer, a conventional computer (for example, a desktop computer, a laptop computer, a tablet computer, etc.), a personal digital processing, or other similar computing devices 400. Hereinafter, in conjunction with Figure 4 the following schematic structural block diagram, a data generation device 400 provided by an embodiment of the disclosure will be described.

[0176] The data generation device 400 includes:

[0177] a text acquisition unit 401 for acquiring a target text;

[0178] a mode determination unit 402 for using a first large model to determine at least one optional Q&A mode related to the target text;

[0179] a data generation unit 403 for using the optional Q&A mode as a target Q&A mode, and using a second large model to generate a Q&A data pair matching the target Q&A mode based on the target text.

[0180] In some alternative embodiments, the mode determination unit 402 is configured to:

[0181] determine a plurality of candidate Q&A modes;

[0182] generate a first prompt message based on the target text and the plurality of candidate Q&A modes;

[0183] use the first large model to determine at least one optional Q&A mode related to the target text from the plurality of candidate Q&A modes based on the first prompt message.

[0184] In some alternative embodiments, the mode determination unit 402 is configured to:

[0185] determine a plurality of original Q&A modes;

[0186] in the case where no mode correction instruction is detected, use the original Q&A mode as a candidate Q&A mode;

[0187] or, in the case where a mode correction instruction is detected, perform a correction operation on the plurality of original Q&A modes based on the mode correction instruction to obtain a plurality of candidate Q&A modes; wherein, the correction operation includes at least one of an amplification operation, a replacement operation, and a deletion operation.

[0188] In some alternative embodiments, the plurality of original Q&A modes include at least one of an inference calculation mode and a discrimination mode.

[0189] In some alternative embodiments, the mode determination unit 402 is configured to:

[0190] Utilize a first large model, based on first prompt information, to determine at least one optional Q&A mode related to the target text from multiple candidate Q&A modes, and the number of Q&A pairs corresponding to the optional Q&A mode.

[0191] In some alternative embodiments, the data generation unit 403 is configured to:

[0192] Generate second prompt information based on the target text, the target Q&A mode, and the number of Q&A pairs corresponding to the target Q&A mode;

[0193] Utilize a second large model, based on the second prompt information, to generate a target number of Q&A data pairs that match the target Q&A mode; wherein the target number is the number of Q&A pairs corresponding to the target Q&A mode.

[0194] In some alternative embodiments, the data generation device 400 further includes a model determination unit, configured to:

[0195] Determine multiple candidate generation models;

[0196] Based on the target Q&A mode, determine the second large model from the multiple candidate generation models.

[0197] In the embodiments of the present disclosure, for the specific functions and examples of each unit in the data generation device 400, reference may be made to the relevant descriptions of the corresponding steps in the foregoing data generation method embodiments, which will not be elaborated herein.

[0198] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0199] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0200] Figure 5 FIG. shows a schematic structural block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure. The electronic device 500 is intended to represent various forms of digital computers, such as in-vehicle computing devices, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device 500 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0201] As shown Figure 5 in FIG. 500, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0202] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of renderers, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0203] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the data generation method. For example, in some embodiments, the data generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the data generation method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured as the data generation method by any other appropriate means (such as by means of firmware).

[0204] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0205] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0206] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0207] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a rendering device for rendering information to the user (e.g., a cathode ray tube (CRT) renderer or a liquid crystal display (LCD) renderer); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0208] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0209] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, or a server of a distributed system, or a server combined with a blockchain.

[0210] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute a data generation method.

[0211] Embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements a data generation method.

[0212] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and this is not limited herein. In addition, in the present disclosure, relational terms such as "first", "second", "third", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In addition, in the present disclosure, "a plurality of" can be understood as at least two.

[0213] The foregoing specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data generation method, comprising: Get the target text; Determine at least one optional question-answering mode related to the target text using the first large model; The optional question-answering model is used as the target question-answering model to generate question-answering data pairs matching the target question-answering model based on the target text using the second largest model.

2. The method according to claim 1, wherein: The step of using the first large model to determine at least one optional question-answering mode related to the target text includes: Determine multiple candidate question-answering models; Based on the target text and the multiple candidate question-answer patterns, generating first prompt information; Using the first large model and based on the first prompt information, at least one optional question-and-answer mode related to the target text is determined from the multiple candidate question-and-answer modes.

3. The method according to claim 2, wherein: The determining of multiple candidate question-answering modes includes: Identify multiple original question-answering patterns; In the case where no mode modification indication is detected, using the original question-answering mode as the candidate question-answering mode; Alternatively, when the mode modification indication is detected, based on the mode modification indication, a modification operation is performed on the multiple original question and answer patterns to obtain the multiple candidate question and answer patterns; wherein the modification operation includes at least one of an expansion operation, a replacement operation, and a deletion operation.

4. The method according to claim 3, wherein: The plurality of original question-answering patterns include at least one of an inference calculation pattern and a determination pattern.

5. The method according to claim 2, wherein: The using the first large model and determining at least one optional question-answering mode related to the target text from the plurality of candidate question-answering modes based on the first prompt information includes: Using the first large model and based on the first prompt information, at least one optional question and answer pattern related to the target text and the number of question and answer pairs corresponding to the optional question and answer pattern are determined from the multiple candidate question and answer patterns.

6. The method according to claim 5, wherein: The step of using the second largest model to generate a question-answer data pair matching the target question-answer mode based on the target text includes: Generate second prompt information based on the target text, the target question-answer mode, and the number of question-answer pairs corresponding to the target question-answer mode; Utilizing the second large model and based on the second prompt information, a target number of question-answer data pairs matching the target question-answer pattern are generated; wherein the target number is the number of question-answer pairs corresponding to the target question-answer pattern.

7. The method according to any one of claims 1 to 6, further comprising: determining a plurality of candidate generative models; Based on the target question-answering model, the second largest model is determined from the multiple candidate generation models.

8. A data generating device, comprising: A text acquisition unit, used for acquiring a target text; A mode determination unit, configured to determine at least one optional question-answer mode related to the target text by using the first large model; A data generating unit is used to use the optional question-answering mode as a target question-answering mode, so as to utilize a second large model and generate a question-answering data pair matching the target question-answering mode based on the target text.

9. An electronic device, comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 7.