Training sample data generation method and device, readable storage medium and program product

By building a sample data generation model and multi-agent collaboration, the problem of insufficient quality and diversity of training sample data in the existing technology is solved, efficient and automated training data generation is achieved, and the adaptability and generalization ability of the model is improved.

CN120277419AActive Publication Date: 2025-07-08LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202510765113.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-quality and diverse programming language processing task training sample data, especially when controlling the difficulty of the problem and verifying test cases, resulting in poor model training results.

Method used

By constructing a sample data generation model, using the language model to modify seed questions, generating a set of test cases covering the constraints of new questions, and performing multiple iterative evolutions to ensure the accuracy and diversity of questions and answers, and combining multi-agent collaboration for question review and answer generation.

Benefits of technology

Efficient and automated training sample data generation is achieved, manual intervention is reduced, the quality and diversity of training data is improved, and the adaptability and generalization ability of the model at different training stages is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277419A_ABST
    Figure CN120277419A_ABST
Patent Text Reader

Abstract

The invention discloses a training sample data generation method and device, a readable storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: inputting a seed topic and a cue word of a programming language processing task into a sample data generation model; and modifying the seed topic in the modification range boundary condition by using the cue word based on the same algorithm thinking, and generating a test case set covering the constraint condition of the new topic according to the type of the new topic and the case generation condition. Answering the new question passing the verification, and outputting an answer passing the verification of the test case set; taking the new question, the test case set and the answer as training sample data; and adjusting and modifying the range boundary condition and / or the cue word and / or the use case generation condition, and generating new training sample data by utilizing the new question. The problem that training data generated by related technologies cannot meet model training requirements can be solved, and high-quality training sample data can be efficiently generated for programming language processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to a method for generating training sample data, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] NLP (Natural Language Processing) can help a computer understand and process human language, enabling the computer to read and understand natural language. With the rapid development of artificial intelligence technology, the accuracy requirements for related language tasks performed using language models are also getting higher and higher, and the demand for high-quality model training sample data is also increasing. Related technologies automatically generate training sample data for programming language processing tasks such as code generation or understanding through multiple agents. The quality of the questions in the training sample data is poor, and the question difficulty is not properly controlled, making it difficult to meet the training requirements of the programming language processing task model. Summary of the Invention

[0003] The present invention provides a method for generating training sample data, an electronic device, a computer-readable storage medium, and a computer program product, which can efficiently generate high-quality training sample data for programming language processing tasks.

[0004] To solve the above technical problems, the present invention provides the following technical solutions: On the one hand, the present invention provides a method for generating training sample data, including: Inputting the seed question and prompt words of the programming language processing task into a sample data generation model constructed based on a language model; Using the prompt words through the sample data generation model, modifying the seed question within the boundary conditions of the modification range based on the same algorithmic thinking, generating a test case set covering the new question constraints according to the new question type and use case generation conditions; verifying the new question and the test case set, outputting the test case set that passes the verification, answering the new question that passes the verification, and outputting the answer verified by the test case set; using the new question, the test case set, and the answer as the first set of training sample data; Inputting the new question into the sample data generation model, adjusting the boundary conditions of the modification range and / or the prompt words and / or the use case generation conditions, and generating the second set of training sample data.

[0005] The present invention also provides an electronic device, including a memory and a processor, and the processor is used to implement the steps of any one of the above methods for generating training sample data when executing a computer program stored in the memory.

[0006] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above training sample data generation methods are implemented.

[0007] Finally, the present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of any of the above training sample data generation methods are implemented.

[0008] The advantages of the technical solution provided by the present invention are as follows: By using a language model to construct a sample data generation model capable of automatically generating training sample data required for programming processing tasks, the automation and intelligence of code-type training sample data generation are realized. There is no need for manual annotation, reducing manual intervention and labor costs, and effectively improving the efficiency of code-type training sample data generation. The newly generated questions are controlled within the same algorithmic thinking as the original questions and within the allowable modified range boundary conditions. This not only makes the difficulty of the newly generated questions controllable but also ensures the accuracy of the new questions, which is conducive to comprehensively covering various programming task scenarios and algorithm requirements, enriching the training data content of programming processing tasks, and is conducive to improving the model training performance of executing programming processing tasks. Further, by reviewing the new questions and test cases, the reliability and accuracy of the generated data are effectively guaranteed. After the new questions are reviewed and passed, answers are provided and the generated answers are verified to ensure the correctness and practicality of the answers and questions. Iterative evolution of the same question can generate training sample data with different levels of complexity, meeting the needs of the language model corresponding to programming language processing tasks for diverse data in different training stages and application scenarios. Through structured training data with multiple difficulty evolution levels including the question-test case-answer triple, it can support supervised learning in the instruction fine-tuning stage and also meet the reward calculation requirements in the reinforcement learning stage, enhancing the adaptability and generalization ability of the language model corresponding to programming language processing tasks. In addition, the present invention also provides corresponding electronic devices, computer-readable storage media, and computer program products for the training sample data generation method, further making the method more practical, and the electronic devices, computer-readable storage media, and computer program products have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1Schematic diagram of the hardware composition framework applicable to the training sample data generation method provided by the present invention; Figure 2 Schematic diagram of the process of a training sample data generation method provided by the present invention; Figure 3 Schematic diagram of a question evolution process provided by the present invention; Figure 4 Schematic diagram of a test case generation process provided by the present invention; Figure 5 Schematic diagram of a question review process provided by the present invention; Figure 6 Schematic diagram of an answer generation process provided by the present invention; Figure 7 Schematic diagram of the process of another training sample data generation method provided by the present invention; Figure 8 Structure framework diagram under an exemplary embodiment of the training sample data generation device provided by the present invention. Detailed implementation manners

[0011] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0012] With the rapid development of artificial intelligence technology, the accuracy requirements for related language tasks performed by language models are also getting higher and higher. For example, when using an LLM (Large Language Model) to perform programming language processing tasks, such as code generation tasks and code understanding tasks, the demand for high-quality model training sample data is also increasing.

[0013] Traditional training sample data is achieved through manual annotation or screening from open-source code libraries. The manual annotation method is not only costly but also difficult to ensure the diversity of sample data. The data obtained from open-source code libraries often has problems such as uneven quality, inconsistent code styles, and lack of systematic difficulty grading. Moreover, the proportion of "difficult problems" in open-source datasets is insufficient. Simple problems cause the model to quickly saturate during training, while there are not enough difficult problems, both of which cannot guarantee the training of a high-performance model. To address the shortcomings of traditional training sample data generation, the current approach is to use automatic generation. However, this method can only generate training sample data in simple scenarios. For programming problems that require in-depth reasoning and global understanding, such as programming competitions, the training sample data generated by this method often has quality problems, such as logical errors and inconsistencies. For example, related technologies generate training data through a code completion model. Due to the lack of strict verification of the generated code, a large number of logical errors or inefficient implementations exist in the data, limiting the training effect of code generation large language models. Magicoder (a large language model) in related technologies uses OSS-INSTRUCT (a data generation method that uses large language models to generate low-prejudice and high-quality code from open-source code snippets) to inspire LLMs to generate instruction data using open-source code snippets. However, it depends on the quality and diversity of open-source code and is difficult to generate high-quality optimized instructions. In addition, the model training sample data generation method usually uses a single model to generate all at once, lacking a systematic problem evolution mechanism and being difficult to cover the complete difficulty spectrum from basic to advanced, all of which cannot meet the requirements for high-quality training sample data.

[0014] In related technologies, when generating training sample data for programming language processing tasks such as automatically generating code or understanding code through multi-agent systems, for example, based on MetaGPT (an open-source multi-agent collaboration framework), the software development process is optimized through multi-agent collaboration. However, it focuses on task decomposition rather than the generation and verification of code training data, and has not been specifically designed for key aspects such as the difficulty evolution of code questions, test case coverage, and answer verification. Based on AGENTVERSE (a multi-agent framework), although multi-agents are used for collaborative processing, it targets general tasks rather than the structured generation of code data. Based on the CMAT (multi-agent collaboration tuning framework), although the collaborative ability of small language models is improved through supervised fine-tuning, it has not dynamically evolved code questions nor generated and verified test cases. In addition, multi-agent systems still have limitations in the field of programming language processing tasks. For example, although a related technology optimizes the communication efficiency between heterogeneous agents through an information navigation mechanism, it solves the problem of information asymmetry and cannot achieve iterative optimization of code data. During the test case generation process, often only the generation of the code itself is concerned, and test cases are generated by relying on simple input-output matching or random test data, lacking in-depth analysis of boundary conditions, abnormal inputs, and algorithm coverage. This neglect of the systematic construction of test cases will result in the generated test cases being unable to effectively verify the robustness of the code, affecting the training effect in the reinforcement learning stage. In addition, related technologies usually cannot simultaneously meet the different requirements of instruction fine-tuning and post-training of reinforcement learning. The generated training sample data has a single use and is difficult to adapt to the requirements of different training stages of large language models.

[0015] Therefore, it can be seen that related technologies are difficult to systematically generate programming questions with gradient difficulty, resulting in training data that is difficult to meet the needs of large language models for algorithmic thinking training. Due to the lack of an effective difficulty assessment mechanism, the complexity of the questions cannot be adjusted through iterative evolution, resulting in simple questions that cannot meet the needs of advanced algorithmic thinking, such as dynamic programming and graph theory optimization, and non-simple questions being obscure and lacking reasonable solution paths. Moreover, in the process of generating training sample data, related technologies ignore boundary conditions, extreme inputs, and algorithm coverage analysis when generating test cases, resulting in test cases being unable to effectively verify the robustness of the code and affecting the training effect in the reinforcement learning stage. In addition, related technologies are difficult to optimize the same question in multiple rounds and cannot generate multiple variants of the same problem, such as dynamic programming problems under different constraint conditions, and thus cannot reproduce the characteristics of tasks with a clear difficulty progression relationship such as programming competition questions. The training sample data is difficult to meet the model's needs for algorithmic thinking training, and thus cannot ensure the training of a high-performance model for processing such programming language processing tasks.

[0016] In view of this, in order to efficiently generate high-quality training sample data for programming language processing tasks and meet the requirements of training sample data in different training stages such as instruction fine-tuning and reinforcement learning, the present invention performs multiple rounds of optimization on the same topic. By generating multiple variants of the same problem, a complete difficulty spectrum is formed, which can effectively enrich the problem data of programming language processing tasks that require in-depth reasoning and global understanding, such as competition-level programming. By iteratively evolving to adjust the problem complexity, the problem quality and difficulty are effectively controlled. It can not only meet the needs of advanced algorithm thinking but also have a reasonable problem-solving path. It not only improves the quality and diversity of training sample data but also divides the data into "problem-answer" pairs (for fine-tuning) and "problem-test case" (for reinforcement learning) through structured storage to meet the requirements of different training stages and solve the problem of the single use of training sample data. Combined with the specific application environment architecture or specific hardware architecture on which the execution of the training sample data generation method depends, the specific application environment architecture or specific hardware architecture is described herein. The following combines Figure 1 Some possible application scenarios related to the technical solution of the present invention are introduced by way of example, which may include the following content: The hardware composition framework may include a first electronic device 11 and a second electronic device 12, which are connected by a network 13. The first electronic device 11 has an interface for invoking multiple language models, such as an API (Application Programming Interface) interface, and deploys a processor for executing the training sample data generation method recorded in any one of the above embodiments. The second electronic device 12 deploys a user terminal for providing a human-computer interaction interface. The user inputs a seed problem and a prompt word through the second electronic device 12. The first electronic device 11 obtains the seed problem and the prompt word through the network 13, completes all or part of the steps in the training sample data generation recorded in the above embodiments, obtains multiple groups of training sample data with a complete difficulty spectrum corresponding to the seed problem, and sends these multiple groups of training sample data to the second electronic device 12. The second electronic device 12 stores them according to the "problem-test case-answer" triple. After the generation of the triple training sample data of a seed problem is completed, the second electronic device 12 may send the next seed problem to the first electronic device 11. Of course, the second electronic device 12 may send the seed problem library to the first electronic device 11. To improve efficiency, save computing resources and storage resources, it may also send the storage address of the seed problem library to the first electronic device 11. After the first electronic device 11 obtains the seed problem library, to improve the efficiency of training sample data generation, and if resources permit, multiple threads may be constructed to simultaneously execute the corresponding training sample data generation tasks for multiple seed problems. Similarly, the storage address of the training sample data of each seed problem may also be sent to the second electronic device 12.

[0017] It should be noted that the above application scenarios are only shown for the convenience of understanding the ideas and principles of the present invention, and the embodiments of the present invention are not restricted in this regard. On the contrary, the embodiments of the present invention can be applied to any applicable scenario.

[0018] After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners. First, please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for generating training sample data provided in this embodiment. This embodiment may include the following content: S201: Input the seed questions and prompt words of the programming language processing task into the sample data generation model constructed based on the language model.

[0019] In this step, the programming language processing task is, for example, a code generation task, a code understanding task, a programming question generation - answering task, an artificial intelligence programming education task. Of course, the training sample data generation method provided by the present invention is applicable to the generation of code - type training sample data, and can also be extended to the generation of training data in other fields, such as question - answering data in natural language processing, annotation data in image recognition, etc. The present invention makes no limitation in this regard. The seed question is the original question of the programming language processing task. This question can be written manually or directly obtained from an existing database, which does not affect the implementation of the present invention. The seed question is the problem corresponding to the programming language processing task, which can be text data, image data, audio - video data. Correspondingly, the sample data generation model has an encoding - decoding processing mechanism in the corresponding format. The prompt word is an instruction for prompting the sample data generation model on how to process the seed question. Similarly, it can be text data, image data, audio - video data, and those skilled in the art can select any data format according to the actual scenario.

[0020] Among them, the sample data generation model can adopt a language model or a combination of multiple language models, which does not affect the implementation of the present invention. The language model can directly adopt any model in related technologies that can help a computer understand and process human language, enabling the computer to read and understand natural language and perform various language tasks. Those skilled in the art can directly adopt existing models that can conduct conversations and perform tasks with certain logical capabilities, such as intelligent chatbot models. Of course, a pre-trained language model can also be adopted and then fine-tuned to achieve the purpose of generating training sample data. A pre-trained language model refers to designing a language model training task based on a large-scale corpus (including, for example, language training materials such as sentences and paragraphs), training a large-scale neural network algorithm structure to learn and implement, and finally obtaining the large-scale neural network algorithm structure and parameters, which are the pre-trained language model. For subsequent other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. The neural network algorithm structure trained by the pre-trained language model can be a convolutional neural network model, a long short-term memory network model, etc., or a model constructed by an attention network, such as Transformer (transformer network model), bert (Bidirectional Encoder Representations from Transformers, bidirectional encoder representations based on Transformer), GPT (Generative Pre-trained Transformer, generative pre-trained transformer), etc.

[0021] S202: Use the prompt words to modify the seed question within the boundary conditions of the modification range through the sample data generation model based on the same algorithmic thinking, and generate a test case set covering the new question constraints according to the new question type and use case generation conditions; verify the new question and the test case set, output the test case set that passes the verification, answer the new question that passes the verification, and output the answer verified by the test case set.

[0022] Among them, the same algorithmic thinking means that the newly generated questions, that is, the new questions in this embodiment, and the algorithms of the seed questions belong to the same type of algorithm, or are algorithms for solving the same problem. For example, if the seed question is to use the linear search method to search for a specific element and its test point is the search algorithm, then the algorithm of the new question can be any search algorithm that can implement the search for a specific element, such as the linear search method or the binary search. Modifying the range boundary condition is used to control the difficulty range of the newly generated questions. It can control the complexity of the new questions from the complexity of the algorithm, the scale of the input data, and the type of data structure. That is, it can limit the complexity of the algorithm and / or the scale of the input data and / or the type of data structure to control the modification range of the new questions to ensure that the generated new questions are both challenging enough and within a reasonable difficulty range. For example, to control the difficulty of the new algorithm through the algorithm complexity, the modification range boundary condition can be set so that the algorithm difficulty level does not exceed 2 levels. If the seed question uses a basic algorithm such as simple search, then the new algorithm can use an intermediate algorithm such as binary search, or a advanced algorithm such as a mathematical optimization method. However, for advanced algorithms and combined algorithms of advanced algorithms that exceed this boundary, they are not allowed. During the process of modifying the seed questions within the modification range boundary condition based on the same algorithmic thinking, it is necessary to always maintain the essential nature of examining the core algorithmic thinking, and at the same time enhance the challenge of the questions through innovative question types. For example, for a traditional sorting algorithm question, its scenario can be extended from simple number sorting to sorting objects with complex attributes, thereby increasing the difficulty and novelty of the question.

[0023] Among them, when the seed problem is modified within the modified range boundary conditions based on the same algorithmic thinking, a new problem is generated, and the type of the new problem is determined according to the test points of the new problem and all the constraints. For example, the new problem is to find the target value in a rotated sorted array, the test point is identified as the variant application of the binary search algorithm, and the constraints include the uniqueness of array elements, the uniqueness of the rotation point, the time complexity requirement, etc. The use case generation conditions at least include the type of use case composition, and the type of use case composition is the type of test case, such as basic use case, stress use case, boundary use case. The test case design scheme is determined according to the test points of the new problem and all the constraints. The test case design scheme may include the test dimensions planned to be covered, the specific use case design ideas for each dimension, and the common error types expected to be found, so as to achieve the purpose of covering all the constraints of the new problem. Taking the new problem as a dynamic programming problem as an example, the test case design scheme may include: dividing the state transition correctness, boundary condition handling, and time efficiency verification into three dimension directions in terms of test dimensions; each dimension at least includes constructing special state transition paths, minimum / maximum scale inputs, illegal input detection, etc.; the expected error types to be found include common problems such as state initialization errors, transition equation errors, and array out-of-bounds. The test case set is a set of multiple test cases generated according to the scheme of this step.

[0024] In this step, after generating new questions and test case sets, to ensure the model training performance and meet the training requirements for programming language processing tasks, it is also necessary to comprehensively evaluate the new questions and test case sets. For example, whether the questions are solvable, whether the difficulty has increased, whether error information or inconsistent information has been introduced, and whether the test cases are appropriate, and specify the reasons in detail. By reviewing and validating the new questions and test cases, the quality of the generated new questions and test cases is ensured. For example, when a certain test case in the test case set cannot cover all possible situations in the question, it will be generated that the test case is inappropriate, and a description of the missing situation will be given. For the new questions or test case sets that do not pass the verification, the prompt words will be adjusted, or the range boundary conditions or the use case generation conditions will be modified to regenerate the new questions and new test case sets until the verification passes. Of course, to avoid entering an infinite loop and ensure efficiency, the maximum number of regenerations can be set. When the limit is exceeded, corresponding error prompt information can be generated. For the new questions that pass the verification, answers are provided for the new questions. The generated data for answering the new questions is the answer. Taking the code generation task as an example, the answer is a piece of generated code data. Taking the code understanding task as an example, the answer is the understanding content data that explains the given code. Taking the competition question setting task as an example, the answer is the competition question and the corresponding answers to the questions. The answer can be text data, image data, audio-visual data, etc., which does not affect the implementation of the present invention. To ensure the accuracy of the final answer, after generating the answer, the generated answer will also be verified, such as whether the answer is accurate, whether it exceeds the constraints of the new question, and the answer needs to pass all or most of the test cases, such as 90%. Similarly, for the questions that do not pass the verification, the prompt words will be adjusted, or the range boundary conditions or the use case generation conditions will be modified to regenerate the new questions until the verification passes. Of course, to avoid entering an infinite loop and ensure efficiency, the maximum number of regenerations can be set. When the limit is exceeded, corresponding error prompt information can be generated. Finally, a corresponding relationship is established among the generated new questions, test case sets, and answers as a set of training sample data, and then it is stored in a structured manner, that is, it can be stored according to the "question - test case - answer" triple, which can not only support the supervised learning in the instruction fine-tuning stage but also meet the reward calculation requirements in the reinforcement learning stage. For example, through "question - answer", it can be applied to the fine-tuning training requirements, and through "question - test case", it can be used for reinforcement learning.

[0025] S203: Input the new questions into the sample data generation model, adjust and modify the range boundary conditions and / or prompt words and / or use case generation conditions to generate the second set of training sample data.

[0026] In this step, the modification range boundary conditions and / or prompt words and / or test case generation conditions can be adjusted according to actual needs. The new problem is used as the seed problem for S201. If the prompt words are adjusted, the new prompt words and the new problem of the first set of training data are used as the seed problem and input into the sample data generation model. When executing S202, if the modification range boundary conditions or test case generation conditions are adjusted, the new modification range boundary conditions or new test case generation conditions are used to replace the old modification range boundary conditions or old test case generation conditions in S202, and the corresponding training sample data is generated according to the method of S202. For the sake of distinction, the new problems, test case sets, and answers generated in the above steps are used as the first set of training sample data, and the training sample data generated in this step is defined as the second set of training sample data. In order to improve the quality and diversity of the generated training sample data, this step will perform multiple rounds of iterative evolution on the same problem. Those skilled in the art can determine the number of iterative evolution rounds according to actual needs. In this embodiment, taking one iteration as an example, if it is iterated twice, the new problem of the second set of training sample data generated by S203 is used as the seed problem for S201 to execute S201-S202, and the cycle continues until the number of iterative evolution rounds is reached.

[0027] In the technical solution provided in this embodiment, a sample data generation model that can automatically generate training sample data required for programming processing tasks is constructed using a language model, realizing the automation and intelligence of code-class training sample data generation, eliminating the need for manual annotation, reducing manual intervention, lowering labor costs, and effectively improving the efficiency of code-class training sample data generation. Controlling the newly generated problems within the same algorithmic thinking as the original problems and within the allowable modification range boundary conditions not only makes the difficulty of the newly generated problems controllable but also ensures the accuracy of the new problems, which is conducive to comprehensively covering various programming task scenarios and algorithm requirements, enriching the training data content of programming processing tasks, and is conducive to improving the training performance of the model for executing programming processing tasks. Further, by reviewing the new problems and test cases, the reliability and accuracy of the generated data are effectively guaranteed. After the new problems pass the review, they are answered, and the generated answers are verified to ensure the correctness and practicality of the answers and problems. Performing iterative evolution on the same problem can generate training sample data of different complexity levels, meeting the needs of the language model corresponding to the programming language processing task for diverse data in different training stages and application scenarios. Through structured training data with multiple difficulty evolution levels containing the "problem - test case - answer" triple, it can not only support supervised learning in the instruction fine-tuning stage but also meet the reward calculation requirements in the reinforcement learning stage, enhancing the adaptability and generalization ability of the language model corresponding to the programming language processing task.

[0028] In the above embodiments, there is no limitation on how to modify the seed questions within the boundary conditions of the modification range based on the same algorithmic thinking. This embodiment also provides various exemplary modification methods for the seed questions, which may include the following: Exemplarily, the examination points of the seed questions can be analyzed to determine the algorithm type and / or data structure type of the seed questions; according to the algorithm type and / or data structure type, various modification methods for the seed questions can be determined; from the random combinations of the seed questions and each modification method, select the modification method that meets the conditions of having the same examination point as the seed questions and at least adding one new technical point as the optimal combined modification method; based on the optimal combined modification method, make corresponding modifications to the seed questions to obtain new questions.

[0029] In this embodiment, first, the algorithm of the seed questions is deconstructed, and the core examination points of the seed questions are deeply analyzed through prompt words, and the algorithm type and data structure type involved are accurately extracted. For example, if the seed question is to search for a specific element, it is clear that the search algorithm it examines may be linear search, binary search, etc., and the data structures involved, such as arrays, linked lists, etc. According to the question type, list the possible directions of question evolution in detail: for algorithmic questions, they can be modified from aspects such as the complexity of the algorithm, the scale of the input data, and the type of data; for data structure questions, they can be modified from aspects such as the organization form of the data structure and the storage method of elements. After determining various modification methods, random combinations can be made, and the optimal modification combination can be selected from numerous modification directions according to the boundary conditions of the modification range. The boundary conditions of the modification range in this embodiment are modification methods that meet the conditions of having the same examination point as the seed questions and at least adding one new technical point. Finally, the final brand-new questions are generated according to the determined optimal combined modification method. For example, combining the previous analysis, generate a question that requires using a specific algorithm to find nodes that meet complex conditions in a large-scale graph data structure. There are multiple optimal combined modification methods, and corresponding modifications can be made to the seed questions respectively based on each optimal combined modification method to obtain multiple initial new questions; screen the candidate new questions that are error-free and non-missing from the initial new questions, compare the algorithm type and / or data structure complexity of the candidate new questions and the seed questions, and use the candidate new questions with a higher complexity than the seed questions as the new questions.

[0030] Exemplarily, the present invention can also simulate the human question-making process through prompt engineering, combine language models to design reasonable prompt words, and achieve the evolution of the original seed questions, such as Figure 3As shown, a large language model with excellent performance (such as DeepSeek V3) can be selected as the topic use case generation sub-model, that is, the sample data generation model at least includes the topic use case generation sub-model, which is also the role assignment. The topic use case generation sub-model completes the modification of the seed topic through four steps: algorithm deconstruction, transformation strategy, strategy combination, and topic generation. During the modification process, it is necessary to always maintain the essence of examining the core algorithm thinking unchanged, and at the same time enhance the challenge of the topic by innovating the question types. For example, if the seed topic is a traditional sorting algorithm topic, its scenario can be extended from simple number sorting to sorting objects with complex attributes, thereby increasing the difficulty and novelty of the topic. Exemplarily, during the algorithm deconstruction process, the topic use case generation sub-model is prompted to deeply analyze the core examination points of the original question and accurately extract the types of algorithms / data structures involved. In the transformation strategy step, the topic use case generation sub-model is prompted to list possible question modification directions in detail according to the question type. In the strategy combination step, the topic use case generation sub-model is prompted to select the optimal modification combination from numerous evolution directions and strictly limit the boundary conditions of the modification range to ensure that the generated new question is both challenging enough and within a reasonable difficulty range. For example, when increasing the algorithm complexity, the realizability and rationality of this complexity in the actual programming scenario should be considered. In the topic generation stage, the topic use case generation sub-model is prompted to generate the final brand-new question. For example, combining the previous analysis, generate a question that requires using a specific algorithm to find nodes that meet complex conditions in a large-scale graph data structure.

[0031] In this embodiment, a question generation prompt template can be pre-constructed. The question generation prompt template at least includes an examination point analysis prompt word, an evolution prompt word, an optimal combination modification method prompt word, and a rewriting prompt word. Input the seed question into the question generation prompt template to obtain question generation prompt words. Input the seed question and the question generation prompt words into the question case generation sub-model. The question case generation sub-model analyzes the examination points of the seed question under the prompt of the examination point analysis prompt word to determine the algorithm type and / or data structure type of the seed question. Under the prompt of the evolution prompt word, according to the algorithm type and / or data structure type, determine various modification methods of the seed question. Under the prompt of the optimal combination modification method prompt word, select a modification method that meets the condition of having the same examination point as the seed question and at least adding one new technical point from the random combination of the seed question and each modification method as the optimal combination modification method. Under the prompt of the rewriting prompt word, modify the seed question accordingly based on the optimal combination modification method to obtain a new question. In the present invention, each prompt word in the prompt word template is data that does not include content in an exemplary scenario. When the exemplary scenario content is input into the prompt word template, a prompt word template dedicated to the exemplary scenario is generated. For the sake of distinction, the template after inputting the exemplary scenario data is defined accordingly. In this embodiment, the exemplary scenario data is the seed question of S201. Input it into the question generation prompt template, and define the question generation prompt template containing the seed question as the question generation prompt word. At this time, the question generation prompt word also includes the examination point analysis prompt word, the evolution prompt word, the optimal combination modification method prompt word, and the rewriting prompt word.

[0032] Among them, in the process of using the question case generation sub-model to modify the seed question within the boundary conditions of the modification range based on the same algorithmic thinking, the question case generation sub-model can be controlled to output sequentially through multiple steps, or the thinking chain process of all steps can be generated at one time. Exemplarily, the question generation prompt template can be: As an expert in processing programming language tasks, rewrite / evolve the given #instruction# into a more complex version. Among them, the given instruction is the seed question.

[0033] Rewrite the given "#instruction#" into a more complex version according to the following steps, which need to meet the following conditions: The new instruction has the same algorithmic core as the original instruction. The solution requires creative thinking rather than direct application and includes at least one important optimization breakthrough.

[0034] Step 1: Please carefully read the "#instruction#", analyze the core examination points of the original #instruction#, and extract the algorithm / data structure type.

[0035] Step 2: Please list 3 possible modification strategies for the #instruction# according to the core examination points of the #instruction#.

[0036] Step 3: According to the listed modification strategies, select the optimal evolutionary combination that meets the requirements of retaining the core examination points of the seed questions and adding at least one new technical point.

[0037] Step 4: Please rewrite the original #Instruction# according to the optimal combination of the evolutionary strategies.

[0038] Step 5: Please carefully read the #Rewritten Instruction#, and find out the unreasonable or missing parts. Ensure that the #Rewritten Instruction# is just a more complex version of the #Instruction#. Just provide the #Final Rewritten Instruction# without any explanation.

[0039] Please reply strictly according to the following format: Step 1 #Core Examination Points#: Step 2 #Evolutionary Strategies#: Step 3 #Optimal Evolutionary Combination#: Step 4 #Rewritten Instruction#: Step 5 #Final Rewritten Instruction#: #Instruction#: {Seed Instruction}.

[0040] Replace the seed questions into the above question generation prompt template, and the question generation prompt words can be obtained. Input the question generation prompt words into the question case generation sub-model, and the question case generation sub-model will generate corresponding data for output according to the format of the above template.

[0041] As can be seen from the above, in this embodiment, through algorithm deconstruction, transformation strategies, strategy combination selection and mutation boundary limitation, the generated new questions have diversity and challenges, can comprehensively cover various programming scenarios and algorithm requirements, and effectively enrich the content of the training sample data for programming language processing tasks. In addition, through the algorithm deconstruction and transformation of the seed questions by the question case generation sub-model, it is ensured that the new questions enhance the challenges while maintaining the examination of the core algorithm thinking, reduce manual intervention, lower the labor cost, and effectively improve the efficiency of generating code class training sample data.

[0042] In the above embodiment, there is no limitation on how to generate a test case set that covers the constraint conditions of the new questions according to the new question type and the test case composition type. This embodiment also provides various exemplary generation methods for the test case set, which may include the following content: Exemplarily, obtain the test case composition type and the use case generation specification, and determine the test case generation scheme according to the algorithm examination points, explicit constraint conditions and implicit constraint conditions of the new questions; the test case generation scheme at least includes the test coverage dimension, the test ideas of each dimension and the error types expected to be found; according to the test case generation scheme, generate multiple test cases that meet the test case composition type and the use case generation specification.

[0043] Among them, the use case generation specification is that each test case has a corresponding relationship between input and output and annotation information of the design intention, ensuring that all problem constraints are covered. This structured prompt can significantly improve the coverage rate and effectiveness of test cases. After determining the algorithm test points and all constraints, a three-dimensional test plan can be dynamically generated according to the characteristics of the problem. According to the test case design scheme, test cases that meet the requirements of the use case composition type and the use case generation specification are generated. For example, the prompt words include that the test cases generated for the problem should at least include: 3 basic use cases, 4 boundary use cases, and 3 stress test cases. The basic use cases are used to verify the correctness for small-scale inputs, the boundary use cases include extreme values, empty inputs, and special conditions, and the stress test cases are used to verify the efficiency for the largest-scale inputs.

[0044] To further improve the practicality and accuracy of test cases, based on the above embodiments, the test cases can also be standardized, and the corresponding data can be formatted and output, such as Figure 4 shown. Correspondingly, according to the test case generation scheme, the implementation process of generating multiple test cases that meet the test case composition type and the use case generation specification can be as follows: The prompt words include automatically supplementing metadata for the test cases and outputting them according to the output format. According to the test case generation scheme, multiple test cases that meet the test case composition type and the use case generation specification are generated, the test case type is added to each test case, and the key inspection points are marked, and the output is carried out according to the input parameter list, expected output, function name, basic / boundary / stress, and test objective description. Taking the binary search problem as an example, the output of the generated test case can be: {"inputs": [[4, 5, 6, 7, 0, 1, 2], 5], "outputs": [1], "fn_name": "search"}. Exemplarily, the output format can be expressed as: {"inputs": [input parameter list], "outputs": [expected output], "fn_name": "function name", "test_type": "basic / boundary / stress", "description": "test objective description"}.

[0045] As can be seen from the above, in this embodiment, by standardizing the test cases, it is not only convenient for automated testing, but also ensures that the test data generated for different problems has a consistent interface. By supplementing metadata, more refined reward signals can be provided in the reinforcement learning stage, improving the quality of training sample data.

[0046] A parallel implementation of the above embodiments. In the test case generation stage, this embodiment can adopt structured prompt engineering to guide the question case generation sub-model to generate a comprehensive coverage test case set, which can include the following content: Pre-construct a case generation prompt template, input the new question into the case generation prompt template to obtain case generation prompt words; input the new question and the case generation prompt words into the question case generation sub-model. The question case generation sub-model obtains the test case composition type and case generation specifications under the prompt of analyzing the test points and extracting the constraint condition prompt words. The case generation specification is that each test case has an input and output correspondence relationship and design intention annotation information; determine the test case generation plan according to the algorithm test points, explicit constraint conditions and implicit constraint conditions of the new question under the prompt of the test case generation plan prompt words; the test case generation plan at least includes the test coverage dimension, the test ideas of each dimension and the error types expected to be found; generate multiple test cases that meet the test case composition type and case generation specifications according to the test case generation plan under the prompt of the case generation prompt words.

[0047] In this embodiment, during the process of using the question case generation sub-model to generate a test case set, the question case generation sub-model can be controlled to output sequentially through multiple steps, or the thinking chain process of all steps can be generated at one time. The case generation prompt template at least includes the test point analysis and constraint condition extraction prompt words, the test case generation plan prompt words and the case generation prompt words. In this embodiment, the exemplary scenario data is the new question, which is input into the case generation prompt template, and the case generation prompt template containing the new question is defined as the case generation prompt words. At this time, the case generation prompt words also include the test point analysis and constraint condition extraction prompt words, the test case generation plan prompt words and the case generation prompt words. Exemplarily, the case generation prompt template can be expressed as: As an expert in processing programming language tasks, please generate a comprehensive coverage test case set for the provided question, with the following requirements: 1. Requirements for case composition: Basic cases (3-5): Small-scale typical inputs to verify the basic correctness of the algorithm; Boundary cases (4-6): Including but not limited to the following types: Input scale boundary (empty input, minimum / maximum scale), Numerical boundary (extreme values, special values), Special condition boundary (special constraints specified in the question); Stress cases (2-3): Maximum allowable scale inputs to verify time efficiency.

[0048] 2. Generation specifications: Each case must include a clear input-output correspondence; Annotate the design intention of each case; Ensure 100% coverage of all constraint conditions described in the question.

[0049] 3. Output format: {"inputs": [input parameter list], "outputs": ["Expected Output"], "fn_name": "Function Name", "test_type": "Basic / Boundary / Stress", "description": "Description of Test Purpose"}

[0050] Step 1: Please carefully read the #question#, analyze the core examination points of the #question#, and extract all the constraints in the question description.

[0051] Step 2: Please generate a test case design plan according to the #core examination points and constraints# of the #question#, and the plan should include: 1) The test dimensions planned to be covered; 2) The specific design ideas for test cases in each dimension; 3) The common error types expected to be discovered.

[0052] Step 3: Generate test cases that meet the requirements of test case composition, generation specifications, and output formats according to the listed #test case design plan#.

[0053] Step 4: Provide the #final test case list# according to the #test cases#, without any explanations.

[0054] Please reply strictly according to the following format: Step 1#Core Examination Points and Constraints#: Step 2#Test Case Design Plan#: Step 3#Test Cases#: Step 4#Final Test Case List#: #Question#: {New Question}.

[0055] Replace the new question into the above test case generation prompt template, and you will get the test case generation prompt words. Input the test case generation prompt words into the question test case generation sub-model, and the question test case generation sub-model will generate corresponding data for output according to the format of the above template.

[0056] As can be seen from the above, in this embodiment, by analyzing the core examination points and constraints to generate a test case design plan, it can effectively expose problems in aspects such as boundary condition handling, algorithm correctness, and execution efficiency of the code, and effectively improve the quality of training sample data. In addition, by generating a test case set through the question test case generation sub-model, it reduces manual intervention, lowers labor costs, and effectively improves the efficiency of generating code-like training sample data.

[0057] In the above embodiment, there is no limitation on how to verify the new question and the test case set. This embodiment also provides various exemplary verification methods for the new question and the test case set, which may include the following content: Generate multiple problem-solving ideas for the new problem, and select at least two target problem-solving ideas for correctly solving the new problem; when the algorithmic difficulty of the new problem is greater than that of the seed problem, the new problem meets the preset semantic conditions without errors and ambiguities, and the new problem matches the test case set, then the new problem passes the verification; if the test case set matches the new problem and the test case set contains at least basic test cases, boundary test cases, and stress test cases, then the test case set passes the verification; when there are not at least two target problem-solving ideas, or the algorithmic difficulty of the new problem is less than or equal to that of the seed problem, or the semantic conditions are not met, then generate problem modification opinions; when the new problem does not match the test case set, then generate problem and test case modification opinions; when the test case set does not contain at least one of the basic test cases, boundary test cases, and stress test cases, then generate test case modification opinions.

[0058] In this embodiment, the problem can be verified by simulating different problem-solving ideas and scenarios to ensure that the generated problem is solvable, the difficulty is increased compared to the seed problem, no error information or inconsistent information is introduced, and the problem matches the test case set. If it is found in the verification that the new problem does not meet the above conditions, it is possible to generate that the new problem is inappropriate and the reasons for the inappropriateness. Among them, the algorithmic difficulty can be divided into basic algorithms, intermediate algorithms, advanced algorithms, and combined algorithms in sequence. Basic algorithms such as simulation, greedy, and simple search are suitable for introductory problems. Intermediate algorithms such as dynamic programming, classic graph theory algorithms, and binary search require a certain amount of algorithm accumulation. Advanced algorithms such as network flow, computational geometry, advanced string processing, and mathematical optimization are usually more difficult. Combined algorithms require combining multiple algorithms or data structures, such as dynamic programming + segment tree. The test case set should match the new problem and be able to cover all possible situations of the new problem. When it is found that a certain test case cannot cover all possible situations in the problem, it is possible to generate that the test case is inappropriate and the feedback information on the missing situation.

[0059] A parallel implementation manner of the above embodiment may use a large language model with excellent performance (such as GPT-4) as the data verification sub-model, that is, the sample data generation model at least includes a generated data verification sub-model, and the data verification sub-model executes the verification process of the questions and test cases. The optimization agent is guided by structured prompt words to comprehensively evaluate the questions. The prompt words at least include verification analysis prompt words and modification opinion prompt words. The data verification sub-model generates multiple problem-solving ideas for the new question under the prompt of the verification analysis prompt words, and selects at least two target problem-solving ideas for correctly solving the new question; when the algorithm difficulty of the new question is greater than the algorithm difficulty of the seed question, the new question meets the preset semantic conditions of no errors and no ambiguity, and the new question matches the test case set, then the new question passes the verification; if the test case set matches the new question, and the test case set at least includes basic test cases, boundary test cases and stress test cases, then the test case set passes the verification and outputs no modification opinions; the data verification sub-model is prompted by the modification opinion prompt words, when there are not at least two target problem-solving ideas, or the algorithm difficulty of the new question is less than or equal to the algorithm difficulty of the seed question, or the semantic conditions are not met, then a question modification opinion is generated and output; when the new question does not match the test case set, then a question and use case modification opinion is generated and output; when the test case set does not include at least one of the basic test cases, boundary test cases and stress test cases, then a use case modification opinion is generated and output.

[0060] Further, in order to improve the verification accuracy, this embodiment adopts a multi-round verification and dynamic feedback mechanism to ensure the quality of the generated questions and test cases, which may include the following content: Pre-construct an audit prompt template, input the seed question, new question and test case set into the audit prompt template to obtain single-round audit prompt words; input the single-round audit prompt words, seed question, new question and test case set into the data verification sub-model multiple times to obtain multiple independent single-round audit results; vote on each single-round audit result. When the voting result does not meet the preset voting passing condition, input each single-round audit result into the audit prompt template to obtain audit result prompt words, and input the audit result prompt words and each single-round audit result into the data verification sub-model. The data verification sub-model modifies the new question or test case according to the modification opinions of each single-round audit result under the prompt of the modification prompt words, and outputs the modified new question or modified test case under the prompt of the audit result output prompt words.

[0061] In this embodiment, as Figure 5As shown, the verification process may include three steps: single-round question verification, multi-round voting, and question modification. Single-round question verification can be achieved through the above embodiments, that is, the data verification sub-model generates multiple problem-solving ideas for the new question under the prompt of the verification analysis prompt word, and selects at least two target problem-solving ideas for correctly solving the new question; when the algorithm difficulty of the new question is greater than the algorithm difficulty of the seed question, the new question meets the preset semantic conditions of no errors and no ambiguity, and the new question matches the test case set, then the new question passes the verification; if the test case set matches the new question, and the test case set contains at least basic test cases, boundary test cases, and stress test cases, then the test case set passes the verification and outputs no modification opinions; under the prompt of the modification opinion prompt word, when there are not at least two target problem-solving ideas, or the algorithm difficulty of the new question is less than or equal to the algorithm difficulty of the seed question, or the semantic conditions are not met, the data verification sub-model generates question modification opinions for output; when the new question does not match the test case set, it generates question and use case modification opinions for output; when the test case set does not contain at least one of basic test cases, boundary test cases, and stress test cases, it generates use case modification opinions for output. The single-round review result will output no modification opinions, or question modification opinions, question and use case modification opinions, and use case modification opinions. After generating the single-round review result, this embodiment adopts a multi-round sampling voting mechanism to summarize the independent single-round review results generated in each round through a voting mechanism. For example, 5 single-round review results are generated for the same question and test case. Only when more than 80% (4 times) of the verifications pass, the new question and the test case set are recognized. Otherwise, the modification opinions are summarized and sent to the question use case generation sub-model as a question modification prompt to modify the new question or the test case set. For example, if the data verification sub-model points out that there is a problem with an unclear condition in the question, the question use case generation sub-model will rephrase the question according to the prompt to make it clearer and more accurate.

[0062] In this embodiment, during the process of validating the generated data using the data validation sub-model, it is possible to control the data validation sub-model to output sequentially through multiple steps, or to generate the thought chain process of all steps at once. The review prompt template includes at least single-round validation prompt words and result output prompt words. The single-round validation prompt words include at least validation analysis prompt words and modification suggestion prompt words; the result output prompt words include at least modification prompt words and review result output prompt words. The review prompt template of this embodiment includes two parts, namely single-round validation prompt words and result output prompt words. For the template corresponding to the single-round validation prompt words in the first part, the exemplary scenario data at this time are seed questions, new questions, and test case sets. Inputting them into the review prompt template, the review prompt template containing these input data is defined as single-round review prompt words. At this time, the single-round review prompt words also include validation analysis prompt words and modification suggestion prompt words. For the template corresponding to the result output prompt words in the second part, the exemplary scenario data at this time are the results of each single-round review. Inputting them into the review prompt template, the review prompt template containing these input data is defined as review result prompt words. At this time, the single-round review prompt words also include modification prompt words and review result output prompt words. Exemplarily, the review prompt template can be expressed as: As an expert in programming language processing tasks, please conduct multi-dimensional validation on the #new question# and provide modification suggestions.

[0063] 1. Note that the #new question# is a more complex version upgraded from the #seed question# through the instruction evolution technology.

[0064] 2. The #test case# is the problem-solving test case for the #new question#.

[0065] 3. The validation dimensions include: 1) Solvability analysis, generating 3 different problem-solving ideas and evaluating whether at least 2 of them can correctly solve the problem; 2) Difficulty assessment: Comparing with the original question, judging whether the algorithm complexity is reasonably improved; 3) Consistency check: Ensuring that the question description is unambiguous and matches the test case; 4) Test case coverage: Verifying whether it includes basic, boundary, and stress tests.

[0066] 4. Generate #modification suggestions# based on the content of #validation analysis#. If all validation dimensions meet the requirements, output "No modification suggestions".

[0067] Please reply strictly in the following format: Step 1 #validation analysis#: Step 2 #modification suggestions#: Among them, #seed question#: {seed question}; #new question#: {new question}; #test case#: {test case}.

[0068] As an expert in processing tasks for programming languages, please strictly modify the #new questions# and #test cases# according to the #modification content# list. Output the modified questions and test cases separately. Please reply strictly in the following format: Step 1 #Modified question#: Step 2 #Modified test case#: Among them, #modification content#: {modification content}, #new question#: {new question}, #test case#: {test case}.

[0069] As can be seen from the above, through multiple verification and voting judgment mechanisms in this embodiment, rigorous question and test case review can be completed, further ensuring the effectiveness of the generation of new questions and test cases, and effectively guaranteeing the reliability and accuracy of the generated data. In addition, by using the data verification sub-model to review and verify the test cases and new questions, manual intervention is reduced, labor costs are lowered, and the efficiency of generating training sample data for code classes is effectively improved.

[0070] In the above embodiment, there is no limitation on how to answer the new questions that have passed the verification. This embodiment also provides various implementation methods for verifying the answers to the new questions, such as Figure 6 As shown, it may include the following contents: Determine the time complexity and space complexity according to the algorithm examination points of the new questions that have passed the verification; determine the corresponding problem-solving idea process and algorithm step annotation information according to the time complexity and space complexity; answer the new questions that have passed the verification according to the problem-solving idea process, annotate the corresponding steps of the answer according to the algorithm step annotation information, and identify whether the answer covers all the constraint conditions; input the test case set and the answer that covers all the constraint conditions into the sandbox verification environment for answer verification, and when the answer verification fails, adjust the answer according to the error information in the answer verification process.

[0071] Among them, the space-time complexity can be calculated according to any relevant technology. The algorithm step annotation information can annotate the steps specified by the user or the steps that play a key role in the answer accuracy, which does not affect the implementation of the present invention. Answer verification can be performed by executing tests in a containerized sandbox environment. The sandbox environment can provide a safe and isolated running space to avoid the impact on external systems when the answer runs, and at the same time ensure the accuracy and reliability of the test. When testing the answer using the sandbox environment, the data segment of the answer can be extracted by string splitting, such as the code segment "```python[code segment]```", and then the extracted data segment and the test case of this question are input into the sandbox environment for execution. Metrics such as memory usage and execution time are monitored in real time, and solutions that exceed the question constraints are directly judged as failed. Whether the verification passes can be measured by the passing rate of the test cases in the sandbox environment. For example, if the passing rate of the answer through the test case set is greater than or equal to 90%, the answer verification passes. When the answer verification fails, that is, if the passing rate of the answer through the test case set is less than 90%, the "answer + error information" can be added to the prompt words, a new answer is regenerated, and the verification is performed again according to the above method. To avoid infinite loops, the maximum number of repetitions can be set for this process, such as limiting the repetition to 3 times. Taking the code generation task as an example, the answer is the generated code data. If an out-of-bounds error occurs during the runtime of the code data generated for the first time, the error information and the original code data are added to the prompt words, and the code generation logic is adjusted based on this information to regenerate new code for verification.

[0072] In a parallel implementation manner with the above embodiments, a large language model (such as Kimi) can be selected as the solution sub-model, that is, the sample data generation model at least includes the solution sub-model. The solution sub-model generates the solution idea and the answer corresponding to the new question according to the question. The solution sub-model will first analyze the question requirements, determine the appropriate algorithm and data structure, and then gradually construct the code logic to finally generate a complete code implementation. For example, for a new question of implementing a specific encryption algorithm, the solution sub-model will elaborate on the principle and steps of the encryption algorithm in detail, and then write the corresponding code according to the algorithm. After the solution sub-model generates the code, the test cases are executed in the sandbox environment to verify the correctness of the code, and the answer generation is optimized through the error feedback loop. The "generated answer + error information" is added to the prompt of the solution sub-model, and the solution sub-model regenerates the answer and verifies it again.

[0073] In this embodiment, a solution hint template can be pre-constructed. The verified new question is input into the solution hint template to obtain an answer generation prompt word. The verified new question and the answer generation prompt word are input into the solution sub-model. The solution sub-model determines the time complexity and space complexity according to the algorithmic test points of the verified new question under the prompt of the answer generation prompt word. According to the time complexity and space complexity, the corresponding problem-solving idea process and algorithm step annotation information are determined. The verified new question is solved according to the problem-solving idea process, and the corresponding steps of the answer are annotated according to the algorithm step annotation information, and it is identified whether the answer covers all the constraints. The answer that covers all the constraints and the error information in the answer verification process are input into the solution hint template to obtain an answer correction prompt word. The answer correction prompt word is input into the solution sub-model, and the solution sub-model regenerates a new answer according to the error information under the prompt of the answer regeneration prompt word.

[0074] In this embodiment, during the process of using the solution sub-model to solve a new question, the solution sub-model can be controlled to output sequentially through multiple steps, or the thinking chain process of all steps can be generated at once. The solution hint template includes at least an answer generation prompt word and an answer regeneration prompt word. The solution hint template of this embodiment includes two parts, namely the answer generation prompt word and the answer regeneration prompt word. For the template corresponding to the former answer generation prompt word, the exemplary scenario data at this time is the verified new question. Input it into the solution hint template, and the solution hint template containing these input data is defined as the answer generation prompt word. At this time, the answer generation prompt word contains the answer generation prompt word in the solution hint template. For the template corresponding to the latter answer regeneration prompt word, the exemplary scenario data at this time is the answer that covers all the constraints and the error information in the answer verification process. Input it into the solution hint template, and the solution hint template containing these input data is defined as the answer correction prompt word. At this time, the answer correction prompt word contains the answer regeneration prompt word. Exemplarily, the solution hint template can be expressed as: As an expert in programming language processing tasks, please generate a #question# solution according to the following steps: 1. Analyze the core algorithm requirements of the question and clarify the time and space complexity goals; 2. Design the problem-solving idea process and mark the key algorithm steps; 3. Write code using the specified programming language, including detailed comments; 4. Self-verify whether the code covers all the constraints of the question; #Question#: {question}.

[0075] Please regenerate a solution for the #new question# according to the #error description# and the #submitted answer#.

[0076] #Submitted Answer#: {Answer that did not pass the verification}; The following problems exist with this answer: #Error Description# {Error description}.

[0077] Please reply strictly in the following format: Step 1#Error Analysis#: Step 2#Code Modification#: Step 3#Self-Check#: #Question#: {Question}.

[0078] After the answer is verified, the generated "Question - Test Case - Answer" can be formatted and saved as a complete piece of data in the database for subsequent training of the language model corresponding to the programming language processing task.

[0079] As can be seen from the above, in the answer verification process of this embodiment, a sandbox is used to perform verification, and those that do not pass the verification are repeatedly answered, ensuring the correctness and practicality of the answer. By adopting a loop verification and answering mechanism, the generated answer solution not only meets the requirements of the question but also has engineering reliability. In addition, by using the answering sub-model to answer and verify the newly generated questions, manual intervention is reduced, the labor cost is lowered, and the efficiency of generating training sample data for code classes is effectively improved.

[0080] Based on the above embodiment, the present invention also provides an implementation method for generating training sample data with different difficulty gradients, which may include the following: Adjust the modification range boundary conditions of the new question to increase the difficulty of the current new question; input the new question and the prompt words into the sample data generation model, modify the new question within the new modification range boundary conditions based on the same algorithm thinking, and generate a new test case set that covers the constraint conditions of the current new question according to the current new question type and use case generation conditions; verify the current new question and the new test case set, output the new test case set that passes the verification, answer the current new question that passes the verification, and output a new answer that passes the verification of the new test case set; use the current new question, the new test case set, and the new answer as the second set of training sample data to generate training sample data with different difficulty levels.

[0081] In this embodiment, in each iteration process, the newly generated question from the previous time is used as the seed question for the current iteration. Referring to the feedback data in the previous generation process, such as the difficulty assessment feedback when generating the new question, the constraint point coverage feedback when generating the test cases, and the error information during the verification process of the new question and test cases, the question modification strategy and parameters are refined and adjusted. For example, in the first iteration, a relatively simple question prototype can be generated according to the basic settings first; in the second iteration, using the same prompt words, according to the feedback of the first generated question in terms of difficulty assessment, knowledge point coverage, etc., the question complexity is increased again, such as adjusting the conditions for applying the algorithm, adding limiting factors, etc. By cycling in this way, a series of questions about specific algorithm applications from simple and basic to complex and advanced can be gradually generated, and the corresponding answers and test cases are equipped synchronously, fully meeting the needs of diverse training data for the large model in different training stages such as pre-training and fine-tuning, as well as different application scenarios such as academic research and industrial applications.

[0082] As can be seen from the above, in this embodiment, through multiple rounds of iterative evolution of the same question, multiple difficulty variants can be generated for the same question, and various types of questions at different difficulty levels can be efficiently collected to form a complete difficulty spectrum, meeting the needs of diverse data in different training stages and application scenarios, and enhancing the adaptability and generalization ability of the programming language processing task model.

[0083] Finally, the present invention also provides a method for generating training sample data based on multi-agent collaboration. In this embodiment, multiple agents collaborate to complete the processes of question evolution, test generation, question review, answer generation, and iterative evolution, as Figure 7 shown, and may include the following contents: Pre - adopt different types of language models as the problem - case generation sub - model, generated - data verification sub - model, and answer sub - model, and combine the problem - case generation sub - model, generated - data verification sub - model, and answer sub - model into a sample - data generation model. Obtain prompt words, which at least include problem - generation prompt words, case - generation prompt words, single - round review prompt words, review - result prompt words, answer - generation prompt words, and answer - correction prompt words. Use different LLMs to respectively undertake the roles of problem evolution, test generation, problem review, and answer generation: The problem - case generation sub - model modifies the seed problem within the boundary conditions of the modification range based on the same algorithmic thinking according to the problem - generation prompt words, and generates a test - case set covering the constraints of the new problem according to the new problem type and case - generation conditions according to the case - generation prompt words; The generated - data verification sub - model verifies the new problem and the test - case set multiple times according to the single - round review prompt words, and when the multi - round voting results do not meet the preset voting - passing conditions, modifies the new problem or the test - case set according to the review - result prompt words and each single - round review result; Here, "or" means that if the new problem fails the verification, the new problem is modified according to the review - result prompt words and each single - round review result, and if the test - case set fails the verification, the test - case set is modified according to the review - result prompt words and each single - round review result. The answer sub - model answers the verified new problem according to the answer - generation prompt words, and modifies the answer according to the answer - correction prompt words when the answer fails the answer - verification process.

[0084] In this embodiment, by constructing a closed - loop process of "problem evolution - test generation - problem review - answer generation", the automatic generation of high - quality programming training data with multiple difficulty levels is realized. Among them, the problem - case generation sub - model is used to realize algorithm deconstruction and problem evolution to ensure the examination of the core algorithmic thinking; the data verification sub - model guarantees the problem quality through a voting and feedback mechanism; the answer sub - model generates code solutions and performs strict tests through execution in a sandbox environment.

[0085] As can be seen from the above, this embodiment can generate training sample data with a complete difficulty spectrum through dynamic difficulty regulation, multi - dimensional test - case generation, and iterative evolution, and finally generate structured training data with multiple difficulty - evolution levels, which can not only support supervised learning in the instruction - fine - tuning stage but also meet the reward - calculation requirements in the reinforcement - learning stage, effectively improving the performance of the language model in code understanding and generation tasks and enhancing its ability to solve complex algorithm problems.

[0086] The present invention also provides a corresponding device for the training - sample - data generation method, further making the method more practical. Among them, the device can be described from the perspective of functional modules and the perspective of hardware respectively. The training - sample - data generation device described below can be mutually referred to with the training - sample - data generation method described above.

[0087] From the perspective of functional modules, refer to Figure 8 , Figure 8 which is a structural diagram of the training sample data generation device provided in this embodiment under a specific implementation manner. The device may include: A data input module 801, configured to input the seed questions and prompt words of the programming language processing task into the sample data generation model constructed based on the language model.

[0088] A data generation module 802, configured to use the prompt words through the sample data generation model to modify the seed questions within the boundary conditions of the modification range based on the same algorithmic thinking, generate a test case set covering the new question constraints according to the new question type and use case generation conditions; verify the new questions and the test case set, output the test case set that passes the verification, answer the new questions that pass the verification, and output the answers that pass the verification of the test case set; use the new questions, the test case set, and the answers as the first group of training sample data. Input the new questions into the sample data generation model, and adjust the boundary conditions of the modification range and / or the prompt words and / or the use case generation conditions to generate the second group of training sample data.

[0089] Exemplarily, in some implementation manners of this embodiment, the above data generation module 802 may further be configured to: analyze the examination points of the seed questions to determine the algorithm type and / or data structure type of the seed questions; determine various modification methods of the seed questions according to the algorithm type and / or data structure type; select, from the random combinations of the seed questions and each modification method, the modification method that meets the condition of having the same examination point as the seed questions and at least adding one new technical point as the optimal combined modification method; perform corresponding modification on the seed questions based on the optimal combined modification method to obtain new questions.

[0090] As an exemplary implementation manner of the above embodiment, the above data generation module 802 may further be configured to: perform corresponding modification on the seed questions respectively based on each optimal combined modification method to obtain multiple initial new questions; screen out candidate new questions without errors and omissions from the initial new questions, compare the algorithm type and / or data structure complexity of the candidate new questions and the seed questions, and use the candidate new questions with a higher complexity than the seed questions as the new questions.

[0091] Exemplarily, in some other embodiments of this embodiment, the above data generation module 802 may further be configured to: obtain the composition types of test cases and the case generation specifications, where the case generation specifications indicate the corresponding relationship between the input and output of each test case and the design intent annotation information; determine a test case generation scheme according to the algorithm test points, explicit constraint conditions, and implicit constraint conditions of the new question; the test case generation scheme at least includes test coverage dimensions, the test ideas for each dimension, and the types of errors expected to be discovered; generate multiple test cases that meet the composition types of test cases and the case generation specifications according to the test case generation scheme.

[0092] Exemplarily, in some other embodiments of this embodiment, the above sample data generation model at least includes a question case generation sub-model, and the data generation module 802 may further be configured to: pre-construct a question generation prompt template, where the question generation prompt template at least includes test point analysis prompt words, evolution prompt words, optimal combination modification method prompt words, and rewrite prompt words; input the seed question into the question generation prompt template to obtain the question generation prompt words; input the seed question and the question generation prompt words into the question case generation sub-model, and the question case generation sub-model analyzes the test points of the seed question under the prompt of the test point analysis prompt words to determine the algorithm type and / or data structure type of the seed question, and determines multiple modification methods of the seed question according to the algorithm type and / or data structure type under the prompt of the evolution prompt words; select, under the prompt of the optimal combination modification method prompt words, from the random combinations of the seed question and each modification method, the modification method that meets the condition of having the same test point as the seed question and at least adding one new technology point as the optimal combination modification method; and perform corresponding modification on the seed question based on the optimal combination modification method under the prompt of the rewrite prompt words to obtain a new question.

[0093] Exemplarily, in some other embodiments of this embodiment, the above sample data generation model at least includes a question use case generation sub-model, and the data generation module 802 can further be used to: pre-construct a use case generation prompt template, where the use case generation prompt template at least includes an analysis of test points and extraction of constraint condition prompt words, a test case generation scheme prompt word, and a use case generation prompt word; input a new question into the use case generation prompt template to obtain the use case generation prompt word; input the new question and the use case generation prompt word into the question use case generation sub-model, and the question use case generation sub-model obtains the test case composition type and the use case generation specification under the prompt of the analysis of test points and extraction of constraint condition prompt words, and the use case generation specification is that each test case has an input and output correspondence relationship and design intention annotation information; determine a test case generation scheme according to the algorithm test points, explicit constraint conditions, and implicit constraint conditions of the new question under the prompt of the test case generation scheme prompt word; the test case generation scheme at least includes a test coverage dimension, the test ideas of each dimension, and the types of errors expected to be found; generate a plurality of test cases that meet the test case composition type and the use case generation specification according to the test case generation scheme under the prompt of the use case generation prompt word.

[0094] Exemplarily, in some other embodiments of this embodiment, the above data generation module 802 can also be used to: generate multiple problem-solving ideas for a new question, and select at least two target problem-solving ideas for correctly solving the new question; when the algorithm difficulty of the new question is greater than the algorithm difficulty of the seed question, the new question meets the preset semantic conditions of no errors and no ambiguity, and the new question matches the test case set, then the new question passes the verification; if the test case set matches the new question and the test case set at least includes basic test cases, boundary test cases, and stress test cases, then the test case set passes the verification; when there are not at least two target problem-solving ideas, or the algorithm difficulty of the new question is less than or equal to the algorithm difficulty of the seed question, or the semantic conditions are not met, then generate question modification opinions; when the new question does not match the test case set, then generate question and use case modification opinions; when the test case set does not include at least one of basic test cases, boundary test cases, and stress test cases, then generate use case modification opinions.

[0095] Exemplarily, in some other embodiments of this embodiment, the above sample data generation model at least includes a generated data verification sub-model, and the above data generation module 802 can further be used to: pre-construct an audit prompt template, where the audit prompt template at least includes a single-round verification prompt word and a result output prompt word. The single-round verification prompt word at least includes a verification analysis prompt word and a modification opinion prompt word; the result output prompt word at least includes a modification prompt word and an audit result output prompt word; input the seed question, the new question, and the test case set into the audit prompt template to obtain a single-round audit prompt word; input the single-round audit prompt word, the seed question, the new question, and the test case set into the data verification sub-model multiple times to obtain multiple independent single-round audit results; vote on each single-round audit result. When the voting result does not meet the preset voting passing condition, input each single-round audit result into the audit prompt template to obtain an audit result prompt word, and input the audit result prompt word and each single-round audit result into the data verification sub-model; wherein, under the prompt of the modification prompt word, the data verification sub-model modifies the new question or the test case according to the modification opinions of each single-round audit result, and outputs the modified new question or the modified test case under the prompt of the audit result output prompt word.

[0096] Exemplarily, in some other embodiments of this embodiment, the above data generation module 802 can also be used to: determine the time complexity and space complexity according to the algorithm test points of the new questions that have passed the verification; determine the corresponding problem-solving idea process and algorithm step annotation information according to the time complexity and space complexity; solve the new questions that have passed the verification according to the problem-solving idea process, annotate the corresponding steps of the answer according to the algorithm step annotation information, and identify whether the answer covers all constraint conditions; input the test case set and the answer that covers all constraint conditions into the sandbox verification environment for answer verification, and when the answer verification fails, adjust the answer according to the error information in the answer verification process.

[0097] Exemplarily, in some other embodiments of this embodiment, the above sample data generation model at least includes an answer sub-model, and the data generation module 802 can further be used to: pre-construct an answer hint template, where the answer hint template at least includes answer hint words and answer re-generation hint words; input the verified new question into the answer hint template to obtain answer generation hint words; input the verified new question and the answer generation hint words into the answer sub-model, and the answer sub-model determines the time complexity and space complexity according to the algorithm test points of the verified new question under the hint of the answer generation hint words; determine the corresponding problem-solving idea process and algorithm step annotation information according to the time complexity and space complexity; answer the verified new question according to the problem-solving idea process, annotate the corresponding steps of the answer according to the algorithm step annotation information, and identify whether the answer covers all constraint conditions; input the answer that covers all constraint conditions and the error information in the answer verification process into the answer hint template to obtain answer correction hint words; input the answer correction hint words into the answer sub-model, and the answer sub-model re-generates a new answer according to the error information under the hint of the answer re-generation hint words.

[0098] Exemplarily, in some other embodiments of this embodiment, the above data generation module 802 can further be used to: adjust the modification range boundary conditions of the new question to increase the difficulty of the current new question; input the new question and hint words into the sample data generation model, modify the new question within the new modification range boundary conditions based on the same algorithm thinking, and generate a new test case set that covers the constraint conditions of the current new question according to the current new question type and use case generation conditions; verify the current new question and the new test case set, output the verified new test case set, answer the verified current new question, and output the new answer verified by the new test case set; use the current new question, the new test case set, and the new answer as the second set of training sample data to generate training sample data with different difficulty levels.

[0099] Exemplarily, in some other embodiments of this embodiment, the above data generation module 802 may further be configured to: pre - adopt different types of language models as the topic use - case generation sub - model, the generated data verification sub - model, and the answer sub - model, and combine the topic use - case generation sub - model, the generated data verification sub - model, and the answer sub - model into a sample data generation model; the topic use - case generation sub - model generates topic generation prompt words, modifies the seed topic within the modification range boundary conditions based on the same algorithmic thinking, and generates a test case set covering the new topic constraint conditions according to the use - case generation prompt words according to the new topic type and use - case generation conditions; the generated data verification sub - model verifies the new topic and the test case set multiple times according to the single - round review prompt words, and when the multi - round voting results do not meet the preset voting - passed conditions, modifies the new topic and the test case set accordingly according to the review result prompt words; the answer sub - model answers the new topic that has passed the verification according to the answer generation prompt words, and modifies the answer according to the answer correction prompt words when the answer does not pass the answer verification process.

[0100] For the description of the features in the corresponding embodiment of the training sample data generation device, reference can be made to the relevant description in the corresponding embodiment of the training sample data generation method, which will not be elaborated here one by one.

[0101] The training sample data generation device mentioned above is described from the perspective of functional modules. Further, the present invention also provides an electronic device, which is described from the hardware perspective. The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above - mentioned embodiments of the training sample data generation method.

[0102] An embodiment of the present application also provides a computer - readable storage medium. A computer program is stored in the computer - readable storage medium, where the computer program is configured to execute the steps in any of the above - mentioned embodiments of the training sample data generation method when running.

[0103] In an exemplary embodiment, the above - mentioned computer - readable storage medium may include, but is not limited to: USB flash drive, read - only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disk, etc., all kinds of media that can store computer programs.

[0104] An embodiment of the present application also provides a computer program product. The above - mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above - mentioned embodiments of the training sample data generation method.

[0105] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the training sample data generation method.

[0106] The above has introduced in detail a training sample data generation method, an electronic device, a computer-readable storage medium, and a computer program product provided by the present invention. Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. Whether the described units and algorithm steps of each example in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A method for generating training sample data, characterized in that, including: Input the seed questions and prompting words for programming language processing tasks into a sample data generation model built based on a language model; Using the prompting words through the sample data generation model, modify the seed questions within the boundary conditions of the modification range based on the same algorithmic thinking, generate a test case set covering the constraints of the new questions according to the new question type and test case generation conditions; verify the new questions and the test case set, output the verified test case set, answer the verified new questions, and output the answers verified by the test case set; Use the new questions, the test case set, and the answers as the first set of training sample data; Input the new questions into the sample data generation model, adjust the boundary conditions of the modification range and / or the prompting words and / or the test case generation conditions to generate a second set of training sample data.

2. The training sample data generation method according to claim 1, wherein Modifying the seed questions within the boundary conditions of the modification range based on the same algorithmic thinking includes: By analyzing the examination points of the seed questions, determine the algorithm type and / or data structure type of the seed questions; According to the algorithm type and / or data structure type, determine multiple modification methods for the seed questions; From the random combinations of the seed questions and each modification method, select the modification method that meets the conditions of having the same examination points as the seed questions and at least adding one new technical point as the optimal combined modification method; Modify the seed questions accordingly based on the optimal combined modification method to obtain new questions.

3. The training sample data generation method according to claim 2, wherein There are multiple optimal combined modification methods. Modifying the seed questions accordingly based on the optimal combined modification methods to obtain new questions includes: Modify the seed questions respectively based on each optimal combined modification method to obtain multiple initial new questions; Screen the candidate new questions with no errors and no omissions from each initial new question, compare the algorithm type and / or data structure complexity of the candidate new questions and the seed questions, and use the candidate new questions with a higher complexity than the seed questions as new questions.

4. The training sample data generation method according to claim 1, wherein Generating a test case set covering the constraints of the new questions according to the new question type and test case composition type includes: Obtain the test case composition type and test case generation specifications, and the test case generation specifications are that each test case includes at least the input and output correspondence and the design intention annotation information; According to the algorithm examination points, explicit constraints, and implicit constraints of the new questions, determine the test case generation plan; the test case generation plan includes at least the test coverage dimension, the test ideas for each dimension, and the error types expected to be found; Generate multiple test cases that meet the test case composition type and the test case generation specifications according to the test case generation plan.

5. The method for generating training sample data according to claim 1, wherein The sample data generation model at least includes a question and test case generation sub-model. Modifying the seed questions within the boundary conditions of the modification range based on the same algorithmic thinking includes: Pre-construct a question generation prompt template, and the question generation prompt template at least includes examination point analysis prompting words, evolution prompting words, optimal combined modification method prompting words, and rewriting prompting words; Input the seed topic into the topic generation prompt template to obtain a topic generation prompt word; Input the seed topic and the topic generation prompt word into the topic use case generation sub-model. The topic use case generation sub-model analyzes the test points of the seed topic under the prompt of the test point analysis prompt word, determines the algorithm type and / or data structure type of the seed topic, and determines various modification methods of the seed topic according to the algorithm type and / or data structure type under the prompt of the evolution prompt word; Under the prompt of the optimal combination modification method prompt word, select a modification method that meets the condition of having the same test point as the seed topic and at least adding one new technical point from the random combination of the seed topic and each modification method as the optimal combination modification method; Based on the optimal combination modification method, make corresponding modifications to the seed topic under the prompt of the rewrite prompt word to obtain a new topic.

6. The method for generating training sample data according to claim 1, wherein The sample data generation model at least includes a topic use case generation sub-model, and generates a test case set covering the constraint conditions of the new topic according to the new topic type and use case generation conditions, including: Pre-construct a use case generation prompt template, which at least includes a test point analysis and constraint condition extraction prompt word, a test case generation scheme prompt word, and a use case generation prompt word; Input the new topic into the use case generation prompt template to obtain a use case generation prompt word; Input the new topic and the use case generation prompt word into the topic use case generation sub-model. The topic use case generation sub-model obtains the test case composition type and use case generation specification under the prompt of the test point analysis and constraint condition extraction prompt word. The use case generation specification is that each test case at least includes the input and output correspondence and the design intention annotation information; Under the prompt of the test case generation scheme prompt word, determine the test case generation scheme according to the algorithm test points, explicit constraint conditions and implicit constraint conditions of the new topic; The test case generation scheme at least includes the test coverage dimension, the test ideas of each dimension, and the error types expected to be found; Under the prompt of the use case generation prompt word, generate multiple test cases that meet the test case composition type and the use case generation specification according to the test case generation scheme.

7. The training sample data generation method according to claim 1, wherein Verify the new topic and the test case set, including: Generate multiple problem-solving ideas for the new topic, and select at least two target problem-solving ideas that correctly solve the new topic; When the algorithm difficulty of the new topic is greater than the algorithm difficulty of the seed topic, the new topic meets the preset semantic conditions of no errors and no ambiguity, and the new topic matches the test case set, then the new topic passes the verification; If the test case set matches the new topic, and the test case set at least includes basic test cases, boundary test cases, and stress test cases, then the test case set passes the verification; When there are not at least two target problem-solving ideas, or the algorithm difficulty of the new topic is less than or equal to the algorithm difficulty of the seed topic, or the semantic conditions are not met, generate topic modification opinions; When the new question does not match the test case set, generate question and use case modification suggestions; When the test case set does not include at least one of the basic test cases, boundary test cases, and stress test cases, generate use case modification suggestions.

8. The training sample data generation method according to claim 1, wherein The sample data generation model at least includes a generated data verification sub-model to verify the new question and the test case set, including: Pre-construct an audit prompt template, the audit prompt template at least includes a single-round verification prompt word and a result output prompt word, the single-round verification prompt word at least includes a verification analysis prompt word and a modification suggestion prompt word; the result output prompt word at least includes a modification prompt word and an audit result output prompt word; Input the seed question, the new question, and the test case set into the audit prompt template to obtain a single-round audit prompt word; Input the single-round audit prompt word, the seed question, the new question, and the test case set into the data verification sub-model multiple times to obtain multiple independent single-round audit results; Vote on each single-round audit result. When the voting result does not meet the preset voting passing condition, input each single-round audit result into the audit prompt template to obtain an audit result prompt word, and input the audit result prompt word and each single-round audit result into the data verification sub-model; Among them, under the prompt of the modification prompt word, the data verification sub-model modifies the new question or the test case according to the modification suggestions of each single-round audit result, and outputs the modified new question or the modified test case under the prompt of the audit result output prompt word.

9. The training sample data generation method according to claim 1, wherein Answer the new question that has passed the verification and output the answer verified by the test case set, including: Determine the time complexity and space complexity according to the algorithm test points of the new question that has passed the verification; Determine the corresponding problem-solving idea process and algorithm step annotation information according to the time complexity and the space complexity; Answer the new question that has passed the verification according to the problem-solving idea process, annotate the corresponding steps of the answer according to the algorithm step annotation information, and identify whether the answer covers all the constraint conditions; Input the test case set and the answer covering all the constraint conditions into the sandbox verification environment for answer verification, and when the answer verification fails, adjust the answer according to the error information in the answer verification process.

10. The method for generating training sample data according to claim 1, wherein The sample data generation model at least includes an answer sub-model to answer the new question that has passed the verification and output the answer verified by the test case set, including: Pre-construct an answer prompt template, the answer prompt template at least includes an answer prompt word and an answer re-generation prompt word; Input the new question that has passed the verification into the answer prompt template to obtain an answer generation prompt word; Generate a prompt based on the verified new question and the answer, and input it into the answer sub-model. The answer sub-model, under the prompt of the answer generation prompt, determines the time complexity and space complexity according to the algorithm test points of the verified new question; according to the time complexity and the space complexity, determine the corresponding problem-solving idea process and algorithm step annotation information; answer the verified new question according to the problem-solving idea process, annotate the corresponding steps of the answer according to the algorithm step annotation information, and identify whether the answer covers all the constraints; Input the answer that covers all the constraints and the error information during the answer verification process into the answer correction template to obtain an answer correction prompt; Input the answer correction prompt into the answer sub-model. The answer sub-model, under the prompt of the answer regeneration prompt, regenerates a new answer according to the error information.

11. The method for generating training sample data according to any one of claims 1 to 10, characterized in that Input the new question into the sample data generation model, and adjust the modified range boundary conditions and / or the prompt and / or the test case generation conditions to generate a second set of training sample data, including: Adjust the modified range boundary conditions of the new question to increase the difficulty of the current new question; Input the new question and the prompt into the sample data generation model, modify the new question within the new modified range boundary conditions based on the same algorithm thinking, and generate a new test case set that covers the constraints of the current new question according to the current new question type and the test case generation conditions; verify the current new question and the new test case set, output the verified new test case set, answer the verified current new question, and output a new answer verified by the new test case set; use the current new question, the new test case set and the new answer as the second set of training sample data to generate training sample data with different difficulty levels.

12. The method for generating training sample data according to any one of claims 1 to 10, characterized in that, The prompt includes a question generation prompt, a test case generation prompt, a single-round review prompt, a review result prompt, an answer generation prompt, and an answer correction prompt; Using the prompt through the sample data generation model includes: Previously use different types of language models as the question test case generation sub-model, the generated data verification sub-model, and the answer sub-model, and combine the question test case generation sub-model, the generated data verification sub-model, and the answer sub-model into a sample data generation model; The question use case generation sub-model generates prompt words according to the question, modifies the seed question within the boundary conditions of the modification range based on the same algorithmic thinking, and generates test case sets that cover the constraints of the new question according to the new question type and use case generation conditions; the generated data verification sub-model verifies the new question and the test case set multiple times according to the single-round review prompt words, and when the multi-round voting results do not meet the preset voting passing conditions, modifies the new question or the test case set accordingly according to the review result prompt words and each single-round review result; the answer sub-model answers the new question that has passed the verification according to the answer generation prompt words, and when the answer does not pass the answer verification process, modifies the answer according to the answer correction prompt words.

13. An electronic device, characterized in that, including: a memory for storing computer programs; a processor for implementing the steps of the training sample data generation method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the training sample data generation method according to any one of claims 1 to 12 are implemented.

15. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the training sample data generation method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Language large model training method, system and device and computer readable storage medium

    CN118210895A

  • Instruction data generation method and device, computer equipment and storage medium

    CN118798216A

  • Customer service work order classification method and device, computer program product and electronic equipment

    CN119003777A

  • Question answering method and device based on large model, training method and device, intelligent agent, equipment and medium

    CN119106123A

  • Large model training method and device, electronic equipment, storage medium and program product

    CN119293500A

Cited By

  • Computing device and method for synthesizing code to generate training data

    CN120524999A

  • Multi-modal model training method and device based on thinking chain prompt pool

    CN121146070A

  • Construction method of character image generation combination library, electronic equipment and storage medium

    CN121330117A

  • Data generation method and electronic equipment

    CN121525766A

  • Method and device for generating problem solving strategy

    CN121981284A