Method and device for generating training data for reasoning large model

By introducing the alignment principle and prompt words of STEM subject information into the inference model to generate training data, the problem of insufficient training data in the STEM field is solved, and the calculation accuracy and problem-solving ability of the model in the STEM field is improved.

CN120450044APending Publication Date: 2025-08-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510541565.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The calculation results of existing inference models in the STEM field are not high, mainly due to the lack of sufficient quantity and quality training data.

Method used

By obtaining prompt words containing alignment principles and science and engineering discipline information, input the inference model to generate response results, and verify them according to the alignment principles, determine the training data that meets the requirements, and adjust the model parameters to improve problem-solving capabilities in the STEM field.

Benefits of technology

It improves the problem solving ability and calculation accuracy of inference models in the STEM field, and reduces the calculation error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450044A_ABST
    Figure CN120450044A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for generating training data for a large reasoning model. The method comprises the steps that a first cue word is obtained, the first cue word comprises an alignment principle and information of the science and engineering subject to which a to-be-generated problem belongs, and the alignment principle is used for constraining a response result generated by a large reasoning model; the first cue word is input into the large reasoning model, a first response is obtained, the first response comprises the first question, the first answer and a reasoning process used for deducing the first question and the first answer, and the first answer is used for answering the first question; and according to an alignment principle, verifying the first response, and if the verification is passed, determining training data for reasoning the large model according to the first question and the first answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular, to a method and apparatus for generating training data for inferring a large model. Background Art

[0002] Large Language Models (LLMs), also known as large models, are deep learning models based on natural language processing, trained on large text corpora and containing hundreds of millions or more parameters. Users can interact with the LLMs by inputting natural language text and receive responses to their questions.

[0003] Compared to conventional large models, inference large models tend to break down a given question into smaller steps (often called reasoning or thought processes) before answering it. This makes them suitable for answering questions that require complex analysis and logical reasoning. Currently, inference large models often suffer from low accuracy when used to answer questions related to STEM (Science, Technology, Engineering, and Mathematics). This is primarily due to a lack of sufficient and high-quality training data in STEM fields. Summary of the Invention

[0004] The embodiments of this specification provide a solution for generating training data for reasoning large models, which can automatically generate training data in the STEM field for reasoning large models.

[0005] In a first aspect, an embodiment of the present specification provides a method for generating training data for an inference large model, comprising: obtaining a first prompt word, the first prompt word including an alignment principle and information about the science and engineering discipline to which the question to be generated belongs, the alignment principle being used to constrain the response result generated by the inference large model; inputting the first prompt word into the inference large model to obtain a first response, the first response including a first question, a first answer, and a reasoning process for deriving the first question and the first answer, the first answer being used to answer the first question; verifying the first response according to the alignment principle, and if the verification passes, determining the training data for the inference large model based on the first question and the first answer.

[0006] In some embodiments, the alignment principle includes at least one of the following: an alignment principle for indicating the modality of the response result; an alignment principle for indicating the language of the response result; an alignment principle for indicating the structure of the question to be generated; an alignment principle for indicating the authenticity of the question to be generated; an alignment principle for indicating the logic of the reasoning process; and an alignment principle for indicating the format of the question to be generated and the answer to be generated.

[0007] In some embodiments, the science and engineering discipline information is multi-level category information.

[0008] In some embodiments, the verifying of the first response according to the alignment principle includes at least one of the following verification methods: evaluating the first response by an evaluation model according to the alignment principle; scoring the first response using a first reward model according to the alignment principle; and obtaining an expert's review result of the first response according to the alignment principle.

[0009] In some embodiments, the first prompt word further includes question difficulty information, where the question difficulty information is used to indicate the difficulty level of the question to be generated.

[0010] In some embodiments, after determining that the first question and the first answer are training data for inferring a large model, the method further includes:

[0011] Based on the training data, the parameters of the inference model are adjusted to obtain the adjusted inference model.

[0012] In some embodiments, the training data includes sample questions and labeled answers, and the parameters of the reasoning model are adjusted based on the training data, including: inputting the sample questions into the reasoning model to obtain the reasoning path and the output answer; using a second reward model to determine the reward score based on the reasoning path, the output answer and the labeled answer, and the reward score is related to at least one of the following indicators: the accuracy of the output answer, the logic of the reasoning path, and the complexity of the reasoning path; according to the reward score, the parameters of the reasoning model are adjusted.

[0013] In some embodiments, after obtaining the adjusted reasoning large model, the method further includes: inputting a second prompt word into the adjusted reasoning large model to obtain a second response, the second prompt word including the alignment principle and the science and engineering discipline information to which the question to be generated belongs, the second response including a second question, a second answer, and a reasoning process for deriving the second question and the second answer, the second answer being used to answer the second question; verifying the second response according to the alignment principle, and if the verification passes, determining that the second question and the second answer are new training data.

[0014] In some embodiments, the difficulty level of the question to be generated indicated by the question difficulty information in the second prompt word is higher than the difficulty level of the question to be generated indicated by the question difficulty information in the first prompt word.

[0015] In a second aspect, an embodiment of the present specification provides a device for generating training data for an inference large model, including: a data acquisition module, used to obtain a first prompt word, the first prompt word including an alignment principle and information about the science and engineering discipline to which the question to be generated belongs, the alignment principle being used to constrain the response result generated by the inference large model; a model inference module, used to input the first prompt word into the inference large model to obtain a first response, the first response including a first question, a first answer and a reasoning process for deriving the first question and the first answer, the first answer being used to answer the first question; a data verification module, used to verify the first response according to the alignment principle, and if the verification passes, determine that the first question and the first answer are training data for the inference large model.

[0016] In a third aspect, an embodiment of this specification provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any implementation manner in the first aspect.

[0017] In the solutions provided in the above embodiments of this specification, questions and high-accuracy answers for STEM fields can be automatically generated by the inference big model as training data, which can be used to adjust the parameters of the inference big model, thereby reducing the calculation error rate of the inference big model for problems in the STEM field and improving the big model's reasoning ability and problem-solving ability for problems in the STEM field. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 is a schematic diagram of a scheme for generating training data for inferring a large model in an embodiment of this specification;

[0020] Figure 2 This is a flow chart of a method for generating training data for inferring a large model in an embodiment of this specification;

[0021] Figure 3 is a schematic diagram of the inference link shown in this specification;

[0022] Figure 4 This is a flowchart for generating training data for inference large models and training inference large models in an embodiment of this specification;

[0023] Figure 5 It is a structural diagram of a device for generating training data for inferring a large model in an embodiment of this specification. DETAILED DESCRIPTION

[0024] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.

[0025] As mentioned above, the lack of high-quality training data in the STEM field limits the ability of large reasoning models to improve their capabilities in this field. Furthermore, current large reasoning models primarily use the following methods for reasoning: First, Direct Question Answering (DQA), where the model directly answers questions without explicitly displaying the reasoning steps. This method cannot handle complex problems requiring multi-step reasoning. Furthermore, due to its opacity, it is difficult to diagnose the causes of errors, making it unsuitable for reasoning and generating problems in the STEM field. Second, Chain-of-Thought (CoT) is a method for guiding large models to perform multi-step reasoning, improving the accuracy of complex problem solving by allowing the model to demonstrate its thought process. However, this method allows for relatively free reasoning and lacks strict structural and logical constraints. For example, invalid or erroneous steps may be generated, and there may be problems such as mixed language usage and modality inconsistencies. The generated answers may be poorly formatted and verifiable, making it difficult to consistently produce high-quality reasoning data that meets all stringent requirements.

[0026] To this end, the present invention proposes a method for generating training data for reasoning large models, such as Figure 1 As shown, first, a first prompt word can be obtained. The first prompt word contains n alignment principles and information about the science and engineering discipline to which the question to be generated belongs. The alignment principles can be related to requirements such as logic, formatting, and structure, etc., to constrain the questions, answers, and reasoning process generated by the reasoning model. Then, the prompt word is input into the reasoning model to obtain a first response. The first response includes the first question, the first answer, and the reasoning process used to derive the first question and the first answer. Next, the first response is verified according to the alignment principles. If the verification passes, the sample question and annotated answer in the training data for the reasoning model are determined based on the first question and the first answer. If the verification fails, the first question and the first answer are not determined as training data.

[0027] This method can guide the model to generate STEM questions by adding science and engineering discipline information to the prompts. Furthermore, by adding alignment principles to the prompts, the questions and answers generated by the model conform to the alignment principles, thus serving as high-quality training data. In this way, the inference model automatically generates STEM-specific questions and highly accurate answers as training data, overcoming the current shortage of high-quality training data in the STEM field. Using the training data obtained by this method to adjust the parameters of the inference model can reduce the inference model's computational error rate for STEM questions and improve the model's reasoning and problem-solving capabilities for STEM problems.

[0028] The following first explains the alignment principle.

[0029] Aligning large models is a core step in ensuring their behavior aligns with human intentions and values, guiding their evolution in line with human needs and expectations. Self-alignment refers to the ability of a model to proactively adjust its behavior to meet pre-defined ethical guidelines, values, or mission objectives during training or inference, without relying on external human feedback. This approach, through built-in mechanisms or intrinsic learning capabilities, enables the model to possess the capabilities of self-reflection and dynamic calibration, reducing reliance on external oversight and improving safety and controllability.

[0030] The alignment principles proposed in the examples of this specification are used to constrain the response results generated by the large inference model. This example does not limit the specific content of the alignment principles. By systematically designing alignment principles, it is possible to achieve comprehensive optimization of model generation capabilities, ensuring that STEM questions have the appropriate difficulty, diversity, and logical rigor, while also improving the accuracy of reference answers. This can cover key aspects such as problem construction, language standardization, modal control, and authenticity verification, significantly improving and strengthening the model's complex reasoning capabilities while overcoming the limitations of conventional thought chain methods.

[0031] In some embodiments, alignment principles may include at least one of the following: an alignment principle for indicating the modality of a response result; an alignment principle for indicating the language of a response result; an alignment principle for indicating the structure of a question to be generated; an alignment principle for indicating the authenticity of a question to be generated; an alignment principle for indicating the logic of a reasoning process; and an alignment principle for indicating the format of a question to be generated and an answer to be generated. Each alignment principle is described below.

[0032] The alignment principle for indicating the modality of the response results, also known as modality consistency alignment, must be strictly limited to pure text questions. It is used to convert descriptions of various complex cross-modal problems that may be involved in STEM disciplines, such as organic chemical molecular structures, biological genetic maps, and other complex problems involving multiple modalities.

[0033] The alignment principle for indicating the language of the response results, also known as language consistency alignment, is used to reduce or prevent the common multi-language confusion reasoning problems during the reasoning process. All aspects of the process, including problem statement, reasoning methods, and answer presentation, must be conducted entirely in English to resolve various problems discovered by the data and reasoning team, such as bad cases (unfavorable situations) caused by the mixing of Chinese and English, leading to model reasoning errors and model verification.

[0034] Alignment principles, also known as problem structure consistency alignment, are used to indicate the structure of the questions to be generated. This prevents logical branching from making the questions unverifiable. Questions can only contain a single question and must not contain sub-questions or derivative questions.

[0035] Alignment principles, also known as question authenticity alignment, are used to indicate the authenticity of questions to be generated. This is used to prevent or reduce false citations or ambiguous statements in the model. Questions must be carefully designed to ensure accuracy and consistency. When citing authoritative sources, direct copying or excerpting is prohibited, and ambiguity, bias, or error must be eliminated. By taking advantage of the inherent flaws of the LLM, the generation of unfounded and meaningless questions can be reduced or prevented from occurring at the source of question generation.

[0036] Alignment principles, also known as logical consistency alignment, are used to indicate the logical integrity of reasoning processes. This principle is used to reduce or prevent errors and confusion within the logical process. Problem-solving must be based entirely on rigorous reasoning or systematic deduction. Empirical methods (such as pattern matching, heuristics, shortcuts, or fabricated data) are prohibited. All intermediate reasoning steps must be validated by STEM theory or logical reasoning to avoid skipping logical connections based on intuitive guesswork.

[0037] The alignment principle used to indicate the format of the questions to be generated and the answers to be generated, also known as question-answer format consistency alignment, is used to prevent or reduce the form of answers that cannot be determined by the model. Questions and answers must follow the format standards specified by the Math-Verify framework. According to the characteristics of STEM subjects, especially biology and other subjects, reduce or avoid situations involving comprehensive liberal arts and science questions that result in long character strings and other undeterminable situations. Moreover, the answer must be a single numerical value that can be objectively verified, and can be in the form of: pure numbers, values with units, ratios, or chemical formulas / equations. The use of formats that are difficult to verify (such as set operations or free text answers) is prohibited. If there are multiple solutions, a summary value (such as a sum or sum of squares) must be provided to ensure the uniqueness of the answer.

[0038] In other embodiments, the alignment principles may also include alignment principles for indicating the complexity of the questions to be generated and the answers to be generated. For example, based on the current traditional single QA question-and-answer questions in the industry, multi-solution problems are constructed, and unique solutions are forced through summary values (such as sum, sum of squares), and more complex question types are added to increase the source of complex questions in the model (such as "calculate the total number of moles of all possible products"); it may also include alignment of question difficulty and topic diversity, which is used to guide the model to generate questions of different difficulty levels and diversity in various disciplines, such as Olympiad-level high-complexity problems and GPQA topic distribution coverage. The complexity of the questions can be improved by referring to the Olympiad BenchMark (benchmark) question construction principles, and the coverage of the questions can be improved by combining the construction of diverse topics, so that the difficulty and diversity of the overall question construction are guaranteed at the same time; it may also include alignment principles for indicating the generation position of the answers to be generated, for example, the final answer should be placed in the \\boxed{{*}} of the response result.

[0039] The detailed process of this method is further explained below. Figure 2A flowchart of a method for generating training data for inferring a large model according to an embodiment of the present specification is shown. The generation process can be performed by any device, platform, or device cluster with computing and processing capabilities, including steps 201-203 as shown below.

[0040] like Figure 2 As shown, in step 201, a first prompt word is obtained.

[0041] The first prompt includes the alignment principle and the science and engineering discipline to which the problem to be generated belongs. The first prompt is used to input the inference model, and the model is expected to return a corresponding response result based on the first prompt. The first prompt can be constructed by the model trainer. In different embodiments, the first prompt can be a different specific prompt, which is not limited in this specification. In different embodiments, the specific method of obtaining the first prompt can vary.

[0042] The alignment principle is used to constrain the response results generated by the inference large model. Specifically, it can be an instruction or constraint that the large model can operate to convey specific alignment requirements (such as question and answer format, inference method, modality requirements, etc.) to the inference large model. For a detailed explanation of the alignment principle, see above.

[0043] Science and engineering discipline information refers to specific STEM fields, such as electrical engineering, chemistry, biology, or sub-disciplines of physics.

[0044] In one implementation, the science and engineering discipline information is multi-level category information, that is, the science and engineering discipline information is classified according to a hierarchical structure (tree structure), and is refined step by step from major categories to minor categories (such as first-level disciplines → second-level disciplines), thereby providing rich and specific contextual information for reasoning large models. This embodiment does not limit the specific levels and formats of multi-level category information. Exemplarily, a STEM discipline knowledge base can be pre-constructed, and the STEM discipline knowledge can be structured and decomposed (Break down) according to preset "three-tier categories". For example, the multi-level category information can be: Chemistry → Organic Chemistry, Spectroscopy → Elimination Reaction, Infrared Spectroscopy, Carbonyl Compounds, Alcohols, Characteristic Infrared Absorption Frequency. For another example, it can be: First-level discipline: [Biology], Second-level discipline: [Molecular Biology, Genetics, Plant Pathology], Specific knowledge points: [Gene knockout, transcription factors, disease resistance, gene interactions, mutant analysis, pathogen resistance, plant-pathogen interactions, double mutant analysis]. This hierarchical structure ensures the topic coverage and depth of the generated questions, allowing for the targeted generation of complex questions in specific segments.

[0045] The current lack of sufficient quantity and quality of training data with Olympiad-level difficulty and disciplinary diversity has limited the ability of large reasoning models to improve their performance in STEM fields. They struggle to meet the challenges of challenging benchmarks such as the Graduate-Level Google-Proof Q&A Benchmark (GPQA). Currently, STEM datasets provided by mainstream open-source platforms (such as Hugging Face and Kaggle) are mostly limited to the elementary level of middle school or undergraduate studies. For example, the STEM dataset on Hugging Face only contains questions on basic concepts, and its knowledge complexity is significantly different from the advanced content covered in GPQA, such as carbonyl compounds and characteristic infrared absorption frequencies. Actual measurements show that approximately 80% of the questions in this dataset are only as difficult as those in a sophomore university course, making it difficult to support the demands of high-level knowledge reasoning.

[0046] In order to solve the problem of severe shortage of high-level complex STEM data, the science and engineering discipline information in this embodiment can be extracted from GPQA. For example, a large model such as the Claude model is used to perform deep semantic analysis and topic extraction on all GPQA queries (questions). After multiple rounds of iterative optimization, a scientific and rigorous three-level label classification system is finally constructed, which specifically includes: 1) first-level disciplines (such as basic disciplines such as chemistry and biology); 2) second-level disciplines (such as professional branches such as organic chemistry under chemistry); 3) specific knowledge points (such as core concepts such as elimination reaction and infrared spectroscopy). The science and engineering discipline information of the third level category can be extracted from this three-level label classification system. For example, in the query related to "hydroboration reaction", the science and engineering discipline information of the third level category extracted is: chemistry (first-level discipline) → organic chemistry, reaction mechanism (second-level discipline) → hydroboration, conjugated diene, stereoselectivity, regioselectivity, borane reagent, Ipc2BH (specific knowledge point). This hierarchical topic labeling scheme not only ensures the comprehensiveness of data synthesis, but also accurately controls the distribution ratio of knowledge points at different difficulty levels. At the same time, a network of associations between labels was established to capture the complex connections between interdisciplinary knowledge and provide reliable contextual information for reasoning in large models. Furthermore, statistical analysis of the distribution of labels and knowledge points within the three-level label classification system revealed an imbalance in the distribution of knowledge points within GPQA. Therefore, each time we extract science and engineering information, we sample based on the distribution of disciplines and corresponding knowledge points. This ensures that the resulting science and engineering information covers topics covered by GPQA while also broadening the distribution of topics involved in complex reasoning.

[0047] In some embodiments, the first prompt word also includes question difficulty information, which is used to indicate the difficulty level of the problem to be generated. For example, the problem difficulty level can be classified according to the education level, including: primary school, junior high school, high school (including technical secondary school), junior college, undergraduate, master's degree, doctoral degree, etc.; it can also be divided according to its complexity, the depth of knowledge involved, and the skills required to solve it. For example, it can be classified as follows: Simple: usually basic knowledge, direct factual questions, the answer can be obtained through simple calculations; Medium difficulty: involves some analysis or reasoning, requires a certain level of understanding and application ability, and the answer usually requires some understanding of the background of the problem and some relatively complex calculations or reasoning; Difficult: requires high theoretical knowledge and problem-solving skills, involves multi-step reasoning or in-depth analysis, and may require interdisciplinary knowledge or more complex calculations; Very difficult: very complex, involving advanced theories, concepts or technologies, requiring deep professional knowledge, and solving these problems usually requires long-term thinking, experimentation and multiple iterations; Extremely difficult: these problems may not have ready-made answers, or the answers require breaking through the limitations of current academic or technical fields. They are usually open-ended problems that experts in academia or industry spend a lot of time and resources to try to solve. It can also be classified according to the difficulty of subject competitions.

[0048] The following example illustrates a first prompt word A obtained in this step. This prompt word provides information about the difficulty of the question by describing the role of the large model. That is, the difficulty of the question is at the level of an Olympic competition. This prompt word A is in English and its Chinese translation is as follows:

[0049] To test the STEM reasoning and complex problem-solving abilities of outstanding graduate students from various biology disciplines, you, a senior biology professor at a world-renowned institution, are creating an Olympiad-level computational problem.

[0050] 1. You must refer to the following resources: Biology; Sub-disciplines: Molecular Biology, Genomics, Epigenetics; Basic Concepts: ChIP-seq (chromatin immunoprecipitation sequencing), DNA (deoxyribonucleic acid)-protein crosslinking, transcription factors, peak calling, fixation methods, IKAROS protein, DNA sequencing, chromatin immunoprecipitation.

[0051] 2. You must randomly select one or more items from the biology sub-discipline, then select several related concepts from the basic concepts based on the sub-discipline to form the outline of the problem. Finally, create a calculation problem.

[0052] Note: Questions must meet the following criteria:

[0053] Alignment Principle 1: All aspects of the process, including question setting, reasoning methods, and answer presentation, must be conducted entirely in English.

[0054] Alignment Principle 2: Place the final generated question between <Question Start> and <Question End>, and do not include the question title.

[0055] Alignment Principle 3: Solutions to problems must rely solely on rigorous reasoning or systematic deduction, avoiding empirical methods such as pattern matching, heuristics, simplification shortcuts, or fabrication. All intermediate reasoning steps should be properly justified to prevent the possibility of skipping logical connections through intuitive guesswork.

[0056] Alignment Principle 4: The answer must be a verifiable number, such as a plain number, a number with units, a ratio, or a biological formula / equation, to allow for objective evaluation. Avoid using difficult-to-verify formats, such as set assignments or free-form text answers. For questions with multiple numerical solutions, require a summary value (such as a sum or sum of squares) to ensure a unique answer.

[0057] Alignment Principle 5: Questions consist entirely of text-based questions.

[0058] Alignment Principle 6: Questions must be carefully crafted to ensure accuracy and consistency. They should cite authoritative sources, without direct copying or quoting, and must be free of ambiguity, bias, or errors.

[0059] Alignment Principle 7: It must contain only one question, without any sub-questions.

[0060] Alignment Principle 8: The final answer should be placed in \boxed{*}.

[0061] Next, in step 202, the first prompt word is input into the inference model to obtain a first response.

[0062] The first response includes a first question, a first answer, and a reasoning process for deriving the first question and the first answer, and the first answer is used to answer the first question.

[0063] A large inference model refers to a large-scale artificial intelligence model with complex logical reasoning capabilities. It is usually based on a large language model (LLM) architecture, but through algorithm optimization, training method improvement, or architecture design, it significantly improves performance in tasks such as mathematical deduction, logical analysis, and causal inference. This embodiment does not limit the large inference model used, and it can be a large model with different neural network architectures. For example, the large inference model in this embodiment can use DeepSeekR1Zero.

[0064] In this step, the first prompt word is input into the reasoning model. The reasoning model will first generate the reasoning process of the first question according to the alignment principle and the instructions of the science and engineering discipline information, and then generate the first question that complies with the alignment principle and belongs to the STEM field indicated by the science and engineering discipline information based on the reasoning process. Then, it will continue to generate the reasoning process of the first answer according to the alignment principle and the instructions of the question. Finally, it will generate an answer that complies with the alignment principle according to the reasoning process, and place the generated question and answer in the corresponding positions in the first response according to the instructions of the alignment principle.

[0065] The following example uses the first prompt word B input into the inference model in this step. The prompt word is in English and its Chinese translation is as follows:

[0066] <Role Beginning> As a senior biology professor at a world-class university, you are designing an Olympic-level computational problem to test the STEM reasoning and complex problem-solving abilities of outstanding graduate students in various biology disciplines. <Role Ending>

[0067] <Begin task description> You must refer to the following resources: Biology; Sub-disciplines: Molecular Biology, Cell Biology, Virology; Basic Concepts: Apoptosis, Viral Proteins, Death Effector Domains, Caspases, DISC Signaling, Protein-Protein Interactions, Cell Death Pathways.

[0068] You must randomly select one or more items from the sub-disciplines of biology and, based on the selected sub-disciplines, select several related concepts from the basic concepts to form an outline for the question. Finally, create a calculation question. <End of Task Description>

[0069] <Mandatory requirements> Note: The question must meet the following conditions:

[0070] Alignment Principle 1: All aspects of the process, including question setting, reasoning methods, and answer presentation, must be conducted entirely in English.

[0071] Alignment Principle 2: Questions consist only of text-based questions.

[0072] Alignment Principle 3: A question must contain only one question and must not have sub-questions.

[0073] Alignment Principle 4: Questions must be carefully crafted to ensure accuracy and consistency. Authoritative sources should be referenced but not directly copied or quoted, and any vague, inaccurate, or biased statements must be avoided.

[0074] Alignment Principle 5: Solutions to problems must rely solely on rigorous reasoning or systematic deduction, avoiding empirical methods such as pattern matching, heuristics, shortcuts, or fabrications. All intermediate reasoning steps should be fully justified to prevent skipping logical connections through intuitive guesswork.

[0075] Alignment Principle 6: Questions and answers follow the format prescribed by the Math-Verify framework.

[0076] Alignment Principle 7: The answer must be a single, verifiable number. This can be a plain number, a number with units, a ratio, or a biological formula / equation to allow for objective evaluation. Avoid using difficult-to-verify formats such as set operations or free-text answers. If a problem has multiple numerical solutions, request an aggregate value (such as a sum or sum of squares) to ensure a unique answer.

[0077] Alignment Principle 8: Place the final generated question between <Question Start> and <Question End>, and do not include the question title.

[0078] Alignment Principle 9: The final answer should be placed in \boxed{*}. <Mandatory end>

[0079] The first response B output by the reasoning model is in English. The first response B shows the detailed reasoning process, as well as the first question and the first answer according to the format requirements of the alignment principle. The Chinese translation is as follows:

[0080] <Thinking> Okay, I need to design an Olympiad-level math problem for biology graduate students, focusing on the following disciplines: molecular biology, cell biology, and virology. This problem will require the application of concepts such as apoptosis, viral proteins, death effector domains (DEDs), caspases, death receptor-activated signaling complexes (DISCs), protein-protein interactions, and cell death pathways. Let me start by selecting content from these sub-disciplines, perhaps combining molecular biology and virology, as the interaction between viral proteins and apoptosis is a good angle.

[0081] Apoptosis involves caspase and DISC signaling. Viral proteins often inhibit apoptosis, aiding viral survival. Therefore, a hypothetical question arises: a hypothetical viral protein interferes with caspase activation. Consider the DISC complex: When death receptors are activated, they form the DISC complex, which in turn activates initiator caspases, such as caspase-8. Caspase-8, in turn, activates executioner caspases, such as caspase-3, ultimately leading to apoptosis. If a viral protein binds to caspase-8, preventing its activation, how would this affect the caspase cascade?

[0082] I need to design a computational problem. Perhaps one that quantifies caspase activation. Suppose the DISC complex normally activates caspase-8 with a certain efficiency. A viral protein binds to caspase-8, reducing its activation rate. The problem might ask how much caspase-3 activity remains under this inhibition. Enzyme kinetics, such as the Michalis-Menden equation, could be used to consider whether the inhibition is competitive or noncompetitive.

[0083] Wait, item 5 requires a rigorous reasoning process, not empirical shortcuts. Therefore, this problem should be deduced step by step. Let's frame it this way: If the viral protein forms a complex with caspase-8, it reduces the amount of active caspase-8 available. Therefore, caspase-8 activates caspase-3. If the initial caspase-8 concentration is known, and the viral protein binds a specific percentage, how much caspase-3 will be activated?

[0084] Let me summarize: assume that each active caspase-8 molecule can cleave and activate 10 caspase-3 molecules. If the viral inhibitor binds 60% of caspase-8, leaving 40% active, the total amount of caspase-3 activated would be (initial caspase-8 concentration) * 0.4 * 10. However, one can get more detailed, using the Hill equation or binding equilibrium. Or perhaps using stoichiometry and binding ratios.

[0085] Alternatively, a simple ratio could be used: if 60% is inhibited, then 40% remains active. If each active caspase-8 activates 10 units of caspase-3, then the total caspase-3 is 10 times the original 40% caspase-8 concentration. However, perhaps the problem requires more steps. For example, given the binding constant Kd, calculate the fraction bound. This might be overly complex. Alternatively, a simple percentage of inhibition could be used.

[0086] Wait, the answer must be a specific number. Let's simplify this. Suppose the question states that the viral protein binds 60% of caspase-8, preventing its activation. Each active caspase-8 activates 10 caspase-3s. If the original caspase-8 concentration is 100 nM, what is the concentration of active caspase-3?

[0087] So, the active caspase-8 concentration is 100nM * (1-0.6) = 40nM. Each active caspase-8 activates 10 caspase-3s, so 40*10 = 400nM caspase-3. But you might want to be more precise. Or use a molar ratio. But make sure the units are clear. Alternatively, a straightforward answer of 400nM would also work. But perhaps add another layer. Perhaps consider that caspase-8 activation is part of a cascade reaction, and there's a time factor involved. But time can complicate things. This question needs to be text-based, without sub-questions. So you might want to make it a stoichiometric question, giving the percentage of inhibition.

[0088] Another approach: DISC complex formation requires certain proteins. Suppose the problem involves calculating the number of DISC complexes formed when a viral protein sequesters a component. For example, if each DISC requires four FADD proteins, and the viral protein binds to FADD with a certain affinity, reducing the available FADD, then calculating the number of DISC complexes formed would be a more complex equilibrium calculation.

[0089] However, the answer needs to be a single number. Let's stick with the caspase inhibition scenario. Let me draft the question:

[0090] "In an experiment studying viral inhibition of apoptosis, a viral protein binds to procaspase-8, preventing its conversion to caspase-8. The DISC signaling complex normally activates procaspase-8 at 50 nM. If the viral protein inhibits procaspase-8 activation by 75%, and each activated caspase-8 molecule activates 25 caspase-3 molecules, calculate the total concentration of activated caspase-3. Please provide your answer in nM."

[0091] So the answer is: 50nM * (1-0.75) = 12.5nM activated caspase-8. 12.5 * 25 = 312.5nM caspase-3. Answer: 312.5nM, boxed.

[0092] But wait, is the inhibitory effect on procaspase-8 or on active caspase-8? The question says "blocks its activation," so the 75% inhibition is on the activation process. So, if 50 nM is normally active and 75% is inhibited, then 25% is activated: 50 * 0.25 = 12.5 nM activated caspase-8. Then, each of these activates 25 caspase-3s, so 12.5 * 25 = 312.5. But the question should clarify whether the 50 nM refers to the total procaspase-8 or the amount that has already been activated. This needs clarification.

[0093] Alternatively, we can rewrite this as follows: The DISC complex activates procaspase-8 to caspase-8. The normal activation efficiency is 50 nM (i.e., under normal conditions, 50 nM of caspase-8 is active). The viral protein reduces this activation by 75%. Therefore, the activated caspase-8 becomes 50 * (1 - 0.75) = 12.5 nM. Each protein then activates 25 caspase-3s. The total caspase-3 level is 12.5 * 25 = 312.5 nM. The answer is 312.5 nM.

[0094] Alternatively, use different numbers to make the answer a whole number. For example, 60% inhibition, 10 activation factors. Then 50 * 0.4 * 10 = 200. The answer is 200 nM. Perhaps this is simpler. Let me adjust the question to make it easier to calculate.

[0095] Question: "In an experiment studying viral inhibition of apoptosis, the death-inducing signaling complex (DISC) activates procaspase-8 to caspase-8, which in turn activates caspase-3. Normally, DISC activation converts 100 nM of procaspase-8 to active caspase-8. A viral protein binds to procaspase-8 and inhibits activation by 60%. If each active caspase-8 molecule activates 5 caspase-3 molecules, what is the total concentration of active caspase-3 produced? Please provide your answer in nM."

[0096] Solution: 100nM*(1-0.6)=40nM active caspase-8. 40*5=200nM caspase-3. Answer: 200nM. Box: 200.

[0097] This seems to tick all the boxes: using apoptosis, DISC, caspases, and viral proteins. The calculations are simple, the answer is unique, and the units are clear. All steps are explained, with no shortcuts. The answer is a single number with units. OK, this should work. <Thinking>

[0098] <Question Start> In an experiment studying viral inhibition of apoptosis, the death-inducing signaling complex (DISC) activates procaspase-8 to caspase-8, which in turn activates caspase-3. Typically, DISC activation converts 100 nM of procaspase-8 into active caspase-8. A viral protein binds to procaspase-8 and inhibits activation by 60%. If each active caspase-8 molecule activates five caspase-3 molecules, what is the total concentration of active caspase-3 produced? Please provide your answer in nanomolar units. <Question End>\boxed{200}

[0099] It can be seen that the above reasoning process is carried out strictly in accordance with the requirements of the alignment principle, step-by-step, clear, and logically coherent. The first question and the first answer generated also meet the requirements of the alignment principle.

[0100] Then, in step 203, the first response is verified according to the alignment principle. If the verification passes, the training data for inferring the large model is determined based on the first question and the first answer.

[0101] In this step, the first question, the first answer, and one or more of the reasoning processes included in the first response can be verified to verify whether they meet the requirements of the alignment principle. For example, the first question and the first case are verified. If the verification passes, the first question is determined as a sample question in the training data, and the first answer is determined as the labeled answer to the sample question, that is, the ground truth. Alternatively, the first answer is reviewed and corrected to obtain the labeled answer to the sample question, thereby obtaining a training set containing a large amount of training data. For another example, in other embodiments, only the first question can be verified, and the answer corresponding to the first question can be obtained by calculating the first question, and the first question and the corresponding answer can be determined as training data.

[0102] This embodiment does not restrict the verification method used. For example, based on more rigorous scientific standards, combined with the alignment principle, further combined with domain experts and large-scale model scoring, the generated first question and first answer pairs can be deeply and comprehensively verified. The purpose is to further filter out poor quality data, such as those with logical errors, unreliable reward signals, and other data that do not meet the requirements of the alignment principle, to ensure that only the highest quality and most reliable data enters the training phase.

[0103] In other embodiments, it is also possible to verify whether the topic of the first question is consistent with the category topic in the science and engineering subject information, to ensure that the first response finally output meets the preset high standards in all dimensions such as difficulty, logic, topic, and format. The first question and answer that have been fully verified and integrated are used to construct the final instance-level self-aligned reasoning data.

[0104] In one implementation, this step may verify the first response using at least one of the following verification methods:

[0105] Based on the alignment principle, the evaluation model evaluates the first response. For example, another powerful evaluation model performs automated evaluation, inputting the first question and the first answer into the evaluation model to check for fluency, consistency, common sense, etc., and outputs an evaluation score.

[0106] Based on the alignment principle, a first-response reward model is used to score the first response. For example, a general or specialized first-response reward model or reward function can be used for scoring to quantitatively evaluate the quality of the inference results, capturing subtle features that are difficult to quantify using LLMs and experts. There are many ways to design and train a first-response reward model, such as using different model architectures for scoring or incorporating a wider range of human feedback or automated evaluation metrics.

[0107] Obtain expert review results for first responses based on alignment principles. For example, human domain experts provide authoritative annotation and review, providing reliable judgment, especially evaluating the accuracy and depth of professional knowledge, and obtaining a review score.

[0108] Furthermore, more diverse verification methods can be integrated, such as involving more domain experts, cross-validation, or more complex automated logic checking tools. Only high-quality first questions and first answers that pass this comprehensive verification process are ultimately confirmed as training data.

[0109] like Figure 3 As shown, taking the alignment principle and science and engineering discipline information shown in the first prompt word B and the first response B as an example, compared with conventional Direct QA and CoT reasoning, the Self-Alignment link designed in this embodiment includes the following core links:

[0110] Starting point (Self-Alignment): The starting point of the entire process is the alignment principle for self-alignment, which indicates that all subsequent steps will strictly follow the alignment principle.

[0111] Question generation: Based on the alignment principle, questions that meet the requirements are generated. This ensures that the questions are not only of high difficulty and topic diversity, but also strictly adhere to the high consistency requirements of language, modality, structure, authenticity, and format.

[0112] Instance-Alignment Reasoning Structure: Compared to the relatively free "reasoning" step in CoT, under the constraints of the alignment principle, this introduces a clear "reasoning structure" definition stage. This step aims to first plan a reasoning framework based on the problem. This reasoning framework can be explicitly output in the first response, as shown in the first paragraph of the "thinking" section in the first response B. It is also implicitly performed within the model and not displayed in the first response. For example, this reasoning framework can be a systematic reasoning framework that meets the requirements of logical consistency and format. It enforces that reasoning must be orderly, planned, and verifiable, rather than arbitrary heuristic deduction.

[0113] Instance-Alignment Reasoning: Guided by the "reasoning structure" defined in the previous step, perform a specific, instanced reasoning process, as shown in the "Thinking" section of the first response B. This step strictly adheres to alignment principles, such as the alignment principle of logical consistency, ensuring that each step of the deduction is supported by STEM theory or logic, eliminating leaps and intuitive guesses, and maintaining language consistency.

[0114] Answer Generation: Finally, based on a rigorous reasoning process, a final answer is generated. This answer must strictly comply with the requirements of the answer-related alignment principles, for example, it must be an objectively verifiable single value and follow the format specification.

[0115] The above embodiment proposes and applies a systematic alignment principle to comprehensively standardize and optimize the quality (difficulty, diversity, logic, authenticity, format, etc.) of STEM reasoning content generated by LLM, introduces a "sample-aligned reasoning structure" and a "sample-aligned reasoning" process, and mandates that the reasoning process be planned, step-by-step, and logically rigorous. This goes beyond the free form of traditional CoT and innovatively integrates structured STEM subject knowledge systems (such as three-level categories) into the data generation process, achieving adaptive control of the depth and breadth of specific sub-fields. A strict multi-stage verification process is also set up to ensure data quality and logical reliability at each stage, ultimately producing high-quality sample-level reasoning data that is conducive to model learning.

[0116] In some embodiments, after obtaining the training data for the inference large model through the above embodiments, the parameters of the inference large model can be adjusted based on the training data to obtain the adjusted inference large model.

[0117] The training data includes sample questions and labeled answers. For example, the first question and the first answer after verification can be determined as the sample question and the labeled answer respectively as training data. This embodiment does not limit the training method used when adjusting the parameters.

[0118] In one embodiment, reinforcement learning can be used for training. When using reinforcement learning for training, sample questions can be input into the inference model to obtain the inference path and output answer. Based on the inference path, the answer is then output and labeled. A second reward model is used to determine a reward score, which is related to at least one of the following indicators: the accuracy of the output answer, the logic of the inference path, and the complexity of the inference path. Finally, the parameters of the inference model are adjusted based on the reward score.

[0119] For example, after inputting a sample question into the large inference model, the corresponding output answer and the reasoning path used to derive the output answer are obtained, and the labeled answer is used as the ground truth. When calculating the reward score, the accuracy of the output answer can be determined by evaluating the match or difference between the output answer generated by the large inference model and the labeled answer. The higher the accuracy, the higher the reward score. The logic of the reasoning path can be determined by judging whether the steps in the reasoning path are logical and whether the reasoning path is reasonable. The stronger the logic, the higher the reward score. The complexity of the reasoning path can be determined by the amount of information in the reasoning path, the number of steps, and the depth of reasoning. The simpler, more intuitive, less information-rich, or fewer steps the reasoning path is, the lower the complexity and the lower the reward score. The deeper the problem-solving depth of the reasoning path, the larger the amount of information, or the more steps, the higher the reward score. After obtaining the reward score, the parameters of the inference model are adjusted according to the reward score. In terms of parameter adjustment, the policy optimization method in reinforcement learning can be adopted, such as the method based on GRPO (Group Relative Policy Optimization) or PPO (Proximal Policy Optimization), to update the model parameters according to the feedback of the reward score.

[0120] This allows the model to learn based on a clear reward score related to the quality of complex reasoning (accuracy, logic, and complexity), inducing the model to continuously produce more accurate, logically rigorous, and complex reasoning. Training is completed when the reward value converges or the number of iterations reaches a preset number. This can more specifically improve the model's performance in reasoning about complex STEM problems than simple supervised learning.

[0121] In other embodiments, supervised learning methods may also be used for training. For example, sample questions are input into the inference model to obtain predicted answers output by the inference model. A loss function is calculated based on the difference between the predicted answers and the labeled answers. Then, the parameters of the inference model are adjusted with the goal of reducing the loss function. When the loss value reaches a preset threshold or the number of iterations reaches a preset number, the training is completed.

[0122] like Figure 4 As shown, in practice, when obtaining the first prompt word, the i-th prompt word template pi can be obtained. The prompt word template contains information such as alignment principles in addition to science and engineering discipline information, and then different science and engineering discipline information is obtained from the science and engineering discipline information database (such as the three-level label classification system mentioned above). According to the prompt word template and science and engineering discipline information, N corresponding first prompt words are constructed, and the first prompt words are input into the reasoning model for reasoning to generate the reasoning process. The first question and the corresponding first answer After verification by the verification module (GeneralVerifier), training data is obtained: including sample questions and proofread answers This training data is reused to train the large model for inference, e.g., reinforcement training or supervised training.

[0123] In this way, the large inference model obtained after adjustment with high-quality training data can generate more accurate and standardized answers to questions in the STEM field.

[0124] In some embodiments, as Figure 4 As shown, after obtaining the adjusted inference model, the adjusted inference model can be used to generate training data again and the process can be repeated. This means that in the next cycle, a more intelligent model will generate inference samples for the next round or subsequent model inference model training. In this way, the improvement of model capabilities directly promotes the generation of higher-quality training data, and higher-quality data further trains stronger models, forming a self-reinforcing virtuous cycle.

[0125] Specifically, the second prompt word can be input into the adjusted reasoning model to obtain a second response. The second prompt word includes the alignment principle and the science and engineering discipline information to which the question to be generated belongs. The second response includes the second question, the second answer, and the reasoning process used to derive the second question and the second answer. The second answer is used to answer the second question. According to the alignment principle, the second response is verified. If the verification passes, the second question and the second answer are determined to be new training data. For a detailed explanation of this process, please refer to Figure 2 The relevant description of the embodiment in will not be repeated here.

[0126] In one implementation, considering that the ability of the large inference model after parameter adjustment is enhanced compared to before the adjustment, it can cope with more difficult prompts and produce higher quality responses. The difficulty level of the problem to be generated indicated by the question difficulty information in the second prompt word can be higher than the difficulty level of the problem to be generated indicated by the question difficulty information in the first prompt word, so as to continuously generate training data with higher difficulty levels.

[0127] Because the closed-loop mechanism ensures that each iteration not only improves the model's average capabilities but also drives the model to explore and generate deeper and more complex reasoning paths, when regenerating training data in the next cycle, the second prompt word used can indicate a higher difficulty level than the first prompt word. In this way, the training data is no longer fixed, but instead adaptively adjusts its difficulty and complexity as the model's capabilities grow. This ensures that the training data remains challenging for the model and continuously drives its progress. Through continuous iterative optimization, the upper limit of the model's ability to handle complex cognitive tasks is constantly challenged and improved.

[0128] For example, by setting the difficulty level of the generated questions in the problem difficulty information to ensure that it reaches the Olympiad level / Master's and Doctoral level, the method in this embodiment can be applied to systematically and large-scale generate training data in the STEM field that meets the depth of Master's and Doctoral levels, the difficulty of Olympiad levels, the diversity of disciplines, rigorous logic, and accurate and verifiable answers. It is expected to greatly improve the performance of the model on high-difficulty reasoning benchmark tests such as GPQA. By using these high-quality data for training (especially in combination with closed-loop optimization), the reasoning accuracy, logic and problem-solving ability of large reasoning models on complex STEM problems can be significantly improved.

[0129] In this process, high-quality, self-aligned training data is used to cold-start large-scale reasoning models. The trained models are then used to generate new batches of more complex and higher-quality training data, and this cycle continues. This closed loop aims to continuously address the inherent shortcomings of conventional large-scale reasoning models in terms of problem complexity, logical rigor, and answer accuracy, continuously raising the upper limit of model intelligence (especially complex STEM reasoning capabilities), thereby achieving better performance on challenging benchmarks such as GPQA.

[0130] Figure 5 This is a schematic diagram of the structure of the apparatus for generating training data for inferring a large model in an embodiment of this specification. The apparatus can be applied to any device, platform, or device cluster with computing and processing capabilities. The apparatus includes:

[0131] Data acquisition module 501, used to obtain a first prompt word, the first prompt word including an alignment principle and information about the science and engineering discipline to which the problem to be generated belongs. The alignment principle is used to constrain the response results generated by the large inference model;

[0132] Model reasoning module 502, configured to input the first prompt word into the reasoning model to obtain a first response, the first response including a first question, a first answer, and a reasoning process for deriving the first question and the first answer, wherein the first answer is used to answer the first question;

[0133] The data verification module 503 is used to verify the first response according to the alignment principle. If the verification passes, the training data for inferring the large model is determined based on the first question and the first answer.

[0134] In some embodiments, the alignment principle includes at least one of the following: an alignment principle for indicating the modality of the response result; an alignment principle for indicating the language of the response result; an alignment principle for indicating the structure of the question to be generated; an alignment principle for indicating the authenticity of the question to be generated; an alignment principle for indicating the logic of the reasoning process; and an alignment principle for indicating the format of the question to be generated and the answer to be generated.

[0135] In some embodiments, the science and engineering discipline information is multi-level category information.

[0136] In some embodiments, the data verification module 503, when used to verify the first response according to the alignment principle, includes at least one of the following verification methods: evaluating the first response by the evaluation model according to the alignment principle; scoring the first response using the first reward model according to the alignment principle; obtaining the expert's review results on the first response according to the alignment principle.

[0137] In some embodiments, the first prompt word further includes question difficulty information, where the question difficulty information is used to indicate the difficulty level of the question to be generated.

[0138] In some embodiments, the device also includes a model training module (not shown in the figure) for adjusting the parameters of the inference big model based on the training data after determining the training data for the inference big model according to the first question and the first answer to obtain an adjusted inference big model.

[0139] In some embodiments, the training data includes sample questions and labeled answers. When the model training module is used to adjust the parameters of the inference model based on the training data, it is specifically used to: input the sample questions into the inference model to obtain the inference path and output the answer; output the answer and label the answer based on the inference path, and use the second reward model to determine the reward score. The reward score is related to at least one of the following indicators: the accuracy of the output answer, the logic of the reasoning path, and the complexity of the reasoning path; and adjust the parameters of the inference model based on the reward score.

[0140] In some embodiments, the device also includes an iterative generation module (not shown in the figure), which is used to input the second prompt word into the adjusted reasoning large model after obtaining the adjusted reasoning large model to obtain a second response. The second prompt word includes the alignment principle and the science and engineering discipline information to which the question to be generated belongs. The second response includes the second question, the second answer, and the reasoning process for deriving the second question and the second answer. The second answer is used to answer the second question; according to the alignment principle, the second response is verified. If the verification passes, the second question and the second answer are determined to be new training data.

[0141] In some embodiments, the difficulty level of the question to be generated indicated by the question difficulty information in the second prompt word is higher than the difficulty level of the question to be generated indicated by the question difficulty information in the first prompt word.

[0142] The embodiment of the present specification also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the following Figure 2 Described method.

[0143] The embodiment of this specification also provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the following is achieved: Figure 2 Described method.

[0144] The embodiments of this specification also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the following Figure 2 Describe the steps of the method.

[0145] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the various embodiments disclosed in this specification may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0146] In some cases, the actions or steps recited in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0147] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.

Claims

1. A method for generating training data for inferring a large model, the method comprising: Obtaining a first prompt word, the first prompt word including an alignment principle and information about the science and engineering discipline to which the problem to be generated belongs, the alignment principle being used to constrain a response result generated by the large inference model; Inputting the first prompt word into the reasoning model to obtain a first response, the first response including a first question, a first answer, and a reasoning process for deriving the first question and the first answer, the first answer being used to answer the first question; According to the alignment principle, the first response is verified. If the verification passes, the training data for the inference large model is determined based on the first question and the first answer.

2. The method according to claim 1, wherein The alignment principle includes at least one of the following: an alignment principle for indicating the modality of the response result; Alignment principles for the language used to indicate the response; an alignment principle for indicating the structure of the problem to be generated; an alignment principle for indicating the authenticity of the problem to be generated; Alignment principles used to indicate the logic of the reasoning process; The alignment principle is used to indicate the format of the question to be generated and the answer to be generated.

3. The method according to claim 1, wherein The science and engineering discipline information is multi-level category information.

4. The method according to claim 1, wherein Verifying the first response according to the alignment principle includes at least one of the following verification methods: According to the alignment principle, the evaluation model evaluates the first response; scoring the first response using a first reward model according to the alignment principle; Obtain an expert's review result of the first response according to the alignment principle.

5. The method according to claim 1, wherein The first prompt word further includes question difficulty information, where the question difficulty information is used to indicate the difficulty level of the question to be generated.

6. The method according to claim 1, wherein After determining that the first question and the first answer are training data for inferring a large model, the method further includes: Based on the training data, the parameters of the inference model are adjusted to obtain the adjusted inference model.

7. The method according to claim 6, wherein: The training data includes sample questions and labeled answers. Adjusting the parameters of the large inference model based on the training data includes: Input the sample question into the reasoning model to obtain a reasoning path and output an answer; Determining a reward score using a second reward model based on the reasoning path, the output answer, and the labeled answer, wherein the reward score is related to at least one of the following indicators: the accuracy of the output answer, the logic of the reasoning path, and the complexity of the reasoning path; According to the reward score, the parameters of the large inference model are adjusted.

8. The method according to claim 6, wherein: After obtaining the adjusted inference model, the method further includes: Inputting a second prompt word into the adjusted reasoning macromodel to obtain a second response, wherein the second prompt word includes the alignment principle and information about the science and engineering discipline to which the question to be generated belongs, and the second response includes a second question, a second answer, and a reasoning process for deriving the second question and the second answer, wherein the second answer is used to answer the second question; According to the alignment principle, the second response is verified, and if the verification passes, the second question and the second answer are determined to be new training data.

9. The method according to claim 8, wherein The difficulty level of the to-be-generated question indicated by the question difficulty information in the second prompt word is higher than the difficulty level of the to-be-generated question indicated by the question difficulty information in the first prompt word.

10. A device for generating training data for inferring a large model, the device comprising: a data acquisition module, configured to acquire a first prompt word, wherein the first prompt word includes an alignment principle and information about the science and engineering discipline to which the problem to be generated belongs, wherein the alignment principle is used to constrain a response result generated by the large inference model; a model reasoning module, configured to input the first prompt word into a large reasoning model to obtain a first response, the first response including a first question, a first answer, and a reasoning process for deriving the first question and the first answer, the first answer being used to answer the first question; A data verification module is used to verify the first response according to the alignment principle. If the verification passes, the first question and the first answer are determined to be training data for inferring the large model.

11. A computing device comprising a memory and a processor, wherein: The memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Model capability alignment method and device, electronic equipment and medium

    CN121457603A

  • Micro-course extraction method, related method and related device

    CN121982695A