Method and device for distillation reinforcement learning data set of intelligent computing center for providing computing power
By constructing reinforcement learning data sets in the intelligent computing center, generating candidate answers, answer ideas and scoring criteria, the diversity and insufficient evaluation of the large model training data set is solved, and the performance and development potential of the large model in complex tasks and real scenarios is improved.
Patent Information
- Application Number
- CN202510670806.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-15
AI Technical Summary
The lack of diversity and complexity of the training data sets of existing large models leads to their limitations in actual business scenarios with no standard answers and insufficient evaluation standards, limiting their performance and development potential in complex tasks and real application scenarios.
By using the big model to generate candidate answers, answer ideas, test points and scoring standards in the intelligent computing center, we build a reinforcement learning data set, simulate diversified problem-solving ideas in complex scenarios, and comprehensively evaluate the ability of the big model through a verification system of answer ideas-test points-scoring standards.
It enhances the performance and development potential of large models in complex tasks and real application scenarios, and improves its performance, especially in the ability to deal with no standard answers and multiple inference ideas.
Smart Images

Figure CN120492933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a method and device for distilling reinforcement learning data sets for intelligent computing centers that provide computing power. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.
[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] Currently, large models are widely used in the field of artificial intelligence. These models already possess a certain level of intelligence. By learning from massive amounts of data, they can generate coherent text, answer complex questions, and simulate different characters to complete command-following tasks with a certain degree of accuracy.
[0008] However, despite the impressive performance of large models in many application scenarios, the following technical challenges remain: First, existing large model training datasets often consist primarily of simple questions and answers, lacking diversity and complexity. In real life, many practical business scenarios lack standardized answers. This limits the models' performance in complex situations or specialized domains, hindering their ability to fully understand and process deeper information needs. Second, current evaluation criteria focus primarily on the model's accuracy and response speed, lacking a comprehensive assessment of its understanding, reasoning, and creativity. This one-sided approach hinders a comprehensive understanding of the true capabilities of large models and limits their development.
[0009] In summary, since the emergence of intelligent computing centers, many problems in actual business scenarios have not had standard answers. The lack of scoring criteria and corresponding training data sets for the answers output by large models has limited their performance and development potential in complex tasks and real-world application scenarios, and has restricted the performance improvement of large models. This problem has become a technical issue that needs to be solved urgently. Summary of the Invention
[0010] The present invention discloses a method and apparatus for distilling reinforcement learning datasets in an intelligent computing center that provides computing power, in order to solve the technical problem that since the emergence of intelligent computing centers, many problems in actual business scenarios have no standard answers, lack of scoring standards and corresponding training datasets for the answers output by large models, which limits their performance and development potential in complex tasks and real application scenarios, and restricts the performance improvement of large models.
[0011] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0012] In a first aspect, the present invention provides a method for distilling a reinforcement learning dataset in an intelligent computing center that provides computing power, the method comprising:
[0013] Step S1: obtaining an original data set, and generating at least one candidate answer corresponding to the original data set based on the original data set and a large model, wherein the large model is deployed on at least one computing node of an intelligent computing center;
[0014] Step S2: Based on the at least one candidate answer and the large model, generating answer ideas corresponding to each candidate answer;
[0015] Step S3: Based on the answering ideas and the large model, generate test points corresponding to the answering ideas;
[0016] Step S4: Based on the answer ideas, the test points and the large model, generating scoring criteria corresponding to the test points;
[0017] Step S5: Merge the original data set, the at least one candidate answer, the answer idea, the test points, and the scoring criteria to obtain a reinforcement learning data set.
[0018] Optionally, step S2 includes:
[0019] Step S21: screening the at least one candidate answer to obtain at least one screened candidate answer, wherein the screening includes at least one of the following: deduplication and quality filtering;
[0020] Step S22: Based on the at least one screened candidate answer and the large model, generate answer ideas corresponding to each candidate answer.
[0021] Optionally, the number of answer ideas is at least one, and step S3 includes:
[0022] Step S31: Processing at least one answer idea to obtain at least one processed answer idea, wherein the processing includes: deduplication processing and atomization processing;
[0023] Step S32: Based on the processed at least one answer idea and the large model, generate a test point corresponding to the answer idea.
[0024] Optionally, step S32 includes any one of the following:
[0025] Step S321: Based on the processed at least one answer idea and the large model, a group of test points corresponding to each answer idea is generated, wherein each group of test points is different;
[0026] Step S322: Based on the processed at least one answer idea and the large model, a group of common test points corresponding to each answer idea is generated.
[0027] Optionally, step S321 includes:
[0028] Step S3211: Based on the processed at least one answer idea and the large model, a set of common test points corresponding to each answer idea is generated;
[0029] Step S3212: Based on the set of common test points, the at least one answer idea and the large model, a set of test points corresponding to each answer idea is generated respectively.
[0030] Optionally, when step S32 includes step S321, each group of test points corresponds to a different scoring standard;
[0031] In the case where step S32 includes step S322, a group of common test points corresponds to the same set of scoring criteria.
[0032] Optionally, each group of test points may have different upper limits for the different scoring criteria.
[0033] Optionally, after step S5, the method further includes:
[0034] Step S6: performing reinforcement learning training on the model based on the reinforcement learning dataset, and scoring the reinforcement learning training of the model based on a preset scoring model;
[0035] In the reinforcement learning training, if the answer to be evaluated generated by the model includes at least one idea to be evaluated, and the idea to be evaluated is an idea that successfully matches the answer idea in the reinforcement learning dataset, then the scoring includes at least one of the following:
[0036] Scoring each of the at least one idea to be evaluated to obtain a scoring result corresponding to each idea to be evaluated, performing a weighted summation on the scoring results, and using the weighted summation result as the scoring result of the answer to be evaluated;
[0037] Each of the at least one idea to be evaluated is scored respectively to obtain a scoring result corresponding to each idea to be evaluated, and the scoring result with the highest score among the scoring results is used as the scoring result of the answer to be evaluated.
[0038] Optionally, after step S5, the method further includes:
[0039] Step S7: Based on the reinforcement learning dataset, a benchmark dataset is constructed.
[0040] Optionally, the reinforcement learning dataset includes:
[0041] The original data set, candidate answers, answer ideas, test points, scoring criteria, and standard answers determined from the candidate answers.
[0042] In a second aspect, the present invention provides a device for distilling a reinforcement learning dataset in an intelligent computing center that provides computing power, the device comprising:
[0043] An acquisition module, configured to execute step S1: acquire an original data set, and generate at least one candidate answer corresponding to the original data set based on the original data set and a large model, wherein the large model is deployed on at least one computing node of an intelligent computing center;
[0044] An execution module, configured to execute step S2: generating an answer idea corresponding to each candidate answer based on the at least one candidate answer and the large model;
[0045] Step S3: Based on the answering ideas and the large model, generate test points corresponding to the answering ideas;
[0046] Step S4: Based on the answer ideas, the test points and the large model, generating scoring criteria corresponding to the test points;
[0047] Step S5: Merge the original data set, the at least one candidate answer, the answer idea, the test points, and the scoring criteria to obtain a reinforcement learning data set.
[0048] In a third aspect, the present invention provides a server comprising: a processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the steps of a method for distilling a reinforcement learning data set in an intelligent computing center providing computing power as described in the first aspect above are implemented.
[0049] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a method for distilling a reinforcement learning data set by an intelligent computing center providing computing power as described in the first aspect above.
[0050] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of a method for distilling a reinforcement learning data set by an intelligent computing center providing computing power as described in the first aspect above.
[0051] In the present invention, by empowering the computing power resources of the intelligent computing center, a large model can be used to automatically distill a reinforcement learning data set containing a variety of candidate answers, answer ideas, test points (test ability points, knowledge points) and scoring standards. Specifically, starting from the questions in the original data set, not only the candidate answers are generated, but also the reasoning logic behind the candidate answers is disassembled, the answer ideas are generated, the test points involved are analyzed, and the scoring standards are automatically formulated to generate a reinforcement learning data set. This reinforcement learning data set can solve two major problems at the same time: on the one hand, by simulating a variety of problem-solving ideas in complex scenarios, the large model can learn how to deal with open questions without standard answers, or learn a variety of reasoning ideas for questions with standard answers; on the other hand, through the "answer ideas-test points-scoring standards" verification system, the large model's depth of understanding, logical reasoning ability, and knowledge application level are comprehensively evaluated, thereby effectively enhancing the performance and development potential of the large model in complex tasks and real application scenarios, and improving the performance of the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0053] Figure 1 A flowchart of a method for distilling reinforcement learning datasets in an intelligent computing center providing computing power disclosed in the present invention;
[0054] Figure 2 A flowchart of a method for distilling reinforcement learning datasets in an intelligent computing center providing computing power disclosed in the present invention;
[0055] Figure 3 A flowchart of a method for distilling reinforcement learning datasets in an intelligent computing center providing computing power disclosed in the present invention;
[0056] Figure 4 This is a structural block diagram of a device disclosed in the present invention that provides computing power for an intelligent computing center to distill reinforcement learning data sets;
[0057] Figure 5 Schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] First, the technical terms involved in the present invention are briefly explained below.
[0060] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0061] The "computing power" (Computational Power, CP) mentioned in the present invention refers to: it is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is about the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP general + CP intelligent + CP super
[0062] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0063] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.
[0064] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.
[0065] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0066] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.
[0067] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0068] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.
[0069] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.
[0070] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0071] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".
[0072] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0073] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0074] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0075] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.
[0076] The "computing power node" mentioned in the present invention refers to the computing resources of the server / container that can process computing tasks.
[0077] The "large model" mentioned in the present invention includes a "large language model".
[0078] The "large language model" mentioned in the present invention refers to a large-scale language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0079] The "reinforcement learning dataset" mentioned in the present invention refers to a data set used for training, evaluation or research of reinforcement learning (RL) algorithms.
[0080] Distillation, as used in this article, refers to distilling natural language reinforcement learning datasets. This involves using a large language model (such as GPT or Llama) as a teacher model to infer and evaluate raw data or candidate outputs, generating a dataset with reinforcement signals such as ratings or preferences. Such datasets are typically used in reinforcement learning or reward modeling to help student models learn better generation strategies.
[0081] A "benchmark dataset," as used in this article, refers to a dataset used to evaluate a particular aspect of a model's performance. It is typically used to evaluate and compare the performance of different algorithms and models on specific tasks. Benchmark datasets provide a consistent way to measure the effectiveness of different methods, enabling researchers and developers to conduct experiments and comparisons on a consistent basis.
[0082] Figure 1 A method for distilling a reinforcement learning dataset using an intelligent computing center that provides computing power is shown, the method comprising:
[0083] Step S1: Obtain an original data set, and generate at least one candidate answer corresponding to the original data set based on the original data set and the large model;
[0084] The large model is deployed on at least one computing node in the intelligent computing center;
[0085] Step S2: Based on at least one candidate answer and the large model, generate answer ideas corresponding to each candidate answer;
[0086] Step S3: Based on the answer ideas and the big model, generate test points corresponding to the answer ideas;
[0087] Step S4: Based on the answer ideas, test points and the big model, generate scoring criteria corresponding to the test points;
[0088] Step S5: Merge the original data set, at least one candidate answer, answer ideas, test points, and scoring criteria to obtain a reinforcement learning data set.
[0089] It should be noted that there are multiple original data sets, that is, multiple original data sets can be obtained and each original data set can be executed separately. Figure 1 The method shown in the figure is used to ensure the richness of the original dataset. In one possible implementation, the reinforcement learning dataset includes: the original dataset, candidate answers, answer ideas, test points, scoring criteria, and standard answers determined from the candidate answers.
[0090] Figure 1 The core of the illustrated method is that the dataset distillation service calls the large model inference service to generate a reinforcement learning dataset through steps S1 to S5. Subsequently, reinforcement learning training and fine-tuning of the large model can be performed based on the computing power of the intelligent computing center and the reinforcement learning dataset.
[0091] It should be noted that through the coordinated processing of the distributed computing resources of the intelligent computing center and the large model, a method for generating a reinforcement learning dataset from the original dataset can be constructed. Specifically, starting with the original dataset, the large model deployed on the computing nodes of the intelligent computing center generates multiple candidate answers, avoiding the data limitations of a single answer. Then, for each candidate answer, the logical chain of its generation process is reversed to form a detailed and parseable answer idea. On this basis, the core knowledge points or ability requirements involved in each answer idea are further extracted, namely the test points described in the present invention. Multi-dimensional scoring criteria are then generated based on the test points. Finally, the original data, candidate answers, answer ideas, test points, and scoring criteria are integrated to distill a reinforcement learning dataset containing multi-level information. Thus, through the "answer generation - idea decomposition - test point extraction - standard setting" processing method, the training data not only contains the corresponding relationship between questions and answers, but also deeply contains the logical reasoning path, test points, and scoring criteria, thereby significantly improving the large model's ability to analyze complex problems and its adaptability to multi-dimensional evaluation.
[0092] And it needs to be explained that, Figure 1The method shown is not only applicable to the generation of reinforcement learning datasets with open answers, but also to the generation of reinforcement learning datasets with standard answers. This is because for questions with standard answers, different answer ideas can also be constructed, and each answer idea will get a standard answer. Corresponding test points can be set for each answer idea.
[0093] Taking a specific application scenario as an example, the relationship between questions, test points, and scoring criteria is explained to assist understanding. Figure 1 The method shown.
[0094] 1. Assume that the original data set consists of 1,000 rental data analysis questions, and each question contains questions and data (including fields such as time, region, and price); one of the questions is "Analyze the quarterly changes in rental demand in a city", and the big model can generate multiple candidate answers, such as Candidate Answer 1: Use a line chart to show the monthly order volume changes; Candidate Answer 2: Use a heat map to show the quarterly rental order volume changes in each region; Candidate Answer 3: Use a line chart to show the correlation between quarterly housing prices and rental order volume; then, the big model can sort out the answer ideas from the candidate answers by clustering similar answers and extracting core ideas. For example, the answer ideas for Candidate Answer 2 can be sorted out as follows: multidimensional statistics and visualization ideas: quarterly-regional order volume statistics → two-dimensional heat map drawing → trend and difference interpretation; then, test points can be defined for the answer ideas, such as the test points corresponding to multidimensional statistics and visualization ideas are multidimensional data statistical ability, two-dimensional heat map drawing ability, and trend and difference analysis ability; and corresponding scoring standards can be formulated for the test points.
[0095] 2. When grading essays or discussion questions, teachers do not give scores based on their feelings. Instead, they formulate scoring rules in advance, breaking down abstract writing into specific test points, with each test point corresponding to a clear score.
[0096] For example, the scoring criteria for the essay "My Dream" (total score 30 points) can be set as follows:
[0097] (1) Degree of relevance to the topic (whether it revolves around "dreams") 0-10 points
[0098] (2) Structural clarity (introduction, main body stratification, and conclusion sublimation) 0-8 points
[0099] (3) Language vividness (use of metaphors, parallelism, and other rhetorical devices) 0-7 points
[0100] (4) Standard writing (no typos, correct punctuation) 0-5 points
[0101] In summary, a reinforcement learning dataset can be distilled from the original dataset, containing at least one candidate answer, answer strategy, test points, and scoring criteria. This enriches the diversity and complexity of the reinforcement learning dataset, allowing subsequent training of large models based on it to effectively enhance their performance and development potential in complex tasks and real-world application scenarios, ultimately improving their performance.
[0102] In one possible implementation, step S2 includes: step S21: screening at least one candidate answer to obtain at least one screened candidate answer, wherein the screening includes at least one of the following: deduplication and quality filtering; step S22: based on the at least one screened candidate answer and the large model, respectively generating answer ideas corresponding to each candidate answer in the at least one screened candidate answer.
[0103] It should be noted that, in the implementation process of step S2, the multiple candidate answers generated in step S1 are first pre-processed to improve the data quality. Specifically, the candidate answers with repeated content can be eliminated through the "de-duplication" operation, and then the logically confusing or obviously wrong answers can be eliminated through "quality filtering" to ensure that the candidate answers processed subsequently are diverse and reliable. Subsequently, based on the set of screened candidate answers, each candidate answer is reversely parsed using a large model to restore the implicit reasoning steps in its generation process, and finally a detailed answer idea corresponding to each candidate answer is formed. In this way, the rationality and richness of the generation basis (candidate answers) of the answer idea are guaranteed, and the answer idea can accurately reflect the differentiated thinking process of different problem-solving paths, providing a basis for the subsequent refinement of test points and the formulation of scoring standards.
[0104] In one possible implementation, the number of answer ideas is at least one, and step S3 includes: step S31: processing at least one answer idea to obtain at least one processed answer idea, wherein the processing includes: deduplication processing and atomization processing; step S32: based on the at least one processed answer idea and the large model, generating test points corresponding to the answer idea.
[0105] It should be noted that, in the implementation process of step S3, the multiple answer ideas generated in step S2 are first structured: duplicate or highly similar idea content is eliminated through "de-duplication processing", and then the complex answer ideas are disassembled into independent units that cannot be further divided through "atomization processing", ensuring that each processed answer idea has minimum integrity and resolvability. Subsequently, based on the set of processed answer ideas, the large model is used to summarize the knowledge points and abstract the ability points of each answer idea, and the test points supporting the idea are extracted (such as the key steps in answering mathematical problems, etc.), and finally a group of test points that strictly correspond to each answer idea are formed (a group of test points includes at least one test point). Therefore, on the basis of processing the answer ideas, test points are generated based on the answer ideas, which can improve the accuracy of test point generation.
[0106] In one possible implementation, step S32 includes any of the following: step S321: based on the processed at least one answer idea and the big model, generate a group of test points corresponding to each answer idea in the at least one answer idea, wherein each group of test points is different; step S322: based on the processed at least one answer idea and the big model, generate a group of general test points corresponding to each answer idea in the at least one answer idea.
[0107] It should be noted that in the specific implementation of step S32, two modes of generating test points can be flexibly selected according to actual needs: the first mode (corresponding to step S321) uses a large model to independently analyze each processed answer idea, and generates an exclusive set of test points for each answer idea. There are differences between the test point groups corresponding to different answer ideas (for example, for two answer ideas for the same mathematical problem, two types of test points are generated, namely "application of properties of geometric figures" and "algebraic equation conversion method"); the second mode (corresponding to step S322) uses a large model to extract the commonalities of all answer ideas, and generates a set of general test points applicable to all answer ideas (for example, general test points such as "correct answer format" and "includes key derivation steps").
[0108] These two modes are respectively applicable to differentiated assessment and standardized assessment scenarios: the former can accurately depict the personalized needs of different answer ideas through the independence of the test point group, while the latter can strengthen the model's unified grasp of core capabilities across scenarios through the universality of the test points. This flexible design allows the generated test points to meet the needs of professional fields for in-depth assessment of subdivided capabilities, while also adapting to the universal training of general thinking methods in basic fields (such as the ability to summarize the main points in reading comprehension), thereby significantly improving the adaptability of subsequent scoring standards. It can also more accurately simulate real-world application scenarios, effectively enhancing the performance and development potential of large models in complex tasks and real-world application scenarios, and improving the performance of large models.
[0109] In one possible implementation, Figure 2 As shown, step S321 includes:
[0110] Step S3211: Based on the processed at least one answer idea and the large model, a set of common test points corresponding to each answer idea is generated;
[0111] Step S3212: Based on a set of common test points, at least one answer idea and a large model, a set of test points corresponding to each answer idea is generated respectively.
[0112] It should be noted that in the implementation process of step S321, an optional implementation method is to first extract the core knowledge points or basic ability requirements that all answer ideas rely on based on the processed answer ideas through the large model, and generate a set of general test points applicable to all answer ideas (such as "formula transformation rules" in mathematical problems); then, based on this general test point group, combined with the specific logical path of each answer idea, further generate its own detailed test points for each answer idea (for example, "geometric auxiliary line construction skills" or "algebraic reasoning methods" for different problem-solving methods). This generation mechanism of general test points to exclusive test points not only ensures the consistency of different answer ideas in basic ability assessment (covering common requirements through general test points), but also supplements personalized test points according to the differences of specific answer ideas (refining special requirements through exclusive test points). As a result, the test points finally generated have both cross-method comparability (general part) and can accurately reflect the unique value of different ideas (exclusive part), thereby significantly improving the compatibility of diversified problem-solving strategies in large model training and the flexibility of the evaluation system. First, general test points are generated, and then based on this, exclusive test points corresponding to each answer idea are generated. The exclusive test points can be checked based on the requirements of the general test points to avoid omissions of test points and ensure the comprehensiveness and accuracy of the test point generation.
[0113] Alternatively, step S321 can be directly executed, i.e., instead of using the general test points, a dedicated test point corresponding to each answer idea can be directly generated based on the processed at least one answer idea and the large model. The test point generation method can be flexibly selected according to actual needs.
[0114] In a possible implementation, when step S32 includes step S321, each group of test points corresponds to a different scoring standard; when step S32 includes step S322, a group of common test points corresponds to the same set of scoring standards.
[0115] It should be noted that in the implementation of step S32, if step S321 is selected (generating an independent test point group for each answer idea), the test point group corresponding to each answer idea will be associated with an exclusive scoring standard (for example, for the "Application of Properties of Geometric Figures" test point group, a scoring rule with a higher weight for graphic analysis can be set, while the "Algebraic Equation Conversion Skills" test point group can focus more on the scoring of the accuracy of formula transformation); if step S322 is selected (generating a general test point group), all answer ideas share the same set of scoring standards (for example, the scoring weights are uniformly formulated according to general test points such as "correct answer format" and "including key derivation steps").
[0116] This differentiated scoring criteria matching mechanism allows for precise quantification of the strengths of different answer strategies through independent scoring criteria in scenarios requiring in-depth assessment of individual capabilities. In scenarios requiring a unified assessment of capabilities, a universal scoring standard ensures consistency in evaluation results. This allows the large model to accommodate both the refined evaluation needs of specialized fields for differentiated capabilities and the standard capabilities required in general scenarios, thereby enhancing the comprehensiveness and adaptability of the evaluation system.
[0117] In a possible implementation, each group of test points has different upper limits for the different scoring standards.
[0118] It is understandable that the upper limit of the scoring criteria for each group of test points can be the same, for example, all are based on a percentage system, and then the total score is calculated based on the weight of each test point; the upper limit of the scoring criteria for each group of test points can also be different, for example: the upper limit of the scoring for the test point corresponding to answer idea A is 60 points, and the upper limit of the scoring for the test point corresponding to answer idea B is 100 points. For example, for the "basic formula application" test point group, the upper limit of the scoring criteria may be set to 60 points (focusing on the achievement of basic capabilities), while the upper limit of the "high-level logical deduction" test point group can be set to 100 points (encouraging breakthroughs in deep innovation capabilities). In this way, differentiated scoring criteria can be formulated for different test point groups themselves, accurately guiding the training direction of the large model between the stability of basic capabilities and the breakthrough of high-level innovation capabilities, and improving the comprehensive performance of the large model.
[0119] In one possible implementation, Figure 3 As shown, after step S5, the method further includes:
[0120] Step S6: Perform reinforcement learning training on the model based on the reinforcement learning dataset, and score the reinforcement learning training of the model based on a preset scoring model.
[0121] It should be noted that the answers generated by the model can be checked using a preset scoring model, and scores are given for different test points. Finally, the final score of the answer is calculated using a preset scoring formula. This scoring method can be encapsulated as the accuracy_reward function of reinforcement learning to perform reinforcement learning on the model.
[0122] Among them, in reinforcement learning training, if the answer to be evaluated generated by the model includes at least one idea to be evaluated, and the idea to be evaluated is an idea that successfully matches the answer idea in the reinforcement learning data set, then the scoring includes at least one of the following: scoring each idea to be evaluated in the at least one idea to be evaluated separately, obtaining the scoring results corresponding to each idea to be evaluated, and performing weighted summation on the scoring results, and using the weighted summation result as the scoring result of the answer to be evaluated; scoring each idea to be evaluated in the at least one idea to be evaluated separately, obtaining the scoring results corresponding to each idea to be evaluated, and using the scoring result with the highest score as the scoring result of the answer to be evaluated.
[0123] It should be noted that during the scoring process of reinforcement learning training, when the answer to be evaluated generated by the model contains multiple ideas to be evaluated, the following scoring method can be used: If a weighted summation mechanism is used, the scoring results of each idea to be evaluated (e.g., idea A gets 80 points, idea B gets 60 points) are calculated according to the preset weights (e.g., idea A weights 70%, idea B weights 30%) to calculate a comprehensive score (80×0.7+60×0.3=74 points); if a highest score selection mechanism is used, the highest score among all ideas to be evaluated (e.g., idea C gets 90 points) is directly selected as the final score. In this way, large model training can both encourage the exploration of multiple ideas (preventing over-reliance on a single idea through a weighted mechanism) and maintain core capability focus (strengthening key advantages through a selection mechanism), thereby further ensuring the performance of large models.
[0124] In a possible implementation, after step S5, the method further includes:
[0125] Step S7: Build a benchmark dataset based on the reinforcement learning dataset.
[0126] This can provide an evaluation benchmark that includes multi-dimensional capability indicators for the capability evaluation of large models, making the evaluation system of large models more complete, and allowing for more comprehensive consideration and training of the understanding, reasoning, and creativity of large models, thereby improving the performance of large models.
[0127] In summary, in the present invention, by empowering the computing power resources of the intelligent computing center, a reinforcement learning data set containing a variety of candidate answers, answer ideas, test points (test ability points, knowledge points) and scoring standards can be automatically generated using a large model. Specifically, starting from the questions in the original data set, not only the candidate answers are generated, but also the reasoning logic behind the candidate answers is disassembled, the answer ideas are generated, the test points involved are analyzed, and the scoring standards are automatically formulated to generate a reinforcement learning data set. This reinforcement learning data set can solve two major problems at the same time: on the one hand, by simulating a variety of problem-solving ideas in complex scenarios, the large model can learn how to deal with open questions without standard answers, or learn a variety of reasoning ideas for questions with standard answers; on the other hand, through the verification system of "answer ideas-test points-scoring standards", the large model's depth of understanding, logical reasoning ability and knowledge application level are comprehensively evaluated. Thereby effectively enhancing the performance and development potential of the large model in complex tasks and real application scenarios.
[0128] In addition, building a benchmark dataset based on the reinforcement learning dataset and then performing reinforcement learning training on the large model to be evaluated, or fine-tuning the large model based on the reinforcement learning dataset, can make the evaluation system of the large model more complete, and can more comprehensively consider and train the large model's understanding ability, reasoning ability and creativity, thereby improving the performance of the large model.
[0129] Figure 4 A device for providing computing power to an intelligent computing center to distill reinforcement learning data sets is shown, such as Figure 4 As shown, the device 40 includes:
[0130] Acquisition module 401 is configured to execute step S1: acquire an original data set, and generate at least one candidate answer corresponding to the original data set based on the original data set and a large model, where the large model is deployed on at least one computing node in an intelligent computing center;
[0131] The execution module 402 is configured to execute step S2: generating a corresponding answer idea for each candidate answer based on at least one candidate answer and the large model;
[0132] Step S3: Based on the answer ideas and the big model, generate test points corresponding to the answer ideas;
[0133] Step S4: Based on the answer ideas, test points and the big model, generate scoring criteria corresponding to the test points;
[0134] Step S5: Merge the original data set, at least one candidate answer, answer ideas, test points, and scoring criteria to obtain a reinforcement learning data set.
[0135] In a possible implementation, step S2 includes:
[0136] Step S21: screening at least one candidate answer to obtain at least one screened candidate answer, wherein the screening includes at least one of the following: deduplication and quality filtering;
[0137] Step S22: Based on the at least one screened candidate answer and the large model, generate answer ideas corresponding to each candidate answer in the at least one screened candidate answer.
[0138] In a possible implementation, the number of answer ideas is at least one, and step S3 includes:
[0139] Step S31: Processing at least one answer idea to obtain at least one processed answer idea, wherein the processing includes: deduplication processing and atomization processing;
[0140] Step S32: Based on the processed at least one answer idea and the large model, generate a test point corresponding to the answer idea.
[0141] In a possible implementation, step S32 includes any one of the following:
[0142] Step S321: Based on the processed at least one answer idea and the large model, a group of test points corresponding to each answer idea in the at least one answer idea is generated, wherein each group of test points is different;
[0143] Step S322: Based on the processed at least one answer idea and the large model, generate a set of common test points corresponding to each answer idea in the at least one answer idea.
[0144] In a possible implementation, step S321 includes:
[0145] Step S3211: Based on the processed at least one answer idea and the large model, generate a set of common test points corresponding to each answer idea in the at least one answer idea;
[0146] Step S3212: Based on a set of common test points, at least one answer idea and the large model, a set of test points corresponding to each answer idea in the at least one answer idea is generated respectively.
[0147] In a possible implementation, when step S32 includes step S321, each group of test points corresponds to a different scoring standard;
[0148] In the case where step S32 includes step S322 , a group of common test points corresponds to the same set of scoring criteria.
[0149] In a possible implementation, each group of test points has different upper limits for the different scoring standards.
[0150] In one possible implementation, the execution module 402 is further configured to, after step S5, execute step S6: perform reinforcement learning training on the model based on the reinforcement learning dataset, and score the reinforcement learning training of the model based on a preset scoring model;
[0151] In reinforcement learning training, if the answer to be evaluated generated by the model includes at least one idea to be evaluated, and the idea to be evaluated is an idea that successfully matches the answer idea in the reinforcement learning dataset, then the scoring includes at least one of the following:
[0152] Scoring each of the at least one idea to be evaluated, obtaining a scoring result corresponding to each idea to be evaluated, performing a weighted summation on the scoring results, and using the weighted summation result as the scoring result of the answer to be evaluated;
[0153] Each of the at least one idea to be evaluated is scored respectively to obtain a scoring result corresponding to each idea to be evaluated, and the scoring result with the highest score among the scoring results is used as the scoring result of the answer to be evaluated.
[0154] In a possible implementation, the execution module 402 is further configured to, after step S5, execute step S7: construct a benchmark dataset based on the reinforcement learning dataset.
[0155] In one possible implementation, the reinforcement learning dataset includes:
[0156] The original data set, candidate answers, answer ideas, test points, scoring criteria, and standard answers determined from the candidate answers.
[0157] In summary, in the present invention, by empowering the computing power resources of the intelligent computing center, a reinforcement learning data set containing a variety of candidate answers, answer ideas, test points (test ability points, knowledge points) and scoring standards can be automatically generated using a large model. Specifically, starting from the questions in the original data set, not only the candidate answers are generated, but also the reasoning logic behind the candidate answers is disassembled, the answer ideas are generated, the test points involved are analyzed, and the scoring standards are automatically formulated to generate a reinforcement learning data set. This reinforcement learning data set can solve two major problems at the same time: on the one hand, by simulating a variety of problem-solving ideas in complex scenarios, the large model can learn how to deal with open questions without standard answers, or learn a variety of reasoning ideas for questions with standard answers; on the other hand, through the verification system of "answer ideas-test points-scoring standards", the large model's depth of understanding, logical reasoning ability and knowledge application level are comprehensively evaluated. Thereby effectively enhancing the performance and development potential of the large model in complex tasks and real application scenarios, and improving the performance of the large model.
[0158] Please refer to Figure 5 The present invention also provides an electronic device 50, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the computer program is executed by the processor 501, the steps of the method for providing computing power for an intelligent computing center to distill a reinforcement learning data set are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.
[0159] The present invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of the above-mentioned method for providing computing power for an intelligent computing center to distill a reinforcement learning dataset, and can achieve the same technical effect. To avoid repetition, the description is omitted here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0160] The present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-mentioned method of providing computing power for an intelligent computing center to distill a reinforcement learning data set, and can achieve the same technical effect. To avoid repetition, they will not be described here.
[0161] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method described in the present invention.
[0163] The present invention is described above with reference to the accompanying drawings, but the present invention is not limited to the above-mentioned specific embodiments. The above-mentioned specific embodiments are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for distilling reinforcement learning datasets in an intelligent computing center that provides computing power, characterized in that: The method comprises: Step S1: obtaining an original data set, and generating at least one candidate answer corresponding to the original data set based on the original data set and a large model, wherein the large model is deployed on at least one computing node of an intelligent computing center; Step S2: Based on the at least one candidate answer and the large model, generating answer ideas corresponding to each candidate answer; Step S3: Based on the answering ideas and the large model, generate test points corresponding to the answering ideas; Step S4: Based on the answer ideas, the test points and the large model, generating scoring criteria corresponding to the test points; Step S5: Merge the original data set, the at least one candidate answer, the answer idea, the test points, and the scoring criteria to obtain a reinforcement learning data set.
2. The method according to claim 1, characterized in that The step S2 comprises: Step S21: screening the at least one candidate answer to obtain at least one screened candidate answer, wherein the screening includes at least one of the following: deduplication and quality filtering; Step S22: Based on the at least one screened candidate answer and the large model, generate answer ideas corresponding to each candidate answer.
3. The method according to claim 1, characterized in that The number of answer ideas is at least one, and step S3 includes: Step S31: Processing at least one answer idea to obtain at least one processed answer idea, wherein the processing includes: deduplication processing and atomization processing; Step S32: Based on the processed at least one answer idea and the large model, generate a test point corresponding to the answer idea.
4. The method according to claim 3, characterized in that The step S32 includes any one of the following: Step S321: Based on the processed at least one answer idea and the large model, a group of test points corresponding to each answer idea is generated, wherein each group of test points is different; Step S322: Based on the processed at least one answer idea and the large model, a group of common test points corresponding to each answer idea is generated.
5. The method according to claim 4, characterized in that The step S321 includes: Step S3211: Based on the processed at least one answer idea and the large model, a set of common test points corresponding to each answer idea is generated; Step S3212: Based on the set of common test points, the at least one answer idea and the large model, a set of test points corresponding to each answer idea is generated respectively.
6. The method according to claim 4, characterized in that In the case where step S32 includes step S321, each group of test points corresponds to a different scoring standard; In the case where step S32 includes step S322, a group of common test points corresponds to the same set of scoring criteria.
7. The method according to claim 6, characterized in that Each group of test points has different corresponding scoring standards and different upper limits for scoring.
8. The method according to claim 1, characterized in that After step S5, the method further includes: Step S6: performing reinforcement learning training on the model based on the reinforcement learning dataset, and scoring the reinforcement learning training of the model based on a preset scoring model; In the reinforcement learning training, if the answer to be evaluated generated by the model includes at least one idea to be evaluated, and the idea to be evaluated is an idea that successfully matches the answer idea in the reinforcement learning dataset, then the scoring includes at least one of the following: Scoring each of the at least one idea to be evaluated to obtain a scoring result corresponding to each idea to be evaluated, performing a weighted summation on the scoring results, and using the weighted summation result as the scoring result of the answer to be evaluated; Each of the at least one idea to be evaluated is scored respectively to obtain a scoring result corresponding to each idea to be evaluated, and the scoring result with the highest score among the scoring results is used as the scoring result of the answer to be evaluated.
9. The method according to claim 1, characterized in that After step S5, the method further includes: Step S7: constructing a benchmark dataset based on the reinforcement learning dataset.
10. The method according to any one of claims 1 to 9, characterized in that The reinforcement learning dataset includes: The original data set, candidate answers, answer ideas, test points, scoring criteria, and standard answers determined from the candidate answers.
11. A device for distilling reinforcement learning datasets in an intelligent computing center that provides computing power, characterized in that: The device comprises: An acquisition module, configured to execute step S1: acquire an original data set, and generate at least one candidate answer corresponding to the original data set based on the original data set and a large model, wherein the large model is deployed on at least one computing node of an intelligent computing center; An execution module, configured to execute step S2: generating an answer idea corresponding to each candidate answer based on the at least one candidate answer and the large model; Step S3: Based on the answering ideas and the large model, generate test points corresponding to the answering ideas; Step S4: Based on the answer ideas, the test points and the large model, generating scoring criteria corresponding to the test points; Step S5: Merge the original data set, the at least one candidate answer, the answer idea, the test points, and the scoring criteria to obtain a reinforcement learning data set.
12. A server, characterized in that: include: A processor, a memory, and a program stored on the memory and executable on the processor, wherein when the program is executed by the processor, the program implements the steps of a method for distilling a reinforcement learning data set in an intelligent computing center providing computing power as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for distilling a reinforcement learning data set in an intelligent computing center providing computing power according to any one of claims 1 to 10.
14. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of a method for distilling a reinforcement learning data set in an intelligent computing center providing computing power as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Method and device for pushing learning resources
CN105787839A
A mathematical problem correction method and device
CN109886851A
Text processing method and device based on machine learning, computer equipment and medium
CN111428021A
Multi-paragraph reading comprehension candidate answer sorting method and device
CN111460089A
Automatic homework correcting method and device, electronic equipment and storage medium
CN111753767A
Cited By
Method and device for realizing automatic construction of reward model through computing power by intelligent computing cloud platform
CN121388002A
Method and device for automatically constructing a reward model by computing power through an intelligent computing cloud platform
CN121388002B
Method and device for realizing closed-loop fine-tuning training of large model by intelligent computing cloud platform through computing power
CN121388607A