Method and apparatus for distilling reinforcement learning dataset by intelligent computing center providing computing power

CN120492933BActive Publication Date: 2026-09-25DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510670806.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2026-09-25
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

[0010]本发明公开了提供算力的智能计算中心蒸馏强化学习数据集的方法及装置,以解决自智能计算中心出现以来,很多实际业务场景的问题并没有标准答案,缺乏对大模型输出答案的评分标准和对应的训练数据集,限制了其在复杂任务和真实应用场景中的表现和发展潜力,制约了大模型的性能提升的技术问题

Benefits of technology

[0051]本发明中,通过加持智能计算中心的算力资源,利用大模型能够自动蒸馏包含多种候选答案、回答思路、考点(考察能力点、知识要点)及评分标准的强化学习数据集。具体来说,从原始数据集的问题出发,不仅生成候选答案,还会拆解出候选答案背后的推理逻辑,生成回答思路,分析其涉及的考点,并自动制定评分标准,进而生成强化学习数据集。这种强化学习数据集能够同时解决两大问题:一方面,通过模拟复杂场景中多样化的解决问题的思路,让大模型学习如何处理无标准答案的开放性问题,或者让大模型学习到有标准答案的问题的多种推理思路;另一方面,通过“回答思路-考点-评分标准”的验证体系,全面评估大模型的理解深度、逻辑推理能力以及知识应用水平,从而有效增强大模型在复杂任务和真实应用场景中的表现和发展潜力,提高大模型的性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492933B_ABST
    Figure CN120492933B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent computing center, wisdom computing center and computing power infrastructure, and discloses a method and device for providing computing power of an intelligent computing center distillation reinforcement learning data set, comprising: obtaining an original data set, generating at least one candidate answer corresponding to the original data set based on the original data set and a large model, and then generating a respective answer idea corresponding to each candidate answer; generating a test point corresponding to the answer idea based on the answer idea and the large model; generating a scoring standard corresponding to the test point based on the answer idea, the test point and the large model; and merging the original data set, the at least one candidate answer, the answer idea, the test point and the scoring standard to obtain a reinforcement learning data set. Thus, the reinforcement learning data set is more diverse and complex, and the large model can be trained according to the reinforcement learning data set subsequently, thereby effectively enhancing the performance and development potential of the large model in complex tasks and real application scenarios, and improving the performance of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent computing centers, smart computing centers, and computing power infrastructure, specifically to a method and apparatus for distilling reinforcement learning datasets in intelligent computing centers that provide computing power. Background Technology

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.

[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".

[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform a certain computing requirement. It is the computing power to achieve the target output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Currently, large-scale models have a wide range of applications in the field of artificial intelligence. These models already possess a certain level of intelligence; through learning from massive amounts of data, they can generate coherent text, answer complex questions, simulate different roles to complete instruction-following tasks, and exhibit a certain degree of accuracy.

[0008] However, despite the good results achieved by large-scale models in many application scenarios, the following technical problems still exist: First, the training datasets of existing large-scale models often consist mainly of simple questions and answers, lacking diversity and complexity. In real life, many practical business scenarios do not have standard answers, which limits the models when faced with complex situations or tasks in specialized domains, making it difficult for them to fully understand and process deeper information needs. Second, current evaluation criteria focus primarily on the model's accuracy and response speed, lacking a comprehensive consideration of its comprehension, reasoning, and creativity. This one-sided evaluation method makes it difficult to fully understand the true capabilities of large-scale models, thus restricting their development.

[0009] In summary, since the emergence of intelligent computing centers, many problems in real-world business scenarios have lacked standard answers. The lack of scoring criteria for the output answers of large models and corresponding training datasets has limited their performance and development potential in complex tasks and real-world application scenarios, and has constrained the performance improvement of large models. This problem has become a technical issue that urgently needs to be solved. Summary of the Invention

[0010] This invention discloses a method and apparatus for providing intelligent computing centers to distill reinforcement learning datasets, in order to address the technical problem that since the emergence of intelligent computing centers, many practical business scenarios have lacked standard answers, and there is a lack of scoring criteria and corresponding training datasets for the output answers of large models, which limits their performance and development potential in complex tasks and real-world application scenarios, and restricts the performance improvement of large models.

[0011] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0012] In a first aspect, the present invention provides a method for providing computing power to distill reinforcement learning datasets in an intelligent computing center, the method comprising:

[0013] Step S1: Obtain the original dataset, and based on the original dataset and the large model, generate at least one candidate answer corresponding to the original dataset. The large model is deployed on at least one computing node in the intelligent computing center.

[0014] Step S2: Based on the at least one candidate answer and the large model, generate a response strategy corresponding to each candidate answer;

[0015] Step S3: Based on the answer approach and the large model, generate the test points corresponding to the answer approach;

[0016] Step S4: Based on the answer approach, the test points, and the large model, generate the scoring criteria corresponding to the test points;

[0017] Step S5: Merge the original dataset, the at least one candidate answer, the answering strategy, the test points, and the scoring criteria to obtain the reinforcement learning dataset.

[0018] Optionally, step S2 includes:

[0019] Step S21: Filter the at least one candidate answer to obtain at least one filtered candidate answer, wherein the filtering includes at least one of the following: deduplication and quality filtering;

[0020] Step S22: Based on at least one candidate answer after screening and the large model, generate a response strategy corresponding to each candidate answer.

[0021] Optionally, the number of answer approaches is at least one, and step S3 includes:

[0022] Step S31: Process at least one answer approach to obtain at least one processed answer approach, wherein the processing includes: deduplication and atomization.

[0023] Step S32: Based on at least one processed answer approach and the large model, generate test points corresponding to the answer approach.

[0024] Optionally, step S32 includes any of the following:

[0025] Step S321: Based on at least one processed answer approach and the large model, generate a set of test points corresponding to each answer approach, wherein each set of test points is different;

[0026] Step S322: Based on at least one processed answer approach and the large model, generate a set of general test points corresponding to each answer approach.

[0027] Optionally, step S321 includes:

[0028] Step S3211: Based on at least one processed answer approach and the large model, generate a set of general test points corresponding to each answer approach;

[0029] Step S3212: Based on the set of general test points, the at least one answer approach, and the large model, generate a set of test points corresponding to each answer approach.

[0030] Optionally, if step S32 includes step S321, each set of test points corresponds to a different scoring standard.

[0031] When step S32 includes step S322, a set of general test points corresponds to the same set of scoring criteria.

[0032] Optionally, the upper limit of the score may vary depending on the different scoring criteria corresponding to each set of test points.

[0033] Optionally, after step S5, the method further includes:

[0034] Step S6: Train the model using reinforcement learning based on the reinforcement learning dataset, and score the reinforcement learning training of the model based on a preset scoring model;

[0035] Wherein, in the reinforcement learning training, if the answer to be evaluated generated by the model includes at least one approach to be evaluated, and the approach to be evaluated is an approach that successfully matches the answer approach in the reinforcement learning dataset, then the scoring includes at least one of the following:

[0036] Each of the at least one approaches to be evaluated is scored to obtain a score result for each approach to be evaluated. The scores are then weighted and summed, and the weighted sum is used as the score result for the answer to be evaluated.

[0037] Each of the at least one approaches to be evaluated is scored to obtain a score result for each approach, and the highest score among the scores is taken as the score result of the answer to be evaluated.

[0038] Optionally, after step S5, the method further includes:

[0039] Step S7: Construct a benchmark dataset based on the reinforcement learning dataset.

[0040] Optionally, the reinforcement learning dataset includes:

[0041] The original dataset, candidate answers, answer strategies, test points, scoring criteria, and the standard answer determined from the candidate answers.

[0042] Secondly, the present invention provides an apparatus for providing computing power for a centrally-controlled intelligent computing system to distill reinforcement learning datasets, the apparatus comprising:

[0043] The acquisition module is used to perform step S1: acquire the original dataset, and generate at least one candidate answer corresponding to the original dataset based on the original dataset and the large model, wherein the large model is deployed on at least one computing node in the intelligent computing center;

[0044] The execution module is used to execute step S2: based on the at least one candidate answer and the large model, generate a response strategy corresponding to each candidate answer;

[0045] Step S3: Based on the answer approach and the large model, generate the test points corresponding to the answer approach;

[0046] Step S4: Based on the answer approach, the test points, and the large model, generate the scoring criteria corresponding to the test points;

[0047] Step S5: Merge the original dataset, the at least one candidate answer, the answering strategy, the test points, and the scoring criteria to obtain the reinforcement learning dataset.

[0048] Thirdly, the present invention provides a server comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for providing computing power in an intelligent computing center for distilling reinforcement learning datasets as described in the first aspect above.

[0049] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for providing computing power in an intelligent computing center for distilling reinforcement learning datasets as described in the first aspect above.

[0050] Fifthly, the present invention provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of a method for providing computing power in an intelligent computing center for distilling reinforcement learning datasets as described in the first aspect above.

[0051] In this invention, by leveraging the computing power of an intelligent computing center, a large-scale model can automatically distill a reinforcement learning dataset containing multiple candidate answers, answer strategies, test points (assessing abilities and knowledge points), and scoring criteria. Specifically, starting from the questions in the original dataset, it not only generates candidate answers but also deconstructs the reasoning logic behind the candidate answers, generates answer strategies, analyzes the test points involved, and automatically formulates scoring criteria, thereby generating a reinforcement learning dataset. This reinforcement learning dataset can simultaneously solve two major problems: on the one hand, by simulating diverse problem-solving approaches in complex scenarios, it allows the large-scale model to learn how to handle open-ended questions without standard answers, or to learn multiple reasoning approaches for questions with standard answers; on the other hand, through a verification system of "answer strategies - test points - scoring criteria," it comprehensively evaluates the large-scale model's depth of understanding, logical reasoning ability, and knowledge application level, thereby effectively enhancing the performance and development potential of the large-scale model in complex tasks and real-world application scenarios, and improving the performance of the large-scale model. Attached Figure Description

[0052] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0053] Figure 1 This is a flowchart of a method for providing computing power to an intelligent computing center for distilling reinforcement learning datasets, as disclosed in this invention.

[0054] Figure 2 This is a flowchart of a method for providing computing power to an intelligent computing center for distilling reinforcement learning datasets, as disclosed in this invention.

[0055] Figure 3 This is a flowchart of a method for providing computing power to an intelligent computing center for distilling reinforcement learning datasets, as disclosed in this invention.

[0056] Figure 4 This is a structural block diagram of an apparatus for providing computing power to a smart computing center for distilling reinforcement learning datasets, as disclosed in this invention.

[0057] Figure 5 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation

[0058] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] The technical terms involved in this invention will be briefly explained below.

[0060] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.

[0061] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values ​​indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = General CP + Intelligent CP + Super CP

[0062] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.

[0063] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices in servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), the commonly used unit of measurement for performance is the number of read / write operations per second per unit capacity (IOPS / TB), and the disaster recovery ratio is an important indicator of security and reliability.

[0064] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.

[0065] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0066] The “computing power” mentioned in this invention includes: general computing power, intelligent computing power, and supercomputing power.

[0067] The "general computing power" mentioned in this invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0068] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovative applications, such as natural language processing and machine vision.

[0069] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0070] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0071] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".

[0072] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0073] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0074] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.

[0075] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.

[0076] The "computing node" mentioned in this invention refers to the computing resources of a server / container capable of processing computing tasks.

[0077] The "large model" mentioned in this invention includes a "large language model".

[0078] The "large language model" mentioned in this invention refers to a large-scale language model (LLM), which is a language model with a large number of parameters. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks, including text summarization, translation, and sentiment analysis.

[0079] The "reinforcement learning dataset" mentioned in this invention refers to a collection of data used for training, evaluating, or researching reinforcement learning (RL) algorithms.

[0080] The "distillation" mentioned in this invention refers to the distillation of natural language reinforcement learning datasets. This involves using large language models (such as GPT, Llama, etc.) as teacher models to infer and evaluate raw data or candidate outputs, generating a dataset with reinforcement signals such as scores or preferences. These datasets are typically used for reinforcement learning or reward modeling to help student models learn better generation strategies.

[0081] The "benchmark dataset" mentioned in this invention refers to a collection of data used to evaluate the performance level of a model in a certain aspect, typically used to assess and compare the performance of different algorithms and models on a specific task. Benchmark datasets provide a consistent method to measure the effectiveness of different methods, enabling researchers and developers to conduct experiments and comparisons on a level playing field.

[0082] Figure 1 This paper presents a method for providing computing power to distill reinforcement learning datasets in an intelligent computing hub, the method comprising:

[0083] Step S1: Obtain the original dataset, and based on the original dataset and the large model, generate at least one candidate answer corresponding to the original dataset;

[0084] The large model is deployed on at least one computing node in the intelligent computing center;

[0085] Step S2: Based on at least one candidate answer and the large model, generate the corresponding answer strategy for each candidate answer;

[0086] Step S3: Based on the answer approach and the overall model, generate the test points corresponding to the answer approach;

[0087] Step S4: Based on the answer approach, key points, and overall model, generate scoring criteria corresponding to the key points;

[0088] Step S5: Merge the original dataset, at least one candidate answer, answer strategy, test points, and scoring criteria to obtain the reinforcement learning dataset.

[0089] It should be noted that the original dataset consists of multiple sets, meaning multiple sets of original datasets can be obtained, and the process can be performed on each set of original datasets separately. Figure 1 The method shown ensures the richness of the original dataset. In one possible implementation, the reinforcement learning dataset includes: the original dataset, candidate answers, answer strategies, test points, scoring criteria, and the standard answer determined from the candidate answers.

[0090] Figure 1 The core of the method shown is as follows: the dataset distillation service calls the large model inference service to generate a reinforcement learning dataset through steps S1 to S5. Furthermore, the large model can be trained and fine-tuned using the computing power of the intelligent computing center and the reinforcement learning dataset.

[0091] It should be noted that by leveraging the distributed computing resources of the intelligent computing center and the collaborative processing of large models, a method for generating reinforcement learning datasets from raw datasets can be constructed. Specifically, starting with the raw dataset, multiple candidate answers are generated by a large model deployed on the computing nodes of the intelligent computing center, avoiding the data limitations of a single answer. Then, for each candidate answer, the logical chain in its generation process is deduced in reverse, forming a detailed and parsable answer strategy. Based on this, the core knowledge points or ability requirements involved in each answer strategy are further extracted, i.e., the test points described in this invention. Then, multi-dimensional scoring criteria are generated by combining the test points. Finally, the raw data, candidate answers, answer strategies, test points, and scoring criteria are integrated to distill a reinforcement learning dataset containing multi-layered information. Thus, through the processing method of "answer generation - strategy decomposition - test point extraction - standard setting," the training data not only contains the correspondence between questions and answers but also deeply includes logical reasoning paths, test points, and scoring criteria, thereby significantly improving the large model's ability to analyze complex problems and its adaptability to multi-dimensional evaluation.

[0092] It should also be noted that, Figure 1The method shown is applicable not only to the generation of reinforcement learning datasets with open-ended answers, but also to the generation of reinforcement learning datasets with standard answers. This is because different answering approaches can be constructed for questions with standard answers, and each answering approach will yield a standard answer. Corresponding test points can be set for each answering approach.

[0093] This section will use a specific application scenario as an example to explain the relationship between the questions, test points, and scoring criteria, in order to aid understanding. Figure 1 The method shown.

[0094] 1. Assume the original dataset contains 1000 rental data analysis questions, each including a question and data (fields such as time, region, and price). One question is "Analyze the quarterly rental demand changes in a certain city." The large model can generate multiple candidate answers, such as: Candidate Answer 1: Use a line chart to show the monthly order volume changes; Candidate Answer 2: Use a heatmap to show the quarterly rental order volume changes in each region; Candidate Answer 3: Use a line chart to show the correlation between quarterly house prices and rental order volume. Then, the large model can organize the answer strategies from the candidate answers by clustering similar answers and extracting core ideas. For example, the answer strategy for Candidate Answer 2 can be organized as: Multidimensional statistics and visualization approach: Quarterly-regional order volume statistics → Two-dimensional heatmap drawing → Trend and difference interpretation. Then, the test points can be defined for the answer strategies, such as the test points corresponding to the multidimensional statistics and visualization approach being multidimensional data statistics ability, two-dimensional heatmap drawing ability, and trend and difference analysis ability. Corresponding scoring criteria can also be developed for the test points.

[0095] 2. When grading essays or discussion questions, teachers do not score based on intuition, but rather develop detailed scoring rules in advance, breaking down the abstract concept of how to write into specific test points, with each test point corresponding to a clear score.

[0096] For example, the scoring criteria for the essay "My Dream" (total score 30 points) could be set as follows:

[0097] (1) Relevance to the topic (whether it revolves around "dreams") 0-10 points

[0098] (2) Structural clarity (introduction, body structure, conclusion) 0-8 points

[0099] (3) Vividness of language (use of rhetorical devices such as metaphor and parallelism) 0-7 points

[0100] (4) Neat handwriting (no misspellings, correct punctuation) 0-5 points

[0101] In summary, a reinforcement learning dataset can be obtained by distillation of the original dataset, containing at least one candidate answer, answer strategy, test points, and scoring criteria. This reinforcement learning dataset is more diverse and complex, and can be used to train large models, thereby effectively enhancing their performance and development potential in complex tasks and real-world applications, and improving their overall performance.

[0102] In one possible implementation, step S2 includes: step S21: filtering at least one candidate answer to obtain at least one filtered candidate answer, wherein the filtering includes at least one of the following: deduplication and quality filtering; step S22: based on the at least one filtered candidate answer and the large model, generating answer ideas corresponding to each candidate answer in the at least one filtered candidate answer respectively.

[0103] It should be noted that in the implementation of step S2, the multiple candidate answers generated in step S1 are first preprocessed to improve data quality. Specifically, duplicate candidate answers are eliminated through a "deduplication" operation, and logically flawed or obviously erroneous answers are removed through "quality filtering," ensuring that the candidate answers processed subsequently are both diverse and reliable. Then, based on the filtered candidate answer set, a large model is used to reverse-analyze each candidate answer, reconstructing the implicit reasoning steps in its generation process, ultimately forming a detailed answer strategy corresponding to each candidate answer. This ensures both the rationality and richness of the foundation for generating answer strategies (candidate answers) and that the answer strategies accurately reflect the differentiated thinking processes of different problem-solving paths, providing a foundation for subsequent key point extraction and the formulation of scoring criteria.

[0104] In one possible implementation, the number of answer ideas is at least one, and step S3 includes: step S31: processing at least one answer idea to obtain at least one processed answer idea, wherein the processing includes: deduplication and atomization; step S32: generating test points corresponding to the answer ideas based on the at least one processed answer idea and the large model.

[0105] It should be noted that in the implementation of step S3, the multiple answer approaches generated in step S2 are first processed in a structured manner: "Duplicate removal" eliminates repetitive or highly similar approaches, and "atomicization" breaks down complex answer approaches into indivisible independent units, ensuring that each processed answer approach possesses minimum completeness and parsingability. Subsequently, based on the processed set of answer approaches, a large model is used to summarize knowledge points and abstract ability points for each approach, extracting the test points supporting the approach (such as key steps in answering mathematical problems), ultimately forming a set of test points strictly corresponding to each answer approach (a set of test points includes at least one test point). Therefore, by processing the answer approaches and then generating test points based on them, the accuracy of test point generation can be improved.

[0106] In one possible implementation, step S32 includes any one of the following: Step S321: Based on the processed at least one answer approach and the large model, generate a set of test points corresponding to each answer approach in the at least one answer approach, wherein each set of test points is different; Step S322: Based on the processed at least one answer approach and the large model, generate a set of general test points corresponding to each answer approach in the at least one answer approach.

[0107] It should be noted that in the specific implementation of step S32, two modes for generating test points can be flexibly selected according to actual needs: The first mode (corresponding to step S321) analyzes each processed answer approach independently through a large model, generating a unique set of test points for each answer approach. There are differences between the test point sets corresponding to different answer approaches (for example, for two answer approaches to the same mathematical problem, two types of test points are generated respectively: "application of geometric properties" and "algebraic equation transformation method"); The second mode (corresponding to step S322) extracts the commonalities of all answer approaches through a large model, generating a set of general test points applicable to all answer approaches (such as "correct answer format" and "includes key derivation steps").

[0108] These two modes are applicable to scenarios of differentiated assessment and standardized assessment, respectively: the former, through the independence of the test point groups, can accurately depict the personalized needs of different answer approaches, while the latter, through the universality of the test points, can strengthen the model's unified mastery of core competencies across scenarios. This flexible design allows the generated test points to not only meet the in-depth assessment needs of specialized fields for specific abilities, but also adapt to the universal training of general thinking methods in basic fields (such as the ability to summarize the main idea in reading comprehension), thereby significantly improving the adaptability of subsequent scoring criteria development. Furthermore, it can more accurately simulate real-world application scenarios, effectively enhancing the performance and development potential of the large model in complex tasks and real-world application scenarios, and improving the overall performance of the large model.

[0109] In one possible implementation, such as Figure 2 As shown, step S321 includes:

[0110] Step S3211: Based on at least one processed answer approach and the overall model, generate a set of general test points corresponding to each answer approach;

[0111] Step S3212: Based on a set of general test points, at least one answer approach, and a large model, generate a set of test points corresponding to each answer approach.

[0112] It should be noted that, in the implementation of step S321, one optional approach is to first extract the core knowledge points or basic ability requirements commonly relied upon by all answer approaches through the large model, generating a set of general test points applicable to all answer approaches (e.g., "formula transformation rules" in mathematical problems); then, based on this set of general test points, and combined with the specific logical path of each answer approach, further generate its own specific detailed test points for each answer approach (e.g., supplementing "geometric auxiliary line construction techniques" or "algebraic reasoning methods" for different problem-solving methods). This mechanism of generating general test points to specific test points ensures consistency in basic ability assessment across different answer approaches (covering common requirements through general test points) while supplementing personalized test points based on the differences in specific answer approaches (refining specific requirements through specific test points). Therefore, the final generated test points possess both cross-method comparability (general part) and accurately reflect the unique value of different approaches (specific part), thereby significantly improving the compatibility of the large model training with diverse problem-solving strategies and the flexibility of the evaluation system. First, general test points are generated, and then, based on these, specific test points corresponding to each answer approach are generated. The specific test points can be checked based on the requirements of the general test points to avoid omissions and ensure the comprehensiveness and accuracy of the test point generation.

[0113] Alternatively, step S321 can be executed directly, which bypasses the general test points and instead generates specific test points for each answer approach based on at least one processed answer strategy and the overall model. The method for generating test points can be flexibly chosen according to actual needs.

[0114] In one possible implementation, if step S32 includes step S321, each set of test points corresponds to a different scoring standard; if step S32 includes step S322, a set of general test points corresponds to the same scoring standard.

[0115] It should be noted that in the implementation of step S32, if step S321 (generating an independent test point group for each answer approach) is selected, then the test point group corresponding to each answer approach will be associated with a unique scoring standard (for example, for the "application of geometric properties" test point group, a scoring rule with higher weight for graphical analysis can be set, while the "algebraic equation transformation skills" test point group can focus more on the scoring of the accuracy of formula transformation); if step S322 (generating a general test point group) is selected, then all answer approaches share the same set of scoring standards (for example, uniformly setting scoring weights according to general test points such as "correct answer format" and "includes key derivation steps").

[0116] This differentiated scoring standard matching mechanism allows for the precise quantification of the advantages of different answer approaches in scenarios requiring in-depth assessment of individualized abilities, while ensuring consistency of evaluation results through universal scoring standards in scenarios requiring standardized evaluation of abilities. Thus, the large model can adapt to the refined evaluation needs of specialized fields regarding differences in subdivided abilities, while also meeting the requirements of standard abilities in general scenarios, thereby enhancing the comprehensiveness and adaptability of the evaluation system.

[0117] In one possible implementation, the upper limit of the score varies for each set of test points and the corresponding different scoring criteria.

[0118] Understandably, the upper limit of the scoring criteria for each group of test points can be the same, for example, all out of 100, and then the total score is calculated by weighting each test point according to its weight. The upper limit of the scoring criteria for each group of test points can also be different. For example, the upper limit of the test point corresponding to answer approach A is 60 points, while the upper limit of the test point corresponding to answer approach B is 100 points. For example, for the "basic formula application" test point group, its upper limit of scoring criteria might be set at 60 points (emphasizing the attainment of basic abilities), while the upper limit of the "advanced logic deduction" test point group could be set at 100 points (encouraging breakthroughs in in-depth innovation capabilities). Therefore, differentiated scoring criteria can be formulated for different test point groups, precisely guiding the training direction of the large model between the stability of basic abilities and the breakthrough of advanced innovation capabilities, thereby improving the overall performance of the large model.

[0119] In one possible implementation, such as Figure 3 As shown, after step S5, the method further includes:

[0120] Step S6: Train the model using reinforcement learning based on the reinforcement learning dataset, and score the model's reinforcement learning training based on a preset scoring model.

[0121] It should be noted that the answers generated by the model can be checked using a pre-defined scoring model, and scores can be given for different test points. Finally, the final score of the answer is calculated using a pre-defined scoring formula. This scoring method can be encapsulated as an accuracy_reward function for reinforcement learning, enabling the model to perform reinforcement learning.

[0122] In reinforcement learning training, if the model-generated answer to be evaluated includes at least one approach to be evaluated, and the approach to be evaluated is an approach that successfully matches the answer approach in the reinforcement learning dataset, then the scoring includes at least one of the following: scoring each approach to be evaluated in the at least one approach to be evaluated, obtaining the scoring result corresponding to each approach to be evaluated, and performing a weighted summation of the scoring results, and using the weighted summation result as the scoring result of the answer to be evaluated; or scoring each approach to be evaluated in the at least one approach to be evaluated, obtaining the scoring result corresponding to each approach to be evaluated, and using the scoring result with the highest score as the scoring result of the answer to be evaluated.

[0123] It should be noted that during the scoring process of reinforcement learning training, when the model-generated answers contain multiple approaches to be evaluated, the following scoring methods can be used: If a weighted summation mechanism is used, the scores of each approach to be evaluated (e.g., approach A gets 80 points, approach B gets 60 points) are weighted according to preset weights (e.g., approach A has a weight of 70%, approach B has a weight of 30%) to calculate a comprehensive score (80 × 0.7 + 60 × 0.3 = 74 points); if a highest score selection mechanism is used, the highest score among all approaches to be evaluated (e.g., approach C gets 90 points) is directly selected as the final score. This allows large-scale model training to both encourage exploration of multiple approaches (preventing over-reliance on a single approach through a weighted mechanism) and maintain focus on core capabilities (strengthening key advantages through a selection mechanism), thereby further ensuring the performance of the large-scale model.

[0124] In one possible implementation, after step S5, the method further includes:

[0125] Step S7: Construct a benchmark dataset based on the reinforcement learning dataset.

[0126] Therefore, we can provide an evaluation benchmark that includes multi-dimensional capability indicators for the assessment of large models, which can make the evaluation system of large models more complete, and can more comprehensively consider and train the understanding, reasoning and creativity of large models, thereby improving the performance of large models.

[0127] In summary, this invention, by leveraging the computing power of an intelligent computing center, utilizes a large-scale model to automatically generate a reinforcement learning dataset containing multiple candidate answers, answer strategies, test points (assessing abilities and knowledge points), and scoring criteria. Specifically, starting from the questions in the original dataset, it not only generates candidate answers but also deconstructs the reasoning logic behind them, generates answer strategies, analyzes the test points involved, and automatically formulates scoring criteria, thereby generating the reinforcement learning dataset. This reinforcement learning dataset can simultaneously solve two major problems: firstly, by simulating diverse problem-solving approaches in complex scenarios, it allows the large-scale model to learn how to handle open-ended questions without standard answers, or to learn multiple reasoning approaches for questions with standard answers; secondly, through a verification system of "answer strategies - test points - scoring criteria," it comprehensively evaluates the large-scale model's depth of understanding, logical reasoning ability, and knowledge application level. This effectively enhances the performance and development potential of the large-scale model in complex tasks and real-world application scenarios.

[0128] In addition, constructing a benchmark dataset based on a reinforcement learning dataset, and then training the large model to be evaluated using reinforcement learning, or fine-tuning the large model based on the reinforcement learning dataset, can make the evaluation system of the large model more complete. This allows for a more comprehensive consideration and training of the large model's understanding, reasoning, and creativity, thereby improving the performance of the large model.

[0129] Figure 4 An apparatus for providing computing power to distill reinforcement learning datasets at an intelligent computing center is shown, such as... Figure 4 As shown, the device 40 includes:

[0130] The module 401 is used to perform step S1: acquire the original dataset, and based on the original dataset and the large model, generate at least one candidate answer corresponding to the original dataset. The large model is deployed on at least one computing node in the intelligent computing center.

[0131] Execution module 402 is used to execute step S2: based on at least one candidate answer and the large model, generate the answer ideas corresponding to each candidate answer respectively;

[0132] Step S3: Based on the answer approach and the overall model, generate the test points corresponding to the answer approach;

[0133] Step S4: Based on the answer approach, key points, and overall model, generate scoring criteria corresponding to the key points;

[0134] Step S5: Merge the original dataset, at least one candidate answer, answer strategy, test points, and scoring criteria to obtain the reinforcement learning dataset.

[0135] In one possible implementation, step S2 includes:

[0136] Step S21: Filter at least one candidate answer to obtain at least one filtered candidate answer, wherein the filtering includes at least one of the following: deduplication and quality filtering;

[0137] Step S22: Based on at least one candidate answer after screening and the large model, generate the corresponding answer ideas for each candidate answer in the at least one candidate answer after screening.

[0138] In one possible implementation, the number of answer ideas is at least one, and step S3 includes:

[0139] Step S31: Process at least one answer approach to obtain at least one processed answer approach, wherein the processing includes: deduplication and atomization.

[0140] Step S32: Based on at least one processed answer approach and large model, generate test points corresponding to the answer approach.

[0141] In one possible implementation, step S32 includes any of the following:

[0142] Step S321: Based on the processed at least one answer approach and the overall model, generate a set of test points corresponding to each answer approach in the at least one answer approach, wherein each set of test points is different;

[0143] Step S322: Based on the processed at least one answer approach and the overall model, generate a set of general test points that correspond to each answer approach in the at least one answer approach.

[0144] In one possible implementation, step S321 includes:

[0145] Step S3211: Based on the processed at least one answer approach and the overall model, generate a set of general test points corresponding to each answer approach in the at least one answer approach;

[0146] Step S3212: Based on a set of general test points, at least one answer approach, and a large model, generate a set of test points corresponding to each answer approach in the at least one answer approach.

[0147] In one possible implementation, if step S32 includes step S321, each set of test points corresponds to a different scoring standard.

[0148] In the case where step S32 includes step S322, a set of general test points corresponds to the same set of scoring criteria.

[0149] In one possible implementation, the upper limit of the score varies for each set of test points and the corresponding different scoring criteria.

[0150] In one possible implementation, the execution module 402 is further configured to execute step S6 after step S5: perform reinforcement learning training on the model based on the reinforcement learning dataset, and score the reinforcement learning training of the model based on a preset scoring model.

[0151] In reinforcement learning training, if the model generates an answer to be evaluated that includes at least one approach to be evaluated, and that approach successfully matches the answer approach in the reinforcement learning dataset, then the scoring includes at least one of the following:

[0152] Each of the at least one approaches to be evaluated is scored to obtain a score result for each approach to be evaluated. The scores are then weighted and summed, and the weighted sum is used as the score result for the answer to be evaluated.

[0153] Each of the at least one approaches to be evaluated is scored separately to obtain a score result for each approach. The highest score among the scores is then taken as the score result for the answer to be evaluated.

[0154] In one possible implementation, the execution module 402 is further configured to perform step S7 after step S5: constructing a benchmark dataset based on the reinforcement learning dataset.

[0155] In one possible implementation, the reinforcement learning dataset includes:

[0156] The original dataset, candidate answers, answer strategies, test points, scoring criteria, and the standard answer determined from the candidate answers.

[0157] In summary, this invention, by leveraging the computing power of an intelligent computing center, utilizes a large-scale model to automatically generate a reinforcement learning dataset containing multiple candidate answers, answer strategies, test points (assessing abilities and knowledge points), and scoring criteria. Specifically, starting from the questions in the original dataset, it not only generates candidate answers but also deconstructs the reasoning logic behind them, generates answer strategies, analyzes the test points involved, and automatically formulates scoring criteria, thereby generating the reinforcement learning dataset. This reinforcement learning dataset can simultaneously solve two major problems: firstly, by simulating diverse problem-solving approaches in complex scenarios, it allows the large-scale model to learn how to handle open-ended questions without standard answers, or to learn multiple reasoning approaches for questions with standard answers; secondly, through a verification system of "answer strategies - test points - scoring criteria," it comprehensively evaluates the large-scale model's depth of understanding, logical reasoning ability, and knowledge application level. This effectively enhances the performance and development potential of the large-scale model in complex tasks and real-world application scenarios, improving its overall performance.

[0158] Please refer to Figure 5 The present invention also provides an electronic device 50, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the computer program is executed by the processor 501, it implements the steps of the above-described method for providing computing power in an intelligent computing center to distill reinforcement learning dataset, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0159] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of the aforementioned method for providing computing power in an intelligent computing center to distill reinforcement learning datasets, achieving the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0160] The present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-described method for providing computing power in an intelligent computing center for distilling reinforcement learning datasets, and achieve the same technical effect. To avoid repetition, the details will not be repeated here.

[0161] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the above methods can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the present invention.

[0163] The present invention has been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other modifications under the guidance of the present invention without departing from the spirit and scope of the claims, and all such modifications are within the protection scope of the present invention.

Claims

1. A method for distilling reinforcement learning datasets using an intelligent computing center that provides computing power, characterized in that, The method includes: Step S1: Obtain the original dataset, and based on the original dataset and the large model, generate at least one candidate answer corresponding to the original dataset. The large model is deployed on at least one computing node in the intelligent computing center. Step S2: Based on the at least one candidate answer and the large model, generate a response strategy corresponding to each candidate answer; Step S3: Based on the answer approach and the large model, generate test points corresponding to the answer approach; wherein, the number of answer approaches is at least one, and step S3 includes: processing the at least one answer approach to obtain the processed at least one answer approach, wherein the processing includes: deduplication and atomization; based on the processed at least one answer approach and the large model, generate a set of general test points corresponding to each answer approach; based on the set of general test points, the at least one answer approach, and the large model, generate a set of test points corresponding to each answer approach respectively. Step S4: Based on the answer approach, the test points, and the large model, generate the scoring criteria corresponding to the test points; Step S5: Merge the original dataset, the at least one candidate answer, the answering strategy, the test points, and the scoring criteria to obtain the reinforcement learning dataset.

2. The method according to claim 1, characterized in that, Step S2 includes: Step S21: Filter the at least one candidate answer to obtain at least one filtered candidate answer, wherein the filtering includes at least one of the following: deduplication and quality filtering; Step S22: Based on at least one candidate answer after screening and the large model, generate a response strategy corresponding to each candidate answer.

3. The method according to claim 1, characterized in that, Each set of test points corresponds to a different scoring standard.

4. The method according to claim 3, characterized in that, Each set of test points has a different scoring standard with a different upper limit for scoring.

5. The method according to claim 1, characterized in that, After step S5, the method further includes: Step S6: Train the model using reinforcement learning based on the reinforcement learning dataset, and score the reinforcement learning training of the model based on a preset scoring model; Wherein, in the reinforcement learning training, if the answer to be evaluated generated by the model includes at least one approach to be evaluated, and the approach to be evaluated is an approach that successfully matches the answer approach in the reinforcement learning dataset, then the scoring includes at least one of the following: Each of the at least one approaches to be evaluated is scored to obtain a score result for each approach to be evaluated. The scores are then weighted and summed, and the weighted sum is used as the score result for the answer to be evaluated. Each of the at least one approaches to be evaluated is scored to obtain a score result for each approach, and the highest score among the scores is taken as the score result of the answer to be evaluated.

6. The method according to claim 1, characterized in that, After step S5, the method further includes: Step S7: Based on the reinforcement learning dataset, construct a benchmark dataset.

7. The method according to any one of claims 1-6, characterized in that, The reinforcement learning dataset includes: The original dataset, candidate answers, answer strategies, test points, scoring criteria, and the standard answer determined from the candidate answers.

8. An apparatus for providing computing power to distill reinforcement learning datasets in an intelligent computing center, characterized in that, The device includes: The acquisition module is used to perform step S1: acquire the original dataset, and generate at least one candidate answer corresponding to the original dataset based on the original dataset and the large model, wherein the large model is deployed on at least one computing node in the intelligent computing center; The execution module is used to execute step S2: based on the at least one candidate answer and the large model, generate a response strategy corresponding to each candidate answer; Step S3: Based on the answer approach and the large model, generate test points corresponding to the answer approach; wherein, the number of answer approaches is at least one, and step S3 includes: processing the at least one answer approach to obtain the processed at least one answer approach, wherein the processing includes: deduplication and atomization; based on the processed at least one answer approach and the large model, generate a set of general test points corresponding to each answer approach; based on the set of general test points, the at least one answer approach, and the large model, generate a set of test points corresponding to each answer approach respectively. Step S4: Based on the answer approach, the test points, and the large model, generate the scoring criteria corresponding to the test points; Step S5: Merge the original dataset, the at least one candidate answer, the answering strategy, the test points, and the scoring criteria to obtain the reinforcement learning dataset.

9. A server, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for providing computing power for a smart computing center distillation reinforcement learning dataset as claimed in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of a method for providing computing power in an intelligent computing center to distill reinforcement learning dataset as described in any one of claims 1-7.

11. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of a method for providing computing power for a centrally-controlled intelligent computing system to distill reinforcement learning datasets as described in any one of claims 1-7.