Method and device for performing large-model self-distillation absorption depth reasoning by intelligent computing center cloud platform through computing power

By building and training a large model self-distillation method to absorb deep reasoning, the problem of large models occupying a large amount of computing resources in the intelligent computing center cloud platform is solved, and more efficient reasoning and service capabilities are achieved.

CN120745848AActive Publication Date: 2025-10-03DATACANVAS LTD

Patent Information

Application Number
CN202511255196.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-03
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In existing technologies, large models occupy a large amount of computing resources in the intelligent computing center cloud platform, resulting in low service efficiency and inability to effectively accelerate the deep reasoning process.

Method used

By obtaining multiple sets of initial sample data, constructing multiple sets of first and second sample data, and training the initial large model, we can obtain the first large model that can self-distill and absorb deep reasoning, so that it can directly generate reasoning chain data and response data without performing deep reasoning steps.

Benefits of technology

It improves reasoning efficiency, reduces computing resource usage, and increases the service capabilities of the intelligent computing center cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745848A_ABST
    Figure CN120745848A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for performing large-model self-distillation absorption depth reasoning by an intelligent computing center cloud platform through computing power, and relates to the technical field of intelligent computing centers, intelligent computing centers, computing power infrastructures and intelligent computing clouds, and the method comprises the following steps: S1, obtaining multiple groups of initial sample data; s2, constructing multiple groups of first sample data and multiple groups of second sample data based on the multiple groups of initial sample data; and S3, training the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model. According to the method, the displayed natural language deep reasoning step is distilled to a continuous hidden space, and the first model can still ensure complete and excellent reasoning ability under the condition of not displaying and generating the middle deep reasoning step, so that resources occupied by a large model can be reduced; therefore, the large model service provided in the intelligent computing center cloud platform is greatly increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of intelligent computing centers, smart computing centers, computing power infrastructure and intelligent computing cloud technology, and specifically to a method and device for an intelligent computing center cloud platform to perform large-scale model self-distillation and absorption deep reasoning through computing power. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.

[0003] An "Intelligent Computing Center" is a facility that utilizes large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, providing a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process parameters. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing parameter data. It is a new type of productivity that integrates parameter computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] The large-scale model inference process on the Intelligent Computing Center cloud platform requires deep reasoning on the question to obtain the inference chain data and generate the corresponding response. However, deep reasoning on large models consumes a considerable amount of time and computing resources. Existing technologies do not provide effective ways to accelerate deep reasoning, nor can they skip or display the deep reasoning generation steps. Consequently, the Intelligent Computing Center cloud platform is unable to provide more services using large models.

[0008] It can be seen that there is a problem in the existing technology that large models occupy a large amount of computing resources, resulting in few large model services being provided in the intelligent computing center cloud platform. Summary of the Invention

[0009] The present invention provides a method and device for an intelligent computing center cloud platform to perform self-distillation and absorb deep reasoning of large models through computing power, so as to solve the problem in the prior art that large models occupy a large amount of computing power resources, resulting in few large model services being provided in the intelligent computing center cloud platform.

[0010] To solve the above problems, the present invention is achieved as follows: In a first aspect, the present invention provides a method for an intelligent computing center cloud platform to perform large-scale model self-distillation and absorption deep reasoning through computing power, comprising: Step S1: Acquire multiple sets of initial sample data, where each set of initial sample data includes a sample question, sample deep reasoning data, sample reasoning chain data, and sample response data; Step S2: constructing multiple sets of first sample data and multiple sets of second sample data based on the multiple sets of initial sample data, wherein each set of first sample data in the multiple sets of first sample data includes the sample question, the sample reasoning chain data, and the sample reply data, and each set of second sample data in the multiple sets of second sample data includes the sample deep reasoning data, the sample reasoning chain data, and the sample reply data; Step S3: Training the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, where the first large model is used to generate reasoning chain data and response data corresponding to the question.

[0011] In one embodiment, step S3 includes: Step S31: training the initial large model based on the multiple sets of first sample data and the multiple sets of second sample data to obtain a first intermediate model; Step S32: calculating a first loss value corresponding to the first intermediate model based on the multiple groups of first sample data; Step S33: calculating a second loss value corresponding to the first intermediate model based on the multiple groups of first sample data and the multiple groups of second sample data; Step S34: Calculate a third loss value, where the third loss value is a weighted value of the first loss value and the second loss value; Step S35: When the third loss value is less than or equal to a preset loss threshold, the first intermediate model is set as the first large model.

[0012] In one embodiment, step S32 includes: Step S321: Calculate a first distribution probability corresponding to each first word-gram in a plurality of first word-grams based on the first intermediate model, and calculate a second distribution probability corresponding to each second word-gram in a plurality of second word-grams, where each first word-gram is a word-gram of the sample inference chain data included in each set of first sample data, and each second word-gram is a word-gram of the sample response data included in each set of first sample data; Step S322: Calculate a plurality of first logarithms and a plurality of second logarithms, wherein the plurality of first logarithms are logarithms of first distribution probabilities corresponding to the plurality of first word-grams, and the plurality of first logarithms have a one-to-one correspondence with the plurality of first word-grams; and the plurality of second logarithms are logarithms of second distribution probabilities corresponding to the plurality of second word-grams, and the plurality of second logarithms have a one-to-one correspondence with the plurality of second word-grams; Step S323: Calculate a first mean and a second mean, where the first mean is the mean of multiple first pairs, and the second mean is the mean of multiple second pairs; Step S324: Calculate the first loss value, where the first loss value is a weighted value of the first mean and the second mean.

[0013] In one embodiment, step S33 includes: Step S331: Calculate a first distribution probability corresponding to each first word-gram in a plurality of first word-grams based on the first intermediate model, and calculate a second distribution probability corresponding to each second word-gram in a plurality of second word-grams; Step S332: Calculate, based on the first intermediate model, a third distribution probability corresponding to each third word-gram in a plurality of third words-grams, and calculate a fourth distribution probability corresponding to each fourth word-gram in a plurality of fourth words-grams, wherein each third word-gram is a word-gram of the sample inference chain data included in each set of second sample data, and each fourth word-gram is a word-gram of the sample response data included in each set of second sample data, wherein the plurality of first words-grams correspond one-to-one to the plurality of third words-grams, and the plurality of second words-grams correspond one-to-one to the plurality of fourth words-grams; Step S333: Calculate a plurality of third logarithms and a plurality of fourth logarithms, wherein the plurality of third logarithms are logarithms of a plurality of first quotient values, each of which is the quotient of a first distribution probability of a corresponding first word-unit and a third distribution probability of a third word-unit; and the plurality of fourth logarithms are logarithms of a plurality of second quotient values, each of which is the quotient of a second distribution probability of a corresponding second word-unit and a fourth distribution probability of a fourth word-unit; Step S334: Calculate a third mean and a fourth mean, where the third mean is the mean of a plurality of first expectations, which are expectations of the plurality of third logarithms, and the plurality of first expectations correspond one-to-one with the plurality of third logarithms; the fourth mean is the mean of a plurality of second expectations, which are expectations of the plurality of fourth logarithms, and the plurality of second expectations correspond one-to-one with the plurality of fourth logarithms; Step S335: Calculate the second loss value, where the second loss value is a weighted value of the third mean and the fourth mean.

[0014] In one embodiment, after step S3, the method further includes: Step S4: receiving the first question; Step S5: Generate reasoning chain data and response data corresponding to the first question based on the first large model; Step S6: Output the reasoning chain data and response data corresponding to the first question.

[0015] In one embodiment, after step S6, the method further includes: Step S7: receiving feedback information, wherein the feedback information includes the modified reasoning chain data and reply data; Step S8: adding the feedback information to the database; Step S9: When the amount of feedback information stored in the database is greater than a set threshold, generating updated training data based on the feedback information, the updated training data including training questions, training reasoning chain data, and training response data; Step S10: Update the first large model based on the updated training data.

[0016] In a second aspect, the present invention further provides a device for an intelligent computing center cloud platform to perform large-scale model self-distillation and absorption deep reasoning through computing power, comprising: An acquisition module, configured to acquire multiple sets of initial sample data, each set of initial sample data including a sample question, sample deep reasoning data, sample reasoning chain data, and sample response data; a construction module, configured to construct multiple sets of first sample data and multiple sets of second sample data based on the multiple sets of initial sample data, wherein each set of first sample data in the multiple sets of first sample data includes the sample question, the sample reasoning chain data, and the sample reply data, and each set of second sample data in the multiple sets of second sample data includes the sample deep reasoning data, the sample reasoning chain data, and the sample reply data; The training module is used to train the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, wherein the first large model is used to generate reasoning chain data and response data corresponding to the question.

[0017] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for the intelligent computing center cloud platform to perform large-model self-distillation and absorption deep reasoning through computing power as described in the first aspect above are implemented.

[0018] In a fourth aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the method of performing large-model self-distillation and absorption deep reasoning through computing power on the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0019] In a fifth aspect, the present invention also provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps in the method of the intelligent computing center cloud platform performing large-model self-distillation and absorption deep reasoning through computing power as described in the first aspect above.

[0020] In the present invention, the method for the intelligent computing center cloud platform to perform self-distillation and absorb deep reasoning of a large model through computing power includes: step S1, obtaining multiple groups of initial sample data, each group of initial sample data in the multiple groups of initial sample data includes sample questions, sample deep reasoning data, sample reasoning chain data and sample reply data; step S2, constructing multiple groups of first sample data and multiple groups of second sample data based on the multiple groups of initial sample data, each group of first sample data in the multiple groups of first sample data includes the sample question, the sample reasoning chain data and the sample reply data, and each group of second sample data in the multiple groups of second sample data includes the sample deep reasoning data, the sample reasoning chain data and the sample reply data; step S3, training the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, and the first large model is used to generate reasoning chain data and reply data corresponding to the problem. In this way, each group of first sample data in multiple groups of first sample data includes sample questions, sample reasoning chain data and sample reply data, and each group of second sample data in multiple groups of second sample data includes sample deep reasoning data, sample reasoning chain data and sample reply data. The initial large model is trained by multiple groups of first sample data and multiple groups of second sample data to obtain the first large model, which distills the displayed natural language deep reasoning steps into a continuous latent space, so that the first model can still ensure complete and excellent reasoning ability without displaying the intermediate deep reasoning steps. The reasoning chain data and reply data corresponding to the question can be directly generated through the first large model, which effectively improves the reasoning efficiency and greatly reduces the occupation of computing resources, thereby greatly increasing the large model services provided in the intelligent computing center cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0022] Figure 1 This is a flow chart of a method provided by the present invention for an intelligent computing center cloud platform to perform large-scale model self-distillation and absorption deep reasoning through computing power; Figure 2 It is a schematic diagram of deep reasoning of self-distillation absorption of a large model provided by the present invention; Figure 3 Schematic diagram of the training process of the initial large model provided by the present invention; Figure 4 This is a structural diagram of an intelligent computing center cloud platform provided by the present invention that uses computing power to perform large-scale model self-distillation and absorb deep reasoning; Figure 5 This is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0024] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0025] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator for measuring the computing power of a data center, including general computing power, supercomputing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0026] The "Network Power" (NP) mentioned in this invention refers to: it is a manifestation of the data transmission capability of computing power facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0027] "Storage Power" (SP) as used in this document refers to the comprehensive capabilities of a data center in four areas: data storage capacity, performance, security and reliability, and environmental friendliness and low-carbon development. It serves as a comprehensive indicator of a data center's data storage capabilities, encompassing both external storage devices like storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second). Disaster recovery ratio is a key indicator of security and reliability.

[0028] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.

[0029] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0030] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0031] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0032] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.

[0033] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.

[0034] The term "intelligent computing center" as used in this document refers to a facility that utilizes large-scale heterogeneous computing resources, including general-purpose computing power (CPUs) and intelligent computing power (GPUs, FPGAs, ASICs, etc.), primarily to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). An intelligent computing center encompasses facilities, hardware, and software, providing a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0035] The "intelligent computing center cloud platform" mentioned in the present invention is referred to as "intelligent computing cloud", which refers to: a cloud computing platform that provides comprehensive services based on the hardware resources and software resources of the intelligent computing center.

[0036] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".

[0037] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0038] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0039] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0040] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water, and electricity.

[0041] The "big model" mentioned in the present invention includes but is not limited to a "big language model" and a "big multimodal model".

[0042] The "large language model" mentioned in this invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0043] The "Multimodal Large Models" mentioned in this invention refer to models that combine multimodal information such as text, images, video, and audio for training, including but not limited to multimodal large language models.

[0044] The "self-distillation absorption of deep reasoning" described in the present invention means: absorbing the deep reasoning process of the large model through training, so that the large model after absorbing the deep reasoning process can omit the deep reasoning process during reasoning, directly obtain the reasoning chain and response data, realize fast reasoning, and reduce the occupation of computing resources.

[0045] See Figure 1 , Figure 1 This is a flowchart of a method provided by the present invention for an intelligent computing center cloud platform to perform large-scale model self-distillation and absorption deep reasoning through computing power, such as Figure 1 As shown, the following steps are included: Step S1: Acquire multiple groups of initial sample data, where each group of initial sample data includes sample questions, sample deep reasoning data, sample reasoning chain data, and sample response data.

[0046] It should be noted that if Figure 2 As shown, in the large model reasoning of the existing technology, after the question (Question) is input into the large model, the large model will perform deep reasoning (Deep Think) to obtain the chain of thought (CoT) data and response data (Answer). The deep reasoning process consumes a lot of time and requires a lot of computing resources in the process. Therefore, in the present invention, the large model is trained by the method of absorbing deep reasoning through self-distillation of the large model to obtain a first large model, so that the first large model can omit the deep reasoning process when reasoning on the question input by the user, and directly obtain the reasoning chain and response data, thereby greatly reducing the reasoning time of the large model for the problem and reducing the occupation of computing resources.

[0047] The above-mentioned multiple sets of initial sample data are obtained by reasoning based on the existing technology model. Specifically, multiple sample questions can be input into the large model in the existing technology, and sample deep reasoning data, sample reasoning chain data and sample response data can be obtained through the large model. It can be understood that since the large model in the existing technology performs deep reasoning during the reasoning process, the accuracy of the reasoning chain data and response data obtained is relatively high. The reasoning chain data output by the large model in the existing technology can be directly used as sample reasoning chain data, the response data output can be used as sample response data, and the deep reasoning data of the large model can be extracted as sample deep reasoning data.

[0048] Among them, each set of initial sample data in the multiple sets of initial sample data includes sample questions, sample deep reasoning data, sample reasoning chain data and sample response data.

[0049] In some embodiments, after obtaining multiple sets of initial sample data, the multiple sets of initial sample data can be preprocessed to delete missing initial sample data. For example, if only a portion of the sample deep reasoning data, sample reasoning chain data, and sample response data is obtained when reasoning on an input sample question using a large model in the prior art, the initial sample data containing the sample data can be deleted, so that each set of initial sample data can effectively implement model training.

[0050] Step S2: construct multiple groups of first sample data and multiple groups of second sample data based on the multiple groups of initial sample data, wherein each group of first sample data in the multiple groups of first sample data includes the sample question, the sample reasoning chain data and the sample reply data, and each group of second sample data in the multiple groups of second sample data includes the sample deep reasoning data, the sample reasoning chain data and the sample reply data.

[0051] The above-mentioned multiple sets of first sample data are sample data that do not include sample depth reasoning data, and the above-mentioned multiple sets of second sample data are sample data that include sample depth reasoning data. Figure 3 As shown, multiple sets of first sample data and multiple sets of second sample data are constructed through multiple sets of initial sample data, and then multiple sets of first sample data and multiple sets of second sample data are constructed through multiple sets of initial sample data to train the large model, so that the trained large model can obtain reasoning chain data and response data without deep reasoning.

[0052] Step S3: Training the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, where the first large model is used to generate reasoning chain data and response data corresponding to the question.

[0053] The above-mentioned initial large model is a large language model or a multimodal model. Multiple sets of first sample data and multiple sets of second sample data are constructed through multiple sets of initial sample data to train the large model to obtain the first large model, and the displayed natural language deep reasoning steps are distilled into a continuous latent space. The first large model does not need to perform deep reasoning, and directly generates the reasoning chain data and response data corresponding to the question, so that the first model can still ensure complete and excellent reasoning capabilities without displaying the intermediate deep reasoning steps, thereby effectively improving the reasoning efficiency and greatly reducing the occupation of computing resources.

[0054] The initial large model may be a large model obtained by obtaining multiple sets of initial sample data. Specifically, in some embodiments, obtaining multiple sets of initial sample data includes: Get multiple sample questions; The multiple sample data are inferred based on the initial large model to obtain sample deep reasoning data, sample reasoning chain data and sample response data corresponding to each sample question.

[0055] In the present invention, the method for the intelligent computing center cloud platform to perform self-distillation and absorb deep reasoning of a large model through computing power includes: step S1, obtaining multiple groups of initial sample data, each group of initial sample data in the multiple groups of initial sample data includes sample questions, sample deep reasoning data, sample reasoning chain data and sample reply data; step S2, constructing multiple groups of first sample data and multiple groups of second sample data based on the multiple groups of initial sample data, each group of first sample data in the multiple groups of first sample data includes the sample question, the sample reasoning chain data and the sample reply data, and each group of second sample data in the multiple groups of second sample data includes the sample deep reasoning data, the sample reasoning chain data and the sample reply data; step S3, training the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, and the first large model is used to generate reasoning chain data and reply data corresponding to the problem. In this way, each group of first sample data in multiple groups of first sample data includes sample questions, sample reasoning chain data and sample reply data, and each group of second sample data in multiple groups of second sample data includes sample deep reasoning data, sample reasoning chain data and sample reply data. The initial large model is trained by multiple groups of first sample data and multiple groups of second sample data to obtain the first large model, which distills the displayed natural language deep reasoning steps into a continuous latent space, so that the first model can still ensure complete and excellent reasoning ability without displaying the intermediate deep reasoning steps. The reasoning chain data and reply data corresponding to the question can be directly generated through the first large model, which effectively improves the reasoning efficiency and greatly reduces the occupation of computing resources, thereby greatly increasing the large model services provided in the intelligent computing center cloud platform.

[0056] In one embodiment, step S3 includes: Step S31: training the initial large model based on the multiple sets of first sample data and the multiple sets of second sample data to obtain a first intermediate model; Step S32: calculating a first loss value corresponding to the first intermediate model based on the multiple groups of first sample data; Step S33: calculating a second loss value corresponding to the first intermediate model based on the multiple groups of first sample data and the multiple groups of second sample data; Step S34: Calculate a third loss value, where the third loss value is a weighted value of the first loss value and the second loss value; Step S35: When the third loss value is less than or equal to a preset loss threshold, the first intermediate model is set as the first large model.

[0057] The above-mentioned first loss value is a loss value calculated based on multiple groups of first sample data. The first loss value can represent the situation where the first intermediate model directly generates reasoning chain data and response data corresponding to the problem; the above-mentioned second loss value is a loss value calculated based on the multiple groups of first sample data and the multiple groups of second sample data. The second loss value can represent the situation where the first intermediate model generates reasoning chain data and response data corresponding to the problem through a deep reasoning process.

[0058] It should be noted that when the first intermediate model directly generates the reasoning chain data and response data corresponding to the question, the reasoning chain data and response data obtained are not necessarily accurate because they have not undergone deep reasoning. During the model training process, the reasoning chain data and response data corresponding to the directly generated question need to be as close as possible to the reasoning chain data and response data corresponding to the question generated through the deep reasoning process, so as to improve the accuracy of the first large model obtained by training, so that the first model can still guarantee complete and excellent reasoning capabilities without displaying the deep reasoning steps in the generation process.

[0059] Therefore, in the present invention, the initial large model is trained based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first intermediate model; the first loss value corresponding to the first intermediate model is calculated based on the multiple groups of first sample data; the second loss value corresponding to the first intermediate model is calculated based on the multiple groups of first sample data and the multiple groups of second sample data; the third loss value is calculated, and the third loss value is a weighted value of the first loss value and the second loss value; when the third loss value is less than or equal to the preset loss threshold, the first intermediate model is set to the first large model. In this way, the first loss value can be used to determine the situation in which the first intermediate model directly generates the reasoning chain data and response data corresponding to the problem, and the second loss value can be used to determine the situation in which the first intermediate model generates the reasoning chain data and response data corresponding to the problem through a deep reasoning process. The third loss value is then calculated using the first loss value and the second loss value. The third loss value can be used to determine whether the effect of the first intermediate model directly generating the reasoning chain data and response data corresponding to the problem is close to the effect of the first intermediate model generating the reasoning chain data and response data corresponding to the problem through a deep reasoning process. When the third loss value is less than or equal to the preset loss threshold, the first intermediate model is set as the first large model, which improves the accuracy of the first large model in directly generating the reasoning chain data and response data corresponding to the problem, so that the first model can still guarantee complete and excellent reasoning capabilities without displaying the deep reasoning steps in the generation process.

[0060] In some embodiments, the third loss value may be calculated using the following formula: ; in the formula is the third loss value, is the first loss value, is the second loss value, is a hyperparameter that balances the first and second loss values.

[0061] The first loss value may be a cross entropy loss value, and the second loss value may be a relative entropy loss value, such as a KL (Kullback-Leibler Divergence) divergence loss value.

[0062] In one embodiment, step S32 includes: Step S321: Calculate, based on the first intermediate model, a first distribution probability corresponding to each first token in a plurality of first tokens, and a second distribution probability corresponding to each second token in a plurality of second tokens, wherein each first token is a token of the sample inference chain data included in each set of first sample data, and each second token is a token of the sample response data included in each set of first sample data; Step S322: Calculate a plurality of first logarithms and a plurality of second logarithms, wherein the plurality of first logarithms are logarithms of first distribution probabilities corresponding to the plurality of first word-grams, and the plurality of first logarithms have a one-to-one correspondence with the plurality of first word-grams; and the plurality of second logarithms are logarithms of second distribution probabilities corresponding to the plurality of second word-grams, and the plurality of second logarithms have a one-to-one correspondence with the plurality of second word-grams; Step S323: Calculate a first mean and a second mean, where the first mean is the mean of multiple first pairs, and the second mean is the mean of multiple second pairs; Step S324: Calculate the first loss value, where the first loss value is a weighted value of the first mean and the second mean.

[0063] The above-mentioned first distribution probability is the distribution probability corresponding to each token in the sample inference chain of multiple groups of first sample data, and the above-mentioned second distribution probability is the distribution probability corresponding to each token in the sample response data of multiple groups of first sample data. The first distribution probability and the second distribution probability are used to determine the distribution of the inference chain data and sample data directly output by the first intermediate model, and then the first loss value for determining the situation where the inference chain data and response data corresponding to the problem directly generated by the first intermediate model can be obtained.

[0064] Furthermore, after the first distribution probability and the second distribution probability are calculated, the first loss value may be obtained by calculating the cross entropy (ie, steps S322 to S324 ).

[0065] In some implementations, the calculation of the first loss value may be expressed by the following formula: ; in the formula Represents the first distribution probability, where x represents the sequence of input questions, c represents the sequence of inference chain data, and c t is the t-th token in the inference chain; represents the second distribution probability, y represents the sequence of reply data, y t Represents the tth word in the reply data sequence; T c represents the number of first words, T y Indicates the number of second tokens.

[0066] In the present invention, based on the first intermediate model, a first distribution probability corresponding to each first word in a plurality of first words is calculated, and a second distribution probability corresponding to each second word in a plurality of second words is calculated, wherein each first word is a word of the sample inference chain data included in each group of first sample data, and each second word is a word of the sample response data included in each group of first sample data; multiple first logarithms and multiple second logarithms are calculated, wherein the multiple first logarithms are the logarithms of the first distribution probabilities corresponding to the multiple first words, and the multiple first logarithms correspond one-to-one to the multiple first words, and the multiple second logarithms are the logarithms of the second distribution probabilities corresponding to the multiple second words, and the multiple second logarithms correspond one-to-one to the multiple second words; a first mean and a second mean are calculated, wherein the first mean is the mean of the multiple first logarithms, and the second mean is the mean of the multiple second logarithms; and the first loss value is calculated, wherein the first loss value is a weighted value of the first mean and the second mean, so as to obtain the first loss value by calculating the first distribution probability and the second distribution probability.

[0067] In one embodiment, step S33 includes: Step S331: Calculate a first distribution probability corresponding to each first word-gram in a plurality of first word-grams based on the first intermediate model, and calculate a second distribution probability corresponding to each second word-gram in a plurality of second word-grams; Step S332: Calculate, based on the first intermediate model, a third distribution probability corresponding to each third word-gram in a plurality of third words-grams, and calculate a fourth distribution probability corresponding to each fourth word-gram in a plurality of fourth words-grams, wherein each third word-gram is a word-gram of the sample inference chain data included in each set of second sample data, and each fourth word-gram is a word-gram of the sample response data included in each set of second sample data, wherein the plurality of first words-grams correspond one-to-one to the plurality of third words-grams, and the plurality of second words-grams correspond one-to-one to the plurality of fourth words-grams; Step S333: Calculate a plurality of third logarithms and a plurality of fourth logarithms, wherein the plurality of third logarithms are logarithms of a plurality of first quotient values, each of which is the quotient of a first distribution probability of a corresponding first word-unit and a third distribution probability of a third word-unit; and the plurality of fourth logarithms are logarithms of a plurality of second quotient values, each of which is the quotient of a second distribution probability of a corresponding second word-unit and a fourth distribution probability of a fourth word-unit; Step S334: Calculate a third mean and a fourth mean, where the third mean is the mean of a plurality of first expectations, which are expectations of the plurality of third logarithms, and the plurality of first expectations correspond one-to-one with the plurality of third logarithms; the fourth mean is the mean of a plurality of second expectations, which are expectations of the plurality of fourth logarithms, and the plurality of second expectations correspond one-to-one with the plurality of fourth logarithms; Step S335: Calculate the second loss value, where the second loss value is a weighted value of the third mean and the fourth mean.

[0068] The third distribution probability is the distribution probability corresponding to each token in the sample inference chains of the multiple sets of second sample data, and the fourth distribution probability is the distribution probability corresponding to each token in the sample response data of the multiple sets of second sample data. It should be noted that the first and second distribution probabilities are the distribution probabilities for directly generating the inference chain data and response data (i.e., without deep inference), while the third and fourth distribution probabilities are the distribution probabilities for generating the inference chain data and response data after deep inference. Thus, the first and second distribution probabilities are used to determine the distribution of the inference chain data and sample data output by the first intermediate model, thereby obtaining a first loss value that characterizes the situation when the first intermediate model directly generates the inference chain data and response data corresponding to the question. Furthermore, the third and fourth distribution probabilities are used to determine the distribution of the inference chain data and sample data output by the first intermediate model after deep inference. Combining the first and second distribution probabilities yields a second loss value, allowing the second loss value to characterize the effectiveness of the first intermediate model in directly generating the inference chain data and response data corresponding to the question, compared to the first intermediate model generating the inference chain data and response data corresponding to the question after deep inference.

[0069] In some embodiments, after the first distribution probability, the second distribution probability, the third distribution probability, and the fourth distribution probability are calculated, the second loss value may be calculated by calculating the KL divergence (ie, steps S332 to S325).

[0070] The calculation process of the second loss value can be expressed by the following formula: ; in the formula represents the first distribution probability, represents the second distribution probability, represents the third distribution probability, represents the fourth distribution probability, Indicates the calculation of expected value.

[0071] Among them, since the KL divergence is defined as , so the negative sign is retained in the formula to ensure that the distributions are closer when the loss function is minimized.

[0072] In the present invention, the first distribution probability corresponding to each first word in the plurality of first words is calculated based on the first intermediate model, and the second distribution probability corresponding to each second word in the plurality of second words is calculated; the third distribution probability corresponding to each third word in the plurality of third words is calculated based on the first intermediate model, and the fourth distribution probability corresponding to each fourth word in the plurality of fourth words is calculated, each third word is a word of the sample inference chain data included in each group of second sample data, and each fourth word is a word of the sample reply data included in each group of second sample data, the plurality of first words correspond one-to-one to the plurality of third words, and the plurality of second words correspond one-to-one to the plurality of fourth words; a plurality of third logarithms and a plurality of fourth logarithms are calculated, the plurality of third logarithms are the logarithms of a plurality of first quotient values, and the plurality of Each of the multiple first quotient values ​​is the quotient of the first distribution probability of the corresponding first word and the third distribution probability of the third word, the multiple fourth logarithms are the logarithms of the multiple second quotient values, and each of the second quotient values ​​is the quotient of the second distribution probability of the corresponding second word and the fourth distribution probability of the fourth word; a third mean and a fourth mean are calculated, the third mean is the mean of the multiple first expectations, the multiple first expectations are the expectations of the multiple third logarithms, the multiple first expectations correspond one-to-one with the multiple third logarithms, the fourth mean is the mean of the multiple second expectations, the multiple second expectations are the expectations of the multiple fourth logarithms, the multiple second expectations correspond one-to-one with the multiple fourth logarithms; the second loss value is calculated, the second loss value is the weighted value of the third mean and the fourth mean. In this way, the second loss value is calculated by the first distribution probability, the second distribution probability, the third distribution probability, and the fourth distribution probability.

[0073] In some embodiments, the first loss value may be calculated using the following formula:

[0074] in, is the standard definition of KL divergence.

[0075] In one embodiment, after step S3, the method further includes: Step S4: receiving the first question; Step S5: Generate reasoning chain data and response data corresponding to the first question based on the first large model; Step S6: Output the reasoning chain data and response data corresponding to the first question.

[0076] In the present invention, a first question is received; reasoning chain data and response data corresponding to the first question are generated based on the first large model; and the reasoning chain data and response data corresponding to the first question are output. In this way, the reasoning chain data and response data corresponding to the first question are directly generated by the trained first large model without the need for a deep reasoning process, effectively improving reasoning speed and reducing computing resource usage.

[0077] In one embodiment, after step S6, the method further includes: Step S7: receiving feedback information, wherein the feedback information includes the modified reasoning chain data and reply data; Step S8: adding the feedback information to the database; Step S9: When the amount of feedback information stored in the database is greater than a set threshold, generating updated training data based on the feedback information, the updated training data including training questions, training reasoning chain data, and training response data; Step S10: Update the first large model based on the updated training data.

[0078] The above feedback information is provided by users or trainers regarding the inference chain data and response data output by the first large model. It should be noted that the inference chain data and response data output by the first model may be inaccurate. In this case, feedback information regarding the inference chain data and response data output by the first model is received and used to further optimize the first large model. The feedback information includes modified inference chain data and response data, which are used to optimize the first large model.

[0079] The database is used to store feedback information. When feedback information is received, it is added to the database. The database obtains the question data corresponding to the feedback information and saves the question data and feedback information as a set of data, so that the training data can be updated based on the feedback information stored in the database.

[0080] The updated training data is training data generated based on feedback information in the database. The first large model can be updated by updating the training data to improve the inference accuracy of the first large model.

[0081] Among them, the feedback information and question data in the database are saved together in the form of a group. When it is necessary to generate updated training data, the reasoning chain data included in the feedback information in each group of data can be extracted as training reasoning chain data, and the reply data included can be extracted as training reply data, and the question data can be used as training questions to obtain updated training data.

[0082] Furthermore, after the updated training data is generated based on the feedback information, the feedback data stored in the database is deleted and the feedback information is saved again, thereby updating the updated training data to achieve continuous optimization of the first large model.

[0083] In the present invention, feedback information is received, including modified inference chain data and response data; the feedback information is added to a database; and when the amount of feedback information stored in the database exceeds a set threshold, updated training data is generated based on the feedback information, including training questions, training inference chain data, and training response data; and the first large model is updated based on the updated training data. In this way, updated training data is generated based on the feedback, and the first large model is updated using the updated training data to further improve the inference accuracy of the first large model.

[0084] See Figure 4 , Figure 4 This is a structural diagram of a device provided by the present invention for an intelligent computing center cloud platform to perform large-scale model self-distillation and absorption deep reasoning through computing power, such as Figure 4 As shown, the device 400 for the intelligent computing center cloud platform to perform large-model self-distillation and absorption deep reasoning through computing power includes: An acquisition module 401 is configured to acquire multiple sets of initial sample data, each set of initial sample data including a sample question, sample deep reasoning data, sample reasoning chain data, and sample response data; A construction module 402 is configured to construct multiple sets of first sample data and multiple sets of second sample data based on the multiple sets of initial sample data, wherein each set of first sample data in the multiple sets of first sample data includes the sample question, the sample reasoning chain data, and the sample response data, and each set of second sample data in the multiple sets of second sample data includes the sample deep reasoning data, the sample reasoning chain data, and the sample response data; The training module 403 is used to train the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, and the first large model is used to generate reasoning chain data and response data corresponding to the question.

[0085] In one embodiment, the training module 403 includes: a training submodule, configured to train the initial large model based on the multiple sets of first sample data and the multiple sets of second sample data to obtain a first intermediate model; A first calculation submodule, configured to calculate a first loss value corresponding to the first intermediate model based on the multiple groups of first sample data; A second calculation submodule, configured to calculate a second loss value corresponding to the first intermediate model based on the multiple groups of first sample data and the multiple groups of second sample data; A third calculation submodule is configured to calculate a third loss value, where the third loss value is a weighted value of the first loss value and the second loss value; A submodule is set, which is used to set the first intermediate model to the first large model when the third loss value is less than or equal to a preset loss threshold.

[0086] In one embodiment, the first calculation submodule includes: a first calculation unit, configured to calculate, based on the first intermediate model, a first distribution probability corresponding to each first word-gram in a plurality of first word-grams, and calculate a second distribution probability corresponding to each second word-gram in a plurality of second word-grams, wherein each first word-gram is a word-gram of the sample inference chain data included in each set of first sample data, and each second word-gram is a word-gram of the sample response data included in each set of first sample data; a second calculation unit, configured to calculate a plurality of first logarithms and a plurality of second logarithms, wherein the plurality of first logarithms are logarithms of first distribution probabilities corresponding to the plurality of first word-units, and the plurality of first logarithms have a one-to-one correspondence with the plurality of first word-units; and the plurality of second logarithms are logarithms of second distribution probabilities corresponding to the plurality of second word-units, and the plurality of second logarithms have a one-to-one correspondence with the plurality of second word-units; a third calculating unit, configured to calculate a first mean and a second mean, wherein the first mean is a mean of a plurality of first pairs, and the second mean is a mean of a plurality of second pairs; The fourth calculation unit is used to calculate the first loss value, where the first loss value is a weighted value of the first mean and the second mean.

[0087] In one embodiment, the second calculation submodule includes: a fifth calculation unit, configured to calculate a first distribution probability corresponding to each first word-gram in the plurality of first word-grams, and calculate a second distribution probability corresponding to each second word-gram in the plurality of second word-grams based on the first intermediate model; a sixth calculation unit, configured to calculate, based on the first intermediate model, a third distribution probability corresponding to each third word-gram in a plurality of third words-grams, and calculate a fourth distribution probability corresponding to each fourth word-gram in a plurality of fourth words-grams, wherein each third word-gram is a word-gram of the sample inference chain data included in each set of second sample data, and each fourth word-gram is a word-gram of the sample response data included in each set of second sample data, wherein the plurality of first words-grams correspond one-to-one to the plurality of third words-grams, and the plurality of second words-grams correspond one-to-one to the plurality of fourth words-grams; a seventh calculation unit, configured to calculate a plurality of third logarithms and a plurality of fourth logarithms, wherein the plurality of third logarithms are logarithms of a plurality of first quotient values, each of which is a quotient of a first distribution probability of a corresponding first word-unit and a third distribution probability of a third word-unit; and the plurality of fourth logarithms are logarithms of a plurality of second quotient values, each of which is a quotient of a second distribution probability of a corresponding second word-unit and a fourth distribution probability of a fourth word-unit; an eighth calculation unit, configured to calculate a third mean and a fourth mean, wherein the third mean is a mean of a plurality of first expectations, which are expectations of the plurality of third logarithms, and the plurality of first expectations correspond one-to-one with the plurality of third logarithms; and the fourth mean is a mean of a plurality of second expectations, which are expectations of the plurality of fourth logarithms, and the plurality of second expectations correspond one-to-one with the plurality of fourth logarithms; A ninth calculation unit is configured to calculate the second loss value, where the second loss value is a weighted value of the third mean and the fourth mean.

[0088] In one embodiment, the apparatus 400 for performing large-model self-distillation absorption deep reasoning through computing power on the intelligent computing center cloud platform further includes: A first receiving module, configured to receive a first question; A first generating module, configured to generate reasoning chain data and response data corresponding to the first question based on the first large model; The output module is used to output the reasoning chain data and response data corresponding to the first question.

[0089] In one embodiment, the apparatus 400 for performing large-model self-distillation absorption deep reasoning through computing power on the intelligent computing center cloud platform further includes: a second receiving module, configured to receive feedback information, wherein the feedback information includes the modified inference chain data and reply data; An adding module, used for adding the feedback information to a database; a second generating module, configured to generate updated training data based on the feedback information when the amount of feedback information stored in the database is greater than a set amount threshold, the updated training data including training questions, training reasoning chain data, and training response data; An updating module is used to update the first large model based on the updated training data.

[0090] The device provided by the present invention for the intelligent computing center cloud platform to perform self-distillation and absorption of deep reasoning on large models through computing power is capable of realizing the various processes of each embodiment of the method for the above-mentioned intelligent computing center cloud platform to perform self-distillation and absorption of deep reasoning on large models through computing power. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be repeated here.

[0091] It should be noted that the device for the intelligent computing center cloud platform in the present invention to perform large-model self-distillation and absorption deep reasoning through computing power can be a device, or a component, integrated circuit, or chip in an electronic device.

[0092] The present invention also provides an electronic device, see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device includes a memory 501, a processor 502, and a program or instruction stored in the memory 501 and running on the memory 501. When the program or instruction is executed by the processor 502, Figure 1 The corresponding intelligent computing center cloud platform uses computing power to perform large-model self-distillation and absorb deep reasoning in any of the steps in the method embodiment and achieve the same beneficial effects, which will not be repeated here.

[0093] The processor 502 may be a CPU, an ASIC, an FPGA, or a GPU.

[0094] Those skilled in the art will understand that all or part of the steps of the method embodiment for implementing the above-mentioned intelligent computing center cloud platform to perform large-model self-distillation and absorption deep reasoning through computing power can be completed through hardware related to program instructions, and the program can be stored in a readable medium.

[0095] The present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above Figure 1 Any steps in the corresponding embodiments of the method for using computing power on a cloud platform to perform large-scale model self-distillation and deep reasoning can achieve the same technical effect. To avoid repetition, they are not described here. The storage medium is, for example, a read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0096] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The corresponding intelligent computing center cloud platform uses computing power to perform large-scale model self-distillation and absorption of deep reasoning in each process of the embodiment of the method, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0097] The terms "first", "second" and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. In addition, the terms "including" and "having" and any variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. In addition, "and / or" is used in this application to represent at least one of the connected objects, for example A and / or B and / or C, which means comprising seven situations including single A, single B, single C, and both A and B exist, both B and C exist, both A and C exist, and both A, B and C exist.

[0098] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of each embodiment of this application.

[0100] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for performing large-scale model self-distillation and deep reasoning on an intelligent computing center cloud platform through computing power, characterized in that: include: Step S1: Acquire multiple sets of initial sample data, where each set of initial sample data includes a sample question, sample deep reasoning data, sample reasoning chain data, and sample response data; Step S2: constructing multiple sets of first sample data and multiple sets of second sample data based on the multiple sets of initial sample data, wherein each set of first sample data in the multiple sets of first sample data includes the sample question, the sample reasoning chain data, and the sample reply data, and each set of second sample data in the multiple sets of second sample data includes the sample deep reasoning data, the sample reasoning chain data, and the sample reply data; Step S3: Training the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, where the first large model is used to generate reasoning chain data and response data corresponding to the question.

2. The method according to claim 1, wherein The step S3 comprises: Step S31: training the initial large model based on the multiple sets of first sample data and the multiple sets of second sample data to obtain a first intermediate model; Step S32: calculating a first loss value corresponding to the first intermediate model based on the multiple groups of first sample data; Step S33: calculating a second loss value corresponding to the first intermediate model based on the multiple groups of first sample data and the multiple groups of second sample data; Step S34: Calculate a third loss value, where the third loss value is a weighted value of the first loss value and the second loss value; Step S35: When the third loss value is less than or equal to a preset loss threshold, the first intermediate model is set as the first large model.

3. The method according to claim 2, wherein The step S32 includes: Step S321: Calculate a first distribution probability corresponding to each first word-gram in a plurality of first word-grams based on the first intermediate model, and calculate a second distribution probability corresponding to each second word-gram in a plurality of second word-grams, where each first word-gram is a word-gram of the sample inference chain data included in each set of first sample data, and each second word-gram is a word-gram of the sample response data included in each set of first sample data; Step S322: Calculate a plurality of first logarithms and a plurality of second logarithms, wherein the plurality of first logarithms are logarithms of first distribution probabilities corresponding to the plurality of first word-grams, and the plurality of first logarithms have a one-to-one correspondence with the plurality of first word-grams; and the plurality of second logarithms are logarithms of second distribution probabilities corresponding to the plurality of second word-grams, and the plurality of second logarithms have a one-to-one correspondence with the plurality of second word-grams; Step S323: Calculate a first mean and a second mean, where the first mean is the mean of multiple first pairs, and the second mean is the mean of multiple second pairs; Step S324: Calculate the first loss value, where the first loss value is a weighted value of the first mean and the second mean.

4. The method according to claim 2, wherein The step S33 includes: Step S331: Calculate a first distribution probability corresponding to each first word-gram in a plurality of first word-grams based on the first intermediate model, and calculate a second distribution probability corresponding to each second word-gram in a plurality of second word-grams; Step S332: Calculate, based on the first intermediate model, a third distribution probability corresponding to each third word-gram in a plurality of third words-grams, and calculate a fourth distribution probability corresponding to each fourth word-gram in a plurality of fourth words-grams, wherein each third word-gram is a word-gram of the sample inference chain data included in each set of second sample data, and each fourth word-gram is a word-gram of the sample response data included in each set of second sample data, wherein the plurality of first words-grams correspond one-to-one to the plurality of third words-grams, and the plurality of second words-grams correspond one-to-one to the plurality of fourth words-grams; Step S333: Calculate a plurality of third logarithms and a plurality of fourth logarithms, wherein the plurality of third logarithms are logarithms of a plurality of first quotient values, each of which is the quotient of a first distribution probability of a corresponding first word-unit and a third distribution probability of a third word-unit; and the plurality of fourth logarithms are logarithms of a plurality of second quotient values, each of which is the quotient of a second distribution probability of a corresponding second word-unit and a fourth distribution probability of a fourth word-unit; Step S334: Calculate a third mean and a fourth mean, where the third mean is the mean of a plurality of first expectations, which are expectations of the plurality of third logarithms, and the plurality of first expectations correspond one-to-one with the plurality of third logarithms; the fourth mean is the mean of a plurality of second expectations, which are expectations of the plurality of fourth logarithms, and the plurality of second expectations correspond one-to-one with the plurality of fourth logarithms; Step S335: Calculate the second loss value, where the second loss value is a weighted value of the third mean and the fourth mean.

5. The method according to any one of claims 1 to 4, characterized in that After step S3, the method further includes: Step S4: receiving the first question; Step S5: Generate reasoning chain data and response data corresponding to the first question based on the first large model; Step S6: Output the reasoning chain data and response data corresponding to the first question.

6. The method according to claim 5, wherein After step S6, the method further includes: Step S7: receiving feedback information, wherein the feedback information includes the modified reasoning chain data and reply data; Step S8: adding the feedback information to the database; Step S9: When the amount of feedback information stored in the database is greater than a set threshold, generating updated training data based on the feedback information, the updated training data including training questions, training reasoning chain data, and training response data; Step S10: Update the first large model based on the updated training data.

7. A device for an intelligent computing center cloud platform that uses computing power to perform large-scale model self-distillation and absorption deep reasoning, characterized in that: include: An acquisition module, configured to acquire multiple sets of initial sample data, each set of initial sample data including a sample question, sample deep reasoning data, sample reasoning chain data, and sample response data; a construction module, configured to construct multiple sets of first sample data and multiple sets of second sample data based on the multiple sets of initial sample data, wherein each set of first sample data in the multiple sets of first sample data includes the sample question, the sample reasoning chain data, and the sample reply data, and each set of second sample data in the multiple sets of second sample data includes the sample deep reasoning data, the sample reasoning chain data, and the sample reply data; The training module is used to train the initial large model based on the multiple groups of first sample data and the multiple groups of second sample data to obtain a first large model, wherein the first large model is used to generate reasoning chain data and response data corresponding to the question.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for performing large-model self-distillation and absorption deep reasoning by computing power on an intelligent computing center cloud platform as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for performing large-model self-distillation and absorption deep reasoning through computing power on an intelligent computing center cloud platform as described in any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for the intelligent computing center cloud platform to perform large-model self-distillation and absorption deep reasoning through computing power as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Large-model end-to-end distillation deployment method, device and equipment for low-computing-power equipment and medium

    CN120066803A

  • Dialogue model training method, artificial intelligence interview method, computing device, storage medium and program product

    CN120067252A

  • Data analysis model intensified training method based on computing power of intelligent computing center

    CN120218234A

  • Question and answer model training method and device, electronic equipment, storage medium and program product

    CN120297414A

  • Dialogue model training method

    US20240412002A1

Cited By

  • Method and device for intelligent computing center cloud platform to carry out self-evolution depth reasoning of group islands through computing power

    CN121352031A