Model fine tuning method for realizing self-adaptive implicit thinking through computing power of intelligent computing center
Through the computing power resources of the intelligent computing center, deep thinking and reasoning are carried out in the hidden space of the big model, combined with the circular hidden thinking module and hybrid expert strategy, the problems of explicit thinking of the big model are solved, and more efficient and accurate answer output is achieved.
Patent Information
- Application Number
- CN202510397242.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
Existing large models take a long time and consume too much resources during explicit thinking, affecting real-time applications and user experience.
Through the computing power of the intelligent computing center, the large model is fine-tuned by the circular hidden thinking module and hybrid expert strategy, the efficient computing power resources of the intelligent computing center are used to deeply think in the hidden space, and reasoning is combined with the thinking chain and hybrid expert strategy.
On the premise of reducing resource consumption, significantly improve thinking efficiency and depth, and can output more accurate answers faster and improve user experience.
Smart Images

Figure CN120338031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method for fine-tuning a model for adaptive implicit thinking through the computing power of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios for artificial intelligence deep learning model development, model fine-tuning, and model inference). An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results through the processing of information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] With the rapid development of artificial intelligence technology, large models have made remarkable progress in multiple fields such as natural language processing, computer vision, and speech recognition. Large models include: "large language models (LLMs)" and "Multimodal Large Models (MLMs)". These models can learn rich knowledge and patterns from massive data through deep learning algorithms, and thus show high accuracy and efficiency in specific tasks. However, current large models still face some technical challenges in practical applications, especially in explicit thinking during the inference process.
[0008] Specifically, although the explicit thinking process can improve the reliability of the output results, its reasoning process often takes a long time, resulting in a poor user experience in actual applications. For example, in scenarios that require real-time response, the long reasoning time may not meet the user's needs, thus affecting the practicality of the large model and user satisfaction. Especially in the handling of complex problems, the time required for explicit thinking may increase significantly, causing the user to wait too long and reducing the interaction efficiency of the large model.
[0009] In addition, the explicit thinking process is usually accompanied by a large consumption of computing resources. The large model needs to perform complex matrix operations and a large number of memory accesses during the reasoning process, which not only increases the computing cost but also may cause it to not operate effectively in resource-constrained environments.
[0010] In summary, although the large model has demonstrated powerful capabilities in multiple fields, its explicit thinking has problems of long thinking time and excessive resource consumption in the output thinking process. Summary of the Invention
[0011] The present invention provides a model fine-tuning method for realizing adaptive implicit thinking through the computing power of an intelligent computing center to solve the problems existing in the prior art that although the large model has demonstrated powerful capabilities in multiple fields, its explicit thinking has problems of long thinking time and excessive resource consumption in the output thinking process.
[0012] To solve the above technical problems, the present invention is implemented as follows:
[0013] In a first aspect, the present invention provides a model fine-tuning method for realizing adaptive implicit thinking through the computing power of an intelligent computing center, and the method includes:
[0014] Step S1: Receive the data to be analyzed input by the user and the problem related to the data to be analyzed;
[0015] Step S2: Based on the data to be analyzed and the problem, fine-tune the large model to be fine-tuned to obtain a target large model; wherein, the large model to be fine-tuned is based on the in-depth thinking mode of the chain of thought and combines the mixture-of-experts strategy to reason about the data to be analyzed and the problem to obtain an output result, and a cyclic implicit thinking module is set in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute an in-depth thinking process based on the chain of thought and the mixture-of-experts strategy in the implicit space of the large model to be fine-tuned.
[0016] Optionally, the architecture of the large model to be fine-tuned is a transformer architecture, and the large model to be fine-tuned of the transformer architecture includes a plurality of Blocks, and a cyclic implicit thinking module is introduced on each Block;
[0017] The looped implicit thinking module includes: a routing module and multiple sub-expert groups. Each sub-expert group contains at least one expert. Each expert adjusts the pre-trained parameters in the corresponding Block in the form of low-rank parameter increments;
[0018] The looped implicit thinking module is used to perform the following steps:
[0019] Step Sa: After splicing the input information of the corresponding Block and the initial state S0, use the current splicing result as the input of the routing module to obtain the output of the routing module. The output of the routing module is a set of weight values, and the set of weight values is used to indicate the probability of each sub-expert group in the multiple sub-expert groups being selected. Among the set of weight values, the sum of all weight values is 1;
[0020] Step Sb: Select N sub-expert groups in order from the largest to the smallest of each weight value, and use the N sub-expert groups together as the target total expert group. The target total expert group includes: the first sub-expert group, the second sub-expert group,..., the Nth sub-expert group. Among them, N is a positive integer, and the weight values of the first sub-expert group to the Nth sub-expert group decrease in turn;
[0021] Step Sc: Use the current splicing result as the input of each sub-expert group in the N sub-expert groups respectively, and obtain the output of each sub-expert group. Perform a weighted summation operation on the output of each sub-expert group according to their respective corresponding probabilities to obtain the output of the target total expert group;
[0022] Step Sd: After splicing the output of the target total expert group and the input information, use the current splicing result as the input of the routing module again to obtain the output of the routing module; the output of the routing module is a set of new weight values, and the set of new weight values is used to indicate the new probability of each sub-expert group in the multiple sub-expert groups being re-selected. Among the set of new weight values, the sum of all new weight values is 1;
[0023] Step Se: Select new N sub-expert groups in order from the largest to the smallest of each new weight value in the set of new weight values output by the routing module, and use the newly selected new N sub-expert groups together as the new target total expert group. The new target total expert group includes: the new first sub-expert group, the new second sub-expert group,..., the new Nth sub-expert group. Among them, N is a positive integer, and the weight values of the new first sub-expert group to the new Nth sub-expert group decrease in turn;
[0024] Step Sf: Use the current splicing result as the input for each new sub-expert group in the new N sub-expert groups respectively, to obtain the output of each new sub-expert group. Perform a weighted summation operation on the outputs of each new sub-expert group according to their respective corresponding probabilities to obtain the output of the new target total expert group;
[0025] Step Sg: Loop and execute Step Sd, Step Se, and Step Sf until the number of loops reaches a preset threshold. Then, splice the output of the current target total expert group with the input information and input it into a multi-layer perceptron to obtain the output result of the corresponding Block;
[0026] Among them, Step Sa to Step Sc correspond to one deep thinking process in the latent space of the large model to be fine-tuned by the cyclic latent thinking module, which performs deep thinking based on the chain of thought and the expert strategy.
[0027] Optionally, Step S2 includes:
[0028] Step S21: Input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned;
[0029] Among them, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned;
[0030] The type is a code type or an answer type. When the type is the code type, the execution action is code; when the type is the answer type, the execution action is an answer generated for the data to be analyzed and the question;
[0031] Step S22: Determine the items to be evaluated based on the output result. The items to be evaluated include at least one of the following: the answer, the format of the output result, and all the codes;
[0032] Step S23: Score based on the items to be evaluated and the corresponding evaluation criteria to obtain a score value;
[0033] Step S24: Fine-tune the parameters in the cyclic latent thinking module based on the score value. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Step S21 to Step S24 until the fine-tuning is completed to obtain the target large model.
[0034] Optionally, the architecture of the large model to be fine-tuned is a transformer architecture, and Step S2 includes:
[0035] Step S25: Based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module in the large model to be fine-tuned, where the parameters in the cyclic implicit thinking module are parameters related to the in-depth thinking process.
[0036] Optionally, step S23 includes:
[0037] Step S231: When the item to be evaluated is the answer, input the preset answer and the answer into a preset reward function for scoring to obtain the score corresponding to the answer;
[0038] Step S232: When the item to be evaluated is the format of the output result, send the output result to a format evaluation executor for execution to obtain a first execution result, and score the first execution result to obtain the score corresponding to the format of the output result;
[0039] Step S233: When the item to be evaluated is all the codes, sequentially execute all the codes in a code sandbox environment to obtain a second execution result, and score the second execution result to obtain the score corresponding to all the codes.
[0040] In a second aspect, the present invention provides a model fine-tuning device for realizing adaptive implicit thinking through the computing power of an intelligent computing center, and the device includes:
[0041] A receiving module, configured to execute step S1: receive the data to be analyzed input by the user and the question related to the data to be analyzed;
[0042] An execution module, configured to execute step S2: based on the data to be analyzed and the question, fine-tune a large model to be fine-tuned to obtain a target large model; wherein, the large model to be fine-tuned is based on a depth thinking mode of a thinking chain, and combines a mixture of experts strategy to reason about the data to be analyzed and the question to obtain an output result, and a cyclic implicit thinking module is arranged in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute an in-depth thinking process based on the thinking chain and the mixture of experts strategy in the implicit space of the large model to be fine-tuned.
[0043] Optionally, the architecture of the large model to be fine-tuned is a transformer architecture, and the large model to be fine-tuned of the transformer architecture includes multiple Blocks, and a cyclic implicit thinking module is introduced on each Block;
[0044] The looped implicit thinking module includes: a routing module and multiple sub-expert groups. Each sub-expert group contains at least one expert, and each expert adjusts the pre-trained parameters in the corresponding Block in the form of low-rank parameter increments;
[0045] The looped implicit thinking module is used to perform the following steps:
[0046] Step Sa: After splicing the input information of the corresponding Block and the initial state S0, use the current splicing result as the input of the routing module to obtain the output of the routing module. The output of the routing module is a set of weight values, and the set of weight values is used to indicate the probability of each sub-expert group in the multiple sub-expert groups being selected. Among the set of weight values, the sum of all weight values is 1;
[0047] Step Sb: Select N sub-expert groups in order from the largest to the smallest according to each weight value, and jointly use the N sub-expert groups as the target total expert group. The target total expert group includes: the first sub-expert group, the second sub-expert group,..., the Nth sub-expert group. Among them, N is a positive integer, and the weight values of the first sub-expert group to the Nth sub-expert group decrease in sequence;
[0048] Step Sc: Use the current splicing result as the input of each sub-expert group in the N sub-expert groups respectively, and obtain the output of each sub-expert group. Perform a weighted sum operation on the outputs of each sub-expert group according to their respective corresponding probabilities to obtain the output of the target total expert group;
[0049] Step Sd: After splicing the output of the target total expert group and the input information, use the current splicing result as the input of the routing module again to obtain the output of the routing module. The output of the routing module is a set of new weight values, and the set of new weight values is used to indicate the new probability of each sub-expert group in the multiple sub-expert groups being re-selected. Among the set of new weight values, the sum of all new weight values is 1;
[0050] Step Se: Re-select new N sub-expert groups in order from the largest to the smallest according to each new weight value in the set of new weight values output by the routing module, and jointly use the re-selected new N sub-expert groups as the new target total expert group. The new target total expert group includes: the new first sub-expert group, the new second sub-expert group,..., the new Nth sub-expert group. Among them, N is a positive integer, and the weight values of the new first sub-expert group to the new Nth sub-expert group decrease in sequence;
[0051] Step Sf: Use the current splicing result as the input of each new sub-expert group in the new N sub-expert groups respectively, obtain the output of each new sub-expert group, and perform a weighted summation operation on the outputs of each new sub-expert group according to their corresponding probabilities to obtain the output of the new target total expert group;
[0052] Step Sg: Loop through Step Sd, Step Se, and Step Sf until the number of loops reaches a preset threshold. Then, splice the output of the current target total expert group with the input information and input it into a multi-layer perceptron to obtain the output result of the corresponding Block;
[0053] Among them, Step Sa to Step Sc correspond to one deep thinking process in the latent space of the large model to be fine-tuned by the cyclic latent thinking module, which is based on the chain of thought and the expert strategy.
[0054] Optionally, Step S2 includes:
[0055] Step S21: Input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned;
[0056] Among them, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned;
[0057] The type is a code type or an answer type. When the type is the code type, the execution action is the code; when the type is the answer type, the execution action is the answer generated for the data to be analyzed and the question;
[0058] Step S22: Determine the item to be evaluated based on the output result. The item to be evaluated includes at least one of the following: the answer, the format of the output result, and all the codes;
[0059] Step S23: Score based on the item to be evaluated and the corresponding evaluation criteria to obtain a score value;
[0060] Step S24: Fine-tune the parameters in the cyclic latent thinking module based on the score value. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Step S21 to Step S24 until the fine-tuning is completed to obtain the target large model.
[0061] Optionally, the architecture of the large model to be fine-tuned is a transformer architecture, and Step S2 includes:
[0062] Step S25: Based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module in the large model to be fine-tuned, where the parameters in the cyclic implicit thinking module are parameters related to the in-depth thinking process.
[0063] Optionally, the step S23 includes:
[0064] Step S231: When the item to be evaluated is the answer, input the preset answer and the answer into a preset reward function for scoring to obtain the score corresponding to the answer;
[0065] Step S232: When the item to be evaluated is the format of the output result, send the output result to a format judgment executor for execution to obtain a first execution result, and score the first execution result to obtain the score corresponding to the format of the output result;
[0066] Step S233: When the item to be evaluated is all the codes, sequentially execute all the codes in a code sandbox environment to obtain a second execution result, and score the second execution result to obtain the score corresponding to all the codes.
[0067] In a third aspect, the present invention provides a server, including: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, it implements the steps of the model fine-tuning method for adaptive implicit thinking through the computing power of the intelligent computing center as described in the first aspect above.
[0068] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the model fine-tuning method for adaptive implicit thinking through the computing power of the intelligent computing center as described in the first aspect above.
[0069] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, it implements the steps of the model fine-tuning method for adaptive implicit thinking through the computing power of the intelligent computing center as described in the first aspect above.
[0070] In the present invention, by leveraging the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the hidden space. Specifically, the intelligent computing center, with its advanced hardware facilities and optimized computing architecture, can support parallel processing and efficient computing resource scheduling, enabling the in-depth thinking and reasoning process of the large model to proceed rapidly in the hidden space. Moreover, the method of hidden space reasoning can not only reduce resource consumption but also significantly improve the efficiency of thinking and deepen the depth of thinking, thereby outputting more accurate answers in a shorter time.
[0071] Furthermore, based on the mixture-of-experts strategy, the large model can flexibly select thinking modules (experts) in the hidden space. Through continuous training of the large model and fine-tuning of low-rank parameters, each expert can gradually specialize and focus on specific subtasks or feature spaces. Based on the mixture-of-experts strategy, the large model can adaptively activate the optimal combination of thinking modules according to the characteristics of the input problem and input data, improving the depth of thinking and the accuracy of problem analysis.
[0072] In summary, based on the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the hidden space. Combined with the mixture-of-experts strategy, the large model can flexibly select thinking modules (experts) in the hidden space, and each thinking module focuses on processing specific subtasks or feature spaces. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth on the premise of reducing resource consumption, output more accurate answers in less time, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0074] Figure 1 is a flowchart of a model fine-tuning method for realizing adaptive hidden thinking through the computing power of the intelligent computing center provided by the present invention;
[0075] Figure 2 is a structural block diagram of a model fine-tuning device for realizing adaptive hidden thinking through the computing power of the intelligent computing center provided by the present invention;
[0076] Figure 3 is a schematic diagram of the working process of a cyclic hidden thinking module provided by the present invention;
[0077] Figure 4 is a flowchart of the working process of a cyclic hidden thinking module provided by the present invention;
[0078] Figure 5 A flowchart showing how the target large model according to the present invention obtains an output through backtracking, reflection, and refinement based on an input;
[0079] Figure 6 A schematic diagram showing the working mechanism of sequence prediction based on implicit thinking of the large model provided by the present invention;
[0080] Figure 7 A structural block diagram of a model fine-tuning device for achieving adaptive implicit thinking through the computing power of an intelligent computing center provided by the present invention;
[0081] Figure 8 A schematic diagram of the structure of the electronic device of the present invention. Detailed implementation manners
[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0083] First, the technical terms related to the present invention will be briefly described below.
[0084] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0085] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive index for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super
[0086] The "Network Power (NP)" described in the present invention refers to: It is an indication of the data transmission capacity of computing power facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0087] The "Storage Power (SP)" described in the present invention refers to: It is the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server-internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0088] The "computing power infrastructure" described in the present invention refers to: A new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize centralized computing, storage, transmission, and application of information.
[0089] The "new type of information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0090] The "computing power" described in the present invention includes: "general computing power", "intelligent computing power", and "super computing power".
[0091] The "general computing power" described in the present invention refers to: The computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0092] The "intelligent computing power" described in the present invention refers to: For various artificial intelligence innovation applications, a computing platform is scaled and deployed based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0093] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0094] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model fine-tuning, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0095] The "intelligent computing center" described in the present invention includes, but is not limited to, the "intelligent computing center".
[0096] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0097] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0098] The "supercomputing center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0099] The "computing power resources" described in the present invention refers to: technologies and facilities required for the development of the digital society with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0100] The "large model" described in the present invention refers to the large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language. Through fine-tuning with a large amount of text data, it can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0101] Figure 1 A model fine-tuning method for realizing adaptive implicit thinking through intelligent computing center computing power according to the present invention is shown. Figure 1 As shown, the method includes:
[0102] Step S1: receiving data to be analyzed and questions related to the data to be analyzed input by a user;
[0103] Step S2: Based on the data and problems to be analyzed, the large model to be fine-tuned is fine-tuned to obtain the target large model; wherein, the large model to be fine-tuned is based on the deep thinking method of the thinking chain, and is combined with the hybrid expert strategy to reason on the data and problems to be analyzed to obtain the output result, and a cyclic implicit thinking module is provided in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute the deep thinking process based on the thinking chain and the hybrid expert strategy in the latent space of the large model to be fine-tuned.
[0104] It should be noted that in step S1, the data to be analyzed and the related questions provided by the user are first received. The data to be analyzed can be any form of information, such as text, images, audio or structured data, and the related questions are the specific information or insights that the user hopes to obtain by analyzing these data. The key to this step is to ensure that the system can accurately understand the user's input, including the format, content and intention of the data. In addition, data preprocessing, such as data cleaning, format conversion and feature extraction, can be performed to prepare for subsequent analysis and reasoning.
[0105] In step S2, the system will fine-tune the large model to be fine-tuned according to the received data to be analyzed and related questions to generate a target large model for a specific task. The fine-tuning process is the process of fine-tuning the existing large model to better adapt it to specific data sets and task requirements. In this process, the large model to be fine-tuned will use its built-in thinking chain mechanism and combine it with the hybrid expert strategy to conduct in-depth thinking and reasoning, analyze the relationship between the data to be analyzed and the questions, and generate preliminary output results.
[0106] It should be noted that a loop implicit thinking module (such as Figure 2 The core function of this module is to perform a deep thinking process based on thought chains and hybrid expert strategies in the latent space. The cyclic implicit thinking module allows the model to iterate multiple times during the reasoning process in order to deeply explore the potential information and logical relationships in the data. Unlike traditional explicit thinking, cyclic implicit thinking can process information in the latent space in a more complex and detailed manner, thereby improving the depth and accuracy of reasoning. In addition, cyclic implicit thinking can also reduce resource consumption because it can efficiently perform calculations in the latent space without frequently calling external resources.
[0107] It should be noted that the Mixture of Experts (MoE) is a machine learning model architecture design method. Its core idea is to dynamically combine multiple specialized sub-models (experts) so that the overall system can flexibly allocate tasks and improve performance according to the characteristics of the input data. The Mixture of Experts strategy consists of the following components: Expert Network (Experts): Multiple independent sub-models (such as neural networks), and each expert focuses on learning a specific pattern, domain, or feature subspace of the input data (such as syntactic analysis, image texture recognition, mathematical reasoning, etc.). Gating Network: A learnable routing mechanism that dynamically calculates the weight distribution based on the current input data to determine which experts to activate and assign task weights.
[0108] In summary, based on the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the latent space. Combining the Mixture of Experts strategy, the large model can flexibly select thinking modules (experts) in the latent space. Each thinking module focuses on processing specific sub-tasks or feature spaces. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth while reducing resource consumption, output more accurate answers in less time, and enhance the user experience.
[0109] Now, taking a cyclic latent thinking module as an example, its working method and data flow are illustrated (as Figure 3 shown).
[0110] First of all, it should be noted that the architecture of the large model to be fine-tuned is the Transformer architecture. The large model to be fine-tuned in the Transformer architecture includes multiple Blocks, and a cyclic latent thinking module is introduced on each Block; the cyclic latent thinking module includes: a routing module and multiple sub-expert groups, each sub-expert group contains at least one expert, and each expert is equivalent to adjusting the pre-trained parameters in the corresponding Block in the form of low-rank parameter increments.
[0111] It should be noted that in the Mixture of Experts strategy, each expert adjusts the pre-trained parameters of the Transformer Block through low-rank parameter increments (Low-Rank Parameter Adaptation). Assuming that the pre-trained parameters of the original Block are R0 and the low-rank increment introduced by the expert is denoted as Δ, the adjusted parameter Ri can be expressed as: Ri = (1 + Δ)R0.
[0112] It should be noted that the routing module has a Softmax-MLP structure. In the mixture-of-experts strategy, the Softmax-MLP of the routing module (Router) is a specifically designed neural network structure for dynamically allocating the weights of input data to different experts. The Softmax-MLP consists of two parts: MLP (Multilayer Perceptron): responsible for mapping the input features to the expert weight space; Softmax activation function: converting the output of the MLP into a probability distribution and ensuring that the sum of all probability values (weight values) is 1.
[0113] The cyclic implicit thinking module is used to perform the following steps, as Figure 4 shown. And it should be noted that the following takes one cyclic implicit thinking module as an example to illustrate its working process. Each cyclic implicit thinking module on each Block will execute the following working process to perform a deep thinking process based on the chain of thought and the mixture-of-experts strategy in the implicit space of the large model to be fine-tuned.
[0114] Step Sa: After concatenating the input information of the corresponding Block and the initial state S0, use the current concatenation result as the input of the routing module to obtain the output of the routing module. The output of the routing module is a set of weight values, and a set of weight values is used to indicate the probability of each sub-expert group among multiple sub-expert groups being selected. In a set of weight values, the sum of all weight values is 1;
[0115] Step Sb: According to the order of each weight value from largest to smallest, sequentially select N sub-expert groups, and use the N sub-expert groups together as the target total expert group. The target total expert group includes: the first sub-expert group, the second sub-expert group,..., the Nth sub-expert group, where N is a positive integer, and the weight values of the first sub-expert group to the Nth sub-expert group decrease in turn;
[0116] Step Sc: Use the current concatenation result as the input of each sub-expert group among the N sub-expert groups respectively, and obtain the output of each sub-expert group. Perform a weighted summation operation on the outputs of each sub-expert group according to their respective corresponding probabilities to obtain the output of the target total expert group;
[0117] Step Sd: After concatenating the output of the target total expert group and the input information, use the current concatenation result as the input of the routing module again to obtain the output of the routing module; the output of the routing module is a new set of weight values, and a new set of weight values is used to indicate the new probability of each sub-expert group among multiple sub-expert groups being reselected. In a new set of weight values, the sum of all new weight values is 1;
[0118] Step Se: According to the order from large to small of a new set of weight values output by the routing module, reselect a new N sub-expert groups in sequence, and use the reselected new N sub-expert groups together as the new target total expert group. The new target total expert group includes: the new first sub-expert group, the new second sub-expert group, ……, the new Nth sub-expert group. Here, N is a positive integer, and the weight values of the new first sub-expert group to the new Nth sub-expert group decrease in sequence;
[0119] Step Sf: Take the current splicing result as the input of each new sub-expert group in the new N sub-expert groups respectively, obtain the output of each new sub-expert group, and perform a weighted summation operation on the outputs of each new sub-expert group according to their respective corresponding probabilities to obtain the output of the new target total expert group;
[0120] Step Sg: Loop and execute Step Sd, Step Se, and Step Sf until the number of loops reaches a preset threshold. Then, splice the output of the current target total expert group with the input information and input it into the multi-layer perceptron to obtain the output result of the corresponding Block;
[0121] Among them, Steps Sa to Sc correspond to one deep thinking process in the hidden space of the large model to be fine-tuned by the loop hidden thinking module during the deep thinking process based on the chain of thought and expert strategy.
[0122] It should be noted that Figure 4 The method shown realizes deep thinking through a dynamic routing and loop optimization mechanism. Specifically, the input information of the Block is spliced with the initial state, and the sub-expert group selection weights are generated by the routing module (the sum of all weight values is 1). Select the Top-N sub-expert groups in descending order of weights; the selected sub-expert groups process the splicing of the initial state and the input in parallel, obtain their respective outputs, and perform a weighted summation on their respective outputs according to the probabilities corresponding to the sub-expert groups to obtain the total output of the Top-N sub-expert groups. Splice the total output with the input information to trigger the routing module to update the weights and reselect the sub-expert groups, and loop to execute the state optimization; after reaching the preset number of loops, splice the final state with the input information and output the result through the MLP. The whole process realizes efficient and low-power deep reasoning in the hidden space through weight dynamic adjustment, state iterative transmission, and expert collaborative optimization.
[0123] It should be noted that assuming there are a total of L sub-experts, they can be divided into M groups (equivalent to the multiple sub-expert groups included in the above loop hidden thinking module), and they can be evenly distributed or not evenly distributed, but it is necessary to ensure that there is at least 1 sub-expert in each sub-expert group. Optionally, an allocation upper limit value can be set, that is, the number of sub-experts in each sub-expert group cannot exceed K, where L, M, and K are all positive integers.
[0124] As shown Figure 3 below, assume that R 1 , R 2 , and R 3 form the first sub-expert group. There are three experts, R 1 , R 2 , and R 3 within the first sub-expert group. The internal data flow is such that the input of the first sub-expert group (the concatenation result of state S0 and the input) is the input of R 1 , the output of R 1 is the input of R 2 , the output of R 2 is the input of R 3 , and the output of R 3 is the output of the first sub-expert group.
[0125] It should be noted that steps Sa to Sc correspond to a single in-depth thinking process in the latent space of the large model to be fine-tuned by the cyclic implicit thinking module, which is based on the chain of thought and expert strategy. And as shown Figure 5 and Figure 6 below, the fine-tuned target large model has the ability to obtain an accurate output based on the input (question and / or data to be analyzed) through an in-depth thinking process (backtracking, reflection, refinement) in the latent space.
[0126] Also, it should be noted that in the initial stage, all sub-expert groups have homogenized capabilities. With the guidance of the routing module and continuous fine-tuning of the low-rank parameter increment (based on LoRA technology) during the training process, each expert gradually focuses on specific sub-tasks or feature spaces. In the cyclic iteration, the routing module dynamically selects the current optimal Top-N expert combination based on the input characteristics, and through continuous cycling, realizes the deepening of progressive thinking. Eventually, within the preset number of cycles, it can adaptively activate the optimal module combination according to the characteristics of the input question and input data, improving the thinking depth and the accuracy of problem analysis.
[0127] In a possible implementation, step S2 includes:
[0128] Step S21: Input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned, and obtain the output result of the large model to be fine-tuned;
[0129] Step S22: Determine the items to be evaluated based on the output result; the items to be evaluated include at least one of the following: the answer, the format of the output result, and all the codes;
[0130] Step S23: Score based on the items to be evaluated and the corresponding evaluation criteria to obtain a score value;
[0131] Step S24: Based on the score, fine-tune the parameters in the cyclic implicit thinking module. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Steps S21 to S24 until the fine-tuning is completed, and then obtain the target large model.
[0132] In Step S21, in addition to the data to be analyzed and the question, the output format also needs to be input into the large model. The output format defines the structure of the output result of the large model, including: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned; and the type is either the code type or the answer type. When the type is the code type, the execution action is the code; when the type is the answer type, the execution action is the answer generated for the data to be analyzed and the question.
[0133] In Steps S22 to S24, it is necessary to determine the items to be evaluated based on the output result; the items to be evaluated include at least one of the following: the answer, the format of the output result, and all the codes. Then, score based on the items to be evaluated and the corresponding evaluation criteria to obtain the score; and then, based on the score, fine-tune the parameters in the cyclic implicit thinking module. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue the fine-tuning until the fine-tuning is completed, and then obtain the target large model.
[0134] In a possible implementation, Step S23 includes: Step S231: When the item to be evaluated is the answer, input the preset answer and the answer into the preset reward function for scoring to obtain the score corresponding to the answer; Step S232: When the item to be evaluated is the format of the output result, send the output result to the format judgment executor for execution to obtain the first execution result, and score the first execution result to obtain the score corresponding to the format of the output result; Step S233: When the item to be evaluated is all the codes, sequentially execute all the codes in the code sandbox environment to obtain the second execution result, and score the second execution result to obtain the score corresponding to all the codes.
[0135] It should be noted that in Step S23, the system carefully scores the output result of the large model to be fine-tuned according to the previously determined items to be evaluated and the corresponding evaluation criteria. The specific scoring process is divided into three sub-steps, which are processed for different items to be evaluated respectively.
[0136] Specifically, when the item to be evaluated is the answer, the preset answer and the answer are input into the preset reward function for scoring to obtain the score corresponding to the answer. The preset correct answer can be the standard answer provided by experts or the reference answer fine-tuned from historical data. The scoring process usually involves the following aspects: Similarity calculation: The system can use text similarity algorithms (such as cosine similarity, etc.) to evaluate the similarity between the generated answer and the preset answer; Reward function: The preset reward function will calculate a score based on the similarity score, usually a normalized score (for example, from 0 to 1 or from 0 to 10), which is used to reflect the quality of the answer. The reward function can be designed according to the requirements of specific tasks to ensure that it can effectively evaluate the accuracy and integrity of the answer. Through this process, the system can assign a quantitative score to the generated answer to reflect its quality.
[0137] When the item to be evaluated is the format of the output result, the output result is sent to the format judgment executor for execution to obtain the first execution result, and the first execution result is scored to obtain the score corresponding to the format of the output result. That is to say, the system sends the output result to a dedicated format judgment executor, which is responsible for verifying the format of the output result. The format of the output result includes its structure, type, and execution actions, etc., to ensure that the output meets the predefined format requirements.
[0138] When the item to be evaluated is all the code, all the code is sequentially executed in the code sandbox environment to obtain the second execution result, and the second execution result is scored to obtain the score corresponding to all the code. That is to say, the system sequentially executes all the generated code in the code sandbox environment to obtain the second execution result, which is used to evaluate the correctness and functionality implementation of the code.
[0139] It should be noted that the system will sequentially execute all the code in a secure code sandbox environment to avoid affecting the main system or causing security issues. Among them, the code sandbox environment provides an isolated execution environment, allowing the system to test the code and functions without affecting other system components.
[0140] Thus, through steps S231, S232, and S233, the system can comprehensively evaluate the output results of the large model to be fine-tuned. This evaluation mechanism not only ensures the accuracy and integrity of the generated answers but also verifies the format of the output results and the effectiveness of the code. Finally, the system will provide a basis for the subsequent fine-tuning process based on these scores to ensure the continuous optimization of the performance of the large model to be fine-tuned.
[0141] In step S24, the parameters in the cyclic implicit thinking module can be fine-tuned based on the scores. After the fine-tuning, it is judged whether the fine-tuning is completed. If the fine-tuning is not completed, steps S21 to S24 are continuously executed until the fine-tuning is completed, and then the target large model is obtained.
[0142] It should be noted that the corresponding output results can be obtained based on each set of sample data in a fine-tuning batch (a set of sample data includes the above-mentioned data to be analyzed, questions, and output formats), and scores can be obtained based on each output result, and then the comprehensive average score can be obtained. The parameters in the cyclic implicit thinking module are fine-tuned based on the comprehensive average score. The condition for the end of the fine-tuning can be that the comprehensive average score reaches a preset threshold, or multiple batches are completed (the fine-tuning is finished).
[0143] And in a specific application scenario, the fine-tuning can be based on reinforcement learning technology. Reinforcement Learning (RL) is a method of learning optimal strategies by interacting with the environment. The parameters in the cyclic implicit thinking module are fine-tuned through reinforcement learning technology (as shown by Reinforcement Learning in Figure 2 ), and the fine-tuning process can be more flexible and efficient, and can adaptively find the optimal parameter adjustment strategy, rather than relying on fixed rules or manual adjustment. This can improve the performance of the model to be fine-tuned and finally obtain a more powerful target large model.
[0144] In a specific application scenario, supervised fine-tuning (SFT) can also be used to fine-tune the parameters in the cyclic implicit thinking module. The present invention does not limit this, and the fine-tuning method can be flexibly selected according to actual needs.
[0145] In a possible implementation manner, the architecture of the large model to be fine-tuned is a transformer architecture. Step S2 includes: Step S25: Based on the LORA technology, freeze the pre-fine-tuning parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module in the large model to be fine-tuned, where the parameters in the cyclic implicit thinking module are parameters related to the in-depth thinking process.
[0146] Among them, the LORA (Low-Rank Adaptation) technology is a technology for fine-tuning large-scale pre-fine-tuned models, aiming to efficiently adjust model parameters in the form of low-rank matrices. The basic idea of LORA is to capture the features of specific tasks by introducing low-rank adapters while keeping the original model parameters unchanged.
[0147] The working principle of LORA is as follows: Freeze pre-fine-tuning parameters: During the fine-tuning process, first freeze the pre-fine-tuning parameters of the large model, which means these parameters do not change during the fine-tuning process. This approach can avoid overfitting when fine-tuning on a small dataset while retaining the knowledge learned by the pre-fine-tuning model on a large-scale dataset; Introduce low-rank adapters: Insert low-rank adapters in some layers of the model. The parameters of these adapters are fine-tunable, and the adaptability of the model is achieved by fine-tuning these parameters; Parameter update: During the fine-tuning process, only update the parameters of the low-rank adapters, rather than the parameters of the entire model. This method greatly reduces the number of parameters that need to be fine-tuned, thereby improving the fine-tuning efficiency and saving computing resources.
[0148] In the present invention, by leveraging the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the latent space. Specifically, the intelligent computing center utilizes advanced hardware facilities and optimized computing architectures, capable of supporting parallel processing and efficient computing resource scheduling, enabling the in-depth thinking and reasoning process of the large model to proceed rapidly in the latent space; moreover, the way of latent space reasoning can not only reduce resource consumption, but also significantly improve the efficiency of thinking and deepen the depth of thinking, thereby outputting more accurate answers in a shorter time.
[0149] And based on the mixture of experts strategy, the large model can flexibly select thinking modules (experts) in the latent space. Through continuous training of the large model and fine-tuning of low-rank parameters, each expert can gradually specialize in focusing on specific subtasks or feature spaces. Based on the mixture of experts strategy, the large model can adaptively activate the optimal module combination according to the characteristics of the input problem and input data, improving the depth of thinking and the accuracy of problem analysis.
[0150] In summary, based on the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the latent space. Combining the mixture of experts strategy, the large model can flexibly select thinking modules (experts) in the latent space, and each thinking module focuses on processing specific subtasks or feature spaces. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth of thinking on the premise of reducing resource consumption, output more accurate answers in less time, and enhance the user experience.
[0151] Figure 7 Shows a model fine-tuning device for realizing adaptive implicit thinking through the computing power of an intelligent computing center according to the present invention. The device 70 includes:
[0152] A receiving module 701, configured to execute step S1: Receive the data to be analyzed input by the user and the question related to the data to be analyzed;
[0153] An execution module 702, configured to execute step S2: fine-tune the large model to be fine-tuned based on the data and questions to be analyzed, to obtain a target large model; wherein, the large model to be fine-tuned is based on the deep thinking mode of the chain of thought, and combines the mixture-of-experts strategy to reason about the data and questions to be analyzed, to obtain an output result, and a cyclic implicit thinking module is set in the large model to be fine-tuned, and the cyclic implicit thinking module is configured to execute a deep thinking process based on the chain of thought and the mixture-of-experts strategy in the implicit space of the large model to be fine-tuned.
[0154] In a possible implementation manner, the architecture of the large model to be fine-tuned is a Transformer architecture, and the large model to be fine-tuned of the Transformer architecture includes a plurality of Blocks, and a cyclic implicit thinking module is introduced on each Block;
[0155] The cyclic implicit thinking module includes: a routing module and a plurality of sub-expert groups, each sub-expert group contains at least one expert, and each expert is equivalent to adjusting the pre-trained parameters in the corresponding Block in the form of low-rank parameter increments;
[0156] The cyclic implicit thinking module is configured to execute the following steps:
[0157] Step Sa: After splicing the input information and the initial state S0 of the corresponding Block, use the current splicing result as the input of the routing module, to obtain the output of the routing module, and the output of the routing module is a set of weight values, and the set of weight values is used to indicate the probability of each sub-expert group in the plurality of sub-expert groups being selected, and in the set of weight values, the sum of all weight values is 1;
[0158] Step Sb: According to the order of the weight values from large to small, sequentially select N sub-expert groups, and use the N sub-expert groups together as the target total expert group, and the target total expert group includes: a first sub-expert group, a second sub-expert group,..., an Nth sub-expert group, wherein, N is a positive integer, and the weight values of the first sub-expert group to the Nth sub-expert group decrease in sequence;
[0159] Step Sc: Use the current splicing result as the input of each sub-expert group in the N sub-expert groups respectively, and obtain the output of each sub-expert group, and perform a weighted summation operation on the output of each sub-expert group according to the corresponding probability, to obtain the output of the target total expert group;
[0160] Step Sd: After concatenating the output of the target general expert group and the input information, use the current concatenation result as the input of the routing module again to obtain the output of the routing module; the output of the routing module is a set of new weight values, and the set of new weight values is used to indicate the new probabilities of each sub-expert group in the multiple sub-expert groups being reselected. In the set of new weight values, the sum of all new weight values is 1;
[0161] Step Se: According to the order of each new weight value in the set of new weight values output by the routing module from large to small, reselect N new sub-expert groups in sequence, and use the reselected N new sub-expert groups together as the new target general expert group. The new target general expert group includes: a new first sub-expert group, a new second sub-expert group,..., a new Nth sub-expert group. Among them, N is a positive integer, and the weight values of the new first sub-expert group to the new Nth sub-expert group decrease in sequence;
[0162] Step Sf: Use the current concatenation result as the input of each new sub-expert group in the new N sub-expert groups respectively to obtain the output of each new sub-expert group, and perform a weighted summation operation on the output of each new sub-expert group according to their respective corresponding probabilities to obtain the output of the new target general expert group;
[0163] Step Sg: Loop to execute Step Sd, Step Se, and Step Sf until the number of loops reaches a preset threshold, then concatenate the output of the current target general expert group with the input information and input it into a multi-layer perceptron to obtain the output result of the corresponding Block;
[0164] Among them, Step Sa to Step Sc correspond to one deep thinking process in the latent space of the large model to be fine-tuned by the cyclic latent thinking module during the deep thinking process based on the chain of thought and the expert strategy.
[0165] In a possible implementation manner, Step S2 includes:
[0166] Step S21: Input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned;
[0167] Among them, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned;
[0168] The type is either a code type or an answer type. When the type is a code type, the execution action is the code; when the type is an answer type, the execution action is the answer generated for the data to be analyzed and the question.
[0169] Step S22: Determine the item to be evaluated based on the output result. The item to be evaluated includes at least one of the following: the answer, the format of the output result, and all the code.
[0170] Step S23: Score based on the item to be evaluated and the corresponding evaluation criteria to obtain a score value.
[0171] Step S24: Based on the score value, fine-tune the parameters in the cyclic implicit thinking module. After the fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Steps S21 to S24 until the fine-tuning is completed to obtain the target large model.
[0172] Optionally, the architecture of the large model to be fine-tuned is the Transformer architecture. Step S2 includes:
[0173] Step S25: Based on the LoRA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module of the large model to be fine-tuned. Among them, the parameters in the cyclic implicit thinking module are the parameters related to the in-depth thinking process.
[0174] In a possible implementation, Step S23 includes:
[0175] Step S231: When the item to be evaluated is the answer, input the preset answer and the answer into the preset reward function for scoring to obtain the score value corresponding to the answer.
[0176] Step S232: When the item to be evaluated is the format of the output result, send the output result to the format judgment executor for execution to obtain the first execution result, and score the first execution result to obtain the score value corresponding to the format of the output result.
[0177] Step S233: When the item to be evaluated is all the code, sequentially execute all the code in the code sandbox environment to obtain the second execution result, and score the second execution result to obtain the score value corresponding to all the code.
[0178] In the present invention, by leveraging the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the latent space. Specifically, the intelligent computing center, with advanced hardware facilities and optimized computing architectures, can support parallel processing and efficient computing resource scheduling, enabling the in-depth thinking and reasoning process of the large model to proceed rapidly in the latent space. Moreover, the method of latent space reasoning can not only reduce resource consumption but also significantly improve the efficiency of thinking and deepen the depth of thinking, thereby outputting more accurate answers in a shorter time.
[0179] Furthermore, based on the mixture-of-experts strategy, the large model can flexibly select thinking modules (experts) in the latent space. Through continuous training of the large model and fine-tuning of low-rank parameters, each expert can gradually specialize in focusing on specific subtasks or feature spaces. Based on the mixture-of-experts strategy, the large model can adaptively activate the optimal module combination according to the characteristics of the input problem and input data, improving the depth of thinking and the accuracy of problem analysis.
[0180] In summary, based on the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the latent space. Combining the mixture-of-experts strategy, the large model can flexibly select thinking modules (experts) in the latent space, with each thinking module focusing on processing specific subtasks or feature spaces. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth on the premise of reducing resource consumption, output more accurate answers in less time, and enhance the user experience.
[0181] Please refer to Figure 8 , the present invention also provides an electronic device 80, including a processor 801, a memory 802, and a computer program stored on the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, it implements the steps of the above-mentioned model fine-tuning method for adaptive implicit thinking through the computing power of the intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0182] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned model fine-tuning method for adaptive implicit thinking through the computing power of the intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium can be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0183] The present invention also provides a computer program product, including computer instructions which, when executed by a processor, implement the steps of the above-mentioned method for adaptively fine-tuning the implicit thinking model through the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0184] It should be noted that, in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method described in the present invention.
[0186] The present invention has been described above in conjunction with the accompanying drawings, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all of them belong to the protection scope of the present invention.
Claims
1. A method for fine-tuning a model that realizes adaptive implicit thinking through the computing power of an intelligent computing center, characterized in that, The method includes: Step S1: Receive the data to be analyzed input by the user and the question related to the data to be analyzed; Step S2: Based on the data to be analyzed and the question, fine-tune the large model to be fine-tuned to obtain a target large model; wherein, the large model to be fine-tuned is based on the deep thinking mode of the chain of thought and combines the mixture-of-experts strategy to reason about the data to be analyzed and the question to obtain an output result. A cyclic hidden thinking module is set in the large model to be fine-tuned, and the cyclic hidden thinking module is used to execute a deep thinking process based on the chain of thought and the mixture-of-experts strategy in the hidden space of the large model to be fine-tuned.
2. The method according to claim 1, wherein the architecture of the large model to be fine-tuned is a Transformer architecture, and the large model to be fine-tuned of the Transformer architecture includes multiple Blocks, and a cyclic hidden thinking module is introduced on each Block; the cyclic hidden thinking module includes: a routing module and multiple sub-expert groups, each sub-expert group contains at least one expert, and each expert is equivalent to adjusting the pre-trained parameters in the corresponding Block in the form of low-rank parameter increments; the cyclic hidden thinking module is used to execute the following steps: Step Sa: After splicing the input information and the initial state S0 of the corresponding Block, use the current splicing result as the input of the routing module to obtain the output of the routing module. The output of the routing module is a set of weight values, and the set of weight values is used to indicate the probability of each sub-expert group in the multiple sub-expert groups being selected. In the set of weight values, the sum of all weight values is 1; Step Sb: According to the order of the weight values from large to small, select N sub-expert groups in sequence, and use the N sub-expert groups together as the target total expert group. The target total expert group includes: the first sub-expert group, the second sub-expert group,..., the Nth sub-expert group, where N is a positive integer, and the weight values of the first sub-expert group to the Nth sub-expert group decrease in sequence; Step Sc: Use the current splicing result as the input of each sub-expert group in the N sub-expert groups respectively, and obtain the output of each sub-expert group. Perform a weighted summation operation on the outputs of each sub-expert group according to their respective corresponding probabilities to obtain the output of the target total expert group; Step Sd: After splicing the output of the target total expert group and the input information, use the current splicing result as the input of the routing module again to obtain the output of the routing module. The output of the routing module is a set of new weight values, and the set of new weight values is used to indicate the new probability of each sub-expert group in the multiple sub-expert groups being reselected. In the set of new weight values, the sum of all new weight values is 1; Step Se: According to the order of each new weight value from large to small in a group of new weight values output by the routing module, reselect N new sub-expert groups in sequence, and use the reselected N new sub-expert groups together as the new target total expert group. The new target total expert group includes: a new first sub-expert group, a new second sub-expert group, …, a new Nth sub-expert group. Among them, N is a positive integer, and the weight values of the new first sub-expert group to the new Nth sub-expert group decrease in sequence; Step Sf: Use the current splicing result as the input of each new sub-expert group in the N new sub-expert groups respectively, obtain the output of each new sub-expert group, and perform a weighted summation operation on the outputs of each new sub-expert group according to their respective corresponding probabilities to obtain the output of the new target total expert group; Step Sg: Loop to execute Step Sd, Step Se, and Step Sf until the number of loops reaches a preset threshold. Then, splice the output of the current target total expert group with the input information and input it into a multi-layer perceptron to obtain the output result of the corresponding Block; Among them, Step Sa to Step Sc correspond to a deep thinking process in the latent space of the large model to be fine-tuned by the looped latent thinking module based on the chain of thought and the expert strategy.
3. The method according to claim 1, wherein Step S2 includes: Step S21: Input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned; Among them, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned; The type is a code type or an answer type. When the type is the code type, the execution action is code; when the type is the answer type, the execution action is an answer generated for the data to be analyzed and the question; Step S22: Determine the item to be evaluated based on the output result. The item to be evaluated includes at least one of the following: the answer, the format of the output result, and all the codes; Step S23: Score based on the item to be evaluated and the corresponding evaluation criteria to obtain a score value; Step S24: Fine-tune the parameters in the looped latent thinking module based on the score value. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Step S21 to Step S24 until the fine-tuning is completed to obtain the target large model.
4. The method according to claim 3, characterized in that, The architecture of the large model to be fine-tuned is a transformer architecture. Step S2 includes: Step S25: Based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the looped latent thinking module of the large model to be fine-tuned. Among them, the parameters in the looped latent thinking module are the parameters related to the deep thinking process.
5. The method according to claim 3, characterized in that, The step S23 includes: Step S231: When the item to be evaluated is the answer, input the preset answer and the answer into a preset reward function for scoring to obtain the score corresponding to the answer; Step S232: When the item to be evaluated is the format of the output result, send the output result to a format judgment executor for execution to obtain a first execution result, and score the first execution result to obtain the score corresponding to the format of the output result; Step S233: When the item to be evaluated is all the codes, sequentially execute all the codes in a code sandbox environment to obtain a second execution result, and score the second execution result to obtain the scores corresponding to all the codes.
6. A model fine-tuning device for realizing adaptive implicit thinking through the computing power of an intelligent computing center, characterized in that, The device includes: A receiving module, configured to execute step S1: Receive the data to be analyzed input by the user and the question related to the data to be analyzed; An execution module, configured to execute step S2: Based on the data to be analyzed and the question, fine-tune the large model to be fine-tuned to obtain a target large model; wherein, the large model to be fine-tuned is based on the deep thinking mode of the chain of thought and combines the mixture-of-experts strategy to reason about the data to be analyzed and the question to obtain an output result, and a cyclic hidden thinking module is set in the large model to be fine-tuned, and the cyclic hidden thinking module is configured to execute a deep thinking process based on the chain of thought and the mixture-of-experts strategy in the hidden space of the large model to be fine-tuned.
7. The device according to claim 6, wherein The architecture of the large model to be fine-tuned is a transformer architecture, and the large model to be fine-tuned of the transformer architecture includes a plurality of Blocks, and a cyclic hidden thinking module is introduced on each Block; The cyclic hidden thinking module includes: a routing module and a plurality of sub-expert groups, each sub-expert group contains at least one expert, and each expert is equivalent to adjusting the pre-trained parameters in the corresponding Block in the form of a low-rank parameter increment; The cyclic hidden thinking module is configured to execute the following steps: Step Sa: After splicing the input information and the initial state S0 of the corresponding Block, use the current splicing result as the input of the routing module to obtain the output of the routing module, and the output of the routing module is a set of weight values, and the set of weight values is used to indicate the probability that each sub-expert group in the plurality of sub-expert groups is selected, and in the set of weight values, the sum of all weight values is 1; Step Sb: According to the order of the weight values from large to small, sequentially select N sub-expert groups, and use the N sub-expert groups together as the target total expert group, and the target total expert group includes: a first sub-expert group, a second sub-expert group,..., an Nth sub-expert group, wherein N is a positive integer, and the weight values of the first sub-expert group to the Nth sub-expert group decrease in sequence. Step Sc: Use the current splicing result as the input for each of the N sub-expert groups, and obtain the output of each sub-expert group. Perform a weighted summation operation on the outputs of each sub-expert group according to their respective corresponding probabilities to obtain the output of the target total expert group; Step Sd: After splicing the output of the target total expert group and the input information, use the current splicing result as the input for the routing module again to obtain the output of the routing module; the output of the routing module is a new set of weight values, and the new set of weight values is used to indicate the new probabilities of each of the multiple sub-expert groups being reselected. Among the new set of weight values, the sum of all the new weight values is 1; Step Se: According to the order of each new weight value from largest to smallest in the new set of weight values output by the routing module, reselect a new N sub-expert groups in sequence, and use the reselected new N sub-expert groups together as the new target total expert group. The new target total expert group includes: a new first sub-expert group, a new second sub-expert group,..., a new Nth sub-expert group. Among them, N is a positive integer, and the weight values of the new first sub-expert group to the new Nth sub-expert group decrease in sequence; Step Sf: Use the current splicing result as the input for each of the new N sub-expert groups, obtain the output of each new sub-expert group, and perform a weighted summation operation on the outputs of each new sub-expert group according to their respective corresponding probabilities to obtain the output of the new target total expert group; Step Sg: Loop and execute Step Sd, Step Se, and Step Sf until the number of loops reaches a preset threshold. Then, splice the output of the current target total expert group with the input information and input it into a multi-layer perceptron to obtain the output result of the corresponding Block; Among them, Step Sa to Step Sc correspond to one deep thinking process in the latent space of the large model to be fine-tuned by the cyclic latent thinking module, which is based on the chain of thought and the expert strategy.
8. A server, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the model fine-tuning method for adaptive latent thinking through the computing power of the intelligent computing center as described in any one of claims 1-5.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the steps of the model fine-tuning method for adaptive latent thinking through the computing power of the intelligent computing center as described in any one of claims 1-5.
10. A computer program product, characterized in that, Including computer instructions. When the computer instructions are executed by the processor, it implements the steps of the model fine-tuning method for adaptive latent thinking through the computing power of the intelligent computing center as described in any one of claims 1-5.