Large model fine tuning method for realizing hidden space thinking based on computing power of intelligent computing center
By introducing a circular hidden thinking module into the intelligent computing center, using hidden space for deep thinking and reasoning, the problems of explicit thinking of large models are solved, and a more efficient and in-depth thinking process is achieved.
Patent Information
- Application Number
- CN202510265268.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
In the explicit thinking process, existing large models have problems such as long thinking and excessive resource consumption, which affects their application efficiency in real-time response and resource-constrained environments.
By introducing a circular hidden thinking module into the intelligent computing center, the computing power resources of the intelligent computing center are used to conduct in-depth thinking and reasoning in the hidden space, and fine-tuning of the large model is achieved to improve efficiency.
Deep thinking and reasoning in hidden space can significantly improve the thinking efficiency and depth of large models, reduce resource consumption, and output more accurate answers in a shorter time and improve user experience.
Smart Images

Figure CN120218235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a large model fine-tuning method for realizing implicit space thinking based on the computing power of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which mainly provides services to society through computing power infrastructure.
[0007] With the rapid development of artificial intelligence technology, large models have made remarkable progress in many fields such as natural language processing, computer vision, and speech recognition. Large models include: "large language model (LLM)" and "Multimodal Large Models (MLM)". These models can learn rich knowledge and patterns from massive data through deep learning algorithms, and thus show high accuracy and efficiency in specific tasks. However, current large models still face some technical challenges in practical applications, especially in explicit thinking during the reasoning process.
[0008] Specifically, although the explicit thinking process can improve the reliability of the output results, its reasoning process often takes a long time, resulting in a poor user experience in practical applications. For example, in scenarios that require real-time response, the long reasoning time may not meet the user's needs, thus affecting the practicality of the large model and user satisfaction. Especially in the handling of complex problems, the time required for explicit thinking may increase significantly, causing the user to wait too long and reducing the interaction efficiency of the large model.
[0009] In addition, the explicit thinking process is usually accompanied by a large consumption of computing resources. The large model needs to perform complex matrix operations and a large number of memory accesses during the reasoning process, which not only increases the computing cost but also may cause it to not run effectively in resource-constrained environments.
[0010] In summary, although the large model has demonstrated powerful capabilities in multiple fields, its explicit thinking has problems of long thinking time and excessive resource consumption in the output thinking process. Summary of the Invention
[0011] The present invention provides a fine-tuning method for a large model that realizes implicit space thinking based on the computing power of an intelligent computing center to solve the problem in the prior art that although the large model has demonstrated powerful capabilities in multiple fields, its explicit thinking has problems of long thinking time and excessive resource consumption in the output thinking process.
[0012] To solve the above technical problems, the present invention is implemented as follows:
[0013] In a first aspect, the present invention provides a fine-tuning method for a large model that realizes implicit space thinking based on the computing power of an intelligent computing center, and the method includes:
[0014] Step S1: Receive the data to be analyzed input by the user and the question related to the data to be analyzed;
[0015] Step S2: Based on the data to be analyzed and the question, fine-tune the large model to be fine-tuned to obtain a target large model; wherein, the large model to be fine-tuned performs reasoning on the data to be analyzed and the question based on the deep thinking mode of the chain of thought, and obtains an output result. A cyclic implicit thinking module is set in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute the deep thinking process based on the chain of thought in the implicit space of the large model to be fine-tuned.
[0016] Optionally, step S2 includes:
[0017] Step S21: Input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned;
[0018] Among them, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned;
[0019] The type is a code type or an answer type. When the type is the code type, the execution action is code; when the type is the answer type, the execution action is the answer generated for the data to be analyzed and the question;
[0020] Step S22: Determine the item to be evaluated based on the output result. The item to be evaluated includes at least one of the following: the answer, the format of the output result, and all the codes;
[0021] Step S23: Score based on the item to be evaluated and the corresponding evaluation criteria to obtain a score value;
[0022] Step S24: Fine-tune the parameters in the cyclic implicit thinking module based on the score value. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Step S21 to Step S24 until the fine-tuning is completed to obtain the target large model.
[0023] Optionally, the architecture of the large model to be fine-tuned is a transformer architecture, and Step S2 includes:
[0024] Step S25: Based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module of the large model to be fine-tuned, where the parameters in the cyclic implicit thinking module are the parameters related to the in-depth thinking process.
[0025] Optionally, the parameters in the cyclic implicit thinking module include the parameters of the low-rank matrix A and the parameters of the low-rank matrix B, and Step S25 includes:
[0026] Step S251: Introduce the cyclic implicit thinking module beside at least one original weight matrix in each layer of the large model to be fine-tuned with the transformer architecture, where the original weight matrix includes at least one of the following: Wq, Wk, Wv, Wd, Wf;
[0027] Step S252: Fine-tune the parameters of the low-rank matrix A and the parameters of the low-rank matrix B in each cyclic implicit thinking module.
[0028] Optionally, the architecture of the large model to be fine-tuned is a Transformer architecture. The large model to be fine-tuned in the Transformer architecture includes multiple layers, and each layer includes a self-attention sub-layer and a feed-forward network sub-layer;
[0029] A recurrent hidden thinking module is introduced beside any one of the original weight matrices in the self-attention sub-layer and the feed-forward network sub-layer;
[0030] The recurrent hidden thinking module includes: a low-rank matrix A, a low-rank matrix B, an adapter, and a gating module;
[0031] The recurrent hidden thinking module is used to perform the following steps:
[0032] Step Sa: Use the input information and the initial state S0 of the corresponding original weight matrix as the input of the adapter to obtain the current output result of the adapter with the same dimension as the input information;
[0033] Step Sb: Use the current output result of the adapter as the input of the low-rank matrix A to obtain the current output result of the low-rank matrix A;
[0034] Step Sc: Use the current output result of the low-rank matrix A as the input of the low-rank matrix B to obtain the current output result of the low-rank matrix B;
[0035] Step Sd: Use the current output result of the low-rank matrix B and the input information as the input of the gating module to obtain the current output result of the gating module;
[0036] Step Se: Determine whether the current output result of the gating module is less than or equal to a preset threshold;
[0037] If so, use the output result of the current output result of the low-rank matrix B and the input information processed by the pre-trained parameter matrix of the large model to be fine-tuned as the output result of the original weight matrix corresponding to the recurrent hidden thinking module;
[0038] If not, use the current output result of the gating module and the input information as the input of the adapter to obtain the current output result of the adapter, and repeat steps Sb to Se until the current output result of the gating module is less than or equal to the preset threshold, then use the output result of the current output result of the low-rank matrix B and the input information processed by the pre-trained parameter matrix as the output result of the original weight matrix corresponding to the recurrent hidden thinking module;
[0039] Among them, the steps from step Sa to step Sc correspond to a deep thinking in the latent space of the large model to be fine-tuned by the cyclic implicit thinking module during the deep thinking process based on the chain of thought.
[0040] Optionally, step S23 includes:
[0041] Step S231: When the item to be evaluated is the answer, input the preset answer and the answer into a preset reward function for scoring to obtain the score corresponding to the answer;
[0042] Step S232: When the item to be evaluated is the format of the output result, send the output result to a format judgment executor for execution to obtain a first execution result, and score the first execution result to obtain the score corresponding to the format of the output result;
[0043] Step S233: When the item to be evaluated is all the codes, sequentially execute all the codes in a code sandbox environment to obtain a second execution result, and score the second execution result to obtain the scores corresponding to all the codes.
[0044] In a second aspect, the present invention provides a large model fine-tuning device for realizing implicit space thinking based on the computing power of an intelligent computing center. The device includes:
[0045] A receiving module, configured to execute step S1: receive the data to be analyzed input by the user and the question related to the data to be analyzed;
[0046] An execution module, configured to execute step S2: fine-tune the large model to be fine-tuned based on the data to be analyzed and the question to obtain a target large model; wherein, the large model to be fine-tuned infers the data to be analyzed and the question based on the deep thinking mode of the chain of thought to obtain an output result, and a cyclic implicit thinking module is provided in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute the deep thinking process based on the chain of thought in the latent space of the large model to be fine-tuned.
[0047] Optionally, the execution module is further configured to execute step S21: input the data to be analyzed, the question, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned;
[0048] Wherein, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned;
[0049] The type is a code type or an answer type. When the type is the code type, the execution action is a code; when the type is the answer type, the execution action is an answer generated for the data to be analyzed and the question;
[0050] Step S22: Determine the item to be evaluated based on the output result. The item to be evaluated includes at least one of the following: the answer, the format of the output result, all the codes;
[0051] Step S23: Score based on the item to be evaluated and the corresponding evaluation criteria to obtain a score value;
[0052] Step S24: Fine-tune the parameters in the cyclic implicit thinking module based on the score value. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute Step S21 to Step S24 until the fine-tuning is completed to obtain the target large model.
[0053] Optionally, the architecture of the large model to be fine-tuned is a Transformer architecture. The execution module is further configured to execute Step S25: Based on the LoRA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module of the large model to be fine-tuned, where the parameters in the cyclic implicit thinking module are parameters related to the in-depth thinking process.
[0054] Optionally, the parameters in the cyclic implicit thinking module include the parameters of low-rank matrix A and the parameters of low-rank matrix B. The execution module is further configured to execute Step S251: Introduce the cyclic implicit thinking module beside at least one original weight matrix in each layer of the large model to be fine-tuned with the Transformer architecture, where the original weight matrix includes at least one of the following: Wq, Wk, Wv, Wd, Wf;
[0055] Step S252: Fine-tune the parameters of low-rank matrix A and the parameters of low-rank matrix B in each cyclic implicit thinking module.
[0056] Optionally, the architecture of the large model to be fine-tuned is a Transformer architecture. The large model to be fine-tuned with the Transformer architecture includes multiple layers, and each layer includes a self-attention sub-layer and a feed-forward network sub-layer;
[0057] A cyclic implicit thinking module is introduced beside any one of the original weight matrices in the self-attention sub-layer and the feed-forward network sub-layer;
[0058] The cyclic implicit thinking module includes: low-rank matrix A, low-rank matrix B, an adapter, and a gating module;
[0059] The looped implicit thinking module is used to perform the following steps:
[0060] Step Sa: Use the input information of the corresponding original weight matrix and the initial state S0 as the input of the adapter to obtain the current output result of the adapter with the same dimension as the input information;
[0061] Step Sb: Use the current output result of the adapter as the input of the low-rank matrix A to obtain the current output result of the low-rank matrix A;
[0062] Step Sc: Use the current output result of the low-rank matrix A as the input of the low-rank matrix B to obtain the current output result of the low-rank matrix B;
[0063] Step Sd: Use the current output result of the low-rank matrix B and the input information as the input of the gating module to obtain the current output result of the gating module;
[0064] Step Se: Determine whether the current output result of the gating module is less than or equal to a preset threshold;
[0065] If so, use the output result after processing the current output result of the low-rank matrix B and the input information through the pre-trained parameter matrix of the large model to be fine-tuned as the output result of the original weight matrix corresponding to the looped implicit thinking module;
[0066] If not, use the current output result of the gating module and the input information as the input of the adapter to obtain the current output result of the adapter, and repeat steps Sb to Se until the current output result of the gating module is less than or equal to the preset threshold, then use the output result after processing the current output result of the low-rank matrix B and the input information through the pre-trained parameter matrix as the output result of the original weight matrix corresponding to the looped implicit thinking module;
[0067] Among them, steps Sa to Sc correspond to a single in-depth thinking in the implicit space of the large model to be fine-tuned by the looped implicit thinking module during the in-depth thinking process based on the chain of thought.
[0068] Optionally, the execution module is further used to perform step S231: When the item to be evaluated is the answer, input the preset answer and the answer into a preset reward function for scoring to obtain the score corresponding to the answer;
[0069] Step S232: When the item to be evaluated is the format of the output result, send the output result to a format evaluation executor for execution to obtain a first execution result, and score the first execution result to obtain the score corresponding to the format of the output result;
[0070] Step S233: When the item to be evaluated is all the said codes, sequentially execute all the said codes in a code sandbox environment to obtain a second execution result, and score the second execution result to obtain the scores corresponding to all the said codes.
[0071] In a third aspect, the present invention provides a server, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the large model fine-tuning method for realizing implicit space thinking based on the computing power of an intelligent computing center as described in the first aspect above.
[0072] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large model fine-tuning method for realizing implicit space thinking based on the computing power of an intelligent computing center as described in the first aspect above.
[0073] In a fifth aspect, the present invention provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps of the large model fine-tuning method for realizing implicit space thinking based on the computing power of an intelligent computing center as described in the first aspect above.
[0074] In the present invention, by leveraging the computing power of an intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the implicit space. Specifically, the intelligent computing center utilizes advanced hardware facilities and optimized computing architectures, which can support parallel processing and efficient computing resource scheduling, enabling the in-depth thinking and reasoning process of the large model to be carried out quickly in the implicit space; and the way of implicit space reasoning can not only reduce resource consumption, but also significantly improve the efficiency of thinking and deepen the depth of thinking, thereby outputting more accurate answers in a shorter time.
[0075] In summary, based on the computing power of an intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the implicit space. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth on the premise of reducing resource consumption, so as to output more accurate answers in less time and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0077] Figure 1 Flow chart of a large model fine-tuning method for realizing implicit space thinking based on the computing power of an intelligent computing center provided by the present invention;
[0078] Figure 2 Structural block diagram of a large model fine-tuning device for realizing implicit space thinking based on the computing power of an intelligent computing center provided by the present invention;
[0079] Figure 3 Flow chart of a large model fine-tuning method for realizing implicit space thinking based on the computing power of an intelligent computing center provided by the present invention;
[0080] Figure 4 Schematic diagram of a sample data provided by the present invention;
[0081] Figure 5 Schematic diagram of the output result of a large model to be fine-tuned provided by the present invention;
[0082] Figure 6 Basic architecture block diagram of a large model provided with a cyclic implicit thinking module by the present invention, and schematic diagram of the working process of the cyclic implicit thinking module;
[0083] Figure 7 Structural block diagram of a large model fine-tuning device for realizing implicit space thinking based on the computing power of an intelligent computing center provided by the present invention;
[0084] Figure 8 Schematic diagram of the structure of the electronic device of the present invention. Detailed implementation manners
[0085] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0086] First, the technical terms related to the present invention will be briefly described below.
[0087] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0088] The "Computational Power (CP)" described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing power of a data center and includes general computing power, supercomputing power, and intelligent computing power. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super
[0089] The "Network Power (NP)" described in the present invention refers to: the manifestation of the data transmission ability of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.
[0090] The "Storage Power (SP)" described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage ability of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0091] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying power, and data storage power, which can realize the centralized computing, storage, transmission, and application of information, and presents characteristics such as multi-element ubiquitous, intelligent and agile, secure and reliable, and green and low-carbon.
[0092] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and popularization of new general technologies, the form of the new type of information infrastructure will be more diverse.
[0093] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and supercomputing power.
[0094] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0095] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled for various artificial intelligence innovation applications and is based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.
[0096] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0097] The "intelligent computing center" described in the present invention refers to a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0098] The "intelligent computing center" described in the present invention includes but is not limited to intelligent computing centers.
[0099] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0100] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, which has computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0101] The "supercomputing center" described in the present invention refers to: namely, a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0102] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0103] The "large model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0104] Figure 1 There is shown a method for fine-tuning a large model that realizes implicit space thinking based on the computing power of an intelligent computing center according to the present invention, as Figure 1 shown, the method includes:
[0105] Step S1: Receive the data to be analyzed input by the user and the question related to the data to be analyzed;
[0106] Step S2: Based on the data to be analyzed and the question, fine-tune the large model to be fine-tuned to obtain a target large model; wherein, the large model to be fine-tuned performs reasoning on the data to be analyzed and the question based on the deep thinking mode of the chain of thought, and obtains an output result. A cyclic implicit thinking module is set in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute the deep thinking process based on the chain of thought in the implicit space of the large model to be fine-tuned.
[0107] It should be noted that in step S1, first, receive the data to be analyzed provided by the user and the related question. The data to be analyzed can be any form of information, such as text, image, audio, or structured data, etc., and the related question is the specific information or insight that the user hopes to obtain by analyzing these data. The key to this step is to ensure that the system can accurately understand the user's input, including the format, content, and intention of the question of the data. In addition, data preprocessing can be performed, such as data cleaning, format conversion, and feature extraction, so as to prepare for subsequent analysis and reasoning.
[0108] In step S2, the system will fine-tune the large model to be fine-tuned according to the received data to be analyzed and related questions to generate a target large model for a specific task. The fine-tuning process is to retrain the existing large model to better adapt to the specific dataset and task requirements. In this process, the large model to be fine-tuned will use its built-in chain of thought mechanism to conduct in-depth thinking and reasoning, analyze the relationship between the data to be analyzed and the questions, and generate preliminary output results.
[0109] It should be noted that a cyclic implicit thinking module (as Figure 2 shown) is set in the large model to be fine-tuned. The core function of this module is to perform in-depth thinking based on the chain of thought in the implicit space. The cyclic implicit thinking module allows the model to perform multiple iterations during the reasoning process to deeply explore the potential information and logical relationships in the data. Different from traditional explicit thinking, cyclic implicit thinking can process information more complexly and meticulously in the implicit space, thus enhancing the depth and accuracy of reasoning. In addition, cyclic implicit thinking can also reduce resource consumption because it can perform calculations efficiently in the implicit space without frequently calling external resources.
[0110] In a possible implementation, as Figure 3 shown, step S2 includes:
[0111] Step S21: Input the data to be analyzed, the questions, and the output format of the large model to be fine-tuned into the large model to be fine-tuned to obtain the output result of the large model to be fine-tuned;
[0112] Step S22: Determine the items to be evaluated based on the output result; the items to be evaluated include at least one of the following: the answer, the format of the output result, and all the codes;
[0113] Step S23: Score based on the items to be evaluated and the corresponding evaluation criteria to obtain a score value;
[0114] Step S24: Fine-tune the parameters in the cyclic implicit thinking module based on the score value. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute steps S21 to S24 until the fine-tuning is completed to obtain the target large model.
[0115] In step S21, in addition to the data to be analyzed and the question, the output format also needs to be input into the large model. The output format defines the structure of the output result of the large model, including: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned; and the type is the code type or the answer type. When the type is the code type, the execution action is the code; when the type is the answer type, the execution action is the answer generated for the data to be analyzed and the question. In an exemplary application scenario, the data to be analyzed, the question, and the output format are:
[0116] You are a senior data analyst and are good at the Python language. The user has provided you with a dataset named `df`. Please write code to solve the user's problem: {question};
[0117] This dataset has {df.shape[0]} rows and {df.shape[1]} columns. The following are 3 sample rows randomly sampled from `df`:
[0118] {df.sample(3).markdown()}
[0119] Please give your answer in the format of a JSON. This JSON contains the following keys:
[0120] · step_id: str, the step to solve, e.g., '1', '2'...
[0121] · type: str, whether it is a code step or an answer step, e.g., 'code', 'answer'
[0122] · description: str, a brief description of this step
[0123] · action: str, the execution action of this step. If type is 'code', it is the code; if type is 'answer', it is the answer to the question.
[0124] Sample data (df) (as Figure 4 shown)
[0125] User question (question): What is the total consumption amount of each customer? And find the customer with the highest consumption amount.
[0126] The large model to be fine-tuned will integrate these inputs, start the in-depth thinking process based on the chain of thought in the latent space, and generate a preliminary output result. The output result is as Figure 5 shown.
[0127] In steps S22 to S24, it is necessary to determine the items to be evaluated based on the output results; the items to be evaluated include at least one of the following: answers, the format of the output results, all codes, and score based on the items to be evaluated and the corresponding evaluation criteria to obtain a score value; then, based on the score value, fine-tune the parameters in the cyclic implicit thinking module. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue the fine-tuning until the fine-tuning is completed to obtain the target large model.
[0128] In a possible implementation, step S23 includes: step S231: when the item to be evaluated is an answer, input the preset answer and the answer into the preset reward function for scoring to obtain the score value corresponding to the answer; step S232: when the item to be evaluated is the format of the output result, send the output result to the format judgment executor for execution to obtain the first execution result, and score the first execution result to obtain the score value corresponding to the format of the output result; step S233: when the item to be evaluated is all codes, sequentially execute all codes in the code sandbox environment to obtain the second execution result, and score the second execution result to obtain the score value corresponding to all codes.
[0129] It should be noted that in step S23, the system carefully scores the output results of the large model to be fine-tuned according to the previously determined items to be evaluated and the corresponding evaluation criteria. The specific scoring process is divided into three sub-steps, which are processed for different items to be evaluated respectively.
[0130] Specifically, when the item to be evaluated is an answer, input the preset answer and the answer into the preset reward function for scoring to obtain the score value corresponding to the answer. The preset correct answer can be the standard answer provided by experts or the reference answer obtained through training with historical data. The scoring process usually involves the following aspects: Similarity calculation: The system can use text similarity algorithms (such as cosine similarity, etc.) to evaluate the similarity between the generated answer and the preset answer; Reward function: The preset reward function will calculate a score value based on the similarity score, usually a normalized score (for example, from 0 to 1 or from 0 to 10), which is used to reflect the quality of the answer. The reward function can be designed according to the requirements of the specific task to ensure that it can effectively evaluate the accuracy and integrity of the answer. Through this process, the system can assign a quantitative score value to the generated answer to reflect its quality.
[0131] When the item to be evaluated is the format of the output result, the output result is sent to the format judgment executor for execution to obtain the first execution result, and the first execution result is scored to obtain the score corresponding to the format of the output result. That is to say, the system sends the output result to a dedicated format judgment executor, which is responsible for verifying the format of the output result. The format of the output result includes its structure, type, execution actions, etc., to ensure that the output meets the predefined format requirements.
[0132] When the item to be evaluated is all the code, all the code is sequentially executed in the code sandbox environment to obtain the second execution result, and the second execution result is scored to obtain the score corresponding to all the code. That is to say, the system sequentially executes all the generated code in the code sandbox environment to obtain the second execution result, which is used to evaluate the correctness and functionality implementation of the code.
[0133] It should be noted that the system will sequentially execute all the code in a secure code sandbox environment to avoid affecting the main system or causing security issues. Among them, the code sandbox environment provides an isolated execution environment, allowing the system to test code and functions without affecting other system components.
[0134] Thus, through steps S231, S232, and S233, the system can comprehensively evaluate the output results of the large model to be fine-tuned. This evaluation mechanism not only ensures the accuracy and integrity of the generated answers but also verifies the format of the output results and the validity of the code. Finally, the system will provide a basis for the subsequent fine-tuning process based on these scores to ensure the continuous optimization of the performance of the large model to be fine-tuned.
[0135] In step S24, the parameters in the cyclic implicit thinking module can be fine-tuned based on the scores. After fine-tuning, it is judged whether the fine-tuning is completed. If the fine-tuning is not completed, steps S21 to S24 are continued until the fine-tuning is completed to obtain the target large model.
[0136] It should be noted that the corresponding output results can be obtained based on each set of sample data in a training batch (a set of sample data includes the above-mentioned data to be analyzed, questions, and output formats), and scores can be obtained based on each output result, and then the comprehensive average score can be obtained. The parameters in the cyclic implicit thinking module are fine-tuned based on the comprehensive average score. The condition for the end of fine-tuning can be that the comprehensive average score reaches the preset threshold, or multiple batches are completed (training is finished).
[0137] And in a specific application scenario, the fine-tuning can be based on reinforcement learning technology. Reinforcement learning (RL) is a method of learning the optimal policy by interacting with the environment. The parameters in the cyclic implicit thinking module are fine-tuned through reinforcement learning technology (such asFigure 2 As shown by Reinforcement Learning in [reference], the fine-tuning process can be more flexible and efficient, capable of adaptively finding the optimal parameter adjustment strategy rather than relying on fixed rules or manual adjustment. This can improve the performance of the model to be fine-tuned and ultimately obtain a more powerful target large model.
[0138] In specific application scenarios, supervised fine-tuning (SFT) can also be used to fine-tune the parameters in the recurrent implicit thinking module. The present invention does not limit this and can flexibly select the fine-tuning method according to actual needs.
[0139] In a possible implementation, the architecture of the large model to be fine-tuned is a transformer architecture, and step S2 includes: step S25: Based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the recurrent implicit thinking module of the large model to be fine-tuned, where the parameters in the recurrent implicit thinking module are parameters related to the in-depth thinking process.
[0140] In a possible implementation, the parameters in the recurrent implicit thinking module include the parameters of the low-rank matrix A and the parameters of the low-rank matrix B. Step S25 includes: step S251: Introduce a recurrent implicit thinking module beside at least one original weight matrix in each layer of the large model to be fine-tuned in the transformer architecture, where the original weight matrix includes at least one of the following: Wq, Wk, Wv, Wd, Wf; step S252: Fine-tune the parameters of the low-rank matrix A and the parameters of the low-rank matrix B in each recurrent implicit thinking module.
[0141] Among them, the meaning of Wq is: linearly transform the input to generate a query vector (Query), and its function is to map the input word embedding (or hidden state) to the query space for calculating the attention score.
[0142] The meaning of Wk is: linearly transform the input to generate a key vector (Key), and its function is to map the input to the key space and cooperate with the query vector to calculate the attention weight.
[0143] The meaning of Wv is: linearly transform the input to generate a value vector (Value), and its function is to map the input to the value space and serve as the carrier for information aggregation in the attention mechanism.
[0144] The meaning of Wd is: linearly transform the output of the multi-head attention to generate the final output, and its function is to map the result after multi-head concatenation back to the original dimension to ensure that the output is consistent with the input dimension.
[0145] The meaning of Wf is: a linear transformation used in the feedforward network, which functions to map the input data to a new dimension for subsequent non-linear transformation or other processing.
[0146] Among them, the LORA (Low-Rank Adaptation) technology is a technique for fine-tuning large-scale pre-trained models, aiming to efficiently adjust model parameters in the form of low-rank matrices. The basic idea of LORA is to capture the features of specific tasks by introducing low-rank adapters while keeping the original model parameters unchanged.
[0147] The working principle of LORA is as follows: Freeze the pre-trained parameters: During the fine-tuning process, first freeze the pre-trained parameters of the large model, which means these parameters do not change during fine-tuning. This approach can avoid overfitting when fine-tuning on small datasets while retaining the knowledge learned by the pre-trained model on large-scale datasets; Introduce low-rank adapters: Insert low-rank adapters in some layers of the model. The parameters of these adapters are trainable, and the adaptability of the model is achieved by fine-tuning these parameters; Parameter update: During the training process, only update the parameters of the low-rank adapters, rather than the parameters of the entire model. This method greatly reduces the number of parameters to be trained, thereby improving the training efficiency and saving computing resources.
[0148] It should be noted that in the large model to be fine-tuned, the recurrent hidden thinking module is a key component responsible for the in-depth thinking process in the hidden space of the large model to be fine-tuned. The parameters of this module are closely related to the inference ability, memory mechanism, and information integration ability of the model when dealing with complex tasks. Based on the LORA technology, the pre-trained parameters of the large model to be fine-tuned can be frozen, and the parameters in the recurrent hidden thinking module of the large model to be fine-tuned can be fine-tuned, which not only improves the fine-tuning efficiency but also reduces the risk of overfitting, enabling the model to better adapt to complex inference and decision-making tasks.
[0149] Now, taking a recurrent hidden thinking module as an example, illustrate its working method and data flow (as Figure 6 shown). The architecture of the large model to be fine-tuned is the transformer architecture. The large model to be fine-tuned with the transformer architecture includes multiple layers, and each layer includes a self-attention sub-layer and a feedforward network sub-layer; A recurrent hidden thinking module is introduced beside any of the original weight matrices in the self-attention sub-layer and the feedforward network sub-layer; The recurrent hidden thinking module includes: low-rank matrix A, low-rank matrix B, an adapter, and a gating module;
[0150] The cyclic implicit thinking module is used to perform the following steps: Step Sa: Use the input information of the corresponding original weight matrix and the initial state S0 as the input of the adapter to obtain the current output result of the adapter with the same dimension as the input information; Step Sb: Use the current output result of the adapter as the input of the low-rank matrix A to obtain the current output result of the low-rank matrix A; Step Sc: Use the current output result of the low-rank matrix A as the input of the low-rank matrix B to obtain the current output result of the low-rank matrix B; Step Sd: Use the current output result of the low-rank matrix B and the input information as the input of the gating module to obtain the current output result of the gating module; Step Se: Determine whether the current output result of the gating module is less than or equal to a preset threshold; if so, use the output result after processing the current output result of the low-rank matrix B and the input information through the pre-trained parameter matrix of the large model to be fine-tuned as the output result of the original weight matrix corresponding to the cyclic implicit thinking module; if not, use the current output result of the gating module and the input information as the input of the adapter to obtain the current output result of the adapter, and repeat Steps Sb to Se until the current output result of the gating module is less than or equal to the preset threshold, then use the output result after processing the current output result of the low-rank matrix B and the input information through the pre-trained parameter matrix as the output result of the original weight matrix corresponding to the cyclic implicit thinking module; where Steps Sa to Sc correspond to a deep thinking process based on the chain of thought in the implicit space of the large model to be fine-tuned during a single deep thinking.
[0151] In summary, in the present invention, by leveraging the computing power of the intelligent computing center, the large model can perform efficient deep thinking and reasoning in the implicit space. Specifically, the intelligent computing center, with its advanced hardware facilities and optimized computing architecture, can support parallel processing and efficient computing resource scheduling, enabling the deep thinking and reasoning process of the large model to proceed rapidly in the implicit space; and the way of implicit space reasoning can not only reduce resource consumption, but also significantly improve the efficiency of thinking and deepen the depth of thinking, thereby outputting more accurate answers in a shorter time.
[0152] In summary, based on the computing power of the intelligent computing center, the large model can perform efficient deep thinking and reasoning in the implicit space. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth of thinking on the premise of reducing resource consumption, thereby outputting more accurate answers in less time and enhancing the user experience.
[0153] Figure 7 Shows a large model fine-tuning device for realizing implicit space thinking based on the computing power of the intelligent computing center according to the present invention, as Figure 7 shown, the device 70 includes:
[0154] A receiving module 701, configured to execute step S1: receive the data to be analyzed input by the user and the questions related to the data to be analyzed;
[0155] An execution module 702, configured to execute step S2: fine-tune the large model to be fine-tuned based on the data to be analyzed and the questions, to obtain a target large model; wherein, the large model to be fine-tuned performs reasoning on the data to be analyzed and the questions based on the depth thinking mode of the chain of thought, to obtain an output result, and a cyclic implicit thinking module is set in the large model to be fine-tuned, and the cyclic implicit thinking module is configured to execute the depth thinking process based on the chain of thought in the implicit space of the large model to be fine-tuned.
[0156] In a possible implementation manner, the execution module 702 is further configured to execute step S21: input the data to be analyzed, the questions, and the output format of the large model to be fine-tuned into the large model to be fine-tuned, to obtain the output result of the large model to be fine-tuned;
[0157] Wherein, the output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned;
[0158] The type is a code type or an answer type. When the type is a code type, the execution action is code; when the type is an answer type, the execution action is an answer generated for the data to be analyzed and the questions;
[0159] Step S22: determine the items to be evaluated based on the output result, and the items to be evaluated include at least one of the following: the answer, the format of the output result, and all the codes;
[0160] Step S23: score based on the items to be evaluated and the corresponding evaluation criteria, to obtain a score value;
[0161] Step S24: fine-tune the parameters in the cyclic implicit thinking module based on the score value. After the fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute steps S21 to S24 until the fine-tuning is completed, to obtain the target large model.
[0162] In a possible implementation manner, the architecture of the large model to be fine-tuned is a transformer architecture, and the execution module 702 is further configured to execute step S25: based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module of the large model to be fine-tuned, wherein the parameters in the cyclic implicit thinking module are the parameters related to the depth thinking process.
[0163] In a possible implementation, the parameters in the cyclic implicit thinking module include the parameters of the low-rank matrix A and the parameters of the low-rank matrix B. The execution module 702 is further configured to execute step S251: beside at least one original weight matrix of each layer of the large model to be fine-tuned in the transformer architecture, a cyclic implicit thinking module is respectively introduced, where the original weight matrix includes at least one of the following: Wq, Wk, Wv, Wd, Wf;
[0164] Step S252: Fine-tune the parameters of the low-rank matrix A and the parameters of the low-rank matrix B in each cyclic implicit thinking module.
[0165] Optionally, the architecture of the large model to be fine-tuned is the transformer architecture. The large model to be fine-tuned in the transformer architecture includes multiple layers, and each layer includes a self-attention sub-layer and a feed-forward network sub-layer;
[0166] A cyclic implicit thinking module is introduced beside any one of the original weight matrices in the self-attention sub-layer and the feed-forward network sub-layer;
[0167] The cyclic implicit thinking module includes: a low-rank matrix A, a low-rank matrix B, an adapter, and a gating module;
[0168] The cyclic implicit thinking module is configured to execute the following steps:
[0169] Step Sa: Use the input information of the corresponding original weight matrix and the initial state S0 as the input of the adapter, and obtain the current output result of the adapter with the same dimension as the input information;
[0170] Step Sb: Use the current output result of the adapter as the input of the low-rank matrix A, and obtain the current output result of the low-rank matrix A;
[0171] Step Sc: Use the current output result of the low-rank matrix A as the input of the low-rank matrix B, and obtain the current output result of the low-rank matrix B;
[0172] Step Sd: Use the current output result of the low-rank matrix B and the input information as the input of the gating module, and obtain the current output result of the gating module;
[0173] Step Se: Determine whether the current output result of the gating module is less than or equal to a preset threshold;
[0174] If so, use the output result after processing the current output result of the low-rank matrix B and the input information through the pre-trained parameter matrix of the large model to be fine-tuned as the output result of the original weight matrix corresponding to the cyclic implicit thinking module;
[0175] If not, use the current output result of the gating module and the input information as the input of the adapter to obtain the current output result of the adapter, and repeat steps Sb to Se until the current output result of the gating module is less than or equal to the preset threshold. Then, use the current output result of the low-rank matrix B and the output result after processing the input information through the pre-trained parameter matrix as the output result of the original weight matrix corresponding to the cyclic implicit thinking module.
[0176] Among them, steps Sa to Sc correspond to a single in-depth thinking process in the implicit space of the large model to be fine-tuned during the in-depth thinking process based on the chain of thought.
[0177] In a possible implementation, the execution module 702 is further configured to execute step S231: when the item to be evaluated is the answer, input the preset answer and the answer into the preset reward function for scoring to obtain the score corresponding to the answer.
[0178] Step S232: when the item to be evaluated is the format of the output result, send the output result to the format judgment executor for execution to obtain the first execution result, and score the first execution result to obtain the score corresponding to the format of the output result.
[0179] Step S233: when the item to be evaluated is all the codes, sequentially execute all the codes in the code sandbox environment to obtain the second execution result, and score the second execution result to obtain the score corresponding to all the codes.
[0180] In summary, in the present invention, by leveraging the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the implicit space. Specifically, the intelligent computing center, with its advanced hardware facilities and optimized computing architecture, can support parallel processing and efficient computing resource scheduling, enabling the in-depth thinking and reasoning process of the large model to proceed rapidly in the implicit space. Moreover, the method of implicit space reasoning can not only reduce resource consumption but also significantly improve the efficiency of thinking and deepen the depth of thinking, thereby outputting more accurate answers in a shorter time.
[0181] In summary, based on the computing power of the intelligent computing center, the large model can perform efficient in-depth thinking and reasoning in the implicit space. Compared with the explicit thinking of existing large models, it can significantly improve the thinking efficiency and depth while reducing resource consumption, thereby outputting more accurate answers in less time and enhancing the user experience.
[0182] Please refer to Figure 8, the present invention also provides an electronic device 80, including a processor 801, a memory 802, and a computer program stored on the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, it implements the steps of the above-mentioned large model fine-tuning method for realizing implicit space thinking based on the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0183] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-mentioned large model fine-tuning method for realizing implicit space thinking based on the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0184] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps of the above-mentioned large model fine-tuning method for realizing implicit space thinking based on the computing power of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0185] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0186] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the method of the present invention.
[0187] The present invention has been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.
Claims
1. A large model fine-tuning method based on the computing power of an intelligent computing center to realize latent space thinking, characterized in that: The method comprises: Step S1: receiving data to be analyzed and questions related to the data to be analyzed input by a user; Step S2: Based on the data to be analyzed and the problem, fine-tune the large model to be fine-tuned to obtain a target large model; wherein the large model to be fine-tuned infers the data to be analyzed and the problem based on a deep thinking method of a thinking chain to obtain an output result, and a cyclic implicit thinking module is provided in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute a deep thinking process based on a thinking chain in the latent space of the large model to be fine-tuned.
2. The method according to claim 1, characterized in that: The step S2 comprises: Step S21: inputting the data to be analyzed, the problem, and the output format of the large model to be fine-tuned into the large model to be fine-tuned, and obtaining the output result of the large model to be fine-tuned; The output format includes: the serial number of the step output by the large model to be fine-tuned, the type of the step output by the large model to be fine-tuned, the description information of the step output by the large model to be fine-tuned, and the execution action indicated by the step output by the large model to be fine-tuned; The type is a code type or an answer type. When the type is the code type, the execution action is a code; when the type is the answer type, the execution action is an answer generated for the data to be analyzed and the question; Step S22: determining an item to be evaluated based on the output result, wherein the item to be evaluated includes at least one of the following: the answer, the format of the output result, and all the codes; Step S23: scoring based on the items to be evaluated and the corresponding evaluation criteria to obtain scores; Step S24: Based on the score, fine-tune the parameters in the cyclic implicit thinking module. After fine-tuning, determine whether the fine-tuning is completed. If the fine-tuning is not completed, continue to execute steps S21 to S24 until the fine-tuning is completed and the target large model is obtained.
3. The method according to claim 2, characterized in that The architecture of the large model to be fine-tuned is a transformer architecture, and step S2 includes: Step S25: Based on the LORA technology, freeze the pre-trained parameters of the large model to be fine-tuned, and fine-tune the parameters in the cyclic implicit thinking module in the large model to be fine-tuned, wherein the parameters in the cyclic implicit thinking module are parameters related to the deep thinking process.
4. The method according to claim 3, characterized in that The parameters in the cyclic implicit thinking module include the parameters of the low-rank matrix A and the parameters of the low-rank matrix B, and the step S25 includes: Step S251: introducing the recurrent implicit thinking module next to at least one original weight matrix of each layer of the large model to be fine-tuned in the transformer architecture, wherein the original weight matrix includes at least one of the following: Wq, Wk, Wv, Wd, Wf; Step S252: fine-tune the parameters of the low-rank matrix A and the parameters of the low-rank matrix B in each of the cyclic implicit thinking modules.
5. The method according to claim 1, characterized in that: The architecture of the large model to be fine-tuned is a transformer architecture, and the large model to be fine-tuned of the transformer architecture includes multiple layers, each of which includes a self-attention sublayer and a feedforward network sublayer; A recurrent implicit thinking module is introduced next to any original weight matrix in the self-attention sublayer and the feedforward network sublayer; The cyclic implicit thinking module includes: a low-rank matrix A, a low-rank matrix B, an adapter and a gating module; The loop implicit thinking module is used to perform the following steps: Step Sa: using the input information of the corresponding original weight matrix and the initial state S0 as the input of the adapter, and obtaining the current output result of the adapter with the same dimension as the input information; Step Sb: using the current output result of the adapter as the input of the low-rank matrix A to obtain the current output result of the low-rank matrix A; Step Sc: using the current output result of the low-rank matrix A as the input of the low-rank matrix B to obtain the current output result of the low-rank matrix B; Step Sd: using the current output result of the low-rank matrix B and the input information as inputs of the gating module to obtain the current output result of the gating module; Step Se: Determine whether the current output result of the gating module is less than or equal to a preset threshold; If yes, the current output result of the low-rank matrix B and the output result of the input information after being processed by the pre-trained parameter matrix of the large model to be fine-tuned are used as the output result of the original weight matrix corresponding to the cyclic implicit thinking module; If not, the current output result of the gating module and the input information are used as the input of the adapter to obtain the current output result of the adapter, and the steps Sb to Se are repeatedly executed until the current output result of the gating module is less than or equal to the preset threshold value, and the current output result of the low-rank matrix B and the output result of the input information after being processed by the pre-trained parameter matrix are used as the output result of the original weight matrix corresponding to the cyclic implicit thinking module; Among them, the steps Sa to Sc correspond to the cyclic implicit thinking module performing a deep thinking process based on the thinking chain in the latent space of the large model to be fine-tuned.
6. The method according to any one of claims 2 to 4, characterized in that: The step S23 comprises: Step S231: when the item to be evaluated is the answer, input the preset answer and the answer into a preset reward function for scoring, and obtain a score corresponding to the answer; Step S232: when the item to be evaluated is the format of the output result, the output result is sent to a format evaluation executor for execution to obtain a first execution result, and the first execution result is scored to obtain a score corresponding to the format of the output result; Step S233: when the items to be evaluated are all the codes, all the codes are sequentially executed in the code sandbox environment to obtain a second execution result, and the second execution result is scored to obtain scores corresponding to all the codes.
7. A large model fine-tuning device based on the computing power of an intelligent computing center to realize latent space thinking, characterized in that: The device comprises: The receiving module is used to execute step S1: receiving the data to be analyzed and questions related to the data to be analyzed input by the user; An execution module is used to execute step S2: based on the data to be analyzed and the problem, fine-tune the large model to be fine-tuned to obtain a target large model; wherein the large model to be fine-tuned reasons the data to be analyzed and the problem based on a deep thinking method of a thinking chain to obtain an output result, and a cyclic implicit thinking module is provided in the large model to be fine-tuned, and the cyclic implicit thinking module is used to execute a deep thinking process based on a thinking chain in the latent space of the large model to be fine-tuned.
8. A server, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a large model fine-tuning method for realizing latent space thinking based on the computing power of an intelligent computing center as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a large model fine-tuning method based on the computing power of an intelligent computing center to realize latent space thinking as described in any one of claims 1 to 6.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of a large model fine-tuning method based on the computing power of an intelligent computing center to realize latent space thinking as described in any one of claims 1-6.