Large model network dynamic expansion method and device of intelligent computing center cloud platform for providing computing power resources
By introducing jump gating modules and recurrent gating modules into large models and dynamically adjusting network layers, the problem of computing power waste during training and inference of large models is solved, achieving more efficient computing and more accurate result output.
Patent Information
- Application Number
- CN202510712428.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Due to the large number of parameters and the design structure of a fixed number of network layers, large models consume a lot of computing resources during inference and training, increasing operating costs and limiting their wide application.
The skip gating module and the recurrent gating module are used to learn the optimal dynamic decision rules during the large model training process, dynamically expand the network layer, and decide whether to skip or recursively execute the network layer according to the problem to be processed.
While ensuring the inference performance of large models, it saves computing resources, avoids waste, and improves the utilization of computing resources.
Smart Images

Figure CN120671816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and specifically to a large-model network dynamic expansion method and device for an intelligent computing center cloud platform that provides computing power resources. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.
[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model fine-tuning, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] In the current field of artificial intelligence, especially in tasks such as natural language processing and computer vision, large models have become the mainstream of research and application. However, the number of parameters in these large models is often very large, resulting in the consumption of a large amount of computing power during deployment and use. In addition, the existing Transformer network architecture has a linear static structure. This means that the number of network layers in large models designed based on the Transformer network architecture is fixed during the training and use phases of the large model, resulting in the consumption of a large amount of computing power during the inference and training processes of large models. This not only wastes computing power resources, but also increases the operating cost of large models, limiting their widespread application.
[0008] In summary, since the emergence of intelligent computing centers, large models have wasted computing resources and increased operating costs in large model inference training due to their large number of parameters and fixed number of network layers. This problem has become a technical issue that needs to be solved urgently. Summary of the Invention
[0009] The present invention provides a large-model network dynamic expansion method and device for an intelligent computing center cloud platform that provides computing resources, in order to solve the technical problem that since the emergence of intelligent computing centers, large models have a large number of parameters and a design structure with a fixed number of network layers, resulting in a waste of computing resources and an increase in operating costs in large-model inference training.
[0010] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0011] In a first aspect, the present invention provides a large-model network dynamic expansion method for an intelligent computing center cloud platform that provides computing resources, the method comprising:
[0012] Step S1: receiving a sample question set input by a user;
[0013] Step S2: Based on the sample problem set, the large model is trained to obtain a trained large model; wherein the large model is provided with a jump gating module and a recurrent gating module;
[0014] The jump gating module and the recurrent gating module are used to learn based on the sample problem set during the training process of the large model, so as to gradually learn the optimal dynamic decision rule;
[0015] The trained large model is used to receive a problem to be processed input by a user, and to reason about the problem to be processed based on the jump gating module, the recurrent gating module, and the optimal dynamic decision rule, and to dynamically expand its own multiple network layers during the reasoning process, and to obtain an output result based on the expanded multiple network layers;
[0016] The optimal dynamic decision rule is used to determine the network layers to be skipped and the network layers to be executed cyclically among the multiple network layers according to the problem to be processed.
[0017] Optionally, step S2 includes:
[0018] Step S21: Based on the sample problem set, the large model is trained in combination with a preset loss function, with the goal of minimizing the loss function, to obtain the trained large model;
[0019] The loss function is used to measure the difference between the output result of the large model and the standard answer, and / or to measure the usage of computing resources during the training process of the large model.
[0020] Optionally, the architecture of the trained large model is a Transformer architecture, and the trained large model of the Transformer architecture includes at least an embedding layer and the multiple network layers;
[0021] The jump gating module is configured to receive the output information of the embedding layer, and output a jump index based on the output information of the embedding layer and the optimal dynamic decision rule, wherein the jump index is used to indicate a jump execution state of each network layer in the multiple network layers, wherein the jump execution state includes skipping the current network layer and activating the current network layer;
[0022] The loop gating module is configured to receive output information of the embedding layer, and output a loop index based on the output information of the embedding layer and the optimal dynamic decision rule, wherein the loop index is used to indicate loop information of each network layer in the multiple network layers;
[0023] The cycle information includes any one of the following: the current network layer belongs to a cycle group and the number of cycles of the cycle group; the current network layer does not belong to a cycle group;
[0024] If the current network layer is independently selected or unselected, the cycle information corresponding to the current network layer is: the current network layer does not belong to the cycle group; if the current network layer is selected and at least one of the network layers adjacent to the current network layer is also selected, the cycle information corresponding to the current network layer is: the cycle group to which the current network layer belongs and the cycle count of the cycle group;
[0025] The plurality of network layers are used to determine whether to jump based on the jump index, and to determine whether to loop based on the loop index.
[0026] Optionally, the number of cycles is related to the complexity of the problem to be processed. The higher the complexity of the problem to be processed, the greater the number of cycles.
[0027] Optionally, if there is a jump network layer in the loop group whose jump execution status is to skip the current network layer, the difference between the total number of network layers in the loop group and the number of the jump network layers is determined. If the difference is greater than or equal to 2, the remaining network layers in the loop group except the jump network layer are looped.
[0028] In a second aspect, the present invention provides a large-scale model network dynamic expansion device for an intelligent computing center cloud platform that provides computing resources, the device comprising:
[0029] The receiving module is configured to execute step S1: receiving a sample question set input by a user;
[0030] An execution module, configured to execute step S2: training the large model based on the sample problem set to obtain a trained large model; wherein the large model is provided with a jump gating module and a recurrent gating module;
[0031] The jump gating module and the recurrent gating module are used to learn based on the sample problem set during the training process of the large model, so as to gradually learn the optimal dynamic decision rule;
[0032] The trained large model is used to receive a problem to be processed input by a user, and to reason about the problem to be processed based on the jump gating module, the recurrent gating module, and the optimal dynamic decision rule, and to dynamically expand its own multiple network layers during the reasoning process, and to obtain an output result based on the expanded multiple network layers;
[0033] The optimal dynamic decision rule is used to determine the network layers to be skipped and the network layers to be executed cyclically among the multiple network layers according to the problem to be processed.
[0034] Optionally, step S2 includes:
[0035] Step S21: Based on the sample problem set, the large model is trained in combination with a preset loss function, with the goal of minimizing the loss function, to obtain the trained large model;
[0036] The loss function is used to measure the difference between the output result of the large model and the standard answer, and / or to measure the usage of computing resources during the training process of the large model.
[0037] In the third aspect, the present invention provides a server comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, the steps of a large-model network dynamic expansion method of an intelligent computing center cloud platform that provides computing resources as described in the first aspect above are implemented.
[0038] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a large-model network dynamic expansion method of an intelligent computing center cloud platform that provides computing resources as described in the first aspect above.
[0039] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of a large-model network dynamic expansion method for an intelligent computing center cloud platform that provides computing resources as described in the first aspect above.
[0040] In the present invention, the large model is provided with a jump gating module and a loop gating module, which can be used to learn based on the sample problem set during the training process of the large model to gradually learn the optimal dynamic decision rule. The jump gating module allows the large model to intelligently skip unnecessary network layers according to the problem to be processed and the optimal dynamic decision rule during the reasoning process, thereby reducing the amount of calculation and the consumption of computing resources, while the loop gating module can repeatedly execute certain key network layers when needed (such as when the problem to be processed is more complex) to enhance the large model's understanding and processing capabilities of complex problems. This ability to dynamically expand the number of layers of the large model's network layer enables the trained large model to adaptively adjust the network layer according to the problem to be processed (actual input situation), thereby saving computing resources, avoiding waste of computing resources, and improving the utilization of computing resources while ensuring the reasoning performance of the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0042] Figure 1 A flowchart of a large-scale model network dynamic expansion method for an intelligent computing center cloud platform providing computing resources provided by the present invention;
[0043] Figure 2 A network hierarchical structure diagram of a large model of an existing Transformer architecture provided by the present invention;
[0044] Figure 3 A network hierarchical structure diagram of a large model of an intelligent computing center cloud platform that provides computing resources provided by the present invention;
[0045] Figure 4 This is a structural block diagram of a large-scale network dynamic expansion device for an intelligent computing center cloud platform that provides computing resources provided by the present invention;
[0046] Figure 5 Schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0048] First, the technical terms involved in the present invention are briefly explained below.
[0049] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0050] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is about the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP general+CP intelligent+CP super
[0051] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0052] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.
[0053] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.
[0054] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0055] The “computing power” mentioned in the present invention includes: “general computing power”, “intelligent computing power” and “super computing power”.
[0056] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0057] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.
[0058] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.
[0059] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model fine-tuning, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0060] The "intelligent computing center cloud platform" mentioned in the present invention refers to: a cloud computing platform that provides comprehensive services based on the hardware resources and software resources of the intelligent computing center.
[0061] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".
[0062] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0063] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0064] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0065] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water, and electricity.
[0066] The "large model" mentioned in this invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is fine-tuned with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0067] The "Poisson distribution" described in the present invention is a discrete probability distribution commonly seen in statistics and probability, and is suitable for describing the probability distribution of the number of random events occurring per unit time.
[0068] The "softmax function" described in the present invention is a mathematical function that normalizes any real-valued vector (usually an unnormalized fraction or logarithmic probability) to a probability distribution.
[0069] Figure 1 The following shows a large model network dynamic expansion method of an intelligent computing center cloud platform providing computing resources according to the present invention, such as Figure 1 As shown, the method includes:
[0070] Step S1: receiving a sample question set input by a user;
[0071] Step S2: Based on the sample problem set, the large model is trained to obtain a trained large model; wherein the large model is provided with a jump gating module and a recurrent gating module;
[0072] It should be noted that the jump gating module and the recurrent gating module are used to learn based on the sample problem set during the training process of the large model, so as to gradually learn the optimal dynamic decision rules; the trained large model is used to receive the problems to be processed input by the user, and to reason about the problems to be processed based on the jump gating module, the recurrent gating module and the optimal dynamic decision rules, and dynamically expand its own multiple network layers during the reasoning process, and obtain output results based on the expanded multiple network layers.
[0073] It should be noted that Figure 1 The core of the technical solution shown is to receive sample problem sets (training data sets) input by users and use these sample problem sets to systematically train the large model to generate an optimized trained large model. The large model integrates a jump gating module and a loop gating module. These two modules play a key role in the training process, enabling the large model to learn from the sample problem set and gradually explore and try and error to eventually master the optimal dynamic decision rule, wherein the optimal dynamic decision rule is used to determine the network layers to be skipped and the network layers to be executed cyclically among multiple network layers according to the problem to be processed. After completing the training, the large model can receive the problem to be processed input by the user, and rely on the jump gating module, the loop gating module and the learned optimal dynamic decision rule to conduct in-depth reasoning on the problem to be processed.
[0074] Moreover, the jump gating module and the loop gating module in the present invention use the softmax function to realize intelligent decision-making. When processing the pending problem input by the user, the jump gating module first scores each network layer, and performs softmax normalization on the scoring result, converts the original score into a probability distribution between 0 and 1, and finally generates a jump index. For example: the optimal dynamic decision rule can be to sort the probabilities of the 10 network layers from high to low, and determine the first 8 network layers as network layers that can be skipped, and the last two network layers as network layers that need to be activated; or, set the threshold to 0.2, the network layer with probability ≥ 0.2 is the network layer to be skipped, and the network layer with probability < 0.2 is the network layer to be activated; or, the network layer with probability ≥ 0.2 is the network layer to be activated, and the network layer with probability < 0.2 is the network layer to be skipped.
[0075] During inference, the large model has the ability to dynamically expand multiple network layers. It can flexibly adjust its network structure based on the complexity and needs of the specific problem, achieving more efficient computation and more accurate output. Specifically, the optimal dynamic decision rule will guide the large model to identify which network layers can be skipped and which network layers need to be executed repeatedly during inference, based on the specific problem being processed. This minimizes unnecessary computation, conserves computing power, and avoids waste while ensuring inference quality.
[0076] In a possible implementation, step S2 includes:
[0077] Step S21: Based on the sample problem set, the large model is trained in combination with a preset loss function, with the goal of minimizing the loss function to obtain a trained large model;
[0078] Among them, the loss function is used to measure the difference between the output results of the large model and the standard answer, and / or to measure the usage of computing resources during the training process of the large model.
[0079] It should be noted that, during the training phase, the present invention uses the sample question set provided by the user to optimize the training of the large model of the preset jump gating module and the loop gating module. The training process is guided by a preset loss function, which simultaneously monitors two key indicators: one is the content difference between the output result of the large model and the standard answer, and the other is the usage of computing resources during the training process (such as the peak value of GPU memory usage, the time taken for a single training, etc.). By continuously adjusting the parameters of the large model (such as the parameters of the jump gating module, the parameters of the loop gating module and the parameters of the network layer), the loss function value is minimized, and finally a large model after training is obtained that can accurately answer questions and intelligently control the consumption of computing resources.
[0080] For ease of understanding, the training process of a large model can be compared to the process of a student reviewing for an exam. During the initial review phase, students randomly select review content and may spend a lot of time on simple questions (similar to the large model's random attempts during training, where layers that should have been skipped were not skipped), or they may skip key knowledge points during review (similar to the large model's mistaken skipping of key layers during training). The result is a poor exam score (similar to the large model's incorrect predictions). When grading the exam, the teacher can provide two types of feedback based on the exam (similar to the loss function): if the answer is incorrect, more points will be deducted (similar to the penalty given if the large model predicts incorrectly); if the answer is correct but lengthy, fewer points will be deducted (similar to the penalty given if the large model predicts correctly but consumes too much computing power). Students can adjust their review strategies accordingly, and during their next review, they may find that some question types require detailed calculations, while others have fixed answer formulas (the process of adjusting review strategies is similar to the parameter update process of the large model). Subsequently, students can learn to perform detailed calculations when encountering certain complex question types, and to directly apply formulas when encountering certain simple question types (similar to how a large model gradually stabilizes in the later stages of training and learns the optimal dynamic decision-making rules).
[0081] In one possible implementation, the architecture of the trained large model is a Transformer architecture, and the trained large model of the Transformer architecture includes at least an embedding layer and multiple network layers;
[0082] a jump gating module, configured to receive output information of the embedding layer, and output a jump index based on the output information of the embedding layer and an optimal dynamic decision rule, wherein the jump index is used to indicate a jump execution state of each network layer in the plurality of network layers, wherein the jump execution state includes skipping the current network layer and activating the current network layer;
[0083] a loop gating module, configured to receive output information of the embedding layer, and output a loop index based on the output information of the embedding layer and an optimal dynamic decision rule, wherein the loop index is used to indicate loop information of each network layer in the plurality of network layers;
[0084] The cycle information includes any of the following: whether the current network layer belongs to a cycle group and the number of cycles of the cycle group; whether the current network layer does not belong to a cycle group; if the current network layer is independently selected or not selected, the cycle information corresponding to the current network layer is: the current network layer does not belong to a cycle group; if the current network layer is selected and at least one of the network layers adjacent to the current network layer is also selected, the cycle information corresponding to the current network layer is: whether the current network layer belongs to a cycle group and the number of cycles of the cycle group;
[0085] Multiple network layers are used to determine whether to jump based on the jump index and whether to loop based on the loop index.
[0086] It should be noted that, please refer to Figure 2 The existing large model using the Transformer architecture has at least an embedding layer (Embed), multiple network layers (Layer 1, Layer 2, Layer N) and a language model head layer (LM Head). It can be seen that its network structure is a linear static structure, and the number of its network layers is fixed during the training and use stages of the large model, resulting in the large model requiring a large amount of computing power resources during the inference and training process.
[0087] The present invention has made improvements to this. The network hierarchy structure diagram of the improved large model can be referred to Figure 3 , Figure 3 The large model using the Transformer architecture shown has at least: an embedding layer (Embed), multiple network layers (Layer 1, Layer i, Layer i+1, Layer j, Layer j+1, Layer j+2, Layer N), a language model head layer (LM Head), a skip gate module (Skip Gate) and a recurrent gate module (Recurrent Gate). First, the embedding layer is responsible for converting the input problem to be processed into a vector representation suitable for processing by the large model, which provides basic information for subsequent network layers. The skip gate module receives the output information of the embedding layer and generates a skip index in combination with the optimal dynamic decision rule. The skip index (skipindex) is used to indicate the jump execution status of each network layer in multiple network layers (that is, if the number of network layers is N, the dimension of the skip index is also N, and N is a positive integer), including whether to skip the current network layer or activate the current network layer, such as Figure 3 In the skip index, blue squares indicate activating the current network layer, while black squares indicate skipping the current network layer. This allows large models to effectively determine which layers are required and which layers can be skipped based on the problem being processed during inference, avoiding unnecessary computation and conserving computing resources while maintaining large model performance.
[0088] At the same time, the recurrent gating module also receives the output information of the embedding layer and generates a recurrent index based on the optimal dynamic decision rule. The recurrent index is used to indicate the recurrent information of each network layer in the multiple network layers. This recurrent information can be whether the current network layer belongs to a recurrent group and the number of recurrent groups, or whether the current network layer does not belong to a recurrent group. Specifically, if the current network layer is independently selected (such as Figure 3The independent orange square in the recurrent index of the Figure 3 If the current network layer is selected and at least one of its adjacent network layers is also selected (such as Figure 3 ), its loop information is "belongs to loop group" and includes the number of loops.
[0089] During the execution of multiple network layers, the large model will decide whether to skip certain network layers based on the jump index, and determine whether to perform a cyclic execution of the network layer based on the loop index.
[0090] In summary, multiple network layers automatically skip unnecessary calculations based on the jump index, and perform repeated calculations a specified number of times on the network layers marked as loop groups based on the loop index. This allows the large model to flexibly adjust its network structure when processing problems to adapt to different input requirements and computing power resources, thereby saving computing power resources while ensuring the performance of the large model.
[0091] In a possible implementation, the number of loops is related to the complexity of the problem to be processed. The higher the complexity of the problem to be processed, the greater the number of loops.
[0092] It should be noted that, in specific implementations, the loop gating module will dynamically adjust the number of loops according to the complexity of the problem to be processed. When dealing with simple problems, it can automatically detect low-complexity features (such as short texts, common vocabulary combinations), and set the number of loops in the relevant network layer to 1 basic calculation; when facing high-complexity problems, the number of loops in the core network layer can be increased by analyzing the structure of long texts, the density of professional terms, and the depth of logical nesting. Thus, by dynamically adjusting the number of loops to adapt to the complexity of the problem to be processed, the large model can perform basic calculations on simple problems, and increase the number of loops in the core network layer on high-complexity problems, thereby saving computing resources while optimizing reasoning efficiency and accuracy.
[0093] It should also be noted that during the training phase of a large model, questions can also be labeled, for example, by dividing them into simple and complex groups. For complex questions, a higher number of iterations can be preset, while for simple questions, a lower number of iterations can be preset, or even no iterations at all. This allows for manual intervention, saving computing resources while maintaining the performance of the large model.
[0094] It should be noted that the number of cycles can be determined in the following two ways: either a fixed value can be manually preset (such as setting all cycle groups to execute 3 times), or the number of cycles can be dynamically controlled based on Poisson distribution.
[0095] In one possible implementation, if there is a jump network layer in the loop group whose jump execution state is to skip the current network layer, the difference between the total number of network layers in the loop group and the number of jump network layers is determined. If the difference is greater than or equal to 2, the remaining network layers in the loop group except the jump network layer are looped.
[0096] It should be noted that when there is a skip network layer in the loop group with the jump execution status of "skip current network layer", the difference between the total number of network layers in the loop group and the number of skip network layers must be calculated first. If the difference is ≥ 2 (for example, the loop group contains 5 network layers, 2 of which are marked as skipped, then 3 remain, 3 ≥ 2), then the specified number of loop calculations will be performed on the remaining 3 network layers that have not been skipped. Or if Figure 3 As shown, the three network layers, Layer j, Layer j+1, and Layer j+2, form a loop group, but the skip execution status of Layer j+1 is "skipping the current network layer." Since the difference between the total number of network layers in the loop group and the number of skipped network layers is 3-1=2, that is, the difference of 2 satisfies the condition of ≥2, the remaining two network layers, Layer j and Layer j+2, are looped. As a result, the large model can effectively identify and perform loop calculations on the network layers that have not been skipped in the loop group, ensuring that the computing power of the network layers in the loop group can still be fully utilized during the reasoning process of complex problems, optimizing reasoning efficiency and improving the accuracy of reasoning results.
[0097] In summary, the large model is equipped with a jump gating module and a loop gating module, which can be used to learn based on the sample problem set during the training process of the large model to gradually learn the optimal dynamic decision rule. The jump gating module allows the large model to intelligently skip unnecessary network layers according to the problem to be processed and the optimal dynamic decision rule during the reasoning process, thereby reducing the amount of calculation and the consumption of computing resources, while the loop gating module can repeatedly execute certain key network layers when needed (such as when the problem to be processed is more complex) to enhance the large model's understanding and processing capabilities of complex problems. This ability to dynamically expand the number of layers of the large model's network layer enables the trained large model to adaptively adjust the network layer according to the problem to be processed (actual input situation), thereby saving computing resources, avoiding waste of computing resources, and improving the utilization of computing resources while ensuring the reasoning performance of the large model.
[0098] Figure 4 The present invention shows a large-scale network dynamic expansion device for an intelligent computing center cloud platform that provides computing resources, such as Figure 4 As shown, the device 40 includes:
[0099] The receiving module 401 is configured to execute step S1: receiving a sample question set input by a user;
[0100] Execution module 402 is used to execute step S2: based on the sample problem set, the large model is trained to obtain a trained large model; wherein the large model is provided with a jump gating module and a recurrent gating module;
[0101] The jump gating module and the recurrent gating module are used to learn based on the sample problem set during the training process of the large model to gradually learn the optimal dynamic decision rules;
[0102] The trained large model is used to receive user input for pending problems and reason about them based on the jump gating module, the recurrent gating module, and the optimal dynamic decision rule. During the reasoning process, it dynamically expands its own multiple network layers and obtains output results based on the expanded multiple network layers.
[0103] The optimal dynamic decision rule is used to determine the network layers to be skipped and the network layers to be executed cyclically among the multiple network layers according to the problem to be processed.
[0104] In a possible implementation, step S2 includes:
[0105] Step S21: Based on the sample problem set, the large model is trained in combination with a preset loss function, with the goal of minimizing the loss function to obtain a trained large model;
[0106] Among them, the loss function is used to measure the difference between the output results of the large model and the standard answer, and / or to measure the usage of computing resources during the training process of the large model.
[0107] In one possible implementation, the architecture of the trained large model is a Transformer architecture, and the trained large model of the Transformer architecture includes at least an embedding layer and multiple network layers.
[0108] a jump gating module, configured to receive output information of the embedding layer, and output a jump index based on the output information of the embedding layer and an optimal dynamic decision rule, wherein the jump index is used to indicate a jump execution state of each network layer in the plurality of network layers, wherein the jump execution state includes skipping the current network layer and activating the current network layer;
[0109] a loop gating module, configured to receive output information of the embedding layer, and output a loop index based on the output information of the embedding layer and an optimal dynamic decision rule, wherein the loop index is used to indicate loop information of each network layer in the plurality of network layers;
[0110] The loop information includes any of the following: whether the current network layer belongs to a loop group and the number of loops in the loop group; whether the current network layer does not belong to a loop group;
[0111] If the current network layer is independently selected or is not selected, the cycle information corresponding to the current network layer is: the current network layer does not belong to the cycle group; if the current network layer is selected and at least one of the network layers adjacent to the current network layer is also selected, the cycle information corresponding to the current network layer is: the cycle group to which the current network layer belongs and the cycle count of the cycle group;
[0112] Multiple network layers are used to determine whether to jump based on the jump index and whether to loop based on the loop index.
[0113] In a possible implementation, the number of loops is related to the complexity of the problem to be processed. The higher the complexity of the problem to be processed, the greater the number of loops.
[0114] In one possible implementation, if there is a jump network layer in the loop group whose jump execution state is to skip the current network layer, the difference between the total number of network layers in the loop group and the number of jump network layers is determined. If the difference is greater than or equal to 2, the remaining network layers in the loop group except the jump network layer are looped.
[0115] In summary, the large model is equipped with a jump gating module and a loop gating module, which can be used to learn based on the sample problem set during the training process of the large model to gradually learn the optimal dynamic decision rule. The jump gating module allows the large model to intelligently skip unnecessary network layers according to the problem to be processed and the optimal dynamic decision rule during the reasoning process, thereby reducing the amount of calculation and the consumption of computing resources, while the loop gating module can repeatedly execute certain key network layers when needed (such as when the problem to be processed is more complex) to enhance the large model's understanding and processing capabilities of complex problems. This ability to dynamically expand the number of layers of the large model's network layer enables the trained large model to adaptively adjust the network layer according to the problem to be processed (actual input situation), thereby saving computing resources, avoiding waste of computing resources, and improving the utilization of computing resources while ensuring the reasoning performance of the large model.
[0116] Please refer to Figure 5 The present invention also provides an electronic device 50, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the computer program is executed by the processor 501, the steps of the large-model network dynamic expansion method of the intelligent computing center cloud platform providing computing resources are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0117] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the large-scale model network dynamic expansion method of the intelligent computing center cloud platform providing computing resources, and can achieve the same technical effect. To avoid repetition, the details are not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0118] The present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the large-model network dynamic expansion method of the intelligent computing center cloud platform that provides computing resources, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0119] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method described in the present invention.
[0121] The present invention is described above with reference to the accompanying drawings, but the present invention is not limited to the above-mentioned specific embodiments. The above-mentioned specific embodiments are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A large-scale network dynamic expansion method for an intelligent computing center cloud platform that provides computing resources, characterized in that: The method comprises: Step S1: receiving a sample question set input by a user; Step S2: Based on the sample problem set, the large model is trained to obtain a trained large model; wherein the large model is provided with a jump gating module and a recurrent gating module; The jump gating module and the recurrent gating module are used to learn based on the sample problem set during the training process of the large model, so as to gradually learn the optimal dynamic decision rule; The trained large model is used to receive a problem to be processed input by a user, and to reason about the problem to be processed based on the jump gating module, the recurrent gating module, and the optimal dynamic decision rule, and to dynamically expand its own multiple network layers during the reasoning process, and to obtain an output result based on the expanded multiple network layers; The optimal dynamic decision rule is used to determine the network layers to be skipped and the network layers to be executed cyclically among the multiple network layers according to the problem to be processed.
2. The method according to claim 1, characterized in that The step S2 comprises: Step S21: Based on the sample problem set, the large model is trained in combination with a preset loss function, with the goal of minimizing the loss function, to obtain the trained large model; The loss function is used to measure the difference between the output result of the large model and the standard answer, and / or to measure the usage of computing resources during the training process of the large model.
3. The method according to claim 1, characterized in that The architecture of the trained large model is a Transformer architecture, and the trained large model of the Transformer architecture includes at least an embedding layer and the multiple network layers; The jump gating module is configured to receive the output information of the embedding layer, and output a jump index based on the output information of the embedding layer and the optimal dynamic decision rule, wherein the jump index is used to indicate a jump execution state of each network layer in the multiple network layers, and the jump execution state includes any one of the following: skipping the current network layer and activating the current network layer; The loop gating module is configured to receive output information of the embedding layer, and output a loop index based on the output information of the embedding layer and the optimal dynamic decision rule, wherein the loop index is used to indicate loop information of each network layer in the multiple network layers; The cycle information includes any one of the following: whether the current network layer belongs to a cycle group and the number of cycles of the cycle group, whether the current network layer does not belong to a cycle group; If the current network layer is independently selected or unselected, the cycle information corresponding to the current network layer is: the current network layer does not belong to the cycle group; if the current network layer is selected and at least one of the network layers adjacent to the current network layer is also selected, the cycle information corresponding to the current network layer is: the cycle group to which the current network layer belongs and the cycle count of the cycle group; The plurality of network layers are used to determine whether to jump based on the jump index, and to determine whether to loop based on the loop index.
4. The method according to claim 3, characterized in that The number of cycles is related to the complexity of the problem to be processed. The higher the complexity of the problem to be processed, the greater the number of cycles.
5. The method according to claim 3, characterized in that If there is a skip network layer in the cyclic group whose jump execution state is skipping the current network layer, determining a difference between the total number of network layers in the cyclic group and the number of the skip network layers; If the difference is greater than or equal to 2, the remaining network layers in the cycle group except the skip network layer are cycled.
6. A large-scale network dynamic expansion device for an intelligent computing center cloud platform that provides computing resources, characterized in that: The device comprises: The receiving module is configured to execute step S1: receiving a sample question set input by a user; An execution module, configured to execute step S2: training the large model based on the sample problem set to obtain a trained large model; wherein the large model is provided with a jump gating module and a recurrent gating module; The jump gating module and the recurrent gating module are used to learn based on the sample problem set during the training process of the large model, so as to gradually learn the optimal dynamic decision rule; The trained large model is used to receive a problem to be processed input by a user, and to reason about the problem to be processed based on the jump gating module, the recurrent gating module, and the optimal dynamic decision rule, and to dynamically expand its own multiple network layers during the reasoning process, and to obtain an output result based on the expanded multiple network layers; The optimal dynamic decision rule is used to determine the network layers to be skipped and the network layers to be executed cyclically among the multiple network layers according to the problem to be processed.
7. The device according to claim 6, characterized in that The step S2 comprises: Step S21: Based on the sample problem set, the large model is trained in combination with a preset loss function, with the goal of minimizing the loss function, to obtain the trained large model; The loss function is used to measure the difference between the output result of the large model and the standard answer, and / or to measure the usage of computing resources during the training process of the large model.
8. A server, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of a large-model network dynamic expansion method for an intelligent computing center cloud platform providing computing resources as described in any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a large-model network dynamic expansion method for an intelligent computing center cloud platform that provides computing resources as described in any one of claims 1-5.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of a large-model network dynamic expansion method for an intelligent computing center cloud platform that provides computing resources as described in any one of claims 1-5.
Citation Information
Patent Citations
Reasoning acceleration method of large language model based on hybrid neural network structure
CN117787410A
Cited By
Method and device for intelligent computing center cloud platform to carry out self-evolution depth reasoning of group islands through computing power
CN121352031A