Mixed depth large model reasoning acceleration method and device based on unified decision
By calculating the similarity of hidden layers, selecting the decision interval and deploying a unified dynamic decision module, the problem of the reasoning speed of the MoD model in actual deployment is solved, and the model inference speed is improved and downstream task performance is maintained.
Patent Information
- Application Number
- CN202510244396.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-11
AI Technical Summary
The existing Mixture of Depths (MoD) models are degraded inference speed due to additional overhead in actual deployment, and downstream task performance is affected.
By calculating the similarity of hidden layers in the target model, select the decision interval with the highest average similarity, and deploy a unified dynamic decision module before this interval, skip the non-critical hidden layers for calculation, and synchronize the parameter update with the target model during the training process.
It significantly improves the model inference speed, maintains the performance of downstream tasks, solves the problem of degradation inference speed caused by additional overhead, and improves the efficiency and performance of the model in practical applications.
Smart Images

Figure CN120297404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model acceleration inference, and particularly to a method and device for accelerating the inference of a hybrid deep large model based on unified decision-making. Background Art
[0002] Large language models (LLMs) perform excellently in text generation tasks, but their huge number of parameters lead to huge computational costs and memory overheads in the inference stage. Especially in resource-constrained local deployment scenarios, this problem is particularly prominent. To address this challenge, researchers have proposed various model compression strategies, including quantization, distillation, and pruning, etc. Among them, pruning improves the inference efficiency by deleting redundant parameters and structures in the model. Further, the Mixture of Depths (MoD) method is proposed, aiming to dynamically adjust the computational resource allocation according to the input context information. By intelligently selecting a subset of model layers for execution, it reduces the floating-point operation volume (FLOPs) and memory occupancy, thereby improving the inference efficiency. However, the existing MoD technology has obvious defects in actual deployment. Although the FLOPs are reduced by skipping some Transformer layers, due to the addition of a dynamic decision-making module before each layer, it leads to additional kernel scheduling and CPU-GPU communication overheads, resulting in a decrease rather than an increase in the actual inference speed. For example, on a single A100 graphics card, the inference speed of the original Llama-2-7b model is 45 tokens / s, while that of the MoD model is only 39.47 tokens / s, and the inference speed has decreased by 12.3%. In addition, the performance of the MoD model in downstream tasks is also affected by the unreasonable selection of the decision interval, resulting in poor model performance. Therefore, there is an urgent need for a new inference acceleration scheme. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method and device for accelerating the inference of a hybrid deep large model based on unified decision-making to eliminate or improve one or more defects existing in the prior art and solve the problem of the decrease in the inference speed of the existing Mixture of Depths (MoD) model due to additional overheads in actual deployment.
[0004] One aspect of the present invention provides a method for accelerating the inference of a hybrid deep large model based on unified decision-making, the method comprising the following steps:
[0005] Randomly sample the preset data set, input the sample into the target model for processing, and calculate the similarity between the input and output of each hidden layer in the target model; based on the preset number of layers as an interval, slide and select consecutive hidden layers in the target model, calculate the similarity between the input and output of each hidden layer, calculate the average similarity of the input and output of each hidden layer in each interval, and take the highest as the decision interval;
[0006] Deploy a unified dynamic decision-making module in front of the decision interval. The unified dynamic decision-making module selects a target hidden layer for participating in the operation of the current input within the decision interval based on the output of the previous hidden layer of the decision layer, so as to skip the remaining hidden layers in the decision interval except the target hidden layer during the calculation process; the unified dynamic decision-making module synchronizes parameter updates with the target model based on a specified task during the training process.
[0007] In some embodiments, the preset data set is consistent with the task content of the target model, is classified according to preset classification attributes, and samples are randomly selected in each classification according to a set ratio to participate in the selection calculation of the decision interval.
[0008] In some embodiments, calculating the similarity between the input and output of each hidden layer includes: calculating the similarity using Euclidean distance, Manhattan distance, Mahalanobis distance, cosine similarity, Pearson correlation coefficient or Jaccard similarity coefficient.
[0009] In some embodiments, the method further includes:
[0010] Based on a preset number of layers as an interval, continuously select hidden layers in the target model, calculate the similarity between the input and output of each hidden layer, calculate the average similarity of the input and output of each hidden layer in each interval, screen out the top set number of candidate decision intervals with the highest average similarity, determine one or more non-overlapping ones of the candidate decision intervals as the decision intervals, and respectively configure the corresponding unified dynamic decision-making modules.
[0011] In some embodiments, the unified dynamic decision-making module consists of a continuous root mean square normalization layer and a multi-layer perceptron.
[0012] In some embodiments, the unified dynamic decision-making module synchronizes parameter updates with the target model based on a specified task during the training process, including:
[0013] Randomly collect samples based on the preset data set;
[0014] Freeze the parameters of the target model, and use the collected samples to update the parameters of the unified dynamic decision-making module until a preset termination condition is reached;
[0015] Freeze the unified dynamic decision-making module, and use the collected samples to update the parameters of the target model until a preset termination condition is reached;
[0016] Use the collected samples to update the parameters of the unified dynamic decision-making module and the target model simultaneously until a preset termination condition is reached.
[0017] In some embodiments, the method further includes:
[0018] Setting an average similarity threshold, and requiring that the average similarity of each hidden layer within the decision interval is higher than the average similarity threshold; otherwise, terminating the inference acceleration and generating a prompt message.
[0019] On the other hand, the present invention further provides a hybrid deep large model inference acceleration device based on unified decision-making, including a processor, a memory, and a computer program / instruction stored on the memory. The processor is used to execute the computer program / instruction, and when the computer program / instruction is executed, the device implements the steps of the above method.
[0020] On the other hand, the present invention further provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the above method are implemented.
[0021] On the other hand, the present invention further provides a computer program product, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the above method are implemented.
[0022] The beneficial effects of the present invention are as follows:
[0023] For the hybrid deep large model inference acceleration method and device based on unified decision-making of the present invention, by processing the samples of a preset data set, calculating the similarity between the input and output of each hidden layer in the target model, then sliding based on a preset number of layers to select consecutive hidden layers, calculating the average similarity of the input and output of the hidden layers within each interval, and selecting the interval with the highest average similarity as the decision interval; deploying a unified dynamic decision module before the decision interval, which based on the output of the previous hidden layer in the decision interval, selects the target hidden layer that participates in the execution of the operation for the current input within the decision interval, so as to skip other hidden layers except the target hidden layer within the decision interval during the calculation process, realizing inference acceleration; the unified dynamic decision module synchronizes parameter updates with the target model based on a specified task during the training process. This method effectively reduces the amount of calculation in the model inference process, significantly improves the inference speed, while maintaining the performance of the model in downstream tasks, solves the problem of the decline in inference speed caused by additional overhead in the prior art, and improves the efficiency and performance of the model in practical applications.
[0024] The additional advantages, objectives, and features of the present invention will be partially elaborated in the following description, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and the drawings.
[0025] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and the above and other objectives achievable with the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:
[0027] Fig. 1(a) is a structural diagram of the original large language model.
[0028] Fig. 1(b) is a structural diagram of the hybrid depth network MoD.
[0029] Figure 2 is a graph of the time required for any layer of the model in Fig. 1 under the cases of selecting to skip, execute, and based on the original structure.
[0030] Figure 3 is a structural diagram of the model of the method for accelerating the inference of the hybrid depth large model based on unified decision according to the present invention.
[0031] Figure 4 is a diagram of the calculation process of the unified dynamic decision module in the method for accelerating the inference of the hybrid depth large model based on unified decision according to the present invention.
[0032] Figure 5 is a graph of the cosine similarity between the input and output of each layer of the Llama-2-7b model.
[0033] Figure 6 is a comparison graph of the inference speeds of the Llama-2-7b model, the MoD model, and the MoD model based on the unified decision of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0035] Here, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0036] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0037] Large language models (LLMs) have demonstrated remarkable capabilities in text generation tasks. Compared with the approach of deploying LLMs on cloud servers and providing remote services, local deployment by users presents many obvious advantages. First, local inference has significant advantages in terms of privacy. Second, local inference does not rely on network connections, improving the robustness of the system. Third, by deploying LLMs locally, enterprises can significantly reduce the cost of cloud computing resources and do not have to configure expensive high-performance computing resources for the operation of LLMs. Finally, local LLMs can provide more personalized services according to the specific usage habits and preferences of users, enabling personalized customization. However, LLM models often contain billions of parameters, requiring expensive computational costs and memory overheads during the inference stage. Therefore, deploying LLMs on resource-constrained platforms is very challenging.
[0038] To address the above issues, researchers have proposed various strategies, including quantization, distillation, and pruning. Among them, pruning is a typical model structure compression method that reduces the size and computational amount of the model by reducing redundant parameters and structures in the model, thereby improving the inference efficiency of the model. Model pruning is divided into two types: structured pruning and unstructured pruning. Structured pruning refers to deleting some redundant structures in the neural network, such as neurons, modules, layers, etc. Unstructured pruning refers to setting some weights in the neural network to zero and using special hardware support for acceleration.
[0039] Pruning permanently removes the weights and structures in the model, but it is unreasonable to allocate the same resources to different generation tasks. During the generation process, the importance of different tokens is not the same, and it is a more reasonable strategy to allocate computational resources according to importance. The Mixture of Depths (MoD) method is an innovative technique that can dynamically adjust the allocation of computational resources based on input context information. Specifically, this method analyzes the context features of the input data and intelligently selects a subset of model layers for execution instead of having all model layers participate in the inference calculation. Through this strategy, MoD can not only significantly reduce the floating-point operation amount (FLOPs) required during model inference but also effectively reduce memory occupancy, thereby improving inference efficiency. In addition, the core idea of this method is to make full use of the sparsity and distribution characteristics of the input information to flexibly regulate the model's computational path while ensuring that the performance and accuracy of the model are not significantly affected. This mechanism of dynamically allocating computational resources is particularly suitable for large-scale models that require efficient inference, especially in the resource-constrained local device environment, showing broad application potential.
[0040] In the prior art, as shown in FIGS. 1(a) and (b), the MoD model places a dynamic decision module in front of each Transformer layer in the original LLM, and through further training (continuous pre-training or fine-tuning), the model is enabled with the ability of dynamic decision-making. During inference, the dynamic decision module makes a decision on whether to execute the corresponding Transformer layer.
[0041] Suppose the input to Transformer layer l is a sequence of token embeddings of length S, that is For a given token embedding, the output corresponding to its dynamic decision module is a scalar calculated by linear projection and passed through the sigmoid activation function, that is where, w θ is the weight of the dynamic decision module.
[0042] Given the token embedding in the input sequence it passes through the dynamic decision module r l before the l-th Transformer layer f l in the large language model. During inference, the dynamic decision module outputs to the l-th Transformer layer, indicating a decision to skip or execute the l-th Transformer layer based on the input token features . The execution result of the l-th layer is defined as:
[0043]
[0044] where, σ is the decision threshold, and generally σ = 0.5.
[0045] Since the MoD adds a dynamic decision module in front of each Transformer layer, the number of additional modules added to the model is relatively large (equivalent to the number of model layers), and these additional modules bring a large amount of additional overhead (kernel scheduling, CPU-GPU communication, etc.). Due to the above disadvantages, although the MoD model reduces the FLOPs during the inference process by skipping some Transformer layers, it cannot achieve a significant acceleration effect in actual deployment inference, and even reduces the inference speed.
[0046] The actual inference speed test on a single A100 graphics card shows such results. The inference speed of the original Llama-2-7b is 45 tokens / s, and that of the MoD is 39.47 tokens / s. The inference speed of the MoD has decreased by 12.3% instead.
[0047] Figure 2 It shows the time overhead of each part of a single Transformer layer in Llama-2-7b and the MoD model when executed using the Huggingface Transformers framework. Although the number of parameters in the dynamic decision-making module is much smaller than that of the Transformer layer, due to the additional overhead, its running time is about 41% of the execution time of a single Llama layer. At the same time, since the CPU needs to select the corresponding branch for execution after obtaining the decision of the dynamic execution module, it brings additional CPU-GPU communication overhead, resulting in more time being required to execute a single Transformer layer in the MoD than in the corresponding Llama layer. The above-mentioned overhead offsets the benefits of skipping some layers, thereby affecting the inference speed of the MoD.
[0048] Specifically, the present invention provides a method for accelerating the inference of a hybrid deep large model based on unified decision-making, and the method includes the following steps S101 to S102:
[0049] Step S101: Randomly sample the preset data set, input the samples into the target model for processing, and calculate the similarity between the input and output of each hidden layer in the target model; based on the preset number of layers as an interval, slide to select consecutive hidden layers in the target model, calculate the similarity between the input and output of each hidden layer, calculate the average similarity of the input and output of each hidden layer in each interval, and take the highest as the decision interval.
[0050] Step S102: Deploy a unified dynamic decision-making module before the decision interval. The unified dynamic decision-making module selects the target hidden layer that participates in the execution operation for the current input within the decision interval based on the output of the previous hidden layer before the decision layer, so as to skip the remaining hidden layers in the decision interval except the target hidden layer during the calculation process; the unified dynamic decision-making module synchronizes parameter updates with the target model based on the specified task during the training process.
[0051] In step S101, samples are randomly collected from the preset data set and input into the target model for processing. Here, the data set is usually a representative data set related to the target model task, and sampling the samples is for subsequent model analysis and optimization based on these samples.
[0052] Next, the similarity between the input and output of each hidden layer in the target model is calculated using the selected samples. The input and output of the hidden layer are the feature representations of the data before and after being processed by this layer, and calculating the similarity is to measure the processing effect of this hidden layer on the input data. If the similarity between the input and output of a certain hidden layer is very high, it means that this layer makes little change to the data in the model inference or there is a certain degree of redundancy.
[0053] Then, based on the preset number of layers, continuously selected hidden layers in the target model are slid as an interval. The hidden layers in the model can be grouped according to a certain number of layers (such as a fixed range of layers) to form multiple consecutive layer intervals. This preset number of layers can be set according to the scale of the target model and the requirements of inference acceleration optimization.
[0054] For the hidden layers in each interval, calculate the similarity between the input and output of each hidden layer, and then calculate the average similarity of the input and output of each hidden layer in this interval. Take the interval with the highest average similarity as the decision interval. This process is similar to finding the most suitable layer interval for simplification processing in a sliding window, reducing unnecessary computational effort in the model inference process.
[0055] In step S102, a unified dynamic decision-making module is deployed before the determined decision interval. The main function of this module is to select the target hidden layer that participates in the execution of operations for the current input within the decision interval based on the output of the hidden layer before the decision layer. The unified dynamic decision-making module can be regarded as an intelligent controller. During model inference, it can dynamically determine which hidden layers need to participate in the current calculation based on the output features of the previous layer, thereby skipping the remaining hidden layers in the decision interval except the target hidden layer. In this way, during the calculation process, the model does not need to execute each hidden layer layer by layer, but only executes the most critical hidden layers through the decision of the unified dynamic decision-making module, greatly reducing the number of calculation steps and resource consumption. In addition, the unified dynamic decision-making module synchronizes parameter updates with the target model based on the specified task during the training process. This means that when training the target model to optimize its performance for a specific task, the parameters of the unified dynamic decision-making module will also be updated accordingly. In this way, the learning of the unified dynamic decision-making module can be more adapted to the task requirements of the target model, and better master how to select appropriate hidden layers for inference in different task scenarios to ensure the balance between model performance improvement and inference acceleration.
[0056] The core of the entire technical solution lies in determining the layer intervals with high redundancy (decision intervals) by analyzing the similarity of the hidden layers, and then using a unified dynamic decision-making module to dynamically select key hidden layers during the inference process and skip the redundant layers, thereby improving the model inference speed. First, in step S101, by calculating the similarity between the input and output of the hidden layer, the processing effect of the hidden layer on the data can be quantified. A hidden layer with a high similarity means that the data does not change much after passing through this layer, and there may be redundancy. These layers can be used as candidate layers for simplification. Then, by means of a sliding window, a continuous layer interval with the highest average similarity is selected as the decision interval. The layers within this interval contribute relatively little to the overall output of the model and can be used as the key area for optimization during inference. Next, in step S102, the unified dynamic decision-making module is responsible for determining which hidden layers should be executed for the current input within the decision interval based on the output features of the previous layer. By learning the relationship between the input data and the output of the hidden layer, this decision-making module can flexibly select the most relevant hidden layer for inference under different input conditions, avoiding the fixed execution of all hidden layers and effectively reducing the computational cost. At the same time, the unified dynamic decision-making module is updated synchronously with the parameters of the target model, ensuring that the decision-making ability of this module can be continuously improved as the target model is optimized, making it more adaptable to the processing requirements of the model for specific tasks, and thus significantly improving the inference efficiency on the premise of ensuring the model performance.
[0057] This technology optimizes the determination of the decision interval and the deployment of the unified dynamic decision-making module, reduces the unnecessary computational amount in the model inference process, and at the same time collaboratively optimizes with the target model during the parameter update process, achieving a balance between efficient inference and model performance improvement, and effectively solving the problem of the decline in inference speed caused by redundant calculations in the existing technology.
[0058] In some embodiments, in step S101, the preset data set is consistent with the task content of the target model, classified according to the preset classification attributes, and samples are randomly selected according to the set ratio in each classification to participate in the selection calculation of the decision interval.
[0059] The purpose of this step is to ensure that the selection of decision intervals is more scientific, reasonable, and representative, thereby improving the stability and reliability of the model inference acceleration effect. Specifically, by making the preset data set consistent with the task content of the target model, it can ensure that the sample data used for selecting decision intervals is highly relevant to the actual application scenario of the model, avoiding the deviation in the selection of decision intervals caused by data mismatch. Classifying according to the preset classification attributes can further refine the sample data, ensuring that suitable decision intervals can be found in different categories and improving the performance of the model in various categories. Randomly selecting samples according to the set ratio in each classification to participate in the calculation of the selection of decision intervals can ensure the diversity and randomness of the samples, avoiding the inaccuracy in the selection of decision intervals caused by the concentration or partiality of sample selection, so as to ensure that the selection of decision intervals can truly reflect the performance of the model on different categories and different data, providing a more reliable basis for subsequent inference acceleration.
[0060] Specifically, it is manifested as follows: First, the selection of decision intervals is more in line with the actual task requirements of the target model, which can effectively avoid the deviation in the selection of decision intervals caused by data mismatch and improve the inference speed and performance of the model in actual applications. Second, by randomly selecting samples in different classifications, it can ensure that the selection of decision intervals has broad representativeness, avoiding the inaccuracy in the selection of decision intervals caused by the limitation of sample selection, thereby improving the inference acceleration effect of the model in various categories and enhancing the generalization ability and adaptability of the model. In addition, this step can also improve the stability and reliability of the selection of decision intervals, reduce the fluctuation in the selection of decision intervals caused by the randomness of sample selection, provide a more stable basis for the inference acceleration of the model, contribute to improving the stability and reliability of the model in actual applications, and enhance the user experience.
[0061] In some embodiments, calculating the similarity between the input and output of each hidden layer includes: calculating the similarity using Euclidean distance, Manhattan distance, Mahalanobis distance, cosine similarity, Pearson correlation coefficient, or Jaccard similarity coefficient.
[0062] In some embodiments, the method further includes: based on the preset number of layers as the target for interval sliding to select consecutive hidden layers in the target model, calculating the similarity between the input and output of each hidden layer, calculating the average similarity of the input and output of each hidden layer within each interval, screening out the top set number of candidate decision intervals with the highest average similarity, determining one or more non-overlapping candidate decision intervals as decision intervals, and respectively configuring corresponding unified dynamic decision modules.
[0063] The purpose of this step is to screen out multiple high-quality candidate decision intervals and determine the final decision interval from them for model inference acceleration. By sliding and selecting consecutive hidden layers in the target model based on a preset number of layers, calculating the average similarity of the input and output of the hidden layers within each interval, and screening out the top several candidate decision intervals with the highest average similarity, it can be ensured that these candidate decision intervals are all relatively important and have less redundancy. Then, one or more non-overlapping decision intervals are determined, and a corresponding unified dynamic decision module is configured for each interval, which can avoid mutual interference between different decision intervals and improve the overall efficiency and performance of model inference acceleration.
[0064] In some embodiments, the unified dynamic decision module consists of a consecutive root mean square normalization layer and a multi-layer perceptron. The root mean square normalization layer (RMSNorm) normalizes the input tensor by calculating the root mean square variance along the channel dimension, which can normalize the output features of the previous hidden layer of the decision area, making their distribution more stable, thus providing better input for the subsequent multi-layer perceptron processing. The multi-layer perceptron (MLP) further performs non-linear transformation on the normalized features through two fully connected layers and an activation function (such as ReLU), so as to more accurately capture the complex relationship between the input features and the hidden layer selection. Specifically, the first fully connected layer maps the normalized features to an intermediate space, the activation function introduces non-linearity, and the second fully connected layer maps the intermediate features to the output space to obtain the selection result of the hidden layer within the decision interval.
[0065] The root mean square normalization layer can more effectively stabilize the training process, accelerate the convergence speed of the model, and improve the training efficiency of the model. Secondly, the multi-layer perceptron can more accurately learn the complex relationship between the input features and the hidden layer selection, so as to more precisely select the key hidden layers for calculation and skip the redundant hidden layers, further improving the inference speed of the model. This structure can also improve the generalization ability of the model, enabling it to maintain good performance in different input data and task scenarios. Through this structure, the unified dynamic decision module can better adapt to different input features and task requirements, achieve more efficient and accurate hidden layer selection, and thus significantly improve the inference speed and efficiency of the model while ensuring the model performance.
[0066] In some embodiments, in step S102, the unified dynamic decision module synchronously updates parameters with the target model during the training process, including steps S201 - S204:
[0067] Step S201: Randomly sample based on a preset data set. Randomly sample based on a preset data set. These samples are usually consistent with the task content of the target model and may be classified according to preset classification attributes to ensure the diversity and representativeness of the samples.
[0068] Step S202: Freeze the parameters of the target model, and use the collected samples to update the parameters of the unified dynamic decision-making module until a preset termination condition is reached. In this stage, the parameters of the target model remain unchanged, while the unified dynamic decision-making module optimizes its own parameters by learning the sample data. On the premise of maintaining the existing capabilities of the target model, it first learns how to allocate appropriate layers for reasoning according to the context information to better adapt to the task requirements of the target model.
[0069] Step S203: Freeze the unified dynamic decision-making module, and use the collected samples to update the parameters of the target model until a preset termination condition is reached. Since the unified dynamic decision-making module already has the ability to judge and select the importance of layers at this time, the target model can further fine-tune its own representation under the fine-grained guidance given by the unified dynamic decision-making module, so that the target model further adapts to the decision-making of the unified dynamic decision-making module while maintaining its original capabilities.
[0070] Step S204: Use the collected samples to update the parameters of both the unified dynamic decision-making module and the target model until a preset termination condition is reached. Use the collected samples to update the parameters of both the unified dynamic decision-making module and the target model until a preset termination condition is reached. In this stage, the parameters of the unified dynamic decision-making module and the target model are updated simultaneously. Through the way of collaborative learning, the two can cooperate better to achieve the balance of model inference acceleration and performance improvement.
[0071] In some embodiments, the method further includes: setting an average similarity threshold, requiring that the average similarity of each hidden layer in the decision interval is higher than the average similarity threshold, otherwise terminate the inference acceleration and generate a prompt message. Determine a suitable average similarity threshold, which can be set according to factors such as the performance requirements of the model, the complexity of the task, and the actual application scenario. When determining the decision interval, calculate the average similarity of each hidden layer in each candidate decision interval and compare it with the set average similarity threshold. If the average similarity of each hidden layer in a certain decision interval is higher than the average similarity threshold, it is considered that the decision interval is valid and can be used for model inference acceleration; otherwise, if the average similarity is lower than the threshold, it is considered that there are many redundant or unimportant layers in the decision interval, which may affect the performance of the model. Therefore, terminate the inference acceleration process and generate the corresponding prompt message to inform the user or the system that it is necessary to reselect the decision interval or adjust the model structure.
[0072] On the other hand, the present invention also provides a hybrid deep large model inference acceleration device based on unified decision-making, including a processor, a memory, and computer programs / instructions stored on the memory. The processor is used to execute the computer programs / instructions, and when the computer programs / instructions are executed, the device implements the steps of the above method.
[0073] On the other hand, the present invention also provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0074] On the other hand, the present invention also provides a computer program product, including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0075] The present invention will be described below in conjunction with a specific embodiment:
[0076] In order to achieve the actual inference acceleration of the MoD model, this embodiment proposes a MoD model structure based on unified decision-making, and uses a unified dynamic decision-making module and a unified decision interval strategy to reduce the number of decision-making modules from the same as the number of model layers to 1; at the same time, a interval selection strategy based on cosine similarity is proposed to ensure the performance of the model on downstream tasks.
[0077] This embodiment improves the model structure of MoD, reduces the number of dynamic decision-making modules from the same as the number of model layers to 1, and at the same time selects continuous l layers in the model as a unified decision interval, and uses each layer in the decision interval of this dynamic decision-making module to decide whether to execute.
[0078] Let represent the i-th layer in the LLM with input X and output X o , and the parameters of this layer are For example, the feed-forward neural network (FeedForward) in the Transformer model can be expressed as Among them respectively represent the input projection matrix and the output projection matrix. Define the MoD model function based on unified decision-making as F U-MoD , which is used to wrap each layer in the unified decision interval of length l in the LLM, so that:
[0079]
[0080] Among them, G(X|W G ) ∈ {0, 1} is the calculation result of the unified dynamic decision-making module, with learnable parameters W G , and X is the output of the previous Transformer layer of the dynamic activation module.Figure 3 Shows the MoD model structure based on unified decision-making.
[0081] Assume the input to the LLM is Since there is generally only one user on the local device, assume Batchsize is 1. This input is a sequence with length T and embedding dimension d. For each token input Then there is:
[0082]
[0083] According to the context, if the unified dynamic decision-making module decides to skip, the input will be directly connected to the output; otherwise, it will go through the calculation logic of this layer.
[0084] The core of this embodiment is the unified dynamic decision-making function G(X|W G ), which only assigns partial layers for inference for each input token. The calculation process is as Figure 4 shown. For the input The output of the dynamic decision-making function is a binary mask matrix:
[0085]
[0086] The dynamic decision-making function first calculates the corresponding importance scores for each layer in layer l for correctly inferring the token at the current position based on the input:
[0087]
[0088] Select the top-K layers with the highest importance scores as the execution layers to participate in model inference. The corresponding output of the dynamic decision-making function is 1, and the remaining l - K layers are non-execution layers that do not participate in model inference, and their dynamic decision-making function output values are 0.
[0089] The decision interval of MoD is all layers of the model, that is, all layers in the model may be skipped. This strategy reduces the performance of the model on downstream tasks. To solve the problem of the model's poor performance on downstream tasks and at the same time be compatible with the above-mentioned MoD model structure based on unified decision-making, this method proposes a unified decision interval selection strategy based on cosine similarity.
[0090] To measure the importance of different layers in the LLM, we randomly select samples from the open-source pre-trained data WikiText. Then, we record the hidden states generated by the LLM for these samples and calculate the cosine similarity between the input and output hidden states of each layer. The calculation of the cosine similarity can be expressed as follows:
[0091]
[0092] Among them, D represents the hidden states of different samples in the record, representing the input and output hidden states of the i-th sample respectively, d represents the hidden size, and L represents the sequence length of each sample.
[0093] Figure 5 Shows the cosine similarity between the inputs and outputs of each layer of the Llama-2-7b model. Generally speaking, as the number of layers increases, the curve gradually rises, indicating that the similarity between the input and output increases layer by layer until it slightly decreases when approaching the later stage. The conclusion is that the cosine similarity between the inputs and outputs of most layers in the model is very high (greater than 0.9), indicating that the contribution of these layers to the final output is small and there is redundancy in a large number of layers; as the number of layers increases, at higher levels of the model, the changes in the input and output gradually tend to be consistent, indicating that in deep networks, the representation of hidden states is more stable. However, when the number of layers approaches the end, the similarity decreases, which reflects that the output is more sensitive or diverse to changes in the input.
[0094] To ensure the performance of the MoD model based on unified decision-making, under the condition of a given decision interval length l, select the l layers with the highest average cosine similarity within the interval as the decision interval, that is:
[0095]
[0096] Among them, S is the starting layer number of the interval.
[0097] This embodiment uses Nvidia series graphics cards as the hardware basis and uses the Python language to implement the algorithm. As Figure 6 shown, the inference speeds of Llama-2-7b, MoD, and the MoD model based on unified decision-making are tested on a single A100 card. This embodiment achieves an acceleration of 1.46 times and realizes the inference acceleration effect.
[0098] As shown in Table 1, on 4 zero-shot benchmarks (BoolQ, OpenBookQA, PIQA, MMLU), the model performance of this solution is 16.2% higher than that of the original MoD model and maintains 93.5% of the performance of the original Llama-2-7b model.
[0099] Table 1 Performance evaluation of three models
[0100]
[0101] Correspondingly, the present invention further provides a device / system, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device / system implements the steps of the method described above.
[0102] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium well-known in the art.
[0103] In summary, for the method and device for accelerating inference of a hybrid deep large model based on unified decision-making according to the present invention, by processing samples of a preset data set, calculating the similarity between the input and output of each hidden layer in the target model, and then slidingly selecting continuous hidden layers based on a preset number of layers, calculating the average similarity of the input and output of the hidden layers in each interval, and selecting the interval with the highest average similarity as the decision interval; a unified dynamic decision-making module is deployed before the decision interval. Based on the output of the previous hidden layer in the decision interval, the module selects the target hidden layer that participates in performing operations on the current input within the decision interval, so as to skip other hidden layers except the target hidden layer within the decision interval during the calculation process, thereby realizing inference acceleration; the unified dynamic decision-making module synchronizes parameter updates with the target model based on a specified task during the training process. By optimizing the selection of the decision interval and the deployment and training of the unified dynamic decision-making module, the method effectively reduces the amount of calculation in the model inference process, significantly improves the inference speed, and at the same time maintains the performance of the model in downstream tasks, solves the problem of the decrease in inference speed caused by additional overhead in the prior art, and improves the efficiency and performance of the model in practical applications.
[0104] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or a communication link.
[0105] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0106] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0107] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for accelerating the inference of a hybrid deep large model based on unified decision-making, characterized in that, The method includes the following steps: Randomly sample a preset data set, input the samples into a target model for processing, and calculate the similarity between the input and output of each hidden layer in the target model; Based on a preset number of layers as an interval, slide to select consecutive hidden layers in the target model, calculate the similarity between the input and output of each hidden layer, calculate the average similarity of the input and output of each hidden layer in each interval, and take the highest as the decision interval; Deploy a unified dynamic decision-making module before the decision interval. The unified dynamic decision-making module selects a target hidden layer for participating in the operation for the current input within the decision interval based on the output of the previous hidden layer before the decision layer, so as to skip the remaining hidden layers other than the target hidden layer in the decision interval during the calculation process; The unified dynamic decision-making module synchronizes parameter updates with the target model based on a specified task during the training process.
2. The method for accelerating inference of a hybrid deep model based on unified decision-making according to claim 1, wherein The preset data set is consistent with the task content of the target model, is classified according to a preset classification attribute, and samples are randomly selected in each classification according to a set ratio to participate in the selection calculation of the decision interval.
3. The method for accelerating inference of a hybrid deep model based on unified decision-making according to claim 1, wherein Calculating the similarity between the input and output of each hidden layer includes: calculating the similarity using Euclidean distance, Manhattan distance, Mahalanobis distance, cosine similarity, Pearson correlation coefficient or Jaccard similarity coefficient.
4. The inference acceleration method of the hybrid deep model based on unified decision-making according to claim 1, wherein, The method further includes: Based on a preset number of layers as an interval, slide to select consecutive hidden layers in the target model, calculate the similarity between the input and output of each hidden layer, calculate the average similarity of the input and output of each hidden layer in each interval, screen out the top set number of candidate decision intervals with the highest average similarity, determine one or more non-overlapping ones among the candidate decision intervals as the decision interval, and respectively configure the corresponding unified dynamic decision-making module.
5. The method for accelerating the inference of a hybrid deep model based on unified decision-making according to claim 1, wherein The unified dynamic decision-making module consists of consecutive root mean square normalization layers and a multi-layer perceptron.
6. The method for accelerating the inference of a hybrid deep model based on unified decision-making according to claim 1, wherein The unified dynamic decision-making module synchronizes parameter updates with the target model based on a specified task during the training process, including: Randomly sample based on the preset data set; Freeze the parameters of the target model, and use the collected samples to update the parameters of the unified dynamic decision-making module until a preset termination condition is reached; Freeze the unified dynamic decision-making module, and use the collected samples to update the parameters of the target model until a preset termination condition is reached; Use the collected samples to update the parameters of the unified dynamic decision-making module and the target model simultaneously until a preset termination condition is reached.
7. The method for accelerating the inference of a hybrid deep model based on unified decision-making according to claim 1, characterized in that The method further includes: Set an average similarity threshold, and require the average similarity of each hidden layer within the decision interval to be higher than the average similarity threshold, otherwise terminate the inference acceleration and generate a prompt message.
8. A hybrid deep model inference acceleration device based on unified decision-making, comprising a processor, a memory, and computer programs / instructions stored on the memory, characterized in that, The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the device implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, The steps of the method according to any one of claims 1 to 7 are implemented when the computer program / instructions are executed by a processor.