Method, device, medium, product and system for reasoning acceleration of pre-training model
By extracting and fusing semantic features at different levels, cutting non-important features based on the importance, combining dynamic cropping and mixed precision training, the problem of slow inference speed of pre-trained models is solved, achieving faster inference speed and lower computational storage requirements.
Patent Information
- Application Number
- CN202510895972.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-30
AI Technical Summary
It is difficult for existing solutions to effectively accelerate the inference speed of large-scale pre-trained models, and it is difficult for the existing technology to achieve the expected results in a single method.
By extracting semantic features at different levels from natural language texts, performing feature fusion, identifying important and non-important features based on the importance of the fusion features, and cropping non-important features, combining dynamic cropping strategies and mixed precision training, the reasoning process of the pre-trained model is optimized.
It improves the expression and generalization capabilities of the model, significantly accelerates the inference speed of pre-trained models, reduces the computing and storage requirements, and is suitable for devices with resource constraints.
Smart Images

Figure CN120408126A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of inference of pre-trained models, and in particular to an inference acceleration method for pre-trained models, an inference acceleration device for pre-trained models, a computer-readable storage medium, a computer program product, and an inference acceleration system for pre-trained models. Background Art
[0002] Large-scale pre-trained models have become the standard for natural language processing (NLP) and other machine learning tasks. However, these large models usually have billions or even more parameters. Although they greatly improve the model performance, they also bring problems of high computational cost and long inference time. In existing solutions, although the inference time of the model can be reduced by pruning, quantization, knowledge distillation (KD), and sparsity to accelerate inference, due to the extremely high complexity of the pre-trained model itself, it is often difficult for a single existing technology to achieve a satisfactory acceleration effect. Summary of the Invention
[0003] This application provides an inference acceleration method for pre-trained models, an inference acceleration device for pre-trained models, a computer-readable storage medium, a computer program product, and an inference acceleration system for pre-trained models, so as to at least solve the problem that the acceleration means for the inference speed of pre-trained models in existing solutions are difficult to achieve the expected effect.
[0004] This application provides an inference acceleration method for pre-trained models, including: obtaining a natural language text, and extracting semantic features at different levels from the natural language text; performing feature fusion on the semantic features at different levels to obtain a plurality of fused features; determining important fused features and unimportant fused features among the plurality of fused features based on the importance degree, where the importance degree represents the influence degree on the output result of the pre-trained model; pruning the unimportant fused features to obtain the pruned fused features, and applying the pruned fused features and the important fused features to the inference of the pre-trained model.
[0005] The present application also provides an inference acceleration device for a pre-trained model, including: a first processing unit configured to obtain a natural language text and extract semantic features at different levels from the natural language text; a second processing unit configured to perform feature fusion on the semantic features at different levels to obtain a plurality of fused features; a third processing unit configured to determine important fused features and unimportant fused features among the plurality of fused features based on the importance degree of the fused features, where the importance degree represents the influence degree on the output result of the pre-trained model; and a fourth processing unit configured to prune the unimportant fused features to obtain pruned fused features, and apply the pruned fused features and the important fused features to the inference of the pre-trained model.
[0006] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-mentioned inference acceleration methods for a pre-trained model.
[0007] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned inference acceleration methods for a pre-trained model.
[0008] The present application also provides an inference acceleration system for a model, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing any of the inference acceleration methods for a pre-trained model.
[0009] Through the present application, semantic features at different levels are extracted from the input text, ensuring that the model can capture information in multiple aspects such as vocabulary, grammar, context, and long-term dependencies, which helps to improve the expression ability and generalization ability of the model. Based on the importance degree of the fused features, important fused features and unimportant fused features among the plurality of fused features are determined, and the unimportant fused features are pruned to obtain pruned fused features, so as to prune those unimportant fused features that have little impact on the final result. Therefore, the problem that the acceleration means for the inference speed of the pre-trained model in the existing solutions is difficult to achieve the expected effect is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To illustrate the embodiments of the present application more clearly, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of an inference acceleration method for a pre-trained model provided by an embodiment of the present application;
[0012] Figure 2 It is a schematic diagram of the relationship among the data loading layer, the model layer, and the output layer provided by the embodiments of the present application;
[0013] Figure 3 It is a schematic flowchart of the inference acceleration process of the pre-trained model provided by the embodiments of the present application;
[0014] Figure 4 It is a schematic structural diagram of an inference acceleration device for a pre-trained model provided by the embodiments of the present application. Specific embodiments
[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0016] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence in time.
[0017] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0018] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the inference acceleration method of the pre-trained model depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0019] The embodiments of the present application provide an inference acceleration method for a pre-trained model. The method will be described in detail in combination with the execution process of the inference acceleration method of the pre-trained model.
[0020] As Figure 1 shown, an inference acceleration method for a pre-trained model includes the following steps:
[0021] Step S101, obtain a natural language text, and extract semantic features at different levels from the natural language text;
[0022] Natural language text refers to the text information composed of language characters used in human daily communication, including written and spoken text forms. In the field of natural language processing (NLP), natural language text is the object of analysis, understanding, and generation, involving analysis at multiple levels such as grammar, semantics, and discourse structure.
[0023] Multi-level feature extraction technology refers to the technology of extracting features at different abstraction levels by designing a multi-layer neural network structure in deep learning models, especially convolutional neural networks (CNNs) and recurrent neural networks (RNNs). The lower layers usually extract simple features from the original signals (such as edges in images, part-of-speech in natural language, etc.), while the higher layers further abstract based on the lower-layer features to form more complex concepts or semantic features, which helps the model understand the deep meaning of the input data.
[0024] Semantic features refer to the features related to meaning that the model can capture when processing text or speech data. In natural language processing, semantic features can refer to the meaning of words, the meaning of sentences, or even the overall idea of a text. These features are crucial for understanding context relationships, inferring intentions and emotions, etc.
[0025] Step S102: Perform feature fusion on semantic features at different levels to obtain multiple fused features;
[0026] In the process of step S102, an adaptive feature fusion technology is required. Adaptive feature fusion technology is a technology that dynamically adjusts the fusion method and weights of features from different sources or levels in the model. It can automatically determine which features are more important and which features can be weakened according to the characteristics of the input data and the requirements of the current task, so as to use information more efficiently and accurately in the model decision-making process.
[0027] Step S103: Determine the important fused features and unimportant fused features among the multiple fused features based on the importance degree of the fused features. The importance degree represents the influence degree on the output result of the pre-trained model;
[0028] The unimportant fused features are the fused features with a relatively smaller influence degree on the output result of the pre-trained model compared to the important fused features.
[0029] Step S104: Clip the unimportant fused features to obtain the clipped fused features, and apply the clipped fused features and the important fused features to the inference of the pre-trained model.
[0030] Step S104 needs to adopt a dynamic pruning strategy. The dynamic pruning strategy selectively shuts down or skips some computational branches or layers based on real-time analysis of the input and output during model operation, so as to reduce unnecessary computational consumption. Different from static pruning, dynamic pruning can flexibly adjust the computational amount according to the actual situation on the premise of ensuring the model performance, and improve the operation efficiency.
[0031] The real-time evaluation mechanism refers to the technology of continuously analyzing the model state, output results or input data during model inference or operation to quickly adjust the model behavior or parameters. In deep learning, this mechanism can be used for dynamic pruning, adaptive learning rate adjustment or model performance monitoring, which helps the model make better decisions in complex and changeable environments.
[0032] In the above steps, extracting semantic features at different levels from the input text ensures that the model can capture information in multiple aspects such as vocabulary, grammar, context, and long-term dependencies, which helps to improve the expression ability and generalization ability of the model. Determining the important and unimportant fusion features among multiple fusion features based on the importance degree of the fusion features, and pruning the unimportant fusion features to obtain the pruned fusion features, so as to prune those unimportant fusion features that have little impact on the final result. Therefore, the problem that the acceleration means of the existing scheme for the inference speed of the pre-trained model is difficult to achieve the expected effect is solved.
[0033] As Figure 2 shown, the above method includes but is not limited to three parts: a data loading layer, a model layer, and an output layer. Among them, the data loading layer is responsible for preparing input samples; the model layer is composed of multiple Transformer encoders and decoders, and each encoder-decoder is further divided into functional modules such as a multi-head self-attention mechanism (Multi-Head Self-Attention, MHA), a feed-forward neural network (Feed-Forward Neural Network, FNN), and a residual connection (Residual Connection); the output layer performs operations such as classification or regression on the finally generated feature vector to obtain a prediction result.
[0034] First, a dynamic padding strategy is adopted. Dynamic padding fills the sequences in each batch dynamically according to the longest sequence length in the batch. By grouping sequences with similar lengths for batch processing and tokenizing the text sequences into vectors as inputs, the padding amount can be reduced, thus accelerating the calculation. This strategy can effectively improve the processing speed and efficiency of data. Then, after multiple rounds of processing by Transformer (a neural network architecture in the field of natural language processing. Its core innovation is to abandon the traditional recurrent neural network structure and instead use the self-attention mechanism to process sequence data. The self-attention mechanism allows the model to consider the context relationship of the entire sequence when processing each element in the sequence, rather than relying only on sequential processing one by one) encoders and decoders, effective capture of long dependencies is achieved; finally, the Softmax function (Softmax is a commonly used activation function in machine learning and deep learning, mainly used in the output layer of multi-classification tasks. It converts a vector containing arbitrary real numbers into a probability distribution. Specifically, the Softmax function maps each element in the input vector to the range (0, 1), and the sum of all outputs is 1, so it can be interpreted as a probability) is used to calculate the probability distributions of various categories, and the category with the maximum probability is selected as the output result. On this basis, an inference acceleration method for a pre-trained model based on hierarchical feature fusion and dynamic pruning strategy is proposed. The former aims to enhance the model expressiveness through multi-level feature interactions, and the latter is to intelligently prune unimportant fusion features without sacrificing accuracy, so as to achieve the purpose of acceleration.
[0035] In a specific embodiment of the present application, pruning of unimportant fusion features includes one of the following: pruning all unimportant fusion features; pruning some of the unimportant fusion features among all unimportant fusion features; pruning at least some of the sub-features of the unimportant fusion features among all unimportant fusion features; dividing the unimportant fusion features into a first part and a second part, pruning at least some of the sub-features of the unimportant fusion features in the first part, and pruning the unimportant fusion features in the second part.
[0036] Different layers or different modules of the model capture semantic information at different granularities:
[0037] Low-level features: part-of-speech, basic phrase structure, local word meaning (usually from the lower layer of the model).
[0038] Middle-level features: syntactic relationship, sentence-level semantics, context-related word meaning (usually from the middle layer of the model).
[0039] High-level features: paragraph / discourse-level semantics, main idea, emotion, intention (usually from the top layer or a specific aggregation layer of the model).
[0040] Here, only an example of cropping some of the unimportant fusion features among all the unimportant fusion features to obtain the cropped fusion features is given (the same applies to other cropping methods, so they will not be elaborated further). For example, if the unimportant fusion features include low-level features, middle-level features, and high-level features, then only the low-level features of the unimportant fusion features can be cropped, or only the low-level features and middle-level features of the unimportant fusion features can be cropped, while retaining the high-level features.
[0041] Another example is that the unimportant fusion features include two-level sub-features, namely the first-level sub-features and the second-level sub-features. Only the first-level sub-features in the unimportant fusion features are cropped, and the second-level sub-features are retained.
[0042] Specifically, cropping all the unimportant fusion features can significantly reduce the computational requirements of the model because the model no longer needs to perform calculations related to unimportant features, which directly reduces the computational time and is particularly important for large-scale models and real-time applications; completely cropping all unimportant features may have a slight impact on the model performance. Cropping some features is a compromise solution that reduces the computational cost while maintaining the prediction accuracy of the model as much as possible; the sub-features may be redundant or low-impact parts of the unimportant features. Cropping them can further reduce the amount of computation while maintaining the integrity of the main features, which is particularly beneficial for accelerating the model inference speed. This is a more refined cropping strategy that not only considers the importance of the entire feature but also the sub-features within the feature. This method can make the model focus more on key information while reducing unnecessary computations and improving the running efficiency. Cropping the unimportant fusion features also means reducing the amount of information that needs to be stored, thereby reducing the memory usage. This is very important for applications deployed on edge devices or with strict memory limitations.
[0043] In a specific embodiment of the present application, cropping some of the unimportant fusion features among all the unimportant fusion features to obtain the cropped fusion features includes: sorting the unimportant fusion features in descending order of their importance levels to obtain an unimportant fusion feature sequence; cropping the unimportant fusion features in the unimportant fusion feature sequence whose importance levels are less than a preset importance level.
[0044] Specifically, each computational branch in the model consumes a certain amount of computational resources and time, and the fusion features that contribute less to the overall performance of the model may lead to computational redundancy. Especially in large-scale data and complex model structures, by sorting and pruning the unimportant fusion features according to their importance, unnecessary computations can be significantly reduced, and the running speed and efficiency of the model can be improved. The storage requirement of the model is directly related to the complexity of the model, including the number of neural network layers, width, and the number of fusion features. Pruning unimportant fusion features reduces the parameters and intermediate results that need to be stored in the model, thereby reducing the storage requirement, which is particularly important for deployment on devices with limited computational resources, such as mobile devices or edge computing devices.
[0045] In a specific embodiment of the present application, semantic features at different levels are extracted from natural language text, including: introducing a self-adjusting activation function and a residual connection for strengthening information transmission in the pre-trained model to extract semantic features at different levels from natural language text, and representing the semantic features using key-value pairs. The self-adjusting activation function is used to provide a non-linear transformation for the pre-trained model;
[0046] Feature fusion is performed on semantic features at different levels to obtain multiple fusion features, including: based on the semantic features, using locality-sensitive hashing to approximately search for the most relevant key-value pairs to determine multiple target key-value pairs; in the case where the type of the target key-value pair is linear, determining the fusion weights of the semantic features at each level corresponding to the target key-value pair according to the task complexity; based on the fusion weights, using a linear fusion function to perform weighted summation on the target key-value pairs to obtain multiple fusion features.
[0047] The key-value pair is a vectorized representation of the semantic feature, and the key-value pair corresponds one-to-one with the semantic feature.
[0048] Specifically, the present application provides a specific usage scenario for obtaining fusion features: for example, developing a real-time translation application where users can input long natural language texts and the system needs to immediately return translation results. In such a scenario, the model not only has to process a large amount of text information but also respond quickly, which poses extremely high requirements on the computational efficiency of the model. When dealing with long texts, the traditional Transformer model involves calculations for all word pairs due to its self-attention mechanism, resulting in a huge amount of computations and making it difficult to achieve real-time response.
[0049] In the feedforward neural network (FNN) of the Transformer, the Swish function (a self-adjusting activation function. Compared with the traditional ReLU function, the Swish function has the following advantages: self-adjustability: The slope of the Swish function can be dynamically adjusted according to the input in the negative region, avoiding the "dead neuron" problem of ReLU, that is, when the input is negative, the output is 0, resulting in gradient disappearance; smoothness: The Swish function is continuous and smooth throughout the domain of definition, which means that its derivative exists at all points, facilitating the propagation of gradients and helping the model to converge during the inference process; non-linearity: The Swish still maintains the characteristics of non-linear transformation, which is very necessary for building deep neural networks because it can help the network learn complex patterns in the data) is used to replace the traditional ReLU activation function (Rectified Linear Unit is a commonly used activation function, and the formula is f(x) = max(0, x). In the positive interval, the ReLU activation function is linear and the derivative is 1, which promotes the rapid transmission of gradients and accelerates the inference process of the neural network. In the negative interval, the output is 0, which helps the network automatically ignore unimportant features and achieve a certain degree of sparsity). The Swish function has a smooth characteristic and can self-adjust during the inference process, providing a better non-linear transformation and helping the model learn more complex feature representations. At the same time, residual connections are introduced to ensure that information can be smoothly transmitted even when the network is deepened, avoiding the gradient disappearance problem and improving the stability and efficiency of the model. In the multi-head self-attention mechanism, instead of comprehensively calculating all positions, Locality Sensitive Hashing (LSH) is used to approximately search for the most relevant key-value pairs, especially in terms of semantic features. For large-scale text sequences, LSH can efficiently find the subset of key-value pairs most relevant to the query word, narrowing the attention calculation scope to the candidate key-value pairs highly relevant to the query word, greatly reducing the computational amount, especially in the processing of long dependencies, and significantly improving the inference speed. Dynamically adjust the fusion weights of the hierarchical features corresponding to the target key-value pairs according to the complexity of the task. For example, in the translation task, the model may rely more on high-level abstract semantic features, and at this time, the weight of the high-level features should be appropriately increased. In this way, the model can focus more on the features crucial to the task, thereby accelerating the inference process without sacrificing the translation quality. After determining the target key-value pairs and their corresponding fusion weights, a linear fusion function (such as weighted average) is used to perform weighted summation on the target key-value pairs to obtain the fused features. This method is simple and computationally efficient and is suitable as a basic fusion strategy to maintain high efficiency in a large number of calculations.)
[0050] Locality-Sensitive Hashing (LSH) is an algorithm for approximate nearest neighbor search, mainly used to process large-scale high-dimensional data sets. The core idea of LSH is to map similar data points to the same hash bucket through hash functions, so that a set of data points closest to the query point can be quickly found without comparing all data points. This method has brought significant improvements in both computational efficiency and storage efficiency. Especially when dealing with large-scale data sets such as text data and image data, LSH can greatly reduce the complexity of distance calculation, making approximate search feasible in practical application scenarios. In natural language processing, LSH is often used to optimize the self-attention mechanism of the Transformer model, accelerating model inference by reducing unnecessary key-value pair calculations.
[0051] Technical advantages of a specific usage scenario for obtaining fused features: The introduction of self-adjusting activation functions, residual connections, and locality-sensitive hashing technology effectively reduces the computational load of the model when processing long texts, greatly improving the inference speed of large-scale language models and meeting the performance requirements of real-time translation applications; it not only reduces the computational intensity but also the required memory resources, enabling the model to run on resource-constrained devices (such as mobile devices), expanding the scope of applications; by dynamically adjusting the fusion weights, the model can focus on using the features most relevant to the task, maintaining high translation quality and semantic understanding accuracy even while optimizing efficiency; these technologies are not only applicable to translation tasks but also widely applicable to other natural language processing fields such as text classification and sentiment analysis, providing a general and effective solution for accelerating the inference of pre-trained models.
[0052] In a specific embodiment of the present application, a residual connection for strengthening information transmission is introduced into the pre-trained model, including: determining target skip layers at multiple levels based on multiple different preset inter-layer residuals, where there is a preset inter-layer residual between any two adjacent levels; performing skip-layer transmission on each level based on the target skip layers.
[0053] Specifically, there are certain deviations. The following are the corrections and supplements to the technical details:
[0054] The skip-layer decision dependence on residuals is a non-standard practice. It is actually fixed by the network structure or dynamically controlled by a gating mechanism. The skip-layer transfer is a preset structure, not a temporary decision. The residuals themselves do not contain jump target information, and the jump path is determined by the architecture design (such as skipping 2 layers in ResNet, where ResNet is short for Residual Network, a deep neural network architecture for image recognition and other tasks). The jump-layer target is predefined (such as jumping from the 3rd layer to the 5th layer), not calculated from the residuals. Types of skip-layer transfer (Skip Connection): Fixed jump: In ResNet, skipping 2 layers to a preset path, regardless of the residual value; Dynamic gating jump: Controlled by learnable parameters to weight whether to skip a layer; Dense jump: Aggregating across multiple layers to a preset path, with the residual as one of the inputs. So whether to skip a layer is determined by the pre-designed network architecture or dynamically calculated by the gating unit, not directly depending on the residual size. In deep neural networks, especially in the Transformer architecture, each layer produces specific output features, which represent the model's understanding at different levels of abstraction. The inter-layer residuals mean calculating the difference information between the output of the current layer and the output of the previous layer. This information reflects the importance or unique contribution of the current layer's features relative to the previous layer's features. Skip-layer transfer is a key feature of the Residual Network (ResNet), which allows the network to directly transfer information between certain layers without passing through all intermediate layers. It can not only accelerate the model's operation but also ensure that the model can pay more attention to and utilize those truly influential deep features when dealing with complex tasks. The mechanism of skip-layer transfer is further refined. When the model determines that the residual information of a certain layer is relatively small or has little impact on the final task, it can skip several layers and directly transfer the information to a farther layer. The target layer of this jump is determined by the size of the residuals. By skipping unnecessary layers, the computational amount is greatly reduced, especially in scenarios where the number of model layers is large, and this improvement in efficiency is particularly significant. Skip-layer transfer reduces the dependence on computing resources, enabling the model to run with lower power consumption or fewer hardware resources, which is very important for edge computing devices or resource-constrained scenarios. By intelligently performing feature selection and information transfer, the model can better generalize to unseen data and reduce the risk of overfitting caused by over-relying on the features of specific layers.
[0055] For dynamic skip-layer based on residuals, introduce a gating mechanism:
[0056] Calculate the L2 norm of the residual (Euclidean norm, which is a way to measure the length of a vector):
[0057] ;
[0058] is the feature result of the output of the l-th layer, The characteristic result output for the (l - 1)-th level compresses the residual tensor into a scalar r for gating decision-making.
[0059] Generate a gating signal:
[0060] ;
[0061] where g is the gating signal strength, σ represents the Sigmoid activation function, W is the weight parameter, r is the residual strength, and b is the bias parameter.
[0062] Decision skip layer:
[0063] ;
[0064] where H output is the characteristic result output by the target skip layer, H l+k is the characteristic result output by the (l + k)-th level, the threshold is a hyperparameter, and k is a preset skip step. The size of the threshold is set using a layer-related threshold, and the formula is:
[0065] ;
[0066] where represents the skip layer threshold of the l-th layer, l is the depth (layer index) of the current layer in the model, usually starting from 0 or 1, L is the total number of layers in the model, is the slope parameter, between 0.5 and 0.7, is the base threshold parameter, between -0.3 and 0.1.
[0067] In a specific embodiment of the present application, semantic features at different levels are fused to obtain multiple fused features, including: in the case where the type of the target key-value pair is non-linear, the weights of each semantic feature are adjusted through the gradient of the backpropagation model to obtain the adjusted weights of each semantic feature; a voting mechanism is used to determine the feature fusion relationship, and the feature fusion relationship represents other semantic features that need to be fused with each semantic feature; based on the adjusted weights and the feature fusion relationship, weighted summation processing is performed on the semantic features at different levels to obtain multiple fused features.
[0068] Feature fusion is to combine the features extracted from different levels (which may also include different sources, such as word vectors, syntactic features, entity recognition results, etc.); common methods in the fusion methods include: concatenation: directly connecting different feature vectors end to end; weighted summation: adding after assigning weights to features at different levels; attention mechanism: allowing the model to dynamically learn the importance weights of different levels of features for the current task, and then performing weighted fusion; gating mechanism: learning to control the information flow and determining which levels of feature information can be passed to the next stage. Creating a new target: creating a more comprehensive, richer unified representation (i.e., the fused feature) that contains information from details to the global level.
[0069] Among them, determining the contribution of each feature to the prediction result of the pre-trained model through the gradient of the backpropagation model specifically includes the following steps:
[0070] Step 1: Forward propagation:
[0071] First, perform a normal forward propagation process on the model, input specific sample data, and obtain the prediction result of the model. In this process, the model performs a series of mathematical operations on the input features according to its weight and bias parameters, and finally generates a prediction value at the output layer.
[0072] Step 2: Calculate the loss:
[0073] Calculate the gap between the model prediction result and the true label, that is, the loss. The choice of the loss function depends on the specific problem. For example, the mean squared error (MSE) is suitable for regression problems, while the cross-entropy loss is usually used for classification tasks.
[0074] Step 3: Backpropagation;
[0075] Then, use the backpropagation algorithm to start from the output layer and calculate the gradient of the loss with respect to each parameter (including weights and biases) layer by layer backward. In a multi-layer neural network, backpropagation decomposes the loss gradient of the output layer to each layer through the chain rule until the input layer.
[0076] Step 4: Calculate the gradient of the feature:
[0077] When backpropagation reaches the input layer, the gradient of the loss function with respect to the input features can be obtained. When the input features change slightly, the change direction and amplitude of the model prediction result. Therefore, a larger gradient means that the feature has a greater impact on the model prediction.
[0078] Step 5: Feature importance evaluation:
[0079] Next, the importance of features can be evaluated in the following ways:
[0080] Absolute gradient value: The larger the absolute value of the gradient of a feature, the greater its contribution to the prediction result.
[0081] Integrated gradient: By accumulating the gradients from a certain feature value to a reference feature value, a more comprehensive estimate of the contribution degree can be provided.
[0082] Gradient input: Multiplying the gradient of a feature by its original input value can measure the relationship between the actual value of the input and its gradient, thereby reflecting the actual impact of the feature on the prediction result.
[0083] Specifically, when the pre-trained model processes non-linear relationships, the gradient contribution of each feature to the prediction result is calculated through backpropagation. The gradient information reflects the importance of each feature in the model prediction, which helps to identify which features are the most critical for the model decision-making; using the calculated gradient contribution to generate a saliency map, which visually shows which feature regions have the greatest impact on the output result of the pre-trained model. Based on the saliency map, the weight allocation in the feature fusion process can be dynamically adjusted to ensure that the most important features receive more attention, while the computational contribution of non-critical features is relatively reduced; a voting mechanism is adopted between different hierarchical features, allowing lower-level features to negotiate information with higher-level features (that is, using a voting mechanism to determine the feature fusion relationship representation and other semantic features that need to be feature-fused for each fusion requirement). This is essentially a self-organizing feature selection process, in which the contribution of lower-level features is determined according to their influence on higher-level features, ensuring that the model's decision-making process fully considers global information and avoiding prediction biases caused by the neglect of local information; for the extracted features, perform depth-wise convolution operations. This operation performs independent convolutions on each feature channel. Compared with traditional convolution operations, this method reduces the cross-channel interaction but can process the information within each channel more finely. This is particularly important for feature extraction in non-linear relationships because it can retain the specificity of each channel while promoting the effective fusion of features.
[0084] When dealing with non-linear relationships, computing resources are allocated more intelligently while ensuring that the most important features are fully utilized, thereby improving the quality and efficiency of model decision-making; Resource optimization and reduced computational costs: By reducing the computational investment in non-critical features, dynamically adjusting the weights in the feature fusion process, and the independent processing of feature channels by depth convolution operations, the computational costs are significantly reduced. This optimization is extremely crucial, especially when dealing with large-scale, high-dimensional data; The generation of saliency maps provides a window into the internal decision-making process of the model, enabling users and developers to better understand why the model makes specific predictions, thus enhancing the interpretability and debuggability of the model. This is particularly important in application scenarios that require human-machine collaboration or trust. The voting mechanism and dynamic weight adjustment strategy enable the model to adaptively respond to different types of input data and changes in task requirements. This flexibility is crucial for dealing with the ever-changing real-world data.
[0085] The non-linear fusion function can use Depth-wise convolution to assign different weights to different channels, use cross-attention layers to introduce learnable attention modules in the feature fusion stage, calculate the weights of features at each level through the Query-Key-Value mechanism, adjust the feature importance using the saliency map of Grad-CAM's gradient backpropagation, and adopt a voting mechanism to enable lower-level features to determine their contribution to higher-level features through negotiation. It can further extract the non-linear relationships between features and enhance the expressive power of features.
[0086] Depth-wise convolution is a convolution operation used in convolutional neural networks (CNNs), especially in mobile deep neural networks. Traditional convolution operations apply a single convolution kernel across all channels of the input, while depth convolution breaks down the convolution process into two steps: First, a convolution kernel is applied separately to each input channel (this is the "depth-wise" part), and then a 1x1 convolution (referred to as "point-wise" convolution) is used to combine the outputs of these channels. The advantage of this approach is that it significantly reduces the amount of computation and the number of parameters, thereby improving the efficiency of the model.
[0087] Grad-CAM is a technique for visualizing the decision-making process of convolutional neural networks. It generates a heatmap by calculating the gradients for a specific class, which are computed with respect to the feature map at a certain layer in the network. Specifically, it uses the gradient information of the target class to weight the feature map in order to highlight the regions in the input image that have the greatest impact on the decision for that class. This visualization method can help understand and interpret the behavior of neural networks and is a powerful tool for analyzing model decisions.
[0088] The Query-Key-Value mechanism is a core concept in the Transformer architecture and is used to implement the self-attention mechanism. It involves three main components:
[0089] Query: This is extracted from the input and is used to "ask" which information is relevant.
[0090] Key: It is also extracted from the input and represents the features of the input data.
[0091] Value: It is the value corresponding to the Key and represents the output information.
[0092] The self-attention mechanism calculates the similarity between the Query and the Key (usually using the dot product) to generate an attention weight, which is then used to weight the Value to generate the final output. This mechanism allows the model to focus on other relevant elements in the input sequence when processing each input element and is the basis for many modern models in natural language processing and computer vision.
[0093] Key: It is a vectorized representation of semantic features (such as text embedding vectors extracted by BERT, Transformer), and as an "index identifier", it stores high-level semantic information of the text (such as topics, intents, entity types).
[0094] Value: Meta-information or task-related parameters associated with the Key: feature weights (such as the importance of this semantics under the current task), labels (such as classification categories), structured data (such as entity attributes in a knowledge graph).
[0095] In a specific embodiment of the present application, in the process of determining multiple target key-value pairs by approximately searching for the most relevant key-value pairs based on semantic features using locality-sensitive hashing, the method further includes:
[0096] According to , determine the hash value of the vector o of the current dimension;
[0097] where h(o) is the hash value of the vector o of the current dimension, a is a random vector of the standard Gaussian distribution, b is an offset uniformly distributed in [0, w], and w is the shard width parameter.
[0098] Specifically, the core advantage of LSH lies in its ability to quickly find similar items in a high-dimensional space. By mapping vectors to hash buckets, vectors with similar Euclidean distances have a higher probability of being mapped to the same bucket. When dealing with large-scale datasets, traditional full-scale calculation methods become extremely slow and resource-intensive, while LSH can significantly accelerate this process. Especially when processing word embeddings or key-value pairs in large-scale language models, it can effectively reduce the computational complexity of the attention mechanism. LSH reduces the computational complexity by decreasing the number of direct comparisons of all vectors, which means that the required computational resources and storage space can be significantly reduced. This is particularly important for models running on resource-constrained devices such as mobile devices and edge servers, as it can effectively improve the running efficiency and deployment feasibility of the model. Since the hash function is designed based on random vectors a and offset b, this method can better adapt to the dynamic changes of data. Even when the data distribution changes or the model continuously learns new information, LSH can still maintain the stability of its search efficiency and effectiveness. In the attention mechanism of the Transformer model, this method allows each query to only calculate with a small number of key-value pairs in the same bucket during the inference process. Compared with the original mechanism that needs to calculate with all key-value pairs, the computational complexity is greatly reduced, significantly improving the inference speed of the model. By using specific hash functions and parameters (such as shard width w), the inference and optimization process of the model can be simplified. For example, instead of making detailed adjustments to each possible key-value pair, the performance of the model can be indirectly optimized by adjusting the hash parameters, reducing the complexity and uncertainty of parameter tuning.
[0099] In a specific embodiment of the present application, based on the fusion weights, a linear fusion function is used to perform weighted summation on the target key-value pairs to obtain the fusion feature, including:
[0100] According to , determine the fusion feature;
[0101] Wherein, is the fusion feature, is the weight of the i-th feature, is the value of the i-th feature, and n is the total number of features.
[0102] Specifically, through weighted summation, the model can highlight those features that contribute the most to the current task according to the weights learned during the inference process. This means that during prediction, the model will pay more attention to these key features, thereby improving the accuracy and reliability of the prediction. The calculation process of the linear fusion function is relatively simple. Compared with complex non-linear fusion, it reduces additional parameter adjustment and function operations, significantly improving the computational efficiency of feature fusion. Especially for large-scale key-value pair processing, it can significantly accelerate the inference process.
[0103] In a specific embodiment of the present application, before determining important and unimportant fusion features among multiple fusion features based on the importance degree of the fusion features, the method further includes: obtaining the importance degree of the fusion features, where obtaining the importance degree of the fusion features includes: obtaining the initial importance degree of the fusion features; allocating dynamic scaling factors to the key attention heads based on a preset importance score to adjust the initial importance degree of the fusion features, and obtaining the adjusted importance degree of the fusion features;
[0104] Determining important and unimportant fusion features among multiple fusion features based on the importance degree of the fusion features includes: determining a preset importance degree based on the current business requirements, where the current business requirements are related to the output result of the pre-trained model; in the case where the adjusted importance degree of the fusion feature is greater than or equal to the preset importance degree, determining the fusion feature as an important fusion feature; in the case where the adjusted importance degree of the fusion feature is less than the preset importance degree, determining the fusion feature as an unimportant fusion feature.
[0105] Among them, a dynamic scaling factor (scaling range 0.5 - 2.0) is applied to the key attention heads, and real-time spectral analysis (principal component contribution rate 85%) is performed on the multi-head attention matrix of the Transformer layer.
[0106] The dynamic scaling factor (Dynamic Scaling Factor) directly adjusts the output of the attention head, but does not directly generate the feature importance degree. The specific steps are as follows:
[0107] Step 1: Evaluate the importance of each attention head in the multi-head self-attention, and generate an importance score {S_head};
[0108] Step 2: Allocate dynamic scaling factors {α_head} to each attention head based on {S_head};
[0109] Step 3: Use {α_head} to scale and weight the output of each attention head to achieve dynamic adjustment;
[0110] The adjusted attention mechanism is more focused on key features, indirectly enhancing the weight of important components in the fusion features.
[0111] In the context of dynamic attention adjustment, key-value pairs (Key-Value): that is, the key-value pair matrix in the attention mechanism, store the projection representation of semantic features. The role of the dynamic scaling factor: strengthen the focus on specific features and enhance the weight of the corresponding semantic features in the fusion. The dynamic scaling factor controls the contribution degree of different semantic features in the fusion through adjustment.
[0112] Specifically, the dynamic scaling factor can adjust the weights of different attention heads in real time according to the characteristics of the input data, which means that the model can more accurately identify and focus on the most relevant and critical information segments in the input sequence, thereby improving the accuracy and quality of the prediction or generation results; by dynamically adjusting the weights of the attention heads, the computational investment in non-critical information can be effectively reduced, especially when dealing with long sequence data, which can significantly reduce the consumption of computing resources, speed up the model inference speed, and ensure the complete processing of important information; the dynamic scaling factor enables the model to adaptively adjust its information processing strategy according to different tasks and scenarios. For example, in some application scenarios, the model may need to pay more attention to local information, while in other scenarios, it needs to focus on global information. The dynamic adjustment mechanism allows the model to flexibly respond to these changes, thereby improving its generalization ability and performance; under certain business requirements, specific types of attention heads may be more important than others. By setting the importance mapping relationship, the model can preferentially process more important information according to the current business needs, achieve personalized feature extraction and processing, and meet diverse requirements. Based on the importance of the fused features and the preset importance of the target, the model can intelligently identify unimportant fused features and adopt appropriate methods (such as pruning unimportant fused features) to reduce redundant calculations, avoid resource waste, and ensure the maximization of computational efficiency; the dynamic adjustment and the setting of importance increase the transparency of the model decision-making process, helping to understand and explain why the model selects certain information for processing, which is extremely important for the auditing, debugging, and optimization of the model, especially in application scenarios that require high interpretability.
[0113] In a specific embodiment of the present application, the method further includes: during the pruning of unimportant fused features and / or the inference process of the pre-trained model, adjusting the hyperparameters of the pre-trained model in real time according to the current business needs.
[0114] Adjust the hyperparameters of batch normalization (BN) according to the characteristics of the input samples to ensure that the best learning rate can be obtained for each iteration, so that the entire model automatically adjusts and optimizes the feature fusion weights and pruning strategies of each layer during the inference process, regularly save checkpoint files during the inference process for subsequent recovery, monitor the system log output, and promptly discover and solve problems.
[0115] Specifically, hyperparameters directly affect the inference process and final performance of the model. Through real-time adjustment, the model can better adapt to the characteristics of specific tasks, such as data distribution, model complexity, etc., thereby improving prediction accuracy and stability while maintaining efficiency. Dynamically adjusting hyperparameters can reasonably allocate computing resources according to the priorities of business requirements. In a resource-constrained environment, this strategy is particularly crucial, which can ensure that the model runs in the most efficient way and avoid resource waste or performance bottlenecks caused by fixed configurations; Different business scenarios may have different performance requirements for the model. For example, some scenarios may value prediction speed more, while others may focus more on prediction accuracy. Real-time adjustment of hyperparameters enables the model to flexibly respond to these changes and adaptively optimize its behavior to meet the needs of specific scenarios.
[0116] In a specific embodiment of the present application, in the process of extracting semantic features at different levels from a natural language text, the method includes: determining the difference in length between different text sequences of the natural language text; performing batch processing on the target text sequences, where the target text sequences are those with the absolute value of the difference in length less than a preset difference.
[0117] Specifically, traditional batch processing padding strategies would extend all sequences to the length of the longest sequence in the batch, which means that a large number of padding bits of shorter sequences would participate in the calculation but would not contribute any meaningful information. By selecting target text sequences with the absolute value of the difference in length less than the preset difference for batch processing, such ineffective calculations can be significantly reduced because these sequences have less padding and are closer to the actual length to be processed; The reduction in computational volume directly leads to an increase in the inference speed of the model. In the inference of deep learning models, the effective utilization rate of computing resources is one of the key factors determining inference efficiency. The batch processing strategy for sequences with similar lengths reduces the unnecessary computational load, thus accelerating the inference progress of the model; Reducing ineffective padding means that less storage space is used to store meaningless numerical values, which not only reduces memory occupancy but also reduces the memory bandwidth pressure on the GPU or CPU because less data needs to be transferred between memory and computing units. This is particularly important for resource-constrained devices, which can ensure that the model runs more smoothly under limited hardware conditions.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0119] The present application also provides a method for accelerating the training of a pre-trained model, specifically including:
[0120] Acquire natural language text and extract semantic features at different levels from the natural language text; fuse the semantic features at different levels to obtain multiple fusion features; determine important fusion features and unimportant fusion features of the multiple fusion features based on the importance of the fusion features; trim the unimportant fusion features to obtain trimmed fusion features, and apply the trimmed fusion features and important fusion features to the training of the pre-trained model.
[0121] Dynamically adjust the floating-point precision of the pre-trained model during the training process of the pre-trained model, including: when the pre-trained model is loaded, adjusting the floating-point precision to a first floating-point precision; when the pre-trained model is predicted, adjusting the floating-point precision to a second floating-point precision, and the second floating-point precision is higher than the first floating-point precision.
[0122] Loading phase (first floating point precision, such as FP16): Using lower precision floating point numbers (such as FP16) can significantly reduce the storage space of model parameters, which is especially beneficial for model loading, because the entire model usually needs to be loaded into RAM during the loading phase. Lower precision means smaller model size, which reduces the demand for memory; in distributed systems, model parameters need to be transmitted over the network. The use of FP16 reduces the amount of data, thereby speeding up the transmission of the model, which is conducive to faster model synchronization and improved efficiency of distributed training; the reduced model size also means increased loading speed, which is crucial for quickly starting model applications or services, especially in environments with limited resources or requiring immediate response.
[0123] Prediction stage (secondary floating-point precision, such as FP32): Using higher precision (such as FP32) during the model prediction stage can avoid numerical overflow or underflow caused by insufficient precision, improving the stability and reliability of the model output. This is crucial for applications that require extremely high prediction accuracy, such as financial analysis and autonomous driving. The prediction stage often requires more accurate calculations to ensure the accuracy of model outputs. Using FP32 can better preserve the details of the model's internal calculations, avoid prediction errors caused by loss of precision, and ensure the high quality and credibility of model predictions. Although FP32 increases storage and computing requirements, this increased resource consumption during the prediction stage is traded for better model performance. Many modern GPUs and accelerators have optimized support for FP32 operations, providing higher throughput and shorter latency, thereby improving overall prediction speed and efficiency.
[0124] In a specific embodiment of the present application, the method further includes: recording various indicators of the pre-training model during the training process.
[0125] Specifically, by recording key metrics during the training process, the training status of the model can be monitored in real time, including changes in the learning rate, whether the loss function converges, the trend of accuracy improvement, etc. This helps to adjust training parameters or strategies in a timely manner and avoid the training getting stuck in local optima or overfitting. The metric records during the training process can assist developers in diagnosing potential problems encountered during model training, such as vanishing or exploding gradients, performance bottlenecks, etc. By analyzing these records, the fault points in the model can be located and corresponding debugging strategies can be adopted to ensure the smooth progress of model training. The recorded metric data can be used as a basis for adjusting hyperparameters, such as the learning rate, regularization strength, batch size, etc. By comparing and analyzing the metrics under different settings, the best combination of hyperparameters suitable for the current task and dataset can be found, thereby optimizing the model performance. The metric records provide a quantitative method for evaluating how the model performance changes with the training cycle. For example, by observing the change in accuracy on the validation set, it can be determined whether the model starts to generalize well. By observing the trend of the loss function, the fitting degree of the model can be understood, and these are all important references for evaluating the final effect of the model.
[0126] As Figure 3 shown, the training acceleration process of the pre-trained model includes:
[0127] Initializing the environment settings to ensure sufficient computing resources are available;
[0128] Preparing the training set and validation set in a preset format;
[0129] Writing or calling corresponding code to implement the model definition;
[0130] End-to-end training of the model;
[0131] Model testing.
[0132] Extracting semantic features at different levels from the input text ensures that the model can capture information in multiple aspects such as vocabulary, grammar, context, and long-term dependencies, which helps to improve the model's expressive ability and generalization ability. The dynamic pruning strategy dynamically determines the importance degree in the fused feature vector based on a real-time evaluation mechanism and prunes those computational branches that have less impact on the final result. In order to accelerate the model training process, an end-to-end mixed-precision training method is adopted to achieve the best balance between speed and accuracy. Therefore, the problem that the existing solutions for accelerating the inference speed of the pre-trained model are difficult to achieve the expected effect is solved. It can be applied to the optimization process of specific deep learning models in the fields of natural language processing, computer vision, speech recognition, etc., broadening the applicable boundary of high-performance AI services, enabling them to serve more diverse demand groups, and promoting the popularization and development of artificial intelligence technology.
[0133] In addition to the above-mentioned approach combining hierarchical feature fusion and dynamic pruning strategies, other potential optimization paths can be explored. For example, introducing lightweight convolution operations to replace traditional matrix multiplication, or adopting more advanced attention mechanism designs (such as linear attention) to replace the standard self-attention mechanism. Dynamic quantization technology can also be explored for secondary optimization, that is, adding a dynamic adjustment factor on the basis of static quantization to better adapt to the unpredictable real-world data distribution. These three improvement measures all contribute to further reducing the computational burden and accelerating the inference process.
[0134] For example, using V100 / A100 GPUs or other higher-performance acceleration cards as the computing platform, and adopting frameworks such as PyTorch (an open-source deep learning framework) or TensorFlow (an open-source framework suitable for machine learning and deep learning for various tasks) as the underlying support tools to build the model structure. Experiments have shown that the solution can greatly improve the inference speed with almost no impact on the model's performance, while reducing the pressure on computing resources. Future work will continue to conduct in-depth research on how to better adapt to different types of tasks.
[0135] Among them, the end-to-end mixed-precision training method uses the final fused features to train the pre-trained model, and dynamically adjusts the floating-point precision of the pre-trained model during the training process. The end-to-end mixed-precision training method is a technique for optimizing inference speed and memory usage, in which part of the model is calculated using lower-precision floating-point numbers (such as FP16, i.e., half-precision), while the key parts use higher precision (such as FP32, i.e., single-precision). In this way, the training process can be faster, while ensuring the stability and accuracy of model training.
[0136] Floating-point numbers are a digital format used in computer science to represent real numbers. They can represent very large and very small numbers, including decimals. Floating-point numbers consist of an integer part, a fractional part, and an exponent part. By adjusting the exponent, different magnitudes of numerical values can be represented. In deep learning, model parameters and calculation results are often represented using floating-point numbers to support high-precision mathematical operations.
[0137] By introducing hierarchical feature fusion and dynamic pruning strategies, significant inference acceleration is achieved while ensuring the quality of the model output, which is especially suitable for processing large-scale complex models. Compared with traditional hardware upgrade solutions, the technical method of the present invention is more cost-effective and easy to deploy, especially suitable for resource-constrained edge computing devices. The feedforward layer structure is reconstructed to improve the operation efficiency of each step. Generally speaking, this helps save more time overhead in the entire inference process. The method of improving the self-attention mechanism with the help of locality-sensitive hashing technology greatly reduces the computational load, especially showing stronger superiority when facing large-scale sequence inputs. The method of dynamically adjusting batch normalization parameters avoids the rigid rules of a one-size-fits-all approach, making the entire training process more intelligent and controllable. Using mixed-precision training effectively alleviates the occurrence probability of overfitting, and at the same time promotes the long-term stable development of the model.
[0138] An embodiment of this application also provides an inference acceleration device for a pre-trained model, as Figure 4 shown. The device includes:
[0139] A first processing unit 41, configured to obtain a natural language text and extract semantic features at different levels from the natural language text; a second processing unit 42, configured to perform feature fusion on the semantic features at different levels to obtain multiple fused features; a third processing unit 43, configured to determine important fused features and unimportant fused features among the multiple fused features based on the importance degree of the fused features, where the importance degree represents the influence degree on the output result of the pre-trained model; a fourth processing unit 44, configured to prune the unimportant fused features to obtain pruned fused features, and apply the pruned fused features and the important fused features to the inference of the pre-trained model.
[0140] In the above device, extracting semantic features at different levels from the input text ensures that the model can capture information in multiple aspects such as vocabulary, grammar, context, and long-term dependencies, which helps improve the expression ability and generalization ability of the model. Determining important fused features and unimportant fused features among the multiple fused features based on the importance degree of the fused features, and pruning the unimportant fused features to obtain pruned fused features, so as to prune those unimportant fused features that have little impact on the final result. Therefore, the problem that the acceleration means for the inference speed of the pre-trained model in the existing solutions is difficult to achieve the expected effect is solved.
[0141] In a specific embodiment of the present application, the fourth processing unit includes: a cropping module configured to perform one of the following steps: cropping off all non-important fusion features; cropping off some of the non-important fusion features among all non-important fusion features; cropping off sub-features of at least some of the non-important fusion features among all non-important fusion features; dividing the non-important fusion features into a first part and a second part, cropping off sub-features of at least some of the non-important fusion features of the first part of the non-important fusion features, and cropping off the non-important fusion features of the second part.
[0142] In a specific embodiment of the present application, the cropping module includes: a first processing sub-module configured to sort the non-important fusion features in descending order of importance to obtain a non-important fusion feature sequence; a second processing sub-module configured to crop off the non-important fusion features in the non-important fusion feature sequence whose importance is less than a preset importance level.
[0143] In a specific embodiment of the present application, the first processing unit includes: a first processing module configured to introduce a self-regulating activation function and a residual connection for strengthening information transmission into a pre-trained model to extract semantic features at different levels from natural language text, and represent the semantic features using key-value pairs, where the self-regulating activation function is used to provide a non-linear transformation for the pre-trained model;
[0144] The first processing module includes: a second processing module configured to determine multiple target key-value pairs by approximately searching for the most relevant key-value pairs based on the semantic features using locality-sensitive hashing; a third processing module configured to determine the fusion weights of the semantic features at each level corresponding to the target key-value pairs according to the task complexity when the type of the target key-value pairs is linear; a fourth processing module configured to perform weighted summation on the target key-value pairs using a linear fusion function based on the fusion weights to obtain multiple fusion features.
[0145] In a specific embodiment of the present application, the first processing module includes: an acquisition sub-module configured to determine multiple levels of target skip layers based on multiple different preset inter-layer residuals, where there is a preset inter-layer residual between any two adjacent levels; a third processing sub-module configured to perform skip-layer transmission on each level based on the target skip layers.
[0146] In a specific embodiment of the present application, the second processing unit includes: a fifth processing module configured to, when the type of the target key-value pair is non-linear, adjust the weight of each semantic feature through the gradient of the backpropagation model to obtain the adjusted weight of each semantic feature; a sixth processing module configured to determine a feature fusion relationship by using a voting mechanism, where the feature fusion relationship represents other semantic features that need to perform feature fusion with each semantic feature; a seventh processing module configured to perform weighted summation processing on semantic features at different levels based on the adjusted weights and the feature fusion relationship to obtain a plurality of fusion features.
[0147] In a specific embodiment of the present application, the second processing module includes a fourth processing sub-module configured to, in the process of determining a plurality of target key-value pairs by approximately searching for the most relevant key-value pairs based on semantic features using locality-sensitive hashing, according to , determine the hash value of the vector o of the current dimension;
[0148] where h(o) is the hash value of the vector o of the current dimension, a is a standard Gaussian distribution random vector, b is an offset uniformly distributed in [0, w], and w is a shard width parameter.
[0149] In a specific embodiment of the present application, the fourth processing module includes:
[0150] a fifth processing sub-module configured to determine a fusion feature according to ;
[0151] where is the fusion feature, is the weight of the i-th feature, is the value of the i-th feature, and n is the total number of features.
[0152] In a specific embodiment of the present application, the apparatus further includes: an acquisition unit configured to acquire the importance degree of the fusion features before determining important fusion features and unimportant fusion features of the plurality of fusion features based on the importance degree of the fusion features, where the acquisition unit includes: an eighth processing module configured to acquire the initial importance degree of the fusion features; a ninth processing module configured to allocate a dynamic scaling factor to a key attention head based on a preset importance score to adjust the initial importance degree of the fusion features to obtain the adjusted importance degree of the fusion features;
[0153] The third processing unit includes: a tenth processing module for determining a preset importance level based on the current service requirement, where the current service requirement is related to the output result of the pre-trained model; an eleventh processing module for determining the fused feature as an important fused feature when the adjusted importance level of the fused feature is greater than or equal to the preset importance level; and a twelfth processing module for determining the fused feature as a non-important fused feature when the adjusted importance level of the fused feature is less than the preset importance level.
[0154] In a specific embodiment of the present application, the device further includes: a fifth processing unit for adjusting the hyperparameters of the pre-trained model in real time according to the current service requirement during the clipping of the non-important fused feature and / or the inference process of the pre-trained model.
[0155] In a specific embodiment of the present application, the first processing unit includes: a thirteenth processing module for determining the difference in length between different text sequences of the natural language text during the extraction of semantic features at different levels from the natural language text; and a fifteenth processing module for performing batch processing on the target text sequence, where the target text sequence is a text sequence whose absolute value of the difference in length is less than the preset difference.
[0156] For the description of the features in the corresponding embodiments of the inference acceleration device of the pre-trained model, reference can be made to the relevant description in the corresponding embodiments of the inference acceleration method of the pre-trained model, which will not be elaborated here one by one.
[0157] The present application also provides an inference acceleration system for a model, including: one or more processors, a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors, and one or more programs include those for executing any inference acceleration method of the pre-trained model.
[0158] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored, and the computer program is set to execute the steps in any of the above-mentioned embodiments of the inference acceleration method of the pre-trained model when running.
[0159] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs, etc., all kinds of media that can store computer programs.
[0160] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the embodiments of the inference acceleration method of the pre-training model are implemented.
[0161] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the embodiments of the inference acceleration method of the pre-training model are implemented.
[0162] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0163] The above has introduced in detail the inference acceleration method of the pre-training model, the inference acceleration device of the pre-training model, the computer-readable storage medium, the computer program product, and the inference acceleration system of the pre-training model provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for accelerating the inference of a pre-trained model, characterized in that Including: Obtain a natural language text, and extract semantic features at different levels from the natural language text; Perform feature fusion on the semantic features at different levels to obtain a plurality of fused features; Determine important fused features and unimportant fused features among the plurality of fused features based on the importance degree of the fused features, where the importance degree represents the influence degree on the output result of the pre-trained model; Crop the unimportant fused features to obtain cropped fused features, and apply the cropped fused features and the important fused features to the inference of the pre-trained model.
2. The inference acceleration method for the pre-trained model according to claim 1, wherein Cropping the unimportant fused features includes one of the following: Crop all of the unimportant fused features; Crop some of the unimportant fused features among all of the unimportant fused features; Crop at least some of the sub-features of the unimportant fused features among all of the unimportant fused features; Divide the unimportant fused features into a first part and a second part, crop at least some of the sub-features of the unimportant fused features in the first part, and crop the unimportant fused features in the second part.
3. The inference acceleration method of the pre-trained model according to claim 2, wherein Cropping some of the unimportant fused features among all of the unimportant fused features to obtain the cropped fused features includes: Sort the unimportant fused features in descending order according to the importance degree of the unimportant fused features to obtain an unimportant fused feature sequence; Crop the unimportant fused features in the unimportant fused feature sequence whose importance degree is less than a preset importance degree.
4. The method for accelerating the inference of the pre-trained model according to claim 1, wherein Extracting semantic features at different levels from the natural language text includes: introducing a self-adjusting activation function and a residual connection for strengthening information transmission in the pre-trained model to extract the semantic features at different levels from the natural language text, and representing the semantic features by using key-value pairs, where the self-adjusting activation function is used to provide a non-linear transformation for the pre-trained model; Performing feature fusion on the semantic features at different levels to obtain a plurality of fused features includes: based on the semantic features, using local sensitive hashing to approximately search for the most relevant key-value pairs to determine a plurality of target key-value pairs; when the type of the target key-value pairs is linear, determining the fusion weights of the semantic features at each level corresponding to the target key-value pairs according to the task complexity; and based on the fusion weights, using a linear fusion function to perform weighted summation on the target key-value pairs to obtain the plurality of fused features.
5. The inference acceleration method of the pre-trained model according to claim 4, wherein Introducing the residual connection for strengthening information transmission in the pre-trained model includes: Determining target skip layers at multiple levels based on a plurality of different preset inter-layer residuals, where there is the preset inter-layer residual between any two adjacent levels; Performing skip layer transmission on each level based on the target skip layer.
6. The inference acceleration method for the pre-trained model according to claim 4, wherein Performing feature fusion on the semantic features at different levels to obtain a plurality of fused features includes: When the type of the target key-value pair is non-linear, adjust the weights of each semantic feature through the gradient of the backpropagation model to obtain the adjusted weights of each semantic feature; Adopt a voting mechanism to determine the feature fusion relationship, where the feature fusion relationship represents other semantic features that need to perform feature fusion with each semantic feature; Based on the adjusted weights and the feature fusion relationship, perform weighted summation processing on semantic features at different levels to obtain multiple fusion features.
7. The inference acceleration method of the pre-trained model according to claim 4, wherein In the process of determining multiple target key-value pairs by using local sensitive hashing to approximately search for the most relevant key-value pairs based on the semantic features, the method further includes: According to , determine the hash value of the vector o in the current dimension; Where h(o) is the hash value of the vector o in the current dimension, a is a random vector of the standard Gaussian distribution, b is an offset uniformly distributed in [0, w], and w is the shard width parameter.
8. The method for accelerating the inference of the pre-trained model according to claim 4, wherein Based on the fusion weights, perform weighted summation on the target key-value pairs by using a linear fusion function to obtain the fusion features, including: According to , determine the fusion feature; Among them, is the fusion feature, is the weight of the i-th feature, is the value of the i-th feature, and n is the total number of the features.
9. The method for accelerating the inference of the pre-trained model according to claim 1, wherein Before determining the important fusion features and unimportant fusion features of the multiple fusion features based on the importance degree of the fusion features, the method further includes: obtaining the importance degree of the fusion features, where obtaining the importance degree of the fusion features includes: obtaining the initial importance degree of the fusion features; based on a preset importance score, allocate a dynamic scaling factor to the key attention heads to adjust the initial importance degree of the fusion features to obtain the adjusted importance degree of the fusion features; Determining the important fusion features and unimportant fusion features of the multiple fusion features based on the importance degree of the fusion features includes: determining a preset importance degree based on the current service requirement, where the current service requirement is related to the output result of the pre-trained model; when the adjusted importance degree of the fusion feature is greater than or equal to the preset importance degree, determining the fusion feature as the important fusion feature; when the adjusted importance degree of the fusion feature is less than the preset importance degree, determining the fusion feature as the unimportant fusion feature.
10. The method for accelerating the inference of the pre-trained model according to claim 1, wherein, The method further includes: During the pruning of the unimportant fusion features and / or the inference process of the pre-trained model, adjust the hyperparameters of the pre-trained model in real time according to the current service requirement.
11. The method for accelerating inference of the pre-trained model according to claim 1, wherein In the process of extracting semantic features at different levels from the natural language text, the method includes: Determine the difference in length between different text sequences of the natural language text; Batch process with the target text sequence, where the target text sequence is the text sequence whose absolute value of the difference in length is less than the preset difference.
12. An inference acceleration device for a pre-trained model, characterized in that, Includes: A first processing unit for obtaining a natural language text and extracting semantic features at different levels from the natural language text; A second processing unit for performing feature fusion on the semantic features at different levels to obtain multiple fusion features; A third processing unit, configured to determine important and unimportant fused features among the multiple fused features based on the importance degree of the fused features, where the importance degree represents the influence degree on the output result of the pre-trained model; A fourth processing unit, configured to prune the unimportant fused features to obtain pruned fused features, and apply the pruned fused features and the important fused features to the inference of the pre-trained model.
13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the inference acceleration method of the pre-trained model according to any one of claims 1 to 11.
14. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the inference acceleration method of the pre-trained model according to any one of claims 1 to 11.
15. An inference acceleration system for a pre-trained model, characterized in that, Comprising: One or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing the inference acceleration method of the pre-trained model according to any one of claims 1 to 11.
Citation Information
Patent Citations
Pre-training model accelerated reasoning method and system based on redundant word deletion
CN113159168A
Visual Transform lightweight method based on Token fusion
CN119992158A
Semi-structured file processing method based on LLMs large language model
CN120218236A
Blower
KR1020230031665A
Cited By
Depth learning model output visualization migration consistency analysis method, system and device, medium and product
CN121119141A
A method, system, device, medium and product for analyzing migration consistency of deep learning model output visualization
CN121119141B
Industrial large model training method, device and system based on efficient fine tuning
CN122088617A
Industry large model training method, device and system based on efficient fine-tuning
CN122088617B