Early exit method based on learnable interlayer integration strategy

By introducing a learnable inter-layer integration strategy into the pre-trained model, combining multi-layer feature information and dynamically adjusting the inference process, the problems of insufficient information sharing and waste of computing resources in the traditional early retreat method are solved, and the robustness and inference speed of the model are improved. It is suitable for the BERT type structural model in the fields of computer vision and multimodality.

CN120277478APending Publication Date: 2025-07-08ANHUI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411872843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The early withdrawal method of traditional pre-trained models has problems such as insufficient information sharing, limited shallow feature extraction capabilities, and unknown expansion effects in other fields in different tasks, resulting in waste of computing resources and insufficient prediction accuracy.

Method used

The early retreat method based on the learning inter-layer integration strategy is adopted. By adding internal learners and classifiers to each layer, combining the feature information of the previous layer and the current layer, using entropy values to judge the model stability, dynamically adjust the inference process, and adding Tanh and Sigmoid activation functions to optimize the learning process.

Benefits of technology

The speed and accuracy trade-offs under different task complexity are realized, the robustness and inference speed of the model are improved, unnecessary computing resource consumption is reduced, and the BERT type structural model is suitable for computer vision and multimodal fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277478A_ABST
    Figure CN120277478A_ABST
Patent Text Reader

Abstract

The invention discloses an early leaving method based on a learnable interlayer integration strategy, and the method comprises the steps: S1, constructing a BERT model in which an internal learner and a classifier are added in n transform layers, and n is a positive integer greater than 1; s2, splicing the hidden vectors of the first i-1 layers, inputting the spliced hidden vectors as input into the ith layer of learner, and connecting the input spliced hidden vectors by the learner to obtain a hyper-parameter lambda i; i is greater than or equal to 1 and less than or equal to n; s3, inputting the logits probability distribution in the step S2 into a classifier to obtain an output Pi; and S4, quantizing the output distribution Pi of the ith layer of the model by using an entropy value to obtain a confidence coefficient e (pi) for outputting the output of the layer, if e (pi) is lower than a preset threshold value tau, the model quits from the current layer, and otherwise, the model continues to go to the next layer for forward propagation. According to the method, the features of the previous branch and the current branch of the model can be effectively combined, and the early-quit mechanism of the model is enabled through the combined features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic data processing, and in particular, to an early exit method based on a learnable inter-layer integration strategy. Background Art

[0002] Today, with the continuous development of pre-trained models, the increasing number of stacked model layers and parameters leads to an increase in the online inference time of the model. As the difficulty coefficients of different tasks vary, it is reasonable to speculate that in tasks with simple scales, the computational resources required by the model are relatively small, and the computational power required by the model for processing this task is relatively small. On the contrary, if a task is more complex, then this task requires more powerful computational power to maintain the accuracy of the task. Traditional pre-trained early exit methods usually have some deficiencies, which we summarize as the following points:

[0003] (1) The features learned by the BERT model in each layer are different. Shallow layers learn more general syntactic features, and deeper layers learn richer semantic features. The multi-branch internal classifiers of traditional early exit methods can only extract information from a single layer and cannot share and combine the information of the internal classifiers of multiple branches.

[0004] (2) The information extraction ability of the pre-trained model in the shallow layer is limited, and the ability to extract shallow features may lead the model to make incorrect decisions.

[0005] (3) The effects of early exit methods based on pre-trained models in other fields are unknown. How to effectively extend the early exit method to BERT-Style models in other fields remains a challenge. Summary of the Invention

[0006] To solve the technical problems in the background art, the present invention proposes an early exit method based on a learnable inter-layer integration strategy.

[0007] An early exit method based on a learnable inter-layer integration strategy proposed by the present invention includes the following steps:

[0008] S1. Construct a BERT model with internal learners and classifiers added to n transformer layers, where n is a positive integer greater than 1, so that each layer can have the ability to predict samples, thus enabling early exit of the pre-trained model and achieving the effect of accelerating model inference;

[0009] S2. Concatenate the hidden vectors of the first i - 1 transformer layers and use them as input to the learner of the i-th transformer layer. The learner connects the input concatenated hidden vectors to obtain a hyperparameter λ i ; According to the hyperparameter λ iThe logits probability distribution of the current layer is obtained by concatenating the feature information of the previous layer and the feature information of the current layer, where i is greater than or equal to 1 and less than or equal to n;

[0010] S3. The logits probability distribution in step S2 is input into a classifier to obtain the output Pi;

[0011] S4. By using entropy to quantify the output distribution Pi of the i-th layer Transformer layer of the model, the confidence e(p i ) of the output of this layer is obtained. Compare e(p i ) with the previously set threshold τ. If e(p i ) is lower than the pre-set threshold τ, the model will exit at the current layer. Otherwise, the model will continue to the next layer for forward propagation. e(p i ) can represent the confidence level of the model at the i-th layer. The lower e(p i ), the more stable the probability distribution of the current layer, and the higher the confidence level of the model, otherwise it is the opposite;

[0012] A linear classifier (classifier) is added after the convolutional layer of each Transformer layer. By means of entropy value, it is judged whether the current prediction result is stable. If the stable threshold is reached, the process exits and the inference process ends. Otherwise, continue the forward propagation of the next layer. The setting of the learner aims to enable the learner to automatically learn the feature information of the previous layer and the current layer, and then combine the two through hyperparameters. The experimental results can prove that under the same acceleration ratio, the method of the present invention can achieve more robust prediction results.

[0013] As a further optimized solution of the present invention, the learner in step S1 includes a layer of Tanh activation function, a linear layer and a final Sigmoid activation function.

[0014] As a further optimized solution of the present invention, the hidden vector of the i-th layer in step S2 is hi, where:

[0015]

[0016] In the formula: x is the input, embedding(x) represents obtaining the word vector of x, and Encoder(h i-1 ) is the existing input h i-1 calculated by the existing encoder.

[0017] As a further optimized solution of the present invention, the logits probability distribution in step S2 is:

[0018] l i =Concat(h1, h2,..., hi-1 )

[0019] r i = Concat(hi)

[0020] λ i = Learner(l i , r i )

[0021] logits i = λ i * l i +(1 - λ i )* r i ;

[0022] Where: li is the aggregation of the hidden vectors of the i - 1 layer transformer layer; ri is the vector aggregation of the current layer transformer layer; Leaner() is the learning by the learning module.

[0023] As a further optimized solution of the present invention, in step S4, e(p i ) = Entropy(p i ), where entropy is the entropy value.

[0024] As a further optimized solution of the present invention, the classifier in step S4 includes a single - layer fully - connected network.

[0025] The learnable inter - layer integration algorithm proposed by the present invention makes the following contributions:

[0026] (1) Abandon the idea that the prediction result of a single classifier represents the global prediction result, combine the features of multiple internal classifiers, and improve the robustness of the prediction result of the model.

[0027] (2) The early - exit algorithm based on the inter - layer integration strategy proposed by the present invention can dynamically maintain the trade - off between the inference speed and the prediction accuracy according to the complexity of the task.

[0028] (3) The algorithm of the present invention can be effectively inserted into the models with BERT - like structures in the fields of computer vision and multi - modality, enabling the improvement of the inference speed of pre - trained models in other fields.

[0029] In the present invention, an early exit method based on a learnable inter-layer integration strategy is proposed. This method can effectively combine the features of the previous branch and the current branch of the model, empower the early exit mechanism of the model through the combined features, and improve the classification accuracy of the model at a shallow layer. Secondly, the method of the present invention can balance the inference speed and the model prediction accuracy of the model, and outperforms other pre-trained model early exit methods in different tasks. Finally, the method proposed by the present invention can be effectively inserted into pre-trained models in other fields with a BERT-type structure, such as the ViT model, and has good performance in the tasks of this field. On the basis of maintaining the trade-off between model speed and accuracy, unnecessary waste of computing resources is reduced.

[0030] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings

[0031] Figure 1 Schematic diagram of the model framework of the present invention; Detailed Description of the Embodiments

[0032] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0033] As Figure 1 shown, an early exit method based on a learnable inter-layer integration strategy includes the following steps:

[0034] S1. Construct a BERT model with internal learners and classifiers added to each of the n transformer layers, where n is a positive integer greater than 1, so that each layer can have the ability to predict samples, which can achieve early exit of the pre-trained model and accelerate the inference of the model. The learner includes a layer of Tanh activation function, a linear layer, and a final Sigmoid activation function, and the classifier includes a single-layer fully connected network;

[0035] S2. Concatenate the hidden vectors of the first i - 1 transformer layers and input them into the learner of the i-th transformer layer. The learner connects the input concatenated hidden vectors to obtain a hyperparameter λ i ; According to the hyperparameter λ i and the feature information of the previous layer and the current layer before concatenation, obtain the logits probability distribution of the current layer, where i is greater than or equal to 1 and less than or equal to n;

[0036] The hidden vector of the i-th layer is hi, where:

[0037]

[0038] In the formula: x is the input, embedding(x) represents obtaining the word vector of x, and Encoder(h i-1 ) is the existing input h i-1 calculated by the existing encoder;

[0039] The logits probability distribution of

[0040] l i = Concat(h1, h2,..., h i-1 )

[0041] r i = Concat(hi)

[0042] λ i = Learner(l i , r i )

[0043] logits i = λ i * l i + (1 - λ i ) * r i ;

[0044] In the formula: li is the aggregation of the hidden vectors of the (i - 1)-th layer transformer layer; ri is the vector aggregation of the current layer transformer layer; Leaner() is the learning by the learning module.

[0045] S3. The logits probability distribution in step S2 is input into the classifier to obtain the output Pi;

[0046] S4. By quantifying the output distribution Pi of the i-th layer transformer layer of the model using the entropy value, the confidence e(p i ) of the output of this layer is obtained, e(p i ) = Entropy(p i );

[0047] Compare e(p i ) with the previously set threshold τ. If e(p i ) is lower than the pre-set threshold τ, the model will exit at the current layer. Otherwise, the model will continue to propagate forward to the next layer. e(p i ) can represent the confidence level of the model at the i-th layer, e(p i)The lower it is, the more stable the probability distribution of the current layer is, and the higher the confidence of the model is. Otherwise, the opposite is true;

[0048] A linear classifier (classifier) is added after the convolutional layer of each transformer layer. The entropy value is used to determine whether the current prediction result is stable. If the stable threshold is reached, the process exits and the inference process ends. Otherwise, the forward propagation of the next layer continues. The setting of the learner aims to enable the learner to automatically learn the feature information of the previous layer and the current layer, and then combine the two through hyperparameters. The experimental results show that under the same acceleration ratio, the method of the present invention can achieve more robust prediction results. There are extremely important syntactic features in the first few layers of the BERT model. If the internal classifier is generally used for global splicing, important features may be covered. Therefore, the present invention proposes a learnable integration strategy (specific step S2), aiming to enable the learner to automatically learn the feature information of the previous layer and the current layer, and then combine the two through hyperparameters (step S2). The experimental results show that under the same acceleration ratio, our method can achieve more robust prediction results.

[0049] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes an equivalent replacement or change, and should be covered within the protection scope of the present invention.

Claims

1. An early exit method based on a learnable inter-layer integration strategy, characterized in that, It includes the following steps: S1. Construct a BERT model with internal learners and classifiers added to each of the n Transformer layers, where n is a positive integer greater than 1; S2. Concatenate the hidden vectors of the first i - 1 transformer layers and use the concatenated result as the input to the learner of the i-th transformer layer. The learner connects the input concatenated hidden vectors to obtain the hyperparameter λ. i ; According to the hyperparameter λ i and the feature information of the previous layer and the current layer before concatenation, obtain the logits probability distribution of the current layer, where i is greater than or equal to 1 and less than or equal to n; S3. The logits probability distribution in step S2 is input into the classifier to obtain the output Pi; S4. By quantifying the output distribution Pi of the i-th Transformer layer of the model using the entropy value, the confidence e(pi) of the output of this layer is obtained. Compare e(pi) with the previously set threshold τ. If e(pi) is lower than the pre-set threshold τ, the model will exit at the current layer. Otherwise, the model will continue to the next layer for forward propagation.

2. The early exit method based on the learnable inter-layer integration strategy according to claim 1, wherein The learner in step S1 includes a layer of Tanh activation function, a linear layer, and a final Sigmoid activation function.

3. The early exit method based on the learnable inter-layer integration strategy according to claim 1, wherein The hidden vector of the i-th layer in step S2 is hi, where: In the formula: x is the input, and embedding(x) represents obtaining the word vector of x.

4. The early exit method based on the learnable inter-layer integration strategy according to claim 1, wherein The logits probability distribution in step S2 is: l i = Concat(h1, h2, …, h i-1 ) r i = Concat(hi) λ i = Learner(l i , r i ) logits i = λ i * l i +(1 - λ i )* r i 5. The early exit method based on the learnable inter-layer integration strategy according to claim 1, wherein In step S4, e(pi)=Entropy(pi).

6. The early exit method based on the learnable inter-layer integration strategy according to claim 1, wherein The classifier in step S4 includes a single-layer fully connected network.