Efficient and safe natural language processing method and system based on mixed attention mechanism
By introducing a hybrid attention mechanism and replacement strategy in the Transformer model, the problem of difficulty in taking into account efficiency and performance in natural language analysis tasks is solved, and better inference speed and model performance are achieved, especially in safe inference scenarios.
Patent Information
- Application Number
- CN202510217331.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-20
AI Technical Summary
The existing Transformer model has problems in both efficiency and model performance in natural language analysis tasks, especially in terms of security inference. The implementation based on MPC has significant computing and communication overhead and high inference delay.
An efficient and safe natural language processing method based on a hybrid attention mechanism is proposed. By combining the hybrid attention mechanism of Softmax Attention and Scaling Attention, two replacement strategies of post-replaced and pre-replaced are designed, the Transformer model is optimized to achieve a better balance of speed and model performance.
Through the mixed attention mechanism and replacement strategy, the inference delay of the natural language processing model is significantly reduced, while improving the model performance, achieving a better balance of speed and performance, especially in tasks such as sentiment analysis, semantic similarity and question-and-answer tasks.
Smart Images

Figure CN120179771A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to an efficient and secure natural language processing method and system based on a hybrid attention mechanism. Background Art
[0002] With the continuous update and iteration of deep learning models, artificial intelligence technology has witnessed rapid development. Especially the emergence of the Transformer model has demonstrated excellent performance and extensive application potential in the field of natural language processing. The rapid development of large language models (LLMs) represented by ChatGPT has further promoted the application and expansion of the Transformer model in more fields. Application scenarios include natural language analysis tasks such as sentiment analysis, intelligent customer service, and text recommendation, which are hereinafter uniformly referred to as inference tasks. However, the problem of user privacy and security that may exist in the natural language analysis provided by large models has become increasingly prominent. The query information input by users when interacting with the model may contain sensitive or personally identifiable information, such as health problems or career information. Frequent queries and interactions may infer more personal information of users through information accumulation, thus revealing personal preferences. Some studies have pointed out that some third-party services or plugins built on ChatGPT have behaviors of over-collecting users' personal information, which brings serious privacy leakage risks.
[0003] To address the privacy and security issues in the natural language analysis service provided by the model, many studies have adopted secure multiparty computation (MPC) to protect the privacy and security of data and models. However, existing neural network models do not consider the impact of MPC on the model during design and optimization, which leads to significant overheads in computation and communication for neural network secure inference implemented based on MPC, and the inference latency is also much higher than that of traditional plaintext inference. Especially in the Transformer model, due to its huge number of parameters and complex operations, this problem is more obvious. For example, when using the same MPC protocol for inference on the CIFAR100 dataset, the ResNet-32 model takes about 16 seconds, while the Transformer model takes about 65 seconds. In contrast, in plaintext inference without using MPC, the inference latency of the Transformer model is usually less than 1s.
[0004] The method of approximate substitution for non-linear functions (refer to Figure 4) Reducing the latency of secure reasoning has been widely studied in CNN networks, mainly in the field of image classification. There is relatively little research on tasks such as natural language analysis, mainly because the Transformer model has more nonlinear functions and is more complex, and approximate substitution often has a greater impact on model performance. Therefore, there is an urgent need to design a better replacement strategy for the Transformer model that takes into account both performance and efficiency to support secure and efficient processing of natural language. Summary of the invention
[0005] The present invention aims to solve the problem that efficiency and model performance cannot be taken into account at the same time when the Transformer model is currently used for natural language analysis tasks. An efficient and safe natural language processing method and system based on a hybrid attention mechanism are proposed. A better balance between the speed and model performance of the Transformer model when safely performing natural language analysis tasks is achieved through a hybrid attention mechanism and an attention mechanism replacement strategy.
[0006] To achieve the above purpose, the technical solution adopted is:
[0007] The present invention provides an efficient and secure natural language processing method based on a hybrid attention mechanism, comprising:
[0008] For the Transformer pre-trained model and task dataset, a hybrid attention mechanism based on Softmax Attention and Scaling Attention is proposed. For the construction of the hybrid attention mechanism model, two selection replacement strategies, post-replaced and pre-replaced, are designed.
[0009] The post-replaced strategy is to calculate the replacement rate through the performance difference rate when you have a Transformer pre-trained model on Scaling Attention for a certain data set, determine the search space, and then quickly search for the key attention head according to the NAS algorithm to restore the key attention head to Softmax Attention;
[0010] The pre-replaced strategy is to set the replacement rate from large to small when there is only a Transformer pre-trained model of a certain data set on Softmax Attention; then, by designing a fast NAS search algorithm to find and retain key attention heads in the Softmax attention mechanism, the key Softmax attention heads are retained and the remaining attention heads are approximately replaced;
[0011] The Transformer model using the MPC protocol is optimized through the above two strategies, and the optimized model is used for natural language processing to achieve a balance between speed and performance.
[0012] According to the efficient and secure natural language processing method based on the hybrid attention mechanism of the present invention, further, calculating the replacement rate through the performance difference rate in the post-replaced strategy includes:
[0013] For the Transformer pre-training models of Softmax Attention and Scaling Attention, calculate the performance difference rate τ:
[0014] τ = Δp / p Softmax =(p Softmax - p Scaling ) / p Softmax
[0015] where p Softmax represents the performance of the Transformer model composed of Softmax Attention, and p Scaling represents the performance of the Transformer model composed of Scaling Attention;
[0016] When the performance difference rate τ is greater than the performance difference rate threshold, it is considered that the model performance drops significantly. By restoring some Softmax attention heads to the ScalingAttention pre-training model to improve the model performance, calculate the number of attention head replacements Head_replaced_n and the replacement rate Head_replaced_p; when the performance difference rate τ is less than the performance difference rate threshold, that is, all are replaced with Scaling attention heads.
[0017] According to the efficient and secure natural language processing method based on the hybrid attention mechanism of the present invention, further, the calculation formula for the number of attention head replacements Head_replaced_n is:
[0018] Head_replaced_n = h * l - 1
[0019] where h represents the number of attention heads in each layer, and l represents the number of layers of the Transformer model; the calculation formula for the replacement rate Head_replaced_p is:
[0020] Head_replaced_p = 1 - 1 / h * l.
[0021] According to the efficient and secure natural language processing method based on the hybrid attention mechanism of the present invention, further, determine the search space, and then quickly search for the key attention heads according to the NAS algorithm, and restoring the key attention heads to SoftmaxAttention includes:
[0022] After determining the replacement rate, for the Transformer pre-trained model with the Scaling attention mechanism, each time select M-Head_replaced_n Scaling attention heads to be restored to Softmax attention heads and used as candidate models, where M represents the number of attention heads of the Scaling Attention pre-trained model; the candidate model search space is M·(M-1)…(Head_repalced_n+1) / (M-Head_repalced_n)!, and then test the performance of all candidate models, and select the model with the best performance as the final optimized model.
[0023] According to the efficient and secure natural language processing method based on the hybrid attention mechanism of the present invention, further, the post-replaced strategy also includes: performing model fine-tuning after the attention heads are selected for replacement, and the fine-tuning is trained using the knowledge distillation method, using the Softmax Attention pre-trained model as the teacher model and the replaced model as the student model, and at the same time freezing the weights of the unreplaced Scaling attention heads, and only fine-tuning the Softmax attention heads.
[0024] According to the efficient and secure natural language processing method based on the hybrid attention mechanism of the present invention, further, in the pre-replaced strategy, the replacement numbers are set from large to small as Head_replaced_n∈{M-1,M-2,…}, where M represents the number of attention heads of the Softmax Attention pre-trained model.
[0025] According to the efficient and secure natural language processing method based on the hybrid attention mechanism of the present invention, further, by designing a fast NAS search algorithm for finding and retaining the key attention heads in the Softmax attention mechanism, retaining the key Softmax attention heads, and approximately replacing the remaining attention heads, including:
[0026] First, replace M-Head_replaced_n heads in the Softmax attention head with Scaling attention heads, where M-Head_replaced_n << M, and test the performance of the model after the replacement; the candidate model search space is M·(M-1)…(Head_repalced_n+1) / (M-Head_repalced_n)!, select the candidate model that has the greatest impact on the model performance after replacement, then the M-Head_replaced_n heads in this model are the key Softmax attention heads in the model, replace the remaining attention heads with Scaling attention heads, and then train the model to obtain the final optimized model.
[0027] Furthermore, the present invention also provides an efficient and secure natural language processing system based on a hybrid attention mechanism for implementing the efficient and secure natural language processing method based on the hybrid attention mechanism as described above, including:
[0028] A post-replaced policy module, which is used to calculate the replacement rate through the performance difference rate, determine the search space, and then quickly search for the key attention heads according to the NAS algorithm and restore the key attention heads to Softmax Attention when there is a Transformer pre-trained model of a certain dataset on Scaling Attention;
[0029] A pre-replaced policy module, which is used to set the replacement rate from large to small when there is only a Transformer pre-trained model of a certain dataset on Softmax Attention; then, by designing a fast NAS search algorithm for finding and retaining the key attention heads in the Softmax attention mechanism, retain the key Softmax attention heads and approximately replace the remaining attention heads.
[0030] Adopting the above technical solutions, the beneficial effects obtained are:
[0031] 1. The present invention explores how to safely and efficiently execute natural language analysis tasks, starting from the internal structure of large language models, exploring the impact of different attention mechanisms and attention heads on the model performance and inference speed under natural language processing models, and through the conclusions given by experiments, explaining the correlation between the internal structure of the model and the model performance on different task datasets.
[0032] 2. The present invention proposes an efficient and secure natural language processing method - AttMA based on a hybrid attention mechanism. Facing different scenario requirements, AttMA provides two different replacement strategies, and designs different NAS fast replacement algorithms according to different natural language processing model conditions and datasets, which can greatly reduce the training cost during the search and replacement process. By adopting the designed hybrid attention mechanism, the accuracy and inference speed of the large language model in securely executing natural language analysis tasks are effectively improved.
[0033] 3. The present invention conducts experiments on various natural language analysis tasks such as sentiment analysis, semantic similarity, question - answering natural language inference, etc. The inference speed and model performance are superior to existing SOTA methods. Compared with MPCFormer and FreeDiv, the average model performance is improved by 3.12% and 5.23% respectively, and at the same time, the inference speed is increased by 1.32× and 1.29×. To more intuitively measure the comprehensive effects of different methods, the present invention defines the inference speed improvement ratio under unit performance loss. Under this index, the AttMA method is improved by 4.15× - 8.97× compared with other methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments of the present invention will be briefly introduced below. Among them, the drawings are only used to show some embodiments of the present invention, rather than limiting all embodiments of the present invention thereto.
[0035] Figure 1 It is a flowchart of the efficient and secure natural language processing method based on the hybrid attention mechanism according to the embodiment of the present invention;
[0036] Figure 2 It is a schematic diagram of the performance of three attention mechanisms on different datasets according to the embodiment of the present invention;
[0037] Figure 3 It is a schematic diagram of the influence of restoring the Softmax Attention attention head on the model performance according to the embodiment of the present invention;
[0038] Figure 4 It is a schematic diagram of the approximation of the attention heads in the Transformer model according to the embodiment of the present invention;
[0039] Figure 5 It is a comparison schematic diagram of the AttMA method and other solutions according to the embodiment of the present invention;
[0040] Figure 6 It is a schematic diagram of the relationship between the number of Softmax heads and the model performance in 4 datasets according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In the following, the exemplary solutions of the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the specific embodiments of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the art.
[0042] From the existing relevant research on how to safely and efficiently perform natural language processing, there is a lack of analysis and general conclusions on the relationship between the internal structure and performance of the model, making it difficult to provide guidance and basis for selecting appropriate attention mechanisms under different scenarios and datasets. Through research and analysis of the internal relationship among the attention heads, attention mechanisms, and network performance of the natural language processing model, the present invention discovers that the relationship between the attention mechanism and the model performance is not only determined by the type of the attention mechanism, but is also closely related to the downstream task dataset of the natural language processing model. On some datasets, the model performance composed of attention mechanisms with fast inference speed may also be better, which is very different from the view obtained from the subconscious. It is precisely this unconventional discovery that provides an opportunity and direction for the present invention to achieve a better balance between the performance and inference speed of the natural language processing model. Further, it is also found that in some attention mechanisms with significantly degraded performance, restoring a small number of attention heads to the Softmax attention mechanism can significantly improve the performance of the natural language processing model. This discovery provides a method route for supporting various task datasets to achieve a better balance. In addition, it is also observed that the distribution rules of different attention mechanisms on the same attention head are consistent. Then, the key attention heads selected under different attention mechanisms should also be the same. This discovery provides a basic basis for designing applicable key attention head search strategies for different natural language processing scenario conditions. Then, through experimental verification, there are the following three conclusions.
[0043] Conclusion 1: The relationship between different attention mechanisms and the performance and inference speed of the natural language processing model is not simply an inverse relationship. On different datasets, the performance of the attention mechanism may vary significantly. Therefore, the task dataset of natural language processing is an important factor to consider when selecting an attention mechanism.
[0044] The attention mechanism mainly identifies the correlation between different parts in the output through attention weights from the internal mechanism. From the perspective of the calculation function of attention weights, the Softmax Attention in the standard Transformer model mainly has three characteristics: ① monotonicity; ② non-negativity; ③ normalized output. Monotonicity helps to maintain the relative magnitude relationship between the input value and the output value; while non-negativity and normalization transform the distribution of input values into probability values and also enhance the contrast between different numerical distributions. From the perspective of the calculation functions of the other two attention mechanisms, 2Quad Attention satisfies characteristics ② and ③, but does not satisfy characteristic ①, while Scaling Attention only satisfies characteristic ①. From the perspective of function characteristics, neither 2QuadAttention nor Scaling Attention can fully satisfy the three characteristics of Softmax Attention, that is, from the perspective of the calculation function alone, the two attention mechanisms cannot accurately approximate Softmax Attention. However, there is no research indicating that the Transformer model using Softmax Attention must have the best performance in natural language processing. Next, the impact of the three different attention mechanisms on the actual performance of the model will be analyzed.
[0045] The experiments were conducted on different task data such as sentiment analysis, semantic similarity, question answering, and natural language inference, aiming to evaluate the performance of three different attention mechanism models. Each dataset was used as training data to train different models.
[0046] It can be observed from Figure 2 that there are significant differences in the performance of these attention mechanisms on different datasets. In datasets such as SST-2, MNLI, QQP, and QNLI, the Transformer models with the three attention mechanisms show similar performance. This indicates that on these datasets, different attention mechanisms have little impact on the model performance.
[0047] However, in the RTE and CoLA datasets, the situation is different. Compared with Softmax Attention, the model performance of 2QuadAttention and Scaling Attention has shown a significant decline. In particular, for the model with ScalingAttention, its performance decline is more significant. In the MRPC dataset, the model performance of Softmax Attention and 2QuadAttention is comparable, but the model performance of Scaling Attention again shows a significant decline. This indicates that on certain datasets, specific attention mechanisms may be more effective. Finally, in the STS-B dataset, compared with SoftmaxAttention, the model performance of both 2Quad Attention and Scaling Attention has declined, but the performance difference between the two is not large. This may mean that in some tasks, the impact of different attention mechanisms on model performance is relatively balanced.
[0048] The above experiments show that the performance and inference speed of natural language processing models under the three attention mechanisms of Softmax Attention, 2Quad Attention, and Scaling Attention are not strictly inversely proportional. Although people may subconsciously think so, the experimental results show that the attention mechanism with faster inference speed may also have the optimal model performance (such as in the SST-2 dataset), which is related to the data to be processed.
[0049] Conclusion 2: Most Softmax attention heads in Transformer-based natural language processing models can be approximately replaced without significantly affecting the overall performance of natural language processing. The optimal number of retained attention heads should be determined according to the characteristics of the dataset. For those datasets with a large performance impact after all replacements, choosing to restore a small number of Softmax attention heads can significantly improve the performance of natural language processing and achieve a better balance between the performance of natural language processing and inference speed.
[0050] To improve the speed of natural language processing, model compression techniques are often used to simplify natural language processing models. Some studies have shown that multiple heads in each layer of the Transformer model are not necessarily better than a single head. During training, multiple attention heads can be used for model training, and during inference, some layers can be simplified to single attention heads without having a great impact on the performance of natural language processing. When compressing the model, some attention heads are directly removed, while attention head approximation is to replace the attention head with other attention mechanisms. The replaced attention head will still have an impact on the model performance. The following is an experimental analysis of the impact of attention head approximation on model performance.
[0051] A series of experiments were conducted to evaluate the impact of replacing the Softmax attention heads on the model performance. The first step of the experiment was to replace all Softmax attention heads in the model with Scaling attention heads. Then, the Softmax attention heads were gradually reintroduced into the model, and the best-performing model after each introduction was selected as the performance evaluation for the number of replaced attention heads. This method can observe how the change in the number of Softmax attention heads affects the overall performance of the model. To study this phenomenon in more detail, the attention heads in these datasets were gradually replaced with the original Softmax type, and the changes in performance were observed when both types of attention heads coexisted in the model.
[0052] Through experiments, it was found that restoring the Softmax attention mechanism significantly improved the performance of natural language processing models on certain datasets. Specifically, in the four datasets of RTE, STS-B, MRPC, and CoLA, increasing the number of Softmax attention heads can significantly improve the performance of natural language processing models, as Figure 3 shown, especially when only 1 attention head was restored, the performance improvement was the most obvious. However, in the four datasets of SST-2, MNLI, QQP, and QNLI, restoring the Softmax attention heads did not significantly improve the performance of natural language processing models, and sometimes even led to a performance decline.
[0053] The experiments show that when making an approximate replacement of the attention mechanism, for those datasets with a significant performance decline after replacement, retaining at least one Softmax attention head may be an effective strategy, which can significantly improve the performance of natural language processing models. In addition, the experimental results also verified that in some cases, the Softmax attention mechanism is not superior to the more efficient Scaling attention mechanism.
[0054] Conclusion 3: In a Transformer-based natural language processing model, different attention mechanisms under the same attention head have similar output distributions. Therefore, when selecting and replacing attention heads, using different attention mechanisms in the model does not affect the search results of key attention heads.
[0055] Conclusion 2 above shows that the performance of natural language processing models can be improved by restoring a small number of key attention heads. Then, under different attention mechanisms, whether the key attention heads are the same is related to how to find and replace key attention heads under different conditions. To address this issue, an analysis experiment was conducted on the output distributions of attention heads at the same position under different attention mechanisms.
[0056] The experimental results show that the attention weight distributions of different attention mechanisms on the same attention heads are similar, but the attention weight distribution of Softmax Attention is clearer. This main advantage is attributed to the role of the Softmax function in amplifying differences. For the input weight values, the Softmax function amplifies the differences between elements through exponentiation operations. This means that larger weight values will become even larger through exponential growth, while smaller weight values will become smaller. In the process of natural language processing, this can allocate more weights to the more important parts of the task, which is also the main reason why the model with Softmax attention mechanism has relatively optimal performance.
[0057] Based on the similarity of the attention head weight distributions of different attention mechanisms at the same position, it can be inferred that the roles of these attention heads in the model are also roughly the same. Therefore, when approximately replacing the attention heads of the model, there are two ideas. One is to use the Softmax Attention standard pre-trained model as a benchmark, and search for the attention head that has the greatest impact on the model performance after replacing a small number of attention heads with Scaling Attention. The other idea is to start from the Scaling Attention pre-trained model and determine the key attention head that can maximize the improvement of the model performance by restoring the Softmax attention head. From the analysis of the experimental results, the positions of the key attention heads determined by the two ideas are basically the same. Therefore, different attention head replacement strategies can be designed according to different scenario conditions.
[0058] Based on the above three experimental conclusions, this embodiment proposes an efficient and secure natural language processing method based on a hybrid attention mechanism - AttMA, which can achieve the rapid search and replacement of key attention heads in a natural language processing model. For two different target requirements in practical applications, two selection and replacement strategies are designed respectively, which can maximize the performance and speed of the model's safe inference when facing different requirements. The AttMA method is mainly designed based on Neural Architecture Search (NAS). NAS is a technology that searches for a better network structure through the automation of neural network structure design. NAS to achieve neural network architecture search includes three steps: defining the search space, determining the search strategy, and determining the performance evaluation strategy. The AttMA method mainly includes three stages: determining the replacement rate, rapid selection and replacement of attention heads based on NAS, and model fine-tuning. The overall framework is as Figure 1 shown. The AttMA method specifically includes:
[0059] The AttMA method addresses the requirements of different natural language processing scenarios. For the Transformer pre-trained model and task datasets, it proposes a hybrid attention mechanism based on Softmax Attention and Scaling Attention. For the construction of a model based on the hybrid attention mechanism, it designs two selection and replacement strategies: post-replaced and pre-replaced. The post-replaced strategy is when there is a Transformer pre-trained model on Scaling Attention for a certain dataset. The replacement rate is calculated through the performance difference rate to determine the search space. Then, according to the NAS algorithm, the key attention heads are quickly searched and restored to Softmax Attention to improve the performance of the Scaling Attention pre-trained model. The pre-replaced strategy is when there is only a Transformer pre-trained model on Softmax Attention for a certain dataset and it is desired to replace it with a more efficient attention mechanism to improve the model inference speed. The replacement rate can be set from large to small, and then a fast NAS search algorithm for finding and retaining key attention heads in the Softmax attention mechanism is designed according to Conclusion 3, which greatly reduces the training overhead brought by the search and replacement strategy. The following details the two strategies separately.
[0060] (1) The post-replaced strategy
[0061] This strategy addresses the scenario requirement of quickly improving the performance of existing pre-trained models. Taking the Softmax Attention and Scaling Attention pre-trained models as examples, the AttMA method can quickly improve the model performance without retraining the model by selecting and restoring the key Softmax attention heads for the Scaling pre-trained model, with basically no increase in inference latency.
[0062] ① Determine the replacement rate
[0063] Current research on different attention mechanisms is relatively rich. There are natural language processing pre-trained models constructed with each common attention mechanism. If there is no pre-trained model for the new attention mechanism, the Softmax Attention of the natural language processing pre-trained model can be all replaced with the new attention mechanism first. Then, based on the parameters of the pre-trained model as the initial model parameters, and using the Transformer pre-trained model as the teacher model for distillation training, a pre-trained model of the new attention mechanism can be quickly trained.
[0064] For two pre-trained models, to eliminate the influence of different model performance evaluation metrics under different datasets, a performance difference rate τ is set, and the calculation formula is as follows:
[0065] τ = Δp / p Softmax = (p Softmax - p Scaling ) / p Softmax
[0066] Among them, p Softmax represents the performance of the Transformer model composed of Softmax Attention, and p Scaling represents the performance of the Transformer model composed of Scaling Attention. According to the calculation formula of the performance difference rate τ, calculate the model performance differences between the two attention mechanisms of Softmax Attention and Scaling Attention under each dataset such as sentiment analysis, semantic similarity, and question answering natural language inference, as shown in Table 1.
[0067] Table 1 Comparison of model performance differences between Softmax Attention and Scaling attention mechanisms on different datasets
[0068]
[0069] Generally, when τ > 0.05, it is considered that the model performance drops significantly, and the performance loss cannot be ignored. It is necessary to restore the model performance. By restoring some Softmax attention heads, the model performance is improved. Calculate the number of replaced attention heads Head_replaced_n:
[0070] Head_replaced_n = h * l - 1
[0071] Among them, h represents the number of attention heads in each layer, and l represents the number of layers of the Transformer model; the calculation formula for the replacement rate Head_replaced_p is:
[0072] Head_replaced_p = 1 - 1 / h * l
[0073] Similarly, when τ < 0.05, Head_replaced_n = h * l, that is, all are replaced with Scaling attention heads.
[0074] ② Fast selection and replacement of attention heads based on NAS
[0075] After determining the replacement rate, for the Transformer pre-trained model of the Scaling attention mechanism, each time select M-Head_replaced_n Scaling attention heads to be restored to Softmax attention heads and used as candidate models. The candidate model search space is M·(M-1)…(Head_repalced_n+1) / (M-Head_repalced_n)! , and then test the performance of all candidate models, and select the model with the best performance as the final optimized model. The search algorithm is shown in Algorithm 1:
[0076]
[0077] ③Model fine-tuning
[0078] After the attention head is replaced, the model needs to be trained and fine-tuned. For example, to make the model parameters more compatible with the new attention mechanism, the model can be fine-tuned. The fine-tuning adopts the Knowledge Distillation (KD) method for training, with the Softmax Attention pre-trained model as the teacher model and the replaced model as the student model. At the same time, the weights of the unreplaced Scaling attention heads are frozen, and only the weights of the Softmax attention heads are fine-tuned. In this way, fine-tuning training can be performed quickly to further restore the model performance.
[0079] (2) Pre-replaced strategy
[0080] This strategy is aimed at natural language processing pre-trained models that only have standard Softmax attention, and hope to choose a more efficient attention mechanism to replace it, while maintaining the model performance as much as possible. Taking the replacement with the Scaling attention mechanism as an example, according to the previous conclusion, most of the Softmax attention heads can be replaced with Scaling attention heads. Therefore, the replacement rate can be set in reverse order from large to small, and only one Softmax attention head can be retained to achieve a better balance between model performance and inference speed on most data sets. That is, the optimization process only requires one training, which greatly reduces the training overhead caused by search and replacement.
[0081] ①Determine the replacement rate
[0082] On the standard Softmax attention natural language processing pre-trained model, consider replacing it with MPC-friendly Scaling Attention to improve the inference speed while maintaining the model performance. According to Conclusions 1 and 2 above, the vast majority of Softmax attention heads in the natural language processing model can be approximately replaced without affecting the model performance, and the optimal number of remaining attention heads needs to be determined according to the dataset situation. Therefore, choosing to decrease the attention head replacement rate from large to small can greatly reduce the search space compared with the traditional method of gradually increasing the replacement rate. Under the pre-replaced strategy, first set the number of replaced attention heads as Head_replaced_n ∈ {M - 1, M - 2, …}, where M represents the number of attention heads of the Softmax Attention pre-trained model, that is, first choose to retain 1 attention head. This strategy performs better on most datasets, thus defining the search space of NAS.
[0083] ② Fast selection and replacement of attention heads based on NAS
[0084] According to the replacement rate, first, Head_replaced_n heads should be selected to be replaced with Scaling attention heads, and then the model performance is tested. However, due to the large number of replaced heads, the model performance drops significantly. Each time, the model performance needs to be restored through re-training before comparison and screening, which still brings a large amount of training overhead. According to Conclusion 3, when screening the key attention heads, it is improved to first replace M - Head_replaced_n heads in the Softmax attention heads with Scaling attention heads. Since M - Head_replaced_n << M, the impact on the model performance after replacement is small, and the performance of the replaced model can be directly tested. The candidate model search space is M·(M - 1)…(Head_repalced_n + 1) / (M - Head_replaced_n)!. Select the candidate model with the greatest impact on the model performance after replacement. Then, the M - Head_replaced_n heads in this model are the key Softmax attention heads in the model. Replace the remaining attention heads with Scaling attention heads, and then train the model to obtain the final optimized model. Adopting this strategy avoids a large amount of repeated training in the process of screening key attention heads and greatly reduces the training overhead. The pre-replaced strategy algorithm is shown in Algorithm 2.
[0085]
[0086] Correspondingly, this embodiment also proposes an efficient and secure natural language processing system based on a hybrid attention mechanism, including:
[0087] The post-replaced policy module is used to calculate the replacement rate through the performance difference rate when having a Transformer pre-trained model on Scaling Attention for a certain dataset, determine the search space, and then quickly search for the key attention heads according to the NAS algorithm, and restore the key attention heads to Softmax Attention.
[0088] The pre-replaced policy module is used to set the replacement rate from large to small when only having a Transformer pre-trained model on Softmax Attention for a certain dataset; then, by designing a fast NAS search algorithm for finding and retaining the key attention heads in the Softmax attention mechanism, retain the key Softmax attention heads and approximately replace the remaining attention heads.
[0089] To verify the effectiveness of the solution in this case, the following further explains with experimental data.
[0090] (I) Experimental settings
[0091] 1) Experimental environment
[0092] For easy comparative analysis, consistent with the existing main methods, the secure inference environment in the AttMA method is mainly based on the secret-sharing-based MPC system built by Crypten, assuming that the participants in the inference are semi-honest. The experimental platform uses two ThinkStation P920 workstations, each equipped with two RTX3090 GPUs, working in an Ethernet with a 10 GbE bandwidth.
[0093] 2) Model structure and dataset
[0094] Datasets of various task types are selected from the GLUE dataset for experiments. The task objectives and basic information of different datasets are shown in Table 2. The model selects the classic Transformer architecture Bert-Base-Cased, with a model depth of 12 layers and 12 attention heads in each layer.
[0095] Table 2 Datasets of different task types selected in the experiment
[0096]
[0097]
[0098] 3) Parameter settings
[0099] In this experiment, after the attention heads were selectively replaced, the models were all trained using the knowledge distillation method, which was specifically carried out in two stages as in MPCFormer. In the first stage, the embedding layer and the Transformer layer were distilled, with the learning rate set to 5e-5, and the number of training epochs was set according to the size of the dataset. In the second stage, the prediction layer was distilled, with the learning rate set to 1e-5, and the number of training epochs was set to 5 rounds.
[0100] During the process of selecting key attention heads, only the model performance was tested. Since selecting a partial dataset for testing would not affect the sorting and selection of key attention heads, to improve the testing efficiency, 1000 groups of data were randomly selected from the test set during testing. It was found in the experiment that when the number of retained Softmax attention heads exceeded 12, the model performance decreased significantly, and the model performance needed to be restored through retraining, which was beyond the scope of this invention. Therefore, in this experiment, the number of replaced attention heads was set between 1 and 12.
[0101] 4) Evaluation metrics
[0102] ① Model performance
[0103] Model performance is an important indicator for measuring the effectiveness of different attention mechanism replacement methods. However, different attention mechanism models perform differently on different datasets. To more intuitively compare the impact of different methods on model performance, the average performance Avg_Perf on each dataset was used as the overall evaluation of the method in terms of model performance. Its calculation formula is:
[0104]
[0105] n represents the number of datasets used in the comparison. In addition, from the perspective of different attention mechanism replacement methods, it is hoped that the impact of the replacement operation on the model performance can be minimized as much as possible. Therefore, on the basis of the average performance, the performance difference ΔAvg_Perf ↓ was added as the comparison of the impact on performance among different methods, as shown in the following formula:
[0106] ΔAvg_Perf ↓ = Avg_Perf Softmax - Avg_Perf *
[0107] ΔAvg_Perf ↓ The smaller it is, the smaller the impact of the method on the performance of the baseline Softmax model after replacement.
[0108] ② Inference speed
[0109] For the convenience of comparative analysis with other methods and to eliminate the influence of different security inference architectures on inference latency, the inference speed is uniformly based on Softmax Attention as the benchmark unit, and the multiple of the inference speed improvement is observed, as shown in the following formula:
[0110]
[0111] ③ Comprehensive evaluation
[0112] The method of the present invention aims to achieve a better balance between model performance and inference speed. To measure the comprehensive ability of the method of the present invention in terms of model performance and inference speed, the improvement ratio of inference speed per unit performance is used as the comprehensive evaluation index for each method, indicating the multiple of the inference speed improvement per unit loss of model performance. For each attention mechanism replacement method, the larger this index is, the greater the inference speed that the method can improve when the performance of the replaced model is equivalent.
[0113] In addition, the search overhead of attention heads is the main overhead in the method of the present invention, and such overhead also exists in many other methods. For the convenience of practical application, in the ablation experiment, the search overhead of replacing different numbers of attention heads is added to the evaluation index. Based on the Scaling Attention pre-trained model, the improvement ratio of the search overhead of replacing different numbers of attention heads on model performance is observed. To accelerate the search speed, the present invention sets the number of search rounds to be proportional to the number of replaced attention heads. Therefore, the improvement ratio of the search overhead of replacing different numbers of attention heads on model performance can be recorded as The larger this index is, the more obvious the improvement effect of the model performance with the same overhead consumption in practical applications.
[0114] (2) Performance of the method of the present invention on different datasets
[0115] Experiments on the AttMA method were conducted on different task datasets, and the benchmark models were the single SoftmaxAttention and Scaling Attention pre-trained models respectively. The influence on model performance and inference speed was observed, and the experimental results are shown in Table 3.
[0116] Table 3 Model performance and inference speed performance of the AttMA method on different task datasets
[0117]
[0118] The table shows the optimal number of key Softmax attention heads found by the AttMA method on different datasets and the performance of the optimized model. The results are tested without retraining and fine-tuning, showing that the AttMA method can adapt to different datasets, replacing more than 90% of the Softmax attention heads with faster ScalingAttention without significantly affecting the model performance.
[0119] From the experimental results, the inference speed of the AttMA method is basically the same as that of the standard Scaling Attention pre-trained model, and the inference speed is about 3× higher than that of the Softmax Attention pre-trained model. However, compared with the Scaling Attention pre-trained model, the AttMA method has a certain improvement in model performance on different datasets, with the improvement range from 0.21 to 4.68, and the average performance improvement is 1.94. This shows that the AttMA method can significantly improve the inference speed of the model while minimizing the impact on the model performance compared with the standard Softmax Attention. Compared with directly replacing the Softmax attention heads with Scaling Attention, it can greatly improve the performance of the Scaling Attention pre-trained model without reducing the inference speed.
[0120] (3) Comparison of the method of the present invention with existing methods
[0121] The AttMA method is compared with existing similar methods including SOTA methods, mainly including MPCFormer, FreeDiv, and single Scaling Attention. The AttMA method shows the model performance of only restoring 1 Softmax attention head on different datasets and the optimal model performance obtained by restoring the Softmax attention heads.
[0122] For convenient comparison, the datasets are selected to be the same as those in the comparison objects. The model performance and inference speed are both based on the standard Softmax attention mechanism model, and the impact on the model performance ΔAvg_Perf ↓ and the inference speed improvement ratio Speed_up after replacing the attention mechanism are compared, and a comprehensive evaluation index is used to uniformly compare and measure each method. The experimental results are shown in Table 4.
[0123] From the experimental results, the AttMA method outperforms other methods in all three datasets. Compared with the MPCFormer scheme, the AttMA method increases the model inference speed by 0.71×, while the model performance increases by 1.53, 1.01, and 6.82 respectively; compared with the FreeDiv scheme, the model inference speed increases by 0.65×, while the model performance increases by 8.03, 1.51, and 6.22 respectively; in addition, compared with the Scaling Attention model with the fastest inference speed, the inference speed of the AttMA method is basically the same, but the model performance increases by 3.98, 0.46, and 4.68 respectively. And when only one Softmax attention head is restored, the model performance also increases by 1.81, 0.23, and 4.29 respectively, showing that the AttMA method can quickly improve the performance of the ScalingAttention pre-trained model without reducing the model inference speed and without any retraining, only by restoring the key Softmax attention heads.
[0124] Table 4 Comparison of the performance of the AttMA method and existing main methods on different datasets
[0125]
[0126]
[0127] To compare the trade-off between the model inference speed and performance of each scheme, a comprehensive evaluation index is used to conduct a comparative analysis of each scheme. The larger this index, the greater the improvement in the inference speed under the condition of the same loss of model performance. The comparison results of this index are as Figure 5 shown. From the experimental results, the AttMA method outperforms other methods. Compared with the MPCFormer and FreeDiv schemes, this index increases by 5.89× and 8.73× respectively, showing that the AttMA method can achieve a better balance between the inference speed and model performance.
[0128] (IV) Ablation experiment
[0129] The search cost of the key attention heads in the method of the present invention is the main cost of the AttMA method. Although this cost is much smaller than the cost of training the model, in order to provide more specific and referenceable implementation details for researchers using the AttMA method, the present invention conducts a comparative analysis of the model performance improvement and the actual cost of the AttMA method. Among them, the model performance improvement is based on the Scaling Attention pre-trained model, and the evaluation index uses the comprehensive evaluation index to represent the improvement effect of the model performance per unit cost consumed during the practice process.
[0130] From the experimental results Figure 6 it can be seen that although the number of Softmax attention heads to be retained for achieving the optimal model performance varies on different datasets, the maximum value of each dataset's metric is to retain only 1 Softmax attention head, indicating that when the implementation cost is limited, directly choosing to retain 1 attention head has the highest cost performance for improving the model performance. This provides a more detailed reference basis for the practical application of the AttMA method.
[0131] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An efficient and secure natural language processing method based on a hybrid attention mechanism, characterized in that: include: For the Transformer pre-trained model and task dataset, a hybrid attention mechanism based on SoftmaxAttention and Scaling Attention is proposed, and two selection replacement strategies, post-replaced and pre-replaced, are designed for the construction of the hybrid attention mechanism model. The post-replaced strategy is to calculate the replacement rate through the performance difference rate when you have a Transformer pre-trained model on Scaling Attention for a certain data set, determine the search space, and then quickly search for the key attention head according to the NAS algorithm to restore the key attention head to Softmax Attention; The pre-replaced strategy is to set the replacement rate from large to small when there is only a Transformer pre-trained model of a certain data set on Softmax Attention; then, by designing a fast NAS search algorithm to find and retain key attention heads in the Softmax attention mechanism, the key Softmax attention heads are retained and the remaining attention heads are approximately replaced; The Transformer model using the MPC protocol is optimized through the above two strategies, and the optimized model is used for natural language processing to achieve a balance between speed and performance.
2. The efficient and secure natural language processing method based on the hybrid attention mechanism according to claim 1 is characterized in that: The post-replaced strategy calculates the replacement rate by using the performance difference rate, including: For the Transformer pre-trained models with two attention mechanisms, Softmax Attention and Scaling Attention, calculate the performance difference rate τ: τ=Δp / p Softmax =(p Softmax -p Scaling ) / p Softmax Among them, p Softmax represents the performance of the Transformer model composed of Softmax Attention, p Scaling Indicates the performance of the Transformer model composed of Scaling Attention; When the performance difference rate τ is greater than the performance difference rate threshold, it is considered that the model performance has declined significantly. The model performance is improved by restoring some Softmax attention heads in the ScalingAttention pre-training model, and calculating the number of attention head replacements Head_replaced_n and the replacement rate Head_replaced_p. When the performance difference rate τ is less than the performance difference rate threshold, all of them are replaced with Scaling attention heads.
3. The efficient and secure natural language processing method based on the hybrid attention mechanism according to claim 2 is characterized in that: The calculation formula for the number of attention head replacements Head_replaced_n is: Head_replaced_n=h*l-1 Among them, h represents the number of attention heads in each layer, l represents the number of layers of the Transformer model; the calculation formula of the replacement rate Head_replaced_p is: Head_replaced_p=1-1 / h*l.
4. The efficient and secure natural language processing method based on the hybrid attention mechanism according to claim 2 is characterized in that: Determine the search space, then quickly search for the key attention head according to the NAS algorithm, and restore the key attention head to Softmax Attention including: After determining the replacement rate, for the Transformer pre-trained model of the Scaling attention mechanism, each time M-Head_replaced_n Scaling attention heads are restored to Softmax attention heads and used as candidate models, where M represents the number of attention heads in the Scaling Attention pre-trained model; the candidate model search space is M·(M-1)…(Head_repalced_n+1) / (M-Head_repalced_n)! , and then the performance of all candidate models is tested, and the model with the best performance is selected as the final optimized model.
5. The efficient and secure natural language processing method based on hybrid attention mechanism according to claim 1, characterized in that: The post-replaced strategy also includes: fine-tuning the model after the attention head is selected and replaced. The fine-tuning is trained using the knowledge distillation method, with the Softmax Attention pre-trained model as the teacher model and the replaced model as the student model. At the same time, the weights of the unreplaced Scaling attention heads are frozen, and only the Softmax attention head is fine-tuned.
6. The efficient and secure natural language processing method based on hybrid attention mechanism according to claim 1, characterized in that: In the pre-replaced strategy, the number of replacements is set from large to small as Head_replaced_n∈{M-1,M-2,…}, where M represents the number of attention heads of the Softmax Attention pre-training model.
7. The efficient and secure natural language processing method based on hybrid attention mechanism according to claim 6, characterized in that: By designing a fast NAS search algorithm to find and retain key attention heads in the Softmax attention mechanism, the key Softmax attention heads are retained and the remaining attention heads are approximately replaced, including: First, replace the M-Head_replaced_n heads in the Softmax attention head with the Scaling attention head, and M-Head_replaced_n<<M, and test the performance of the replaced model; the candidate model search space is M·(M-1)…(Head_repalced_n+1) / (M-Head_repalced_n)!, select the candidate model that has the greatest impact on the model performance after replacement, then the M-Head_replaced_n heads in the model are the key Softmax attention heads in the model, replace the remaining attention heads with the Scaling attention heads, and then train the model to obtain the final optimized model.
8. An efficient and secure natural language processing system based on a hybrid attention mechanism, characterized in that: The method for implementing an efficient and secure natural language processing method based on a hybrid attention mechanism as described in any one of claims 1 to 7 comprises: The post-replaced strategy module is used to calculate the replacement rate through the performance difference rate when a Transformer pre-trained model of a certain data set is available on Scaling Attention, determine the search space, and then quickly search for the key attention head according to the NAS algorithm to restore the key attention head to Softmax Attention; The pre-replaced strategy module is used when there is only a certain data set in the Transformer pre-training model on Softmax Attention. The replacement rate is set from large to small; then a fast NAS search algorithm is designed to find and retain the key attention heads in the Softmax attention mechanism, retain the key Softmax attention heads, and approximately replace the remaining attention heads.
9. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.