Attention dilution optimization method combining wavelet feature analysis and anisotropy loss
By combining wavelet feature analysis and anisotropic loss attention dilution optimization method, the Transformer attention mechanism has solved the problem of high computational complexity and poor robustness in long-sequence data processing, and efficient attention calculation and model robustness improvement are achieved.
Patent Information
- Application Number
- CN202510416785.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing Transformer attention mechanism has high computational complexity when processing long sequence data, resulting in increased computing power and memory overhead, and is poorly robust, making it difficult to adapt to different downstream tasks.
Combining wavelet characteristic analysis and anisotropic loss, an attention dilution optimization method is proposed. By presetting multiple quantized vocabulary lists, using wavelet transformation and quantized loss calculation, the sparseness of the attention score matrix is achieved, and attention calculation is adopted using the Query-Key sharing mechanism.
With almost no loss of model accuracy, the computational complexity is significantly reduced, the computational efficiency is improved, and the model is suitable for larger context lengths, and the model is enhanced.
Smart Images

Figure CN119940449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large models, and in particular to an attention dilution optimization method combining wavelet feature analysis and anisotropic loss, and a corresponding computer program product and natural language processing device. Background Art
[0002] In recent years, large model technology has risen rapidly. With its powerful capabilities and generalization, it has set off a wave of changes in many fields, especially in natural language processing (NLP) and computer vision (CV). Existing large models are mainly based on Transformer. Once released, Transformer attracted the attention of most researchers in the field of artificial intelligence. After several years of development, models based on this architecture have replaced the traditional recurrent neural network (RNN) and convolutional neural network (CNN) in many tasks. This is due to its high parallelism and its unique self-attention mechanism (Self-Attention). Figure 1 (a) in the figure shows the calculation process of attention. By calculating the dot product of the query vector (query) and the key vector (key), scaling and applying the softmax function, the probability distribution corresponding to each key is obtained. Finally, the value vector is weighted and summed with the distribution, and the output represents the weighted value of attention. It can effectively capture the dependency between any positions in the input sequence and model global information. Figure 1 (b) in the figure shows a multi-head attention layer. This structure concatenates the weight representation vectors obtained by passing the vector through several attention layers. It can extract information with different focuses by training different weights. At the same time, it does not need to input sequence data step by step, so it can greatly improve the training speed and the efficiency of model parallel computing. With its unique advantages, Transformer has achieved excellent results in many tasks since its inception. However, although the attention mechanism has given it a leading position in many fields, as the sequence length increases, the computational overhead of the attention score is also facing increasingly severe computational complexity challenges: if the length of the two input sequences is n, then the computational complexity of the attention score is O(n 2). Therefore, when the input sequence increases by an order of magnitude, the computational overhead of the self-attention mechanism will increase quadratically, which not only consumes a lot of computing resources, but also brings additional memory bottleneck challenges. These problems are particularly evident when processing large-scale data such as long documents, multimodal data, and image sequences, and have become an important obstacle to the advancement of large-scale model technology. However, the attention score matrix is usually sparse, and the inner product between most queries and keys in the matrix is small. After softmax processing, these vectors account for a negligible proportion of the final weight representation. Therefore, the rational use of the sparsity of the attention score matrix can not only reduce unnecessary computational overhead, but also reduce the proportion of weights that need to be saved, thereby effectively alleviating the memory bottleneck and computing power overhead problems.
[0003] One of the mainstream solutions for Transformer attention sparsity is to reduce the computational complexity of the attention score matrix while ensuring excellent model performance by applying a specific sparsity pattern. The research on sparse patterns mainly focuses on four mainstream strategies: 1. Local Sliding Window Attention This strategy is based on the locality principle in NLP, that is, the correlation between a token and its adjacent tokens is usually higher than that between distant tokens. The researchers effectively reduced the computational complexity by limiting each query to focus only on the keys within its fixed window.
[0004] 2. Global Sparse Attention This strategy selects a fixed number of tokens so that it can focus on information at all locations, thereby reducing the amount of computation while improving the ability to capture global information. A typical model that applies this strategy is ETC, which combines local sliding windows with global sparsity to enable the model to simultaneously learn long-range dependencies and neighboring context expressions.
[0005] 3. Random Sparse Attention In order to introduce the expression of distant tokens, the random sparsity strategy selects attention pairs for calculation through random sampling. Although random sparsity can take tokens at distant locations into consideration while reducing the amount of calculation, for example, the Big Bird model (Manzil Zaheer et al., 2020) combines sliding windows, global tokens, and random sparsity to achieve excellent performance in long text processing tasks, it also introduces randomness, which reduces the stability and robustness of the model.
[0006] 4. Dilated Sparse Attention In order to expand the receptive field of the sliding window, researchers introduced the idea of dilated convolution. They introduced a dilation factor in the local window so that the model can focus on tokens at farther positions and expand the receptive field of the local sliding window. Longformer is a typical example, which can maintain good performance even when the context length is 4096.
[0007] The above methods are widely used in natural language understanding, time series analysis and other fields, but these four sparse modes are usually only applicable to a certain data type and have poor robustness on other downstream tasks; the existing Transformer attention mechanism still has problems such as losing global dependency information and introducing additional errors. Summary of the invention
[0008] In order to solve the problems of low efficiency and poor robustness of the existing Transformer attention mechanism, the present invention provides an attention dilution optimization method combining wavelet feature analysis and anisotropic loss, and its corresponding computer program product and natural language processing device.
[0009] The technical solution provided by the present invention is as follows: An attention dilution optimization method combining wavelet feature analysis and anisotropic loss comprises the following steps: S1: Preset in order from low to high accuracy n Initialize the quantization vocabulary. For each layer of input features X emb Perform separation wavelet transform to obtain the low-frequency approximation coefficient A and high-frequency detail coefficient D of the input feature. Calculate the energy distribution corresponding to A and D E low and E high ; Then calculate it by the following formula S The value of S Select the initialization quantization vocabulary of the corresponding sequence number: , in, Indicates a floor operation.
[0010] S2: Calculate the score-aware quantization loss between each query and each quantization point in the initial quantization vocabulary, and use it as the distance between the two. Assign each query to the quantization point closest to it, and then obtain the updated quantization vocabulary. The updated quantization vocabulary contains multiple independent quantization partitions. C j, each quantization partition consists of a quantization point and its corresponding at least one query.
[0011] S3: Use the updated quantization vocabulary to quantize all queries into corresponding quantization points ; Build a lookup table by calculating the inner product of the quantization point and the key; Use the quantization points in the lookup table that are related to each other Data pairs with key vector k Approximate representation of the required query vector q With key vector k Data pair , and then realize attention calculation.
[0012] As a further improvement of the present invention, in the attention calculation process of S3, a Query-Key sharing mechanism is adopted to make the query and key share weights.
[0013] As a further improvement of the present invention, the score-aware quantization loss of any query and quantization point The calculation formula is as follows: , In the above formula, q i Indicates i query; Indicates the first j Quantitative points; express The horizontal component of express The vertical component of represents the weight coefficient of the vertical component; Represents the weight coefficient of the horizontal component.
[0014] As a further improvement of the present invention, the weight coefficients of the vertical component and the horizontal component are and The calculation formula is: ; In the above formula, p Represents the dimension of the data; w represents a weight function for balancing the parallel quantization error and the vertical quantization error. By selecting this function, we can better focus on the larger inner product term, thereby giving a larger weight to the parallel quantization error and penalizing the vertical error component; is an angle parameter that satisfies: .
[0015] As a further improvement of the present invention, anisotropic loss LInstead of score-aware quantization loss, it is used to measure the distance between any query and the quantization point; anisotropic loss The calculation formula is: .
[0016] As a further improvement of the present invention, the updated quantization word list contains multiple quantization partitions, each quantization partition C j All queries in the same quantization point belong to the same quantization point; the quantization point in the quantization vocabulary The iteration formula is as follows: , In the above formula, I represents an identity matrix.
[0017] As a further improvement of the present invention, the input feature X emb The low-frequency approximation coefficient A and high-frequency detail coefficient D and their energy distribution E low and E high The calculation formula is as follows: , In the above formula, Represents a separable wavelet transform operation.
[0018] The present invention also includes an application of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as mentioned above to replace the self-attention module in a Transformer-based neural network to realize attention calculation.
[0019] The present invention also includes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as mentioned above, and then completes the attention calculation of the input features.
[0020] The present invention also includes a natural language processing device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, a Transformer-based data processing model is used to process input data; wherein the self-attention module in the data processing model uses the aforementioned attention dilution optimization method combining wavelet feature analysis and anisotropic loss to realize attention calculation.
[0021] The present invention has the following beneficial effects: The present invention applies technologies such as anisotropic quantization loss, wavelet feature extraction, and QK weight sharing to the dilution optimization of the Transformer's self-attention mechanism. These new attention sparse strategies can reduce computational complexity and improve computational efficiency without losing model accuracy. In the inference stage, the present invention uses a lookup table to quickly obtain approximate attention scores, avoid redundant calculations, and improve inference speed.
[0022] The sparse attention calculation method of the present invention can reduce the original quadratic computational complexity, allowing the Transformer to run efficiently on a larger context length and be suitable for resource-constrained devices. The scheme of the present invention sets up a variety of quantization vocabulary with different precision levels, and flexibly selects the best quantization vocabulary in combination with wavelet feature analysis technology, which makes the scheme of the present invention applicable to processing different types of data types and various downstream tasks, significantly enhancing the robustness of the scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a schematic diagram of the existing self-attention mechanism used by the Transformer introduced in the background technology. Part (a) of the figure is the schematic diagram of a simple dot product attention mechanism, and part (b) is the schematic diagram of a multi-head attention mechanism.
[0024] Figure 2 Schematic diagram of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided in Example 1 of the present invention.
[0025] Figure 3 This is a curve showing how the performance of the model using the attention dilution scheme provided by the present invention changes with the number of quantized vocabularies in a test experiment. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.
[0028] Example 1 This embodiment provides an attention dilution optimization method that combines wavelet feature analysis and anisotropic loss. The scheme applies technologies such as anisotropic quantization loss, wavelet feature extraction, and QK weight sharing to the self-attention mechanism of Transformer, and then designs a more efficient attention sparse optimization method. This method can greatly reduce the computational overhead of Transformer in long sequence tasks with almost no loss of model accuracy. Therefore, this scheme can be easily combined with new technologies such as distillation and further improve the performance of the model, providing new possibilities for the application of large models in resource-constrained environments, thereby promoting the further development of the Transformer architecture in NLP, CV and cross-modal tasks.
[0029] Specifically, Figure 2 As shown, the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided in this embodiment includes the following steps: S1: Generate multiple sparse quantization word lists, and use wavelet feature analysis technology to analyze the input features, and then adaptively select the most appropriate quantization word list in the space to quantize the query according to the frequency distribution of the input. The process includes: S11: First, preset the number of n ) initialize the quantization vocabulary.
[0030] In this embodiment, the benefit of quantizing the query into quantization points in the quantization vocabulary is that the number of calculations can be reduced. However, the input features of each layer of the network are only a small part of the data, and a single vocabulary is difficult to fully represent all queries. To address this problem, this embodiment first generates an original quantization vocabulary with the widest range and the most quantization points based on the distribution of input features, and then generates a quantization vocabulary based on the original quantization vocabulary in order of accuracy from low to high. n For example, assuming that the number of the initialized quantization word lists is 8, the first initialized quantization word list is a quantization word list with coarser feature changes and lower quantization accuracy, while the eighth initialized quantization word list is a quantization word list with finer feature changes and higher quantization accuracy.
[0031] S12: Select the most appropriate initialization quantization vocabulary according to the frequency distribution of the current input features.
[0032] In practical applications, this embodiment first performs an input feature X emb Perform separation wavelet transform to obtain low-frequency approximate coefficients of input features A and high frequency detail factor D , the process is expressed as: , in, Represents a separable wavelet transform operation.
[0033] Next, the low-frequency approximation coefficients are calculated by the following formula: A and high frequency detail factor D The corresponding energy distribution E low and E high :in, , .
[0034] Finally, according to the known E low and E high Calculate a S value, and according to S The value of selects the initialization quantization vocabulary corresponding to the sequence number: , in, Indicates a floor operation.
[0035] Combining the formula, we can see that: In this embodiment, when the high-frequency detail coefficient of the input feature D Energy distribution E high The higher the value, the higher the precision level of the quantization vocabulary will be selected to quantize the query; conversely, the lower the precision level of the quantization vocabulary will be selected to quantize the query; this can significantly improve the robustness of the selected quantization vocabulary under different input samples.
[0036] S2: Iteratively update the quantization vocabulary to obtain each quantization point and all its corresponding queries to form a quantization partition C j .
[0037] Specifically, at this stage, the present embodiment first calculates the score-perceived quantization loss between each query and each quantization point in the initial quantization vocabulary, and uses it as the distance between the two. Then each query is assigned to the quantization point closest to it, thereby obtaining an updated quantization vocabulary. The updated quantization vocabulary contains multiple independent quantization partitions. C j , each quantization partition consists of a quantization point and its corresponding at least one query.
[0038] The traditional self-attention mechanism uses reconstruction error to measure the distance between the quantization point and the query, but the reconstruction error is not sensitive to the inner product and is not the optimal measurement indicator. This embodiment uses score-aware quantization loss to measure the distance between the query and all quantization points, which can give a greater weight to the parallel component while penalizing the vertical component, thereby better quantizing the query from the perspective of a larger inner product value.
[0039] Among them, the score-aware quantization loss of any query and quantization point The calculation formula is as follows: , In the above formula, q i Indicates the first i query; Indicates the first j Quantitative points; express The horizontal component of express The vertical component of express q i and The vector difference between represents the weight coefficient of the vertical component; Represents the weight coefficient of the horizontal component.
[0040] Among them, the horizontal component and the vertical component and The calculation formula is as follows: , Weight coefficients for vertical and horizontal components and The calculation formula is as follows: , In the above formula, p Represents the dimension of the data; w represents a weight function for balancing the parallel quantization error and the vertical quantization error. By selecting this function, we can better focus on the larger inner product term, thereby giving a larger weight to the parallel quantization error and penalizing the vertical error component; is an angle parameter that satisfies: .
[0041] Through the above method, the loss distance between all queries and each quantization point can be calculated, and the quantization point with the smallest distance is selected as the quantization point of the current query. By quantizing all queries, several quantization partitions can be obtained. Cj , all queries in each quantization partition belong to the same quantization point. Under this condition, the update formula of the quantization vocabulary in the iteration process is as follows: , In the above formula, Indicates the updated quantization point in the quantization vocabulary; I represents an identity matrix.
[0042] S3: Use the updated quantization vocabulary to quantize all queries into corresponding quantization points ; Build a lookup table by calculating the inner product of the quantization point and the key; Use the quantization points in the lookup table that are related to each other Data pairs with key vector k Approximate representation of the required query vector q With key vector k Data pair , and then realize attention calculation.
[0043] Traditional Transformers need to perform inner product calculations on all query-key pairs when calculating the attention matrix. In the solution of this embodiment, a larger number of queries are quantized into a smaller number of quantization points in the quantization vocabulary, thereby achieving dimensionality reduction of the data. On this basis, this embodiment further constructs a lookup table of the inner product between each quantization point in the quantization vocabulary and the key, and uses the lookup table to calculate the inner product between the query and the key. and k The approximate inner product of replace q Vector Sum k Exact value of vector inner product , which further reduces the number of attention score calculations to reduce computational complexity.
[0044] In the S2 stage of this scheme, the score-aware quantization loss is calculated The purpose of is to measure the similarity between each query and each quantization point, and accordingly q i Compare the values of the score-aware quantization loss of all quantization points in the quantization vocabulary to find the closest quantization point, thereby achieving q i Assign to the corresponding quantization point Belong to the quantitative partition C j It can be seen that the above scheme does not require accurate calculation of each , but only need to determine the value of each The size relationship is sufficient.
[0045] Based on the above conclusions, in the further optimized solution of this embodiment, Under the premise that t Represents the query vector q and health vector k The inner product of T 0 Indicates the preset t The threshold value; I represents the unit matrix; it can be simplified by the following method The calculation formula is: Because there are: , So, in T 0 =0.2, we can conclude .
[0046] consider Represents the weight coefficient of the horizontal component Weight coefficient with vertical component The ratio of , so: .
[0047] On this basis, this embodiment defines anisotropic loss The calculation formula is: .
[0048] Therefore, in further optimization, anisotropic loss can be used in step S2. L Replace score-aware quantization loss ,Will L Used to measure the distance between any query and the quantized point. Since the anisotropic loss L does not need to calculate the weight coefficient of the horizontal component Weight coefficient with vertical component This can further reduce the computational complexity of attention.
[0049] Finally, in a further optimized solution of this embodiment, the Query-Key sharing mechanism can be used to make the query and key share weights during the attention calculation process in step S3, thereby further reducing unnecessary calculations while ensuring that the expressiveness of the model is not seriously affected.
[0050] The above content introduces the complete content of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided in this embodiment. Combined with the above content, it can be seen that this scheme introduces an attention dilution optimization method based on anisotropic loss and wavelet feature extraction in the self-attention mechanism of Transformer, and combines the Query-Key sharing mechanism to achieve more efficient sparseness of attention, which provides a new sparse strategy.
[0051] In practical applications, the new sparse strategy provided in this embodiment can be applied to all data processing tasks that use Transformer and its self-attention mechanism (such as NLP, CV and cross-modal tasks), thereby replacing the existing self-attention mechanism to implement attention calculation. This significantly reduces the computational overhead while ensuring that the model accuracy does not decrease significantly, improves the accuracy of attention calculation, and enhances the adaptability and robustness of the Transformer-based network model to long sequences and different downstream tasks.
[0052] Example 2 Based on the attention dilution optimization method combining wavelet feature analysis and anisotropic loss and its application provided in Example 1, this embodiment further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as mentioned above, and then completes the attention calculation of the input features.
[0053] This embodiment also provides a storage medium in which a computer program is stored. When the computer program is executed by a processor, the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as mentioned above is implemented to complete the attention calculation of the input features.
[0054] This embodiment also provides a natural language processing / computer vision processing / multimodal data processing device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, a Transformer-based data processing model is used to process the input data; wherein the self-attention module in the data processing model uses the aforementioned attention dilution optimization method combining wavelet feature analysis and anisotropic loss to realize attention calculation.
[0055] In this embodiment, the natural language processing / computer vision processing / multimodal data processing device is essentially a computer device. In actual application, the computer device can adopt an embedded device and be deployed in various terminal devices to support data processing and interaction. It is also used as an independent computer device to support data processing requirements in certain scenarios. This non-embedded computer device can adopt a medium or large computer device such as a laptop, a tablet computer, a desktop computer, or a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers) that can execute a computer program.
[0056] Specifically, the computer device of this embodiment at least includes but is not limited to: a memory and a processor that can be connected to each other through a system bus. In this embodiment, the memory (i.e., a readable storage medium) includes a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory may be an internal storage unit of a computer device, such as a hard disk or a memory of the computer device. In other embodiments, the memory may also be an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Of course, the memory may also include both an internal storage unit of a computer device and an external storage device thereof. In this embodiment, the memory is generally used to store an operating system and various application software installed on the computer device. In addition, the memory may also be used to temporarily store various types of data that have been output or are to be output.
[0057] The processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor is generally used to control the overall operation of a computer device.
[0058] Verification experiment In order to verify the effectiveness of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided by the present invention, this experiment uses the BERT-small model as the test object to compare the performance of the model before and after adopting the solution of the present invention.
[0059] 1. Experimental Content This experiment set up a control group and an experimental group (referred to as the present invention). Among them, the control group selected the solution provided in the prior art that did not adopt the attention sparse strategy. In the training phase, the control group scheme used the strategy of pre-training first and then distillation fine-tuning to adjust the BERT-small model, which means that the effect of the model will be better than the strategy of simple pre-training and fine-tuning.
[0060] The experimental group selected the BERT-small model that adopted the attention dilution optimization strategy provided by the present invention. Among them, the model of the experimental group was trained on a 4*NVIDIA RTX 4090 platform, the hidden_size of the model was 512 (here it can be directly changed to the input token length of the model is 512, and each token is represented by a 512-dimensional vector), and the number of layers was 4. First, in the pre-training stage, the experimental group used the Wikipedia corpus to pre-train the BERT-small model in combination with the attention sparse strategy provided by the present invention, and compared it with the control group scheme.
[0061] The experimental group scheme pre-initializes eight sets of quantized vocabulary lists. The entire training process totals 100,000 steps, using a two-stage training protocol. The initial 10,000 steps are the warm-up stage. Each iteration simultaneously updates the model gradient and quantized vocabulary list. This method enables the vocabulary list to converge quickly to a good state. In the subsequent 90,000 steps, in order to reduce the training overhead caused by frequent codebook updates, the vocabulary list is updated using a gradient accumulation strategy, that is, it is updated once every k steps, while other model parameters are updated in real time to maintain the responsiveness of the training. This training can roughly achieve a balance between computational efficiency and optimization stability. The entire training process uses the default hyperparameters. The Adam optimizer ensures stable and efficient optimization of model parameters and quantization vocabulary.
[0062] After pre-training, the model of the experimental group solution was fine-tuned using the General Language Understanding Evaluation (GLUE) benchmark on datasets such as SST-2, QQP, and MNLI. During fine-tuning, the weights and quantized vocabulary obtained from pre-training were used as the basis, and the model was trained for 4 epochs using hyperparameters of a batch size of 32 and a learning rate of 1e-4. Through such a pre-training and fine-tuning process, while effectively reducing the computational complexity of the model, the model performance is maintained to the greatest extent, improving its adaptability and accuracy in different natural language processing tasks, and providing strong support for model applications in resource-constrained scenarios.
[0063] 2. Experimental Results and Analysis 2.1 Performance Loss When the vocabulary size is 128, this experiment first uses the GLUE test benchmark to evaluate the GLUE test scores of the experimental group and the control group on the SST-2, QQP, QNLI and MNLI datasets. The results are shown in Table 1: Table 1: GLUE scores of different schemes on various datasets
[0064] From the experimental data in the analysis table, it can be found that the performance difference between the two schemes is roughly between 1% and 3%. This shows that compared with the scheme without attention dilution, the attention dilution optimization measurement provided by the present invention does not cause serious loss of performance such as model accuracy, and can therefore replace the existing self-attention mechanism.
[0065] 2.2 Model speed and number of parameters Based on the above experiments, the attention calculation efficiency of the two schemes was further compared and evaluated. It was found that the attention calculation effect of the model using the attention dilution optimization of the present invention can be improved by 3.5 times compared with the existing model. At the same time, the overall model speed and model parameter quantity of the two schemes using different attention mechanisms were compared, and the results are shown in Table 2.
[0066] Table 2: Comparison of model speed and parameter size in the two solutions
[0067] By analyzing the data in the above table, we can find that the solution of the present invention not only achieves more efficient attention calculation performance, but also can achieve a 1.9-fold improvement in overall model speed, while the number of model parameters can be reduced by 8.7%. This shows that the present invention has achieved the expected attention sparse effect, greatly reduced the amount of data calculation for attention calculation, and helps to alleviate the memory bottleneck and computing power overhead problems of existing large models when processing large-scale data.
[0068] 2.3 Model Performance with Different Numbers of Quantized Spatial Vocabularies In order to verify the performance of the dynamic vocabulary selection based on wavelet feature analysis designed in the solution provided by the present invention, this experiment tested the model performance under different configurations of 1, 2, 4 and 8 initial quantized vocabularies on the SST-2 dataset. The curve of the model performance changing with the number of quantized vocabularies is shown in Figure 2. Figure 3 The benchmark scores for these four configurations are 76.3, 85.2, 87.5, and 88.2 respectively.
[0069] Combination Figure 3It can be seen from the data in that: when the number of the initial quantization vocabulary set increases, it helps to improve the final performance of the model. This is because the present invention adopts dynamic selection of quantization vocabulary based on wavelet feature analysis to expand the dynamic range of the vocabulary, thereby improving the robustness of the quantization model.
[0070] The above-described embodiment only expresses one implementation mode of the present invention, and its description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the attached claims.
Claims
1. An attention dilution optimization method combining wavelet feature analysis and anisotropic loss, characterized in that: It includes: Preset in order from low to high accuracy level n Initialize the quantization vocabulary; The input features of each layer X emb Perform separation wavelet transform to obtain the low-frequency approximation coefficient A and high-frequency detail coefficient D of the input feature; calculate the energy distribution corresponding to A and D E low and E high ; Calculated by the following formula S The value of S Select the initial quantization vocabulary corresponding to the precision level: , in, Indicates a round-down operation; Calculate the score-aware quantization loss between each query and each quantization point in the initial quantization vocabulary and use it as the distance between the two; assign each query to the quantization point closest to it, and then obtain an updated quantization vocabulary; the updated quantization vocabulary contains multiple independent quantization partitions C j , each quantization partition consists of a quantization point and its corresponding at least one query; Use the updated quantization vocabulary to quantize all queries to corresponding quantization points ; Build a lookup table by calculating the inner product of the quantization point and the key; Use the quantization points in the lookup table that are related to each other Data pairs with key vector k Approximate representation of the required query vector q With key vector k Data pair , and then realize attention calculation.
2. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 1, characterized in that: In the attention calculation process, the Query-Key sharing mechanism is used to make the query and key share weights.
3. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 1, characterized in that: Score-aware quantization loss for any query and quantization point The calculation formula is as follows: , In the above formula, q i Indicates i query; Indicates the first j Quantitative points; express The horizontal component of express The vertical component of represents the weight coefficient of the vertical component; Represents the weight coefficient of the horizontal component.
4. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 3, characterized in that: Weight coefficients for vertical and horizontal components and The calculation formula is: ; In the above formula, p Represents the dimension of the data; w represents the weight function used to balance the parallel quantization error and the vertical quantization error, is an angle parameter that satisfies: .
5. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 4, characterized in that: Using anisotropic loss L Instead of score-aware quantization loss, it is used to measure the distance between any query and the quantization point; anisotropic loss The calculation formula is: 。 6. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 5, characterized in that: The updated quantization vocabulary contains multiple quantization partitions, each of which C j All queries in the same quantization point belong to the same quantization point; the quantization point in the quantization vocabulary The iteration formula is as follows: , In the above formula, I represents an identity matrix.
7. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 1, characterized in that: Input Features X emb The low-frequency approximation coefficient A and high-frequency detail coefficient D and their energy distribution E low and E high The calculation formula is as follows: , In the above formula, Represents a separable wavelet transform operation.
8. An application of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described in any one of claims 1 to 7 to replace the self-attention module in a Transformer-based neural network to realize attention calculation.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described in any one of claims 1 to 8 is implemented, thereby completing the attention calculation of the input features.
10. A natural language processing device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, a Transformer-based data processing model is used to process the input data, wherein the self-attention module in the data processing model uses the attention dilution optimization method combined with wavelet feature analysis and anisotropic loss as described in any one of claims 1 to 8 to realize attention calculation.
Citation Information
Patent Citations
Image restoration method based on wavelet transform attention model
CN111047541A
Quantization-based long text question and answer reasoning method and device and medium
CN114817500A
Transform network video restoration method based on sparse self-attention of most correlation area
CN118469872A
Emotion classification algorithm combining double attention mechanism and Bi-LSTM
CN119538005A
Systems and methods for retrieving patient information using large language models
US12254005B1