Attention Dilution Optimization Method Combining Wavelet Feature Analysis and Anisotropic Loss

By combining wavelet feature analysis and anisotropic loss attention dilution optimization method, the Transformer attention mechanism has solved the problem of high computational complexity and poor robustness in long-sequence data processing, and efficient sparse calculation and robustness enhancement are achieved.

CN119940449BActive Publication Date: 2025-07-04ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510416785.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-04
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing Transformer attention mechanism has high computational complexity when processing long sequence data, resulting in high computing resources consumption and poor robustness, making it difficult to adapt to different types of downstream tasks.

Method used

Combining wavelet feature analysis and attention dilution optimization method of anisotropic loss, through the quantized vocabulary list and the Query-Key sharing mechanism, the calculation complexity is reduced and robustness is improved, including wavelet transformation, quantized vocabulary list update and the application of anisotropic loss.

Benefits of technology

It significantly reduces the computational complexity without losing model accuracy, improves the computational efficiency, and enhances the model's adaptability in long sequences and different downstream tasks, and is suitable for resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940449B_ABST
    Figure CN119940449B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of large models, and particularly to an attention dilution optimization method combining wavelet feature analysis and anisotropic loss, and its corresponding computer program product and natural language processing device. The solution first presets multiple initial quantization vocabularies according to accuracy levels; and flexibly selects an initial quantization vocabulary in combination with the wavelet feature analysis result of the input features. Then, the score-aware quantization loss or anisotropic loss is used to measure the distance between each query and the quantization points horizontally, so as to partition and update the quantization vocabulary. Finally, all queries are quantized into corresponding quantization points by using the updated quantization vocabulary; a lookup table is constructed by calculating the inner product of the quantization points and the keys; and the approximate representation result in the lookup table is used to implement attention calculation. The present invention solves the problems of low efficiency and poor robustness existing in the attention mechanism of the existing Transformer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large models, and in particular to an attention dilution optimization method combining wavelet feature analysis and anisotropic loss, and its corresponding computer program product and natural language processing device. Background Art

[0002] In recent years, large model technology has risen rapidly. With its powerful capabilities and generalization ability, it has set off a wave of changes in many fields, which is particularly remarkable in the fields of natural language processing (NLP) and computer vision (CV). Existing large models are mainly implemented based on Transformer. Since its release, Transformer has attracted the attention of the vast majority of researchers in the field of artificial intelligence. After several years of development, models based on this architecture have replaced the traditional recurrent neural network (RNN) and convolutional neural network (CNN) in the dominant position in many tasks. This is due to its high parallelism and its unique self-attention mechanism. Figure 1 Figure (a) in [reference] shows the calculation process of attention. By calculating the dot product of the query vector and the key vector, performing scale scaling and applying the softmax function, the probability distribution corresponding to each key is obtained. Finally, the value vector is weighted and summed using this distribution to output the value representing the attention weight. It can effectively capture the dependencies between any positions in the input sequence and model global information. Figure 1 Figure (b) in [reference] shows the multi-head attention layer. This structure concatenates the weight representation vectors obtained by passing the vector through several attention layers. Different weights can be trained to extract information with different focuses. At the same time, it does not require input sequence data step by step in time, so it can greatly improve the training speed and the efficiency of parallel computing of the model. Thanks to its unique advantages, Transformer has achieved excellent results in many tasks since its inception. However, although the attention mechanism has given it a leading position in many fields, with the increase in the sequence length, the computational cost of calculating the attention score also faces an increasingly severe challenge of computational complexity: if the lengths of two input sequences are both n, then the computational complexity of calculating the attention score is O(n 2) Therefore, when the input sequence increases by an order of magnitude, the computational cost of the self-attention mechanism will increase exponentially, which not only consumes a large amount of computing power resources but also brings additional memory bottleneck challenges. These problems are particularly evident when processing large-scale data such as long documents, multi-modal data, and image sequences, and have become an important obstacle restricting the progress of large model technology. However, the attention score matrix usually exhibits sparse characteristics, where the inner products between most queries and keys in the matrix are small. After softmax processing, the proportion of these vectors in the final weight representation is negligible. Therefore, reasonably utilizing the sparsity of the attention score matrix can not only reduce unnecessary computational costs but also decrease the proportion of weights that need to be saved, thus effectively alleviating the memory bottleneck and computational power overhead problems.

[0003] One of the mainstream solutions for Transformer attention sparsity is to apply specific sparse patterns to reduce the computational amount of the attention score matrix while ensuring excellent model performance. The research on sparse patterns mainly focuses on four mainstream strategies:

[0004] 1. Local Sliding Window Attention

[0005] This strategy is based on the locality principle in NLP, that is, the correlation between a token and its adjacent tokens is usually higher than that with distant tokens. Researchers effectively reduce the computational complexity by restricting each query to only focus on the keys within its fixed window.

[0006] 2. Global Sparse Attention

[0007] This strategy fixedly selects several tokens to enable them to attend to information at all positions, thereby reducing the computational amount while enhancing the global information capture ability. A typical model applying this strategy is ETC, which combines local sliding window and global sparsity to enable the model to simultaneously have the ability to learn long-range dependencies and express adjacent contexts.

[0008] 3. Random Sparse Attention

[0009] To introduce the representation of tokens at relatively long distances, the random sparse strategy selects attention pairs for calculation through random sampling. Although random sparsity can take tokens at relatively long distances into consideration while reducing the computational cost, for example, the Big Bird model (Manzil Zaheer et al., 2020) combines the sliding window, global token, and random sparse modes, showing excellent performance in long text processing tasks, it also introduces randomness, resulting in reduced stability and robustness of the model.

[0010] 4. Dilated Sparse Attention

[0011] To expand the receptive field of the sliding window, researchers introduced the idea of dilated convolution, introducing a dilation factor in the local window, enabling the model to focus on tokens at farther positions and expanding the receptive field of the local sliding window. Longformer is a typical example, which can even maintain good performance when the context length is 4096.

[0012] The above methods are widely used in fields such as natural language understanding and time series analysis. However, these four sparse modes usually only apply to a certain type of data and have poor robustness in other downstream tasks; the existing Transformer attention mechanism still has problems such as losing global dependency information and introducing additional errors. Summary of the Invention

[0013] To solve the problems of low efficiency and poor robustness existing in the existing Transformer attention mechanism, the present invention provides an attention dilution optimization method combining wavelet feature analysis and anisotropic loss, as well as its corresponding computer program product and natural language processing device.

[0014] The technical solution provided by the present invention is as follows:

[0015] An attention dilution optimization method combining wavelet feature analysis and anisotropic loss, which includes the following steps:

[0016] S1: Preset n initial quantization vocabularies in ascending order of precision. Perform discrete wavelet transform on the input features of each layer X emb to obtain the low-frequency approximation coefficient A and high-frequency detail coefficient D of the input features. Calculate the corresponding energy distributions of A and D E low and E high ; then calculate the value of S according to the following formula and based onS Select the initialization quantization word table corresponding to the serial number:

[0017] ,

[0018] Among them, represents the floor operation.

[0019] S2: Calculate the score-aware quantization loss between each query and each quantization point in the initialization quantization word table, and use it as the distance between the two. Assign each query to the quantization point with the closest distance, and then obtain the updated quantization word table. The updated quantization word table contains multiple independent quantization partitions C j , and each quantization partition is composed of a quantization point and at least one query corresponding to it.

[0020] S3: Quantize all queries into corresponding quantization points using the updated quantization word table ; construct a lookup table by calculating the inner product of the quantization points and the keys; use the correlated quantization points in the lookup table and the data pair of the key vector k to approximately represent the required query vector q and the key vector k data pair , and then implement the attention calculation.

[0021] As a further improvement of the present invention, in the attention calculation process of S3, a Query-Key sharing mechanism is adopted to make the queries and keys share weights.

[0022] As a further improvement of the present invention, the score-aware quantization loss between any query and the quantization point is calculated as follows:

[0023] ,

[0024] In the above formula, q i represents the i th query; represents the j th quantization point in the quantization word table; represents the horizontal component of; represents the vertical component of; represents the weight coefficient of the vertical component; represents the weight coefficient of the horizontal component.

[0025] As a further improvement of the present invention, the weight coefficients of the vertical component and the horizontal component and The calculation formula for is as follows:

[0026] ,

[0027] In the above formula, p represents the dimension of the data; w represents a weight function used to balance the parallel quantization error and the vertical quantization error. By selecting this function, one can better focus on larger inner product terms, and thus give a greater weight to the parallel quantization error and penalize the vertical error component; is an angle parameter, which satisfies: .

[0028] As a further improvement of the present invention, the anisotropic loss L is used instead of the score-aware quantization loss to measure the distance between any query and the quantization point; the calculation formula for the anisotropic loss is as follows:

[0029] .

[0030] As a further improvement of the present invention, the updated quantization vocabulary contains multiple quantization partitions, and all queries within each quantization partition C j belong to the same quantization point; the iterative formula for the quantization point in the quantization vocabulary is as follows:

[0031] ,

[0032] In the above formula, I represents an identity matrix.

[0033] As a further improvement of the present invention, the low-frequency approximation coefficient A and the high-frequency detail coefficient D of the input feature X emb and their energy distributions E low and E high are calculated as follows:

[0034] ,

[0035] In the above formula, represents the discrete wavelet transform operation.

[0036] The present invention also includes an application of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described above to replace the self-attention module in a Transformer-based neural network to implement attention calculation.

[0037] The present invention also includes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described above, and further completes the attention calculation for the input features.

[0038] The present invention also includes a natural language processing device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it processes the input data using a data processing model based on Transformer; wherein, the self-attention module in the data processing model implements attention calculation using the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described above.

[0039] The present invention has the following beneficial effects:

[0040] The present invention applies technologies such as anisotropic quantization loss, wavelet feature extraction, and Q-K weight sharing to the dilution optimization of the self-attention mechanism of Transformer. These new attention sparsity strategies can reduce the computational complexity and improve the computational efficiency with almost no loss of model accuracy. Among them, in the inference stage, the present invention uses a lookup table to quickly obtain approximate attention scores, avoiding redundant calculations and improving the inference speed.

[0041] The sparse attention calculation method of the present invention can reduce the original square-level computational complexity, enabling Transformer to operate efficiently on a larger context length and being applicable to resource-constrained devices. The solution of the present invention sets up quantization vocabularies with multiple different precision levels, and flexibly selects the best quantization vocabulary in combination with wavelet feature analysis technology, which enables the solution of the present invention to be applicable to processing different types of data and various different downstream tasks, significantly enhancing the robustness of the solution. Description of the Drawings

[0042] Figure 1 It is a schematic diagram of the existing self-attention mechanism adopted by Transformer introduced in the background technology. In the figure, part (a) is the schematic diagram of a simple dot-product attention mechanism, and part (b) is the schematic diagram of a multi-head attention mechanism.

[0043] Figure 2 It is a schematic diagram of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided in Embodiment 1 of the present invention.

[0044] Figure 3 It is a curve showing the change of the performance of the model adopting the attention dilution scheme provided by the present invention with the number of quantization vocabularies in the test experiment. Detailed Embodiments

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.

[0047] Embodiment 1

[0048] This embodiment provides an attention dilution optimization method combining wavelet feature analysis and anisotropic loss. This solution applies technologies such as anisotropic quantization loss, wavelet feature extraction, and Q-K weight sharing to the self-attention mechanism in Transformer, and then designs a more efficient attention sparse optimization method. This method can greatly reduce the computational overhead of Transformer in long sequence tasks with almost no loss of model accuracy. Therefore, this solution can be easily combined with new technologies such as distillation and further improve the performance of the model, providing new possibilities for the application of large models in resource-constrained environments, thereby promoting the further development of the Transformer architecture in NLP, CV, and cross-modal tasks.

[0049] Specifically, as Figure 2 shown, the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided in this embodiment includes the following steps:

[0050] S1: Generate multiple sparsified quantization vocabularies, and analyze the input features using wavelet feature analysis technology, and then adaptively select the most suitable quantization vocabulary in the space according to the input frequency distribution to quantize the query. The process includes:

[0051] S11: First, preset multiple (assumed to be n ) initialization quantization vocabularies in ascending order of accuracy.

[0052] In this embodiment, quantizing the query into quantization points in the quantization vocabulary has the advantage of reducing the number of calculations. However, the features input to each layer of the network are only a small part of the data, and a single vocabulary is difficult to comprehensively represent all queries. To address this issue, in this embodiment, a raw quantization vocabulary with the widest range and the most quantization points is first generated in combination with the distribution of the input features, and then, in the order of increasing precision, n initial quantization vocabularies are generated based on the raw quantization vocabulary. For example, assuming the number of initial quantization vocabularies is 8, the first initial quantization vocabulary is a vocabulary with coarser feature changes and lower quantization precision, while the eighth initial quantization vocabulary is a vocabulary with finer feature changes and higher quantization precision.

[0053] S12: Select the most suitable initial quantization vocabulary according to the frequency distribution of the currently input features.

[0054] In practical applications, in this embodiment, the input features of each layer are first X emb subjected to discrete wavelet transform to obtain the low-frequency approximation coefficients A and high-frequency detail coefficients D of the input features. The process is expressed as:

[0055] ,

[0056] where represents the discrete wavelet transform operation.

[0057] Next, the energy distributions A corresponding to the low-frequency approximation coefficients D and the high-frequency detail coefficients E low and E high are calculated respectively through the following formula: where , .

[0058] Finally, based on the known E low and E high a S value is calculated through the following formula, and the initial quantization vocabulary corresponding to the serial number is selected according to the S value:

[0059] ,

[0060] where represents the floor operation.

[0061] Combined with the formula, it can be seen that in this embodiment, when the high-frequency detail coefficients of the input featuresD Energy distribution E high The higher it is, the higher-precision quantization table will be selected to quantize the query; conversely, the lower-precision quantization table will be selected to quantize the query; and this can significantly improve the robustness of the selected quantization table under different input samples.

[0062] S2: Iteratively update the quantization table to obtain each quantization point and all the corresponding queries to form a quantization partition C j .

[0063] Specifically, at this stage, this embodiment first calculates the score-aware quantization loss between each query and each quantization point in the initialized quantization table and uses it as the distance between the two. Then each query is assigned to the quantization point with the closest distance, and then the updated quantization table is obtained. The updated quantization table contains multiple independent quantization partitions C j , and each quantization partition is composed of a quantization point and at least one corresponding query.

[0064] The traditional self-attention mechanism uses the reconstruction error to measure the distance between the quantization point and the query, but the reconstruction error is insensitive to the inner product and is not an optimal measurement index. This embodiment uses the score-aware quantization loss to measure the distance between the query and all quantization points, which can give a greater weight to the parallel component while punishing the vertical component, so as to better quantize the query from the perspective of a larger inner product value.

[0065] Among them, the score-aware quantization loss between any query and the quantization point has the following calculation formula:

[0066] ,

[0067] In the above formula, q i represents the i th query to be quantized; represents the j th quantization point in the quantization table; represents 's horizontal component; represents 's vertical component; represents q i and 's vector difference; represents the weight coefficient of the vertical component; represents the weight coefficient of the horizontal component.

[0068] Among them, the horizontal component and the vertical component and The calculation formulas are as follows:

[0069] ,

[0070] The weight coefficients of the vertical component and the horizontal component and The calculation formulas are as follows:

[0071] ,

[0072] In the above formula, p represents the dimension of the data; w represents a weight function used to balance the parallel quantization error and the vertical quantization error. By selecting this function, it is possible to pay more attention to larger inner product terms, and thus give a greater weight to the parallel quantization error and penalize the vertical error component; is an angle parameter, which satisfies: .

[0073] Through the above method, the loss distance between all queries and each quantization point can be calculated. Select the quantization point with the smallest distance as the quantization point of the current query. By quantizing all queries, several quantization partitions C j can be obtained. All queries within each quantization partition belong to the same quantization point. Under this condition, the update formula of the quantization vocabulary during the iteration process is as follows:

[0074] ,

[0075] In the above formula, represents the updated quantization point in the quantization vocabulary; I represents an identity matrix.

[0076] S3: Quantize all queries into corresponding quantization points using the updated quantization vocabulary ; Construct a lookup table by calculating the inner product of the quantization point and the key; Use the correlated quantization points in the lookup table and the data pair of the key vector k to approximately represent the required query vector q and the key vector k data pair , thereby realizing the attention calculation.

[0077] Traditional Transformers need to calculate the inner product of all query-key pairs when computing the attention matrix. In the solution of this embodiment, a larger number of queries are quantized into a smaller number of quantization points in the quantization vocabulary, thereby achieving dimensionality reduction of the data. On this basis, this embodiment further constructs a lookup table for the inner product between each quantization point in the quantization vocabulary and the keys, and uses the approximate inner product of and k in the lookup table to replace q the exact value of the inner product of the vector sum k and the vector , thereby further reducing the number of attention score calculations to reduce the computational complexity.

[0078] In the S2 stage of this solution, calculating the score-aware quantization loss is to measure the similarity between each query and each quantization point, and based on this, compare the values of the score-aware quantization loss between the specified q i and all quantization points in the quantization vocabulary to find the closest quantization point, thereby achieving the assignment of q i to the quantization partition to which the corresponding quantization point C j belongs. It can be seen that the above solution does not need to accurately calculate each value, but only needs to determine the magnitude relationship of each .

[0079] Based on the above conclusion, in the further optimized solution of this embodiment, on the premise of , where t represents the inner product of the query vector q and the key vector k ; T 0 represents the preset t threshold; I represents the identity matrix; then the calculation formula of can be simplified by the following method:

[0080] Because there is:

[0081] ,

[0082] So, when T 0 = 0.2, it can be obtained that .

[0083] Considering represents the weight coefficient of the horizontal component The ratio with the weight coefficient of the vertical component Thus, there is

[0084] .

[0085] On this basis, this embodiment defines the calculation formula of the anisotropic loss as

[0086] .

[0087] Therefore, in the further optimization, in step S2, the anisotropic loss L can be used to replace the score-aware quantization loss , and L is used to measure the distance between any query and the quantization point. Since the anisotropic loss L does not need to calculate the exact values of the weight coefficient of the horizontal component and the weight coefficient of the vertical component, this can further reduce the computational complexity of the attention

[0088] Finally, in the further optimized solution of this embodiment, in the attention calculation process of step S3, a Query-Key sharing mechanism can be adopted to make the query and the key share weights. Further, on the premise of ensuring that the expression ability of the model is not severely affected, unnecessary calculations can be reduced

[0089] The above content introduces the complete content of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided by this embodiment. From the above content, it can be seen that this solution introduces an attention dilution optimization method based on anisotropic loss and wavelet feature extraction into the self-attention mechanism of the Transformer, and combines the Query-Key sharing mechanism to achieve more efficient sparsity of the attention, which provides a new sparsity strategy

[0090] In practical applications, the new sparsity strategy provided by this embodiment can be applied to all data processing tasks that adopt the Transformer and its self-attention mechanism (such as NLP, CV, and cross-modal tasks), and then replace the existing self-attention mechanism to implement attention calculation. Thus, on the premise of ensuring that the model accuracy does not decrease significantly, the computational overhead can be significantly reduced, the accuracy of attention calculation can be improved, and the adaptability and robustness of the Transformer-based network model to long sequences and different downstream tasks can be enhanced

[0091] Embodiment 2

[0092] Based on the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided in Embodiment 1 and its applications, this embodiment further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described above, and further completes the attention calculation for the input features.

[0093] This embodiment also provides a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described above, and further completes the attention calculation for the input features.

[0094] This embodiment also provides a natural language processing / computer vision processing / multimodal data processing device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it processes the input data using a data processing model based on Transformer; among them, the self-attention module in the data processing model implements attention calculation using the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described above.

[0095] In this embodiment, the natural language processing / computer vision processing / multimodal data processing device is essentially a computer device. In actual application processes, this computer device can be an embedded device and be deployed in various terminal devices to support data processing and interaction. It can also be applied as an independent computer device to support data processing requirements in certain scenarios. Such a non-embedded computer device can be a laptop computer, a tablet computer, a desktop computer, or a medium and large-sized computer device such as a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers) that can execute computer programs.

[0096] Specifically, the computer device of this embodiment includes at least, but is not limited to, a memory and a processor that can be communicatively connected to each other through a system bus. In this embodiment, the memory (i.e., the readable storage medium) includes flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory may also be used to temporarily store various data that have been output or will be output.

[0097] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device.

[0098] Verification experiment

[0099] To verify the effectiveness of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss provided by the present invention, this experiment uses the BERT-small model as the test object to compare the performance of the model before and after adopting the solution of the present invention.

[0100] I. Experimental content

[0101] This experiment sets up a control group and an experimental group (denoted as the present invention). Among them, the control group selects the solution that does not adopt the attention sparsity strategy provided in the prior art. In the training stage, the control group solution adjusts the BERT-small model by using the strategy of pre-training first and then distillation fine-tuning, which means that the effect of the model will be better than that of the strategy of simple pre-training plus fine-tuning.

[0102] The experimental group selected the BERT-small model that adopted the attention dilution optimization strategy provided by the present invention. Among them, the model of the experimental group was trained on a 4*NVIDIA RTX 4090 platform. The hidden_size of the model was 512 (here it can be directly changed to the input token length of the model being 512, and each token was represented by a 512-dimensional vector), and the number of layers was 4. First, in the pre-training stage, the experimental group selected the Wikipedia corpus to pre-train the BERT-small model in combination with the attention sparsity strategy provided by the present invention, and at the same time compared it with the control group's solution.

[0103] The experimental group's solution pre-initialized eight groups of quantization vocabularies. The entire training process totaled 100,000 steps and adopted a two-stage training protocol. The initial 10,000 steps were the warm-up stage. In each iteration, the model gradients and quantization vocabularies were updated simultaneously. This method could make the vocabulary converge to a good state quickly. In the subsequent 90,000 steps, to reduce the training overhead caused by frequent update of the codebook, the gradient accumulation strategy was adopted to update the vocabulary, that is, it was updated every k steps, while other model parameters were updated in real time to maintain the responsiveness of the training. This training could roughly achieve a balance between computational efficiency and optimization stability. The entire training process used the default hyperparameters of the Adam optimizer to ensure the stable and efficient optimization of model parameters and quantization vocabularies.

[0104] After completing the pre-training, the model of the experimental group's solution was fine-tuned using the General Language Understanding Evaluation (GLUE) benchmark on datasets such as SST-2, QQP, and MNLI. During fine-tuning, based on the weights and quantization vocabularies obtained from the pre-training, with hyperparameters of a batch size of 32 and a learning rate of 1e-4, the model was trained for 4 epochs. Through such pre-training and fine-tuning processes, while effectively reducing the computational complexity of the model, the model performance was maintained to the greatest extent, improving its adaptability and accuracy in different natural language processing tasks, providing strong support for the application of the model in resource-constrained scenarios.

[0105] II. Experimental Results and Analysis

[0106] 2.1. Performance Loss

[0107] In the case where the vocabulary size was 128, in this experiment, the GLUE test scores obtained by the experimental group and the control group's solutions were first evaluated using the GLUE test benchmark on datasets such as SST-2, QQP, QNLI, and MNLI. The results obtained are shown in Table 1:

[0108] Table 1: GLUE Scores of Different Solutions on Each Dataset

[0109]

[0110] By analyzing the experimental data in the table, it can be found that the performance difference between the two schemes is roughly between 1% and 3%. This indicates that compared with the scheme without attention dilution, the attention dilution optimization measurement provided by the present invention does not cause serious losses to the performance such as the model accuracy, so it can replace the existing self-attention mechanism for use.

[0111] 2.2. Model Speed and Number of Parameters

[0112] Based on the above-mentioned experiment, further evaluation of the attention calculation efficiency of the two schemes was carried out and it was found that: compared with the existing model, the model using the attention dilution optimization of the present invention can achieve a 3.5-fold improvement in the attention calculation effect. At the same time, comparing the overall model speed and the number of model parameters of the two schemes using different attention mechanisms, the results are shown in Table 2.

[0113] Table 2: Comparison of Model Speed (speed) and Number of Parameters (Parameter) in Two Schemes

[0114]

[0115] By analyzing the data in the above table, it can be found that: the scheme of the present invention not only achieves more efficient attention calculation performance, but also can achieve a 1.9-fold improvement in terms of the overall model speed. At the same time, the number of model parameters can be reduced by 8.7%. This shows that the present invention has achieved the expected attention sparsity effect, realized a significant reduction in the data operation volume of attention calculation, and helps to alleviate the memory bottleneck and computing power overhead problems existing in the existing large models when processing large-scale data.

[0116] 2.3. Model Performance with Different Numbers of Quantized Vocabulary Tables

[0117] To verify the performance of the dynamic vocabulary table selection based on wavelet feature analysis designed in the scheme provided by the present invention, this experiment tested the model performance of the model under different configurations with 1, 2, 4, and 8 initial quantized vocabulary tables on the SST-2 dataset. Among them, the curve of the model performance changing with the number of quantized vocabulary tables is as Figure 3 shown. The benchmark scores under these 4 configurations are 76.3, 85.2, 87.5, and 88.2 respectively.

[0118] Combined with Figure 3 the data in it, it can be known that: when the number of set initial quantized vocabulary tables increases, it helps to improve the final performance of the model. This is because the dynamic selection of the vocabulary table based on wavelet feature analysis adopted by the present invention can expand the dynamic range of the vocabulary table, thereby improving the robustness of the quantization model.

[0119] The above-described embodiments merely represent one implementation mode of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.

Claims

1. An attention dilution optimization method combining wavelet feature analysis and anisotropic loss, characterized in that It includes: Preset in the order from low to high precision level n initial quantization tables; For the input features of each layer X emb Perform a discrete wavelet transform to obtain the low-frequency approximation coefficient A and the high-frequency detail coefficient D of the input features; calculate the corresponding energy distributions of A and D E low and E high ; calculate the value of S through the following formula and select the initialization quantization vocabulary of the corresponding precision level according to S : , Among them, represents the floor operation; Calculate the score-aware quantization loss between each query and each quantization point in the initialized quantization vocabulary and use it as the distance between the two; assign each query to the quantization point with the closest distance, and then obtain the updated quantization vocabulary; the updated quantization vocabulary contains multiple independent quantization partitions C j , and each quantization partition consists of a quantization point and at least one corresponding query; Quantize all queries into corresponding quantization points using the updated quantization table ; construct a lookup table by calculating the inner product of the quantization points and the keys; use the correlated quantization points in the lookup table with the key vector k to approximately represent the required query vector q with the key vector k data pair , thereby realizing attention calculation.

2. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 1, wherein: During the attention calculation process, a Query-Key sharing mechanism is adopted to make the query and key share weights.

3. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 1, characterized in that Score-Aware Quantization Loss of Any Query and Quantization Points The calculation formula is as follows: , In the above formula, q i represents the i th query; represents the j th quantization point in the quantization table; represents the horizontal component of; represents the vertical component of; represents the weight coefficient of the vertical component; represents the weight coefficient of the horizontal component.

4. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 3, characterized in that: Weight coefficients of the vertical component and the horizontal component and The calculation formula is as follows: , In the above formula, p represents the dimension of the data; w represents a weight function used to balance the parallel quantization error and the vertical quantization error, is an angle parameter, which satisfies: .

5. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 4, wherein Adopt anisotropic loss L instead of the score-aware quantization loss to measure the distance between any query and the quantization points; the anisotropic loss has the following calculation formula: 。 6. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 5, characterized in that The updated quantization table contains multiple quantization partitions, and all queries within each quantization partition C j belong to the same quantization point; the iterative formula for the quantization point in the quantization table is as follows: , In the above formula, I represents an identity matrix.

7. The attention dilution optimization method combining wavelet feature analysis and anisotropic loss according to claim 1, characterized in that Input features X emb The low-frequency approximation coefficients A and high-frequency detail coefficients D thereof and their energy distributions E low and E high The calculation formulas thereof are as follows: , In the above formula, represents the discrete wavelet transform operation.

8. An application of the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described in any one of claims 1-7 in replacing the self-attention module in a Transformer-based neural network to implement attention calculation.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described in any one of claims 1-8, and further completes the attention calculation for the input features.

10. A natural language processing device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that: When the processor executes the computer program, it processes the input data using a data processing model based on Transformer, wherein the self-attention module in the data processing model implements attention calculation using the attention dilution optimization method combining wavelet feature analysis and anisotropic loss as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Image restoration method based on wavelet transform attention model

    CN111047541A

  • Transform network video restoration method based on sparse self-attention of most correlation area

    CN118469872A