Hierarchical auxiliary sparse attention method for long text large language model
By introducing a hierarchically assisted sparse attention method in the long text large language model, the problem of the square-level growth of the first word element response time in long text processing is solved, and the inference efficiency is improved and user experience is improved.
Patent Information
- Application Number
- CN202510003045.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
When long text large language model processes long text, the first word element response time increases squarely, resulting in a decline in user experience, and the existing technology lacks effective solutions.
The hierarchically assisted sparse attention method is adopted. By adding offset branches of parameter sharing in each layer of large language model, slicing the context into multiple fragments, pooling and low-resolution representation extraction, and retrieving low-resolution information based on local attention, combined with offset masking to reduce calculation costs.
It effectively reduces the response time of the first word element, improves the inference efficiency, improves the user experience, and maintains or improves the model performance.
Smart Images

Figure CN119990363A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a hierarchical assisted sparse attention method for a long text and large language model. Background Art
[0002] In recent years, the application of large language models with long contexts has become increasingly popular. However, the traditional causal attention mechanism faces significant performance bottlenecks when processing long texts. The memory usage and inference latency of the causal attention mechanism usually increase quadratically with the increase of the input sequence, which may cause users to face significant delays when generating the first word, affecting the interactive experience.
[0003] Although existing methods can alleviate the memory bottleneck caused by the causal attention mechanism, the quadratic growth of the first word response time still exists in long text reasoning, resulting in a degraded user experience. There is currently a lack of effective solutions to deal with this delay problem, especially when processing long sequences, where users may experience delays of more than one minute. Therefore, optimizing the reasoning efficiency of long texts has become an urgent problem to be solved. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a hierarchical assisted sparse attention method for long text and large language model, so as to improve the reasoning efficiency, reduce the first word response time, and ensure that the model performance is not reduced.
[0005] The present invention is implemented as follows: a hierarchical assisted sparse attention method for a long text large language model, the method comprising:
[0006] Step S1: Add a parameter-sharing offset branch in each large language model layer to obtain a new large language model;
[0007] Step S2: Divide the context into multiple segments and input them into the large language model layer to obtain local feature output. At the same time, each segment is pooled and input into the offset branch of the same layer to obtain a low-resolution representation.
[0008] Step S3: splice the low-resolution representation output by the offset branch of the previous layer to the local features of the large language model layer of the current layer, and output it to the large language model layer of the next layer;
[0009] Step S4: fine-tune the new large language model, and connect a language modeling head after the last large language model layer to output the processing results of the downstream tasks.
[0010] Furthermore, the step S3 also includes adding retrieval of the low-resolution information on the basis of local attention.
[0011] Furthermore, the branches of each layer of the large language model include a normalization module, a shifted attention mechanism, a normalization module and a multi-layer perceptron arranged in sequence, the branch parameters of each layer are consistent with the layer parameters of the large language model, and an offset mask is introduced in the shifted attention mechanism, and the offset mask is obtained by removing the elements on the diagonal based on the original mask.
[0012] Furthermore, the downstream task is a question-answering system, and the language modeling head is a linear transformation layer for mapping the representation space into the output space.
[0013] The present invention has the following advantages: by introducing a hierarchical assisted sparse attention mechanism, the quadratic growth problem of the first word unit response time caused by the long text model in the pre-filling stage is solved, the pre-filling process is accelerated, the reasoning efficiency is improved, and the first word unit response time is effectively reduced; the present invention combines local attention and global information extraction, which can maintain or enhance model performance while improving reasoning efficiency and improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The present invention will be further described below in conjunction with embodiments with reference to the accompanying drawings.
[0015] Figure 1 The present invention is an execution flow chart of a hierarchical assisted sparse attention method for a long text and large language model.
[0016] Figure 2 It is a schematic diagram of the large language model framework of the present invention. DETAILED DESCRIPTION
[0017] The technical solution in the embodiment of the present application has the following overall idea: based on the existing large language model, a hierarchical auxiliary sparse attention mechanism is proposed to optimize the information recovery and reasoning efficiency in long text processing. First, the context is divided into multiple segments, and a low-resolution representation is extracted for each segment. The retrieval of these low-resolution information is added on the basis of local attention, and the lost global information can be supplemented by the low-resolution representation, which plays a role in calibrating the entire attention. Therefore, the entire attention score can be expressed as the sum of local attention and low-resolution global attention. It is calculated by introducing an offset mask and removing the elements on the main diagonal in the original mask. The present invention can reduce the computational cost to a certain extent while maintaining model performance.
[0018] In order to better understand the above technical solution, the technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0019] See also Figure 1 and 2As shown, the present invention provides a hierarchical assisted sparse attention method for a long text large language model, the method comprising:
[0020] Step S1: Add a parameter-sharing offset branch in each large language model layer to obtain a new large language model;
[0021] Step S2: Divide the context into multiple segments and input them into the large language model layer to obtain local feature output. At the same time, each segment is pooled and input into the offset branch of the same layer to obtain a low-resolution representation.
[0022] Step S3: splice the low-resolution representation output by the offset branch of the previous layer to the local features of the large language model layer of the current layer, and output it to the next large language model layer as the input of the next large language model layer;
[0023] Step S4: Fine-tune the new large language model and connect a language modeling head after its last large language model layer to output the processing results of downstream tasks. We fine-tune the projection matrices of all self-attention and multi-layer perceptron modules to ensure the optimization of model performance without comprehensive retraining. Experimental evidence shows that the large oracle model enhanced by our method shows excellent performance and efficiency on various tasks.
[0024] Preferably, the step S3 further includes adding a retrieval function for the low-resolution information on the basis of the local attention. T .
[0025] Preferably, the branches of each layer of the large language model include a standardization module, a shifted attention mechanism, a standardization module and a multi-layer perceptron arranged in sequence, and the branch parameters of each layer are consistent with the layer parameters of the large language model. In order to prevent future word-unit information from leaking in the causal attention mechanism, an offset mask is introduced into the offset attention mechanism, and the offset mask is obtained by removing the elements on the diagonal on the basis of the original mask.
[0026] Preferably, the downstream task is a question answering system, and the language modeling head is a linear transformation layer for mapping the representation space into the output space.
[0027] The following is combined with Figure 2 The above technical solution of the present invention is described as follows:
[0028] This invention introduces the attention mechanism into the traditional language model framework by proposing a dedicated branch to process the captured global information. In order to make the representation output by this branch capture the global information more accurately and be easier to be used by the model, it is processed in the following way:
[0029] (1) We divide the context into multiple segments and input them into the original large language model layer with causal attention mechanism for processing to obtain the corresponding processing results. During the processing of each segment, some caches will be created in each layer, which is called key-value cache. At the same time, multiple segments are pooled and input into the branch with offset attention mechanism in the same layer for low-resolution feature extraction to obtain low-resolution representation. The low-resolution representation is then spliced into the processing result as the output feature of the current layer.
[0030] Since the original model has 32 layers, when the next layer is processed, the output feature X from the previous layer is -1 As the input of the current layer, it is input into the current large language model layer for processing to obtain the corresponding processing result, and the output feature X -1 After pooling, the data is input to the branch with offset attention mechanism at the same layer for low-resolution feature extraction to obtain a low-resolution representation of each word in the text segment. Then aggregate them to get the low-resolution representation sequence of text segment words
[0031]
[0032] Wherein, M represents the number of blocks, that is, the total number of text segments, and the calculation formula for each low-resolution representation is as follows:
[0033]
[0034] (2) Use the large language model layer of the previous layer to calculate the global representation
[0035]
[0036] (3) In this way, the lost global information can be supplemented by low-resolution representation, which plays a role in calibrating the entire attention. Therefore, the entire attention score can be expressed as the sum of local attention and low-resolution global attention:
[0037]
[0038] in It's R T , which is calculated in the previous layer, represents a global, low-resolution information for compensation, which is input into the next layer for calculation and processing after splicing. A global representation A fragment of It can be expressed as When processing the current fragment T, T-1 fragments 1 to T-1 are needed. All fragments before this fragment have a low-resolution representation, and only this fragment T is a high-resolution representation that is directly used for calculation. Since the calculation of the global representation of each layer is used in the next layer, the first layer does not add global information and does not need to calculate global information. Similarly, the last layer does not need to calculate global information because it will not be used in the subsequent layer.
[0039] (4) In order to prevent future word-unit information from leaking in the causal attention mechanism, the present invention introduces an offset mask, which is obtained by removing the elements on the main diagonal in the original mask.
[0040] like Figure 2 As shown, the network structure of the present invention has two branches, the left branch is used to process local information in the attention mechanism, and the right branch is used to process global information. The purpose of setting up two branches is to allocate less computation to global information that has less impact on generation, and allocate more computation to local information that has greater impact on generation. In terms of details, for each layer, at the input level, the two branches share the input, where the right branch needs to perform a pooling operation on the blocks before input to reduce the resolution of the sequence, and the left branch needs to merge the low-resolution representation of the output of the right branch of the previous layer, and in the attention part, the global information is absorbed by paying attention to these low-resolution representations. The right side uses offset attention in the attention part, that is, the original causal attention mask is offset by one unit to the lower left corner. Such a structure can reasonably allocate important computational amounts to local information that plays a key role, thereby greatly saving computational overhead as a whole.
[0041] like Figure 2 In the square area in the lower left corner, in the query tag and key value pair in the offset attention mechanism, the diagonal line is white, that is, the unit block that does not need to be calculated. The splicing result shows that part of the lower left corner is a white area, that is, a block that does not need to be calculated. The method of the present invention can improve the reasoning efficiency and reduce the calculation unit at the same time. Because the introduction of low-resolution representation supplements the lost global information and ensures the model performance.
[0042] The LlaMA-2-7B model is used as the original large language model of the present invention. Experiments are conducted on the original large language model and the improved large language model of the present invention on the data set PG19 (a language modeling benchmark data set proposed by DeepMind), ProofPile (mathematical reasoning), and CodeParrot (code completion), and the experimental data shown in Table 1 are obtained:
[0043] Table 1
[0044]
[0045] The above experimental evidence shows that the large language model enhanced by our method shows excellent performance and efficiency. By optimizing the large language model and improving the efficiency of long-context processing and reasoning, the present invention can effectively reduce the reasoning delay of the long text model. The implementation of this technology will promote the widespread adoption of long-context large language models in practical applications.
[0046] Although the specific implementation modes of the present invention are described above, those skilled in the art should understand that the specific implementation modes described are only illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. A hierarchical assisted sparse attention method for long text and large language model, characterized by: The method comprises: Step S1: Add a parameter-sharing offset branch in each large language model layer to obtain a new large language model; Step S2: Divide the context into multiple segments and input them into the large language model layer to obtain local feature output. At the same time, each segment is pooled and input into the offset branch of the same layer to obtain a low-resolution representation. Step S3: splice the low-resolution representation output by the offset branch of the previous layer to the local features of the large language model layer of the current layer, and output it to the large language model layer of the next layer; Step S4: fine-tune the new large language model, and connect a language modeling head after the last large language model layer to output the processing results of the downstream tasks.
2. The hierarchical assisted sparse attention method for long text and large language model according to claim 1, characterized in that: The step S3 also includes adding the retrieval of the low-resolution information on the basis of the local attention.
3. The hierarchical assisted sparse attention method for long text and large language model according to claim 1, characterized in that: The branches of each layer of the large language model include a standardization module, a shifted attention mechanism, a standardization module and a multi-layer perceptron arranged in sequence. The branch parameters of each layer are consistent with the layer parameters of the large language model, and an offset mask is introduced in the shifted attention mechanism. The offset mask is obtained by removing the elements on the diagonal on the basis of the original mask.
4. The hierarchical assisted sparse attention method for long text and large language model according to claim 1, characterized in that: The downstream task is a question answering system, and the language modeling head is a linear transformation layer for mapping the representation space into the output space.
Citation Information
Cited By
Long text processing method and device based on large language model, equipment and storage medium
CN120705288A