Rotation position coding design method suitable for large model long context understanding
By introducing the imaginary part calculation result of rotation position encoding into the large language model as a self-attention head, and computing it in parallel with the real part attention, the problem of imaginary part information loss is solved, and the ability and efficiency of long text comprehension are improved.
Patent Information
- Application Number
- CN202511868645.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-20
AI Technical Summary
Existing rotational position encoding suffers from the loss of imaginary part information and insufficient analysis of computational form in large language models, which affects the ability to understand long texts.
The imaginary part of the rotation position encoding is reintroduced into the self-attention calculation as another self-attention head to be calculated in parallel with the real part attention. By combining real and imaginary part attention, the long text comprehension ability of large language models is enhanced.
It reduces information loss, improves the ability to understand long texts, reduces attention parameters and key-value caching, and improves the efficiency and effectiveness of long text tasks.
Smart Images

Figure CN121705412A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of artificial intelligence, and in particular to a rotation position coding design method suitable for long context understanding of a large model. BACKGROUND
[0002] The existing large language model mainly uses a Transformer architecture based on an attention mechanism, and the position coding is a very key module. Among a plurality of position coding design schemes, the rotation position coding (RoPE) is the most mainstream position coding method in the current large language model. The rotation position coding unifies the absolute position, that is, the index position of each word symbol in the input sequence, and the relative position information, that is, the relative distance between each two word symbols, through a rotation matrix and complex multiplication. The effect of the relative position coding is realized in the form of the absolute position coding. Meanwhile, due to the property of the trigonometric function in the rotation position coding, the rotation position coding also has the ability of semantic aggregation and long context decay.
[0003] However, the rotation position coding still faces challenges including length extrapolation, multi-modal input and data perception, thereby also promoting a large amount of improvement work. Among them, the length extrapolation is the main research direction, and the existing work makes the model process the context far beyond the training window by scaling the rotation angle base, interpolation or compressing the index range, or combining the rotation position coding with sparse attention. Some other work extends the rotation position coding to heterogeneous, cross-modal input, especially text-video sequence input, and achieves the effect of compatible heterogeneous input by assigning different dimensions to depict the space-time position. In addition, some research designs a parameterized method to code the context information, refines or replaces the rotation position coding, and realizes data dependency.
[0004] However, very few works re-examine the calculation form of the rotation position coding or analyze its inherent limitations. Especially noteworthy is that the rotation position coding has a loss of imaginary part information in the rotation format compared with the complex multiplication format, which has been ignored by other works. Although some existing research attempts to introduce complete complex calculation into the self-attention mechanism or neural network, the analysis of the characteristics and functions of the imaginary part quantity in the rotation position coding is still blank. SUMMARY
[0005] The application aims to provide a rotation position coding design method suitable for long context understanding of a large model, which improves the long text understanding capability.
[0006] The application can be achieved by the following technical solutions. A rotation position coding design method suitable for long context understanding of a large model, comprising the following steps: The hidden state sequence in the large language model is mapped to a query vector sequence, a key vector sequence and a value vector sequence through a linear mapping, and a copy of the query vector sequence is obtained, wherein the large language model is provided with a self-attention layer comprising multiple attention heads; The query vector sequence and the key vector sequence are subjected to traditional rotary position encoding to obtain a query vector sequence with position encoding and a key vector sequence with position encoding; The copied query vector sequence is rotated and then subjected to traditional rotary position encoding to obtain a rotated and position-encoded query vector sequence; Based on the rotated and position-encoded query vector sequence, the query vector sequence with position encoding, the key vector sequence and the value vector sequence, self-attention operation is performed to obtain the output result of the self-attention layer, and the design process is completed.
[0007] Further, the traditional rotary position encoding is as follows: The feature vectors are grouped in pairs in the dimension direction, and each group is rotated at a rotation frequency to encode the relative position into the self-attention calculation, wherein different groups are rotated at different rotation frequencies .
[0008] Further, the query vectors and key vectors with position encoding in the query vector sequence with position encoding and the key vector sequence with position encoding are respectively represented as: , , wherein, is the element of the 2nth dimension of the query vector at position t, is the element of the 2nth dimension of the key vector at position s, is the rotation frequency; wherein, in the complex form, vector rotation is equivalent to the corresponding complex rotation in the complex plane, i.e. , wherein, is the element of the nth dimension of the query vector at position t in the complex form, is the element of the nth dimension of the key vector at position s in the complex form, is the imaginary unit; The above two forms satisfy the following equivalence relation: , wherein, is the conjugate of the nth dimension element of the key vector at position s in the complex form.
[0009] Further, the rotation of the copied query vector sequence is a clockwise rotation of 90 degrees.
[0010] Furthermore, the query vector in the rotated position-encoded query vector sequence is represented as: , In the formula, Let t be the element of the 2n-th dimension of the query vector used for imaginary part attention calculation. Let t be the element of the 2nth dimension of the query vector at position t. is the rotation frequency.
[0011] Furthermore, for the rotated and copied query vector sequence, the corresponding self-attention calculation result is equivalent to the calculation result of the imaginary part discarded in the complex form of the rotated position encoding, expressed as: , In the formula, Let t be the element of the 2n-th dimension of the query vector used for imaginary part attention calculation. Let be the element of the 2n-th dimension of the key vector at position s. Let be the element of the nth dimension of the query vector at position t in complex form. It is the conjugate of the nth dimension element of the key vector at position s in complex form.
[0012] Furthermore, the step of obtaining the output result of the self-attention layer includes: The rotated position-encoded query vector sequence and the position-encoded query vector sequence are concatenated to obtain a concatenated query vector sequence, which serves as a new query vector sequence for the attention head. This new query vector sequence is used for attention calculation of the real and imaginary parts. Perform self-attention operation between the concatenated query vector sequence and the key vector sequence and value vector sequence to obtain the self-attention calculation result; The self-attention calculation results are linearly mapped to obtain the output results of the self-attention layer.
[0013] Furthermore, the self-attention calculation of the real part corresponds to the real part result of complex multiplication, and the self-attention calculation of the imaginary part corresponds to the imaginary part result of complex multiplication, wherein the specific calculation formula is as follows: , In the formula, The complete attention fraction in complex form, The real part attention score, The imaginary part of the attention score. The number of attention dimensions. Let be the element of the nth dimension of the query vector at position t in complex form. Let be the conjugate of the nth dimension element of the key vector at position s in complex form. For rotation matrix, Let t be the query vector used for calculating the imaginary part attention. Let s be the key vector at position s.
[0014] Furthermore, the self-attention calculations for the real and imaginary parts are performed in parallel as two sets of self-attention in a multi-head attention process.
[0015] Furthermore, while maintaining the same attention head, the output linear mapping is such that the input dimension is consistent with the original self-attention; while maintaining the same key-value cache, the output linear mapping is such that the input dimension is twice that of the original self-attention.
[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) Starting from the complex form of traditional rotation position encoding, this invention reintroduces the originally discarded imaginary part calculation result into the self-attention calculation, as another self-attention head to be calculated simultaneously with the real part attention. By combining the real and imaginary part attention, the characteristics of the two attentions in mathematically converging local semantics and characterizing long-range dependencies are brought into play, thereby enhancing the long text understanding ability of the large language model.
[0017] (2) The present invention introduces the imaginary part result of the complex form of the rotation position code into the self-attention calculation, which reduces the information loss caused by taking the real part in the traditional rotation position code implementation.
[0018] (3) This invention achieves imaginary part attention calculation by retaining the key vector and value vector and rotating the query vector by 90 degrees, thus preserving the mathematical properties of traditional rotation position encoding and being compatible with existing self-attention acceleration algorithms.
[0019] (4) When the number of self-attention heads is the same, the present invention introduces half of the attention heads to perform virtual part attention calculation, thereby reducing half of the attention parameters and key value cache, which is beneficial to the efficiency of long text training and inference.
[0020] (5) Under the same number of attention heads and the same cache setting, the present invention achieves better results in both long and short tasks, surpassing other positional encoding schemes; especially in long text tasks, it has a more obvious advantage compared with other methods. Attached Figure Description
[0021] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram (a) of the dual attention of the real and imaginary parts of the present invention; Figure 3This is a schematic diagram (II) of the dual attention of the real and imaginary parts of the present invention. Detailed Implementation
[0022] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0023] Example 1 This embodiment provides a rotation position encoding design method suitable for large model long context understanding, combined with Figure 1 The method includes the following steps: 101. Input linear mapping to obtain the query key-value vector.
[0024] This embodiment takes a large language model with 376M parameters (12 self-attention layers, 12 attention heads, attention head dimension 128, hidden state dimension 1536) as an example.
[0025] This step primarily involves mapping the hidden state sequence through input linearity to obtain the query vector sequence, key vector sequence, and value vector sequence. Specifically: Hidden state sequences of large language models The query vector sequence is obtained after input linear mapping. Key vector sequence Value vector sequence While maintaining the same attention heads (EH), the query vector sequence, key vector sequence, and value vector sequence obtained by linear mapping are reduced by half compared to the original self-attention, i.e., there are only 6 attention heads, corresponding to a reduction of half in the input mapping matrix; while maintaining the same cache (EC), the above sequences are of the same size.
[0026] 102. Perform traditional RoPE on the query and key vector.
[0027] This step primarily involves copying the original query vector sequence and performing traditional rotational position encoding on both the original query vector and key vector sequences. Specifically: In step 102, a copy of the original query vector sequence is made, and the original query vector and key vector sequences are subjected to traditional rotation position encoding. That is, the feature vectors are grouped pairwise along the dimensional direction, and each group is rotated. The rotation frequency of different groups is... The differences result in a query vector sequence and a key vector sequence with positional encoding: , , In the formula, Let t be the element of the 2nth dimension of the query vector at position t. Let be the element of the 2n-th dimension of the key vector at position s. is the rotation frequency.
[0028] In complex form, a vector rotation is equivalent to a rotation of the corresponding complex number in the complex plane, that is: , In the formula, Let be the element of the nth dimension of the query vector at position t in complex form. Let be the element of the nth dimension of the key vector at position s in complex form. It is the imaginary unit.
[0029] The two forms above satisfy the following equivalence relation: the product of the rotated vectors is equal to the real part of the complex multiplication. .
[0030] In the formula, It is the conjugate of the nth dimension element of the key vector at position s in complex form.
[0031] 103. Perform RoPE on the query vector that has been rotated 90 degrees.
[0032] This step primarily involves rotating the copied query vector sequence 90 degrees clockwise before performing traditional rotation position encoding. Specifically: In step 103, the copied query vector is rotated 90 degrees clockwise (i.e., (radians), then perform traditional rotation position encoding to obtain the query vector sequence of the rotated position encoding, i.e.: , In the formula, Let t be the element of the 2n-th dimension of the query vector used for imaginary part attention calculation. This is the element of the 2nth dimension of the query vector at position t.
[0033] like Figure 2 As shown, the self-attention calculation result corresponding to the query vector sequence after rotation by 90 degrees is equivalent to the calculation result of the discarded imaginary part (negative imaginary part) in the complex form of the rotation position encoding: , In the formula, Let t be the element of the 2n-th dimension of the query vector used for imaginary part attention calculation. Let be the element of the 2n-th dimension of the key vector at position s. Let be the element of the nth dimension of the query vector at position t in complex form. It is the conjugate of the nth dimension element of the key vector at position s in complex form.
[0034] If the rotation matrix is simplified as Therefore, the position encoding process in steps 102 and 103 can be summarized as follows: .
[0035] in, Let t be the query vector used for real part attention calculation. Let t be the query vector used for calculating the imaginary part attention. Let s be the key vector at position s.
[0036] 104. Concatenate the query vectors encoded twice.
[0037] This step mainly involves concatenating the rotated, position-encoded query vector sequence from step 103 with the position-encoded query vector sequence obtained through traditional position encoding in step 102, resulting in a concatenated query vector sequence, which serves as a new query vector for the attention head. Specifically: In step 104, the calculation result is concatenated to the original encoded query vector to form a new query vector for the attention heads. At this point, while maintaining the same attention heads (EH), six heads are responsible for calculating the real and imaginary self-attention functions respectively; while maintaining the same cache (EC), twelve heads are responsible for calculating the real and imaginary self-attention functions respectively. The real self-attention calculation corresponds to the real part result of complex multiplication, and the imaginary self-attention calculation corresponds to the imaginary part result of complex multiplication, as shown in the following formulas.
[0038] , .
[0039] In the formula, The complete attention fraction in complex form, The real part attention score, The imaginary part of the attention score. The number of attention dimensions. Let be the element of the nth dimension of the query vector at position t in complex form. Let be the conjugate of the nth dimension element of the key vector at position s in complex form. For rotation matrix, Let t be the query vector used for calculating the imaginary part attention. Let s be the key vector at position s.
[0040] 105. Perform self-attention on the query key-value vector sequence.
[0041] This step primarily involves performing self-attention operations between the concatenated query vector sequence and the key-value vector sequence. Specifically: In step 105, self-attention operations are performed between the concatenated query vector sequence and the key-value vector sequence. It should be noted that the self-attention calculations for the real and imaginary parts are performed in parallel as two groups of self-attention in multi-head attention. They can be regarded as two different self-attention groups, which are compatible with existing self-attention acceleration algorithms, including grouped self-attention in the architecture and FlashAttention, PagedAttention, and other framework-based solutions.
[0042] 106. Perform an output linear mapping on the self-attention results.
[0043] This step primarily involves performing an output linear mapping on the self-attention calculation results to obtain the output of the self-attention layer. Specifically: In step 106, the output linear mapping is performed on the self-attention calculation results to obtain the output results of the self-attention layer. Under the condition of keeping the same attention head, the input dimension of the output mapping is consistent with the original self-attention, that is, 1536. Under the condition of keeping the same key-value cache, the input dimension of the output mapping is twice that of the original self-attention, that is, 3072.
[0044] like Figure 2 As shown, this invention proposes an overall structure for rotational positional encoding design suitable for large-scale, long-context understanding models, RoPE++. It re-incorporates the imaginary part self-attention results from traditional positional encoding into the large language model as a new set of attention heads for computation. Specifically, as follows... Figure 3 As shown, with the same number of attention heads (EH) and the same cache (EC), RoPE++ performs self-attention operations on the key-value cache and the query vector corresponding to the real and imaginary part calculation results, respectively, to obtain two attention results: one for the real part and one for the imaginary part. Leveraging the characteristic that the imaginary part self-attention is better at capturing semantic dependencies in long texts, RoPE++ improves the long text understanding capabilities of large language models. Furthermore, by reducing the number of attention parameters and key-value cache by half while maintaining the same total number of attention heads, it facilitates efficient long text understanding. On a 376M large language model, as shown in Table 1, RoPE++ outperforms the baseline (RoPE) and other positional encoding designs on short text tasks, regardless of the same number of attention heads (EH) and the same cache (EC) settings. And as shown in Table 2, RoPE++ also outperforms the baseline on long text tasks, showing even greater advantages in longer contexts.
[0045] Table 1. Comparison of short text performance between the method of this invention and other methods. Table 2. Comparison of long-text performance between the method of this invention and the baseline method. If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A rotation position encoding design method suitable for large model long context understanding, characterized in that, Includes the following steps: The hidden state sequence in the large language model is linearly mapped to obtain the query vector sequence, key vector sequence and value vector sequence, and a copy of the query vector sequence is made. The large language model has a self-attention layer containing a multi-head attention head. The query vector sequence and key vector sequence are subjected to traditional rotation position encoding to obtain a query vector sequence and key vector sequence with position encoding; Based on the copied query vector sequence, it is rotated and then subjected to traditional rotation position encoding to obtain the rotated position encoded query vector sequence; Based on the rotated position-encoded query vector sequence, the position-encoded query vector sequence, the position-encoded key vector sequence, and the value vector sequence, self-attention operation is performed to obtain the output result of the self-attention layer, thus completing the design process.
2. The rotation position encoding design method suitable for large model long context understanding according to claim 1, characterized in that, Traditional rotational position encoding is as follows: The feature vectors are grouped pairwise along the dimensional direction, and each group is sorted according to the rotation frequency. Rotation is performed to encode the relative positions into the self-attention computation, where different groups rotate at different frequencies. Rotate.
3. The rotation position encoding design method suitable for large model long context understanding according to claim 1, characterized in that, The position-encoded query vector and key vector in the query vector sequence and key vector sequence are respectively represented as follows: , , In the formula, Let t be the element of the 2nth dimension of the query vector at position t. Let be the element of the 2n-th dimension of the key vector at position s. The rotational frequency; In the complex form, vector rotation is equivalent to the corresponding complex rotation in the complex plane, i.e.: , In the formula, Let be the element of the nth dimension of the query vector at position t in complex form. Let be the element of the nth dimension of the key vector at position s in complex form. The imaginary unit; The two forms described above satisfy the following equivalence relation: , In the formula, It is the conjugate of the nth dimension element of the key vector at position s in complex form.
4. The rotation position encoding design method suitable for large model long context understanding according to claim 1, characterized in that, The copied query vector sequence is rotated 90 degrees clockwise.
5. The rotation position encoding design method suitable for large model long context understanding according to claim 1, characterized in that, The query vector in the query vector sequence after rotational position encoding is represented as: , In the formula, Let t be the element of the 2n-th dimension of the query vector used for imaginary part attention calculation. Let t be the element of the 2nth dimension of the query vector at position t. is the rotation frequency.
6. The rotation position encoding design method suitable for large model long context understanding according to claim 1, characterized in that, For the rotated copy of the query vector sequence, the corresponding self-attention calculation result is equivalent to the calculation result of the imaginary part discarded in the complex form of the rotated position encoding, expressed as: , In the formula, Let t be the element of the 2n-th dimension of the query vector used for imaginary part attention calculation. Let be the element of the 2n-th dimension of the key vector at position s. Let be the element of the nth dimension of the query vector at position t in complex form. It is the conjugate of the nth dimension element of the key vector at position s in complex form.
7. The rotation position encoding design method suitable for large model long context understanding according to claim 1, characterized in that, The steps for obtaining the output of the self-attention layer include: The rotated position-encoded query vector sequence and the position-encoded query vector sequence are concatenated to obtain a concatenated query vector sequence, which serves as a new query vector sequence for the attention head. This new query vector sequence is used for attention calculation of the real and imaginary parts. Perform self-attention operation between the concatenated query vector sequence and the key vector sequence and value vector sequence with position encoding to obtain the self-attention calculation result; The self-attention calculation results are linearly mapped to obtain the output results of the self-attention layer.
8. A rotation position encoding design method suitable for large model long context understanding according to claim 7, characterized in that, The self-attention calculation of the real part corresponds to the real part result of complex multiplication, and the self-attention calculation of the imaginary part corresponds to the imaginary part result of complex multiplication, wherein the specific calculation formula is as follows: , In the formula, For the complete attention result in complex form, For real attention results, For the imaginary part of attention, The number of attention dimensions. Let be the element of the nth dimension of the query vector at position t in complex form. Let be the conjugate of the nth dimension element of the key vector at position s in complex form. Let be a rotation matrix. Let t be the query vector used for calculating the imaginary part attention. Let s be the key vector at position s.
9. A rotation position encoding design method suitable for large model long context understanding according to claim 7, characterized in that, The self-attention calculations for the real and imaginary parts are performed in parallel as two sets of self-attention in a multi-head attention process.
10. A rotation position encoding design method suitable for large model long context understanding according to claim 7, characterized in that, While maintaining the same attention head, the output linear mapping is such that the input dimension is consistent with the original self-attention; while maintaining the same key-value cache, the output linear mapping is such that the input dimension is twice that of the original self-attention.
Citation Information
Cited By
Large language model training method and device based on range position coding, and medium
CN122153661A