Large language model long text extrapolation method and device, electronic equipment and storage medium
By synchronously expanding the initial window size and position coding of the sliding window attention mechanism, the problem of the sharp increase in computing resource requirements when processing long text is solved, achieving more efficient long text processing and reducing inference costs.
Patent Information
- Application Number
- CN202510055799.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
When existing large language models process long text, the demand for computing resources has increased dramatically, resulting in high inference costs and limited extrapolated text lengths of the sliding window attention mechanism.
By synchronously expanding the initial window size and initial position encoding of the sliding window attention mechanism, the target large language model is obtained, thereby processing longer text. The specific method includes determining the search scores of multiple attention heads, selecting the search head and local head, and expanding their window size and position codes respectively.
The large language model effectively handles longer texts, reduces inference costs, and avoids the problem that the processing effect is limited by the preset length of the text to be processed.
Smart Images

Figure CN120012924A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method, device, electronic device and storage medium for extrapolating long texts using a large language model. Background Art
[0002] With the breakthroughs in deep learning, reinforcement learning and other technologies and the accumulation of high-quality texts on the Chinese Internet, the field of natural language processing has developed rapidly, and large language models based on generative decoding have achieved advanced results in various natural language processing tasks. Deep learning usually relies on a large amount of labeled data, but in recent years, with the introduction of various self-supervised pre-training models, semi-supervised and unsupervised algorithms, large language models can use a large amount of unlabeled corpus for learning, which greatly reduces the number of labeled corpora required by large language models for various downstream tasks, and reduces the technical and resource threshold for migrating them to downstream business areas.
[0003] At present, with the continuous improvement of positional encoding technology and the increasing maturity of training technology, the training and application of large language models have become a field that is mainly limited by computing resources. When processing shorter texts, large language models can achieve excellent results due to the superiority of the model architecture. However, when computing resources are limited, processing long texts usually requires a large amount of computing resources due to the complexity of the attention mechanism algorithm, and the resource requirements will increase sharply as the length of the text increases.
[0004] Existing solutions for expanding the extrapolation capabilities of large language models for long texts mainly include two categories: global attention mechanism and sliding window attention mechanism. The global attention mechanism enhances the ability of large language models to handle text lengths that have not been seen during the training phase through methods such as linear interpolation, so that the perplexity of the model will not increase significantly when the global attention mechanism is used to expand the model context window; the sliding window mechanism uses a multi-layer receptive field mechanism to expand the model's receptive field capabilities without increasing computing and video memory resources.
[0005] However, for large language models that use global attention mechanisms to implement extrapolation, the inference cost increases dramatically with the length, while for models that use sliding window attention mechanisms to implement extrapolation, the extrapolated text length is limited. Summary of the invention
[0006] The present invention provides a large language model long text extrapolation method, device, electronic device and storage medium to solve the defects existing in the related art.
[0007] The present invention provides a large language model long text extrapolation method, comprising: Get the text to be processed of preset length; If the preset length is greater than the sequence length of the training text of the initial large language model, based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, the initial position encoding of the sliding window attention mechanism is expanded, and the initial window size is expanded to obtain a target large language model; The text to be processed is processed based on the target large language model.
[0008] According to a large language model long text extrapolation method provided by the present invention, the sliding window attention mechanism includes multiple attention heads; The step of expanding the initial window size comprises: Determine a test set, wherein the test set includes decoding results corresponding to a plurality of texts to be decoded; Based on the decoding result, determining the number of successful searches of each word in the to-be-decoded text by the multiple attention heads, selecting a search head from the multiple attention heads based on the number of successful searches, and determining a local head other than the search head from the multiple attention heads; The initial window size of the retrieval header is extended to the preset length, and the initial window size of the local header remains unchanged.
[0009] According to a large language model long text extrapolation method provided by the present invention, the cache information corresponding to the retrieval head and the local head are respectively stored in different cache spaces, and the retrieval head and the local head perform dot product attention calculations respectively.
[0010] According to a large language model long text extrapolation method provided by the present invention, the selecting a retrieval head from the multiple attention heads based on the number of successful retrievals includes: For any attention head, based on the number of successful retrievals of the attention head and the sequence length and number of the decoding results in the test set, calculate the retrieval score of the attention head; Based on the retrieval score of any of the attention heads, determine whether any of the attention heads is the retrieval head.
[0011] According to a large language model long text extrapolation method provided by the present invention, the initial position encoding of the sliding window attention mechanism based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model is expanded, including: Determining initial encoding parameters for each position in the initial position encoding; Calculating a target coding parameter based on a ratio of the preset length to the initial window size and the initial coding parameter; Based on the target coding parameters, the extended position coding is determined.
[0012] According to a large language model long text extrapolation method provided by the present invention, the expanding the initial window size comprises: The initial window sizes of the pre-filling phase and the decoding phase in the decoding process of the initial large language model are respectively expanded.
[0013] According to a large language model long text extrapolation method provided by the present invention, the processing of the to-be-processed text based on the target large language model includes: Based on the target large language model, generating target position codes for the to-be-processed text segments in a fixed size according to the expanded position codes; Based on the target position code, the text to be processed is processed.
[0014] The present invention also provides a large language model long text extrapolation device, comprising: A text acquisition module is used to acquire a text to be processed of a preset length; A size expansion module, configured to expand the initial position encoding of the sliding window attention mechanism based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, and expand the initial window size to obtain a target large language model if the preset length is greater than the sequence length of the training text of the initial large language model; A processing module is used to process the text to be processed based on the target large language model.
[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the large language model long text extrapolation method as described in any one of the above is implemented.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the large language model long text extrapolation method as described in any one of the above.
[0017] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the large language model long text extrapolation methods described above.
[0018] The large language model long text extrapolation method, device, electronic device and storage medium provided by the present invention first obtain a text to be processed of a preset length; then if the preset length is greater than the sequence length of the training text of the initial large language model, based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, the initial position encoding of the sliding window attention mechanism is expanded, and the initial window size is expanded to obtain a target large language model; finally, based on the target large language model, the text to be processed is processed. The method synchronously expands the initial window size and the initial position encoding of the sliding window attention mechanism, so that the target large language model has the ability to process longer texts. Furthermore, by processing the text to be processed by the target large language model, the processing effect can be guaranteed, the inference cost can be reduced, and the processing effect is not limited by the preset length of the text to be processed. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 It is a schematic diagram of the network structure of the global attention mechanism provided by the present invention.
[0021] Figure 2 It is a flowchart of the long text extrapolation method of a large language model provided by the present invention.
[0022] Figure 3 It is an expanded schematic diagram of the initial window size and initial position encoding of the sliding window attention mechanism provided by the present invention.
[0023] Figure 4 It is a schematic diagram of the visualization results of retrieval scores in the large language model long text extrapolation method provided by the present invention.
[0024] Figure 5 It is a schematic diagram of the separate attention method in the large language model long text extrapolation method provided by the present invention.
[0025] Figure 6 It is a schematic diagram of the cache and calculation of each attention head in the long text extrapolation method of a large language model provided by the present invention.
[0026] Figure 7 It is a structural schematic diagram of the large language model long text extrapolation device provided by the present invention.
[0027] Figure 8It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] Figure 1 This is a schematic diagram of the network structure of the global attention mechanism. The window size of the global attention mechanism is the window size of each layer of the attention network, which is the same as the sequence length of the training text. Figure 1 The window size of the global attention mechanism is 6 words. This allows us to see the information of all the words before the word when decoding words at different positions in the input text, and has good long-distance modeling capabilities.
[0030] In the global attention mechanism, each word in the input text needs to calculate the attention weight of other words in the input text, and the complexity is O(n 2 ), where n is the sequence length of the input text. During inference, the cache is performed, and the cost of the cache is proportional to the sequence length of the input text. Therefore, when the sequence length processed by the global attention mechanism is four times the original, its computational cost increases to sixteen times the original, and the cache cost increases to four times the original. This shows that the technology is limited by computing and video memory resources and cannot process ultra-long texts of up to 200,000 or even one million words.
[0031] To improve this situation, when processing input text with a sequence length far exceeding that of the training phase, the position encoding is linearly interpolated to shorten the distance between relative positions, thereby enabling the large language model to process input text of this sequence length. However, the inference cost of this method increases dramatically as the sequence length of the input text increases.
[0032] When the sliding window attention mechanism is used to process long texts, the receptive field is only the text fragment at the end of the sequence of the long text that is less than the sequence length of the long text, that is, the window is only the sequence length of the text fragment.
[0033] In the autoregressive decoding process, when decoding a word at each position, only a fixed number of words before it can be seen. Due to the multi-layer attention network of the sliding window attention mechanism, the actual receptive field of the words in the upper attention network can be further expanded through the sliding window attention mechanism, and the words in the bottom layer far beyond the window can be seen.
[0034] The sliding window attention mechanism reduces the computational complexity by limiting the attention range to a fixed-size window. Each word only pays attention to a fixed number of words that are close to it in the sequence. This method reduces the complexity from O(n 2 ) is reduced to O(n*w), where w is the window size and n is the sequence length. Nevertheless, the cost of this reduction in complexity is that the large language model's perception of long-distance dependencies is weakened. Although the window of the multi-layer attention network can expand the receptive field of the large language model to a certain extent, this ability will drop rapidly as the distance beyond the window increases. And in the decoding process, since it is impossible to predict where the information required for subsequent questions is in the original text, this part of information cannot be put into the cache in advance. Therefore, when the large language model performs long text extrapolation, that is, when processing input text that is longer than the training text in the inference stage, the performance of the large language model will drop significantly. As a result, the length of extrapolated text to which this method can be applied is limited.
[0035] Based on this, an embodiment of the present invention provides a long text extrapolation method for a large language model.
[0036] Figure 2 FIG. 1 is a flow chart of a large language model long text extrapolation method provided in an embodiment of the present invention. Figure 2 As shown, the method includes: S1, obtaining a text to be processed of a preset length; S2, if the preset length is greater than the sequence length of the training text of the initial large language model, based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, the initial position encoding of the sliding window attention mechanism is expanded, and the initial window size is expanded to obtain a target large language model; S3: Process the text to be processed based on the target large language model.
[0037] Specifically, the large language model long text extrapolation method provided in the embodiment of the present invention is executed by a large language model long text extrapolation device, which can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.
[0038] First, execute step S1 to obtain a text to be processed of a preset length. The text to be processed may be a long text, that is, the preset length is longer than the sequence length of the training text of the initial large language model, and the training text is the text sample used by the initial large language model during training. Therefore, the processing result obtained by directly using the initial large language model to process the text to be processed is not ideal, and the initial large language model needs to be updated with the preset length so as to use the updated target large language model to process the long text. In addition, the text to be processed may also be a short text, that is, the preset length is shorter than or equal to the sequence length of the training text of the initial large language model, and the target large language model provided in the embodiment of the present invention may also be used to process the short text.
[0039] Among them, the initial large language model can include a multi-layer attention network to implement a sliding window attention mechanism. The sliding window attention mechanism, because it has a multi-layer receptive field, makes the actual receptive field of the upper network larger than the initial window size, and the local information captured by the underlying network can be gathered at the top layer. However, since this information transmission process is lossy, the initial large language model cannot know in advance which part of the information needs to be used later, so its actual receptive field is not much larger than the initial window size. This also causes the performance of the initial large language model to be severely degraded when processing training texts that far exceed the window length. In order to enable the initial large language model to have the ability to process longer texts, the initial window size and initial position encoding of the sliding window attention mechanism need to be expanded synchronously.
[0040] Then, step S2 is executed. When the preset length is greater than the sequence length of the training text of the initial large language model, the preset length and the initial window size of the sliding window attention mechanism of the initial large language model can be used to expand the initial position encoding of the sliding window attention mechanism and expand the initial window size to obtain the target large language model.
[0041] Here, the initial window size and initial position encoding in each layer of the attention network of the initial large language model need to be expanded. By expanding the initial window size, the window can capture a wider range of text information and realize the extrapolation ability of the target large language model on longer texts. Since the length of text seen in the window during training is short, the window does not have the ability to directly process more text, that is, the initial position encoding in the window is inconsistent during reasoning and training. By expanding the initial position encoding, the ability of the target large language model to process longer texts can be improved.
[0042] The initial position encoding of the sliding window attention mechanism can be obtained by encoding through relative position encoding or by rotating position encoding (Rotary Position Embedding, RoPE), which is not specifically limited here. Among them, RoPE is a position encoding method that can integrate the relative position information of words at different positions in the sequence into the attention calculation and improve the structural performance of mainstream large language models such as Transformer. Compared with relative position encoding, RoPE has better extrapolation. Although Rope has a certain degree of extrapolation, the effect of this extrapolation will be significantly reduced when processing long texts.
[0043] In order to expand the position encoding to adapt to longer text sequences, one existing method is to perform linear interpolation on the word position, which will cause the performance of the large language model to degrade significantly. Another method is to increase the initial encoding parameters by using the magnification factor of the preset length compared to the training text. For example, the magnification factor of the preset length compared to the initial large language model training text is k, and the initial encoding parameters are , then the increased initial encoding parameters are This method works well on large language models that use a global attention mechanism, but cannot be applied to large language models that use a sliding window attention mechanism.
[0044] In an embodiment of the present invention, when the preset length is greater than the sequence length of the training text of the initial large language model, the initial position encoding of the sliding window attention mechanism can be expanded using the preset length and the initial window size of the sliding window attention mechanism of the initial large language model.
[0045] Here, expanding the initial position code means increasing the initial code parameters of each position in the initial position code, so that the expanded target position code can adapt to a longer text sequence. For example, a mathematical operation can be performed on the preset length and the initial window size, and then the result of the mathematical operation is multiplied by the initial position code to obtain the expanded target position code. The mathematical operation can include a single operation such as addition, subtraction, multiplication, division, logarithm, exponentiation, or a combination of at least two single operations, which is not specifically limited here.
[0046] At the same time, the initial window size also needs to be expanded, that is, the initial window size needs to be increased. For example, it can be directly increased to a preset length. At this time, the sliding window attention mechanism evolves into a global attention mechanism. The window size to be increased can also be specified as needed. No specific limitation is given here.
[0047] It is understandable that if only the initial position encoding is expanded or only the initial window size is expanded, the ability to extrapolate long texts cannot be improved, but the long text field of view of the initial large language model will be degraded to the window of the sliding window attention mechanism. In the embodiment of the present invention, by synchronously expanding the initial window size and the initial position encoding of the sliding window attention mechanism, the target large language model has the ability to process longer texts.
[0048] like Figure 3 The figure shows the initial window size of the sliding window attention mechanism and the expansion diagram of the initial position encoding. In the initial large language model before expansion, the window of the sliding window attention mechanism only contains 9 attention network nodes, and the sliding window attention mechanism can process 4 words in the input sequence at the same time. In the target large language model after expansion, the sliding window attention mechanism contains 18 attention network nodes, and the sliding window attention mechanism can process 6 words in the input sequence at the same time.
[0049] Finally, step S3 is executed to process the text to be processed using the target large language model. Here, the text to be processed can be input into the target large language model, and the target large language model processes the text to be processed and gives a reply to the text to be processed. For example, if the text to be processed is a specific question of the user, the reply given by the target large language model can be the answer to the specific question.
[0050] The large language model long text extrapolation method provided in the embodiment of the present invention first obtains a text to be processed of a preset length; then if the preset length is greater than the sequence length of the training text of the initial large language model, based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, the initial position encoding of the sliding window attention mechanism is expanded, and the initial window size is expanded to obtain a target large language model; finally, based on the target large language model, the text to be processed is processed. The method synchronously expands the initial window size and the initial position encoding of the sliding window attention mechanism, so that the target large language model has the ability to process longer texts. Furthermore, by processing the text to be processed through the target large language model, the processing effect can be guaranteed, the inference cost can be reduced, and the processing effect is not limited by the preset length of the text to be processed.
[0051] Based on the above embodiment, the sliding window attention mechanism includes multiple attention heads; The step of expanding the initial window size comprises: Determine a test set, wherein the test set includes decoding results corresponding to a plurality of texts to be decoded; Based on the decoding result, determining the number of successful searches of each word in the to-be-decoded text by the multiple attention heads, selecting a search head from the multiple attention heads based on the number of successful searches, and determining a local head other than the search head from the multiple attention heads; The initial window size of the retrieval header is extended to the preset length, and the initial window size of the local header remains unchanged.
[0052] Specifically, the sliding window attention mechanism is a multi-head attention mechanism, which may include multiple attention heads. Based on the interpretability of the multi-head attention mechanism, the retrieval capability of each attention head can be analyzed and calculated, and each attention head can be divided into a retrieval head and a local head according to its retrieval capability. Among them, the retrieval capability of the attention head can be represented by the number of successful retrievals of each word in the long text by the attention head. The attention head that can capture key information from the long text is the retrieval head, and the retrieval head has the ability to model long distances. The remaining attention heads except the retrieval head are local heads, and the local heads do not have the ability to model long distances.
[0053] In order to distinguish the retrieval head from the local head, the test set is introduced , O is the number of decoding results in the test set, the oth decoding result Can be the text to be decoded The corresponding decoding result is , l Is the text to be decoded The sequence length, that is, the text to be decoded The number of words contained in , n is the decoding result The sequence length, that is, the decoding result The number of words contained in .
[0054] When using multiple attention heads to determine the decoding result, if the decoding result The kth word in The attention weight corresponding to the i-th attention head is , and take The position j corresponding to the maximum value in is , j is the position that the i-th attention head pays most attention to. If the word at position j in the text to be decoded X and If they are consistent, it is determined that the retrieval of the i-th attention head is successful.
[0055] Using the above method, the number of successful retrievals of each attention head can be summarized. Furthermore, the number of successful retrievals can be used to select a retrieval head from multiple attention heads. For example, the attention head with a number of successful retrievals greater than a threshold can be directly selected from each attention head as the retrieval head, and the remaining attention heads can be used as local heads.
[0056] In the embodiment of the present invention, the initial window of the search head can be extended to the full text of the text to be processed, that is, the initial window size of the search head can be extended to the preset length of the text to be processed. The initial window of the local head is not extended, that is, the initial window size of the local head remains unchanged. Compared with extending the initial window size of all attention heads, a large amount of computing and video memory resources can be saved, and video memory overhead is greatly reduced.
[0057] Here, the operation of classifying multiple attention heads into retrieval heads and local heads can be named as the split-head attention method. If there are 5 attention heads in total, 2 of which are retrieval heads and 3 are local heads, the split-head attention method is as follows: Figure 5 shown. Figure 5 In , each local head window includes 12 attention nodes and can process 4 words at the same time. Each search head window includes 18 attention nodes and can process 6 words at the same time.
[0058] In an embodiment of the present invention, starting from the receptive field length of each attention head in the multi-head attention mechanism, combined with the expansion of position encoding, different window size strategies are adopted for different attention heads, the retrieval head adopts the global attention mechanism, and the local head adopts the sliding window attention mechanism, thereby achieving an effect close to the global attention mechanism with a very small computing resource growth rate.
[0059] In addition, in the embodiment of the present invention, the calculation of the search head is universal, does not require dynamic calculation during reasoning, and is compatible with multi-head attention and group attention methods. The split-head attention method not only works on the sliding window attention mechanism, but also on the global attention mechanism. The split-head attention method not only extends the extrapolation length of the sliding window attention mechanism, but also brings very little video memory overhead.
[0060] On the basis of the above embodiment, the cache information corresponding to the retrieval header and the local header are respectively stored in different cache spaces, and the retrieval header and the local header perform dot product attention calculations respectively.
[0061] Specifically, in the conventional multi-head attention mechanism, since the calculation and cache strategies of all attention heads are consistent, they are usually spliced into a matrix for batch matrix multiplication, and the cache space (KV-Cache) is also allocated together. However, in the separate attention method, the window sizes of the retrieval head and the local head are different. The retrieval head needs to use a larger cache space and window size, and the retrieval head and the local head cannot be cached and calculated uniformly. Therefore, in the embodiment of the present invention, the cache information corresponding to the retrieval head and the local head are respectively stored in different cache spaces, and the retrieval head and the local head perform point multiplication attention calculations respectively, thereby improving the calculation efficiency.
[0062] like Figure 6 As shown, attention heads 1, 3, and 5 are retrieval heads, and attention heads 2 and 4 are local heads. In the cache space, the cache information corresponding to attention heads 1, 3, and 5 is stored in the same area, and the cache information corresponding to attention heads 2 and 4 is stored in the same area.
[0063] The input of the multi-head attention mechanism includes the query vector q, the key vector K, and the value vector V. The key vector K and the value vector V need to be cached, while the query vector q does not need to be cached, that is, the cache information corresponding to the attention heads 1, 2, 3, 4, and 5 all include the key vector K and the value vector V.
[0064] When performing attention calculation, the query vector q, key vector K and value vector V are divided into retrieval head group and local head group: q=[ , ],K=[ , ],V=[ , ].in f Indicates the search header group: , , , are the query vectors of attention heads 1, 3, and 5 respectively. are the key vectors of attention heads 1, 3, and 5 respectively, are the value vectors of attention heads 1, 3, and 5 respectively. l Represents a local header group: , , , are the query vectors of attention heads 2 and 4 respectively, are the key vectors of attention heads 2 and 4 respectively, are the value vectors of attention heads 2 and 4 respectively.
[0065] Calculate according to the dot product attention calculation method, that is, .in , is the attention calculation result of the i-th attention head. , a is the final calculation result of the split attention method. In the specific calculation, the calculation of the same group can be spliced into the same matrix for parallel calculation. Figure 6 shown.
[0066] Based on the above embodiment, selecting a search head from the multiple attention heads based on the number of successful searches includes: For any attention head, based on the number of successful retrievals of the attention head and the sequence length and number of the decoding results in the test set, calculate the retrieval score of the attention head; Based on the retrieval score of any of the attention heads, determine whether any of the attention heads is the retrieval head.
[0067] Specifically, in the embodiment of the present invention, the retrieval score of each attention head can be calculated according to the number of successful retrievals of each attention head. That is, in, Indicates the decoding result The sequence length is n, and O is the number of decoding results.
[0068] After that, the retrieval score of each attention head is compared with the score threshold. If the retrieval score of an attention head is greater than the score threshold, the attention head is determined to be a retrieval head, otherwise it is a local head. The score threshold can be set as needed, for example, it can be 0.5 or other values.
[0069] Taking the initial large language model including 40 layers of attention network as an example, each layer of attention network has a multi-head attention neural network with 40 attention heads. The above method is used to visualize the retrieval score, such as Figure 4 shown. Figure 4 The horizontal axis is the layer number (layer_id) of the attention network, and the vertical axis is the attention head number (head_id) in the attention network. The darker the color, the lower the score, and the brighter the color, the higher the score. The attention heads with retrieval scores greater than 0.5 account for about 5% of the total number of attention heads, that is, about 80 attention heads are identified as retrieval heads.
[0070] On the basis of the above embodiment, the initial window size of the sliding window attention mechanism based on the preset length and the initial large language model is used to expand the initial position encoding of the sliding window attention mechanism, including: Determining initial encoding parameters for each position in the initial position encoding; Calculating a target coding parameter based on a ratio of the preset length to the initial window size and the initial coding parameter; Based on the target coding parameters, the extended position coding is determined.
[0071] Specifically, if the position encoding method is a rotation position encoding method, the method uses a sine and cosine function to generate an initial position encoding, and then the initial position encoding can be expanded by modifying the period of the sine and cosine function.
[0072] The initial position encoding can be expressed as: ; in, , B is the initial encoding parameter, and n is the position of the word. The period of the initial position encoding, that is, the period of the sine and cosine functions, will increase as the dimension d increases.
[0073] Therefore, when the initial position encoding of the sliding window attention mechanism is expanded, the initial encoding parameters of each position in the initial position encoding can be determined first. The initial encoding parameter can be 10000 or other values, which are not specifically limited here.
[0074] Thereafter, the target coding parameter can be calculated using the ratio of the preset length to the initial window size and the initial coding parameter. For example, the target coding parameter can be expressed as: ; in, is the target encoding parameter, is the initial encoding parameter, L is the preset length, is the initial window size. S is much smaller than the sequence length of the training text, and L is larger than the sequence length of the training text.
[0075] When the preset length L is 3 times the initial window size S, .
[0076] Finally, the denominator of the expanded sine and cosine functions can be determined by the target encoding parameters: , and then the expanded position code can be determined. The expanded position code can be expressed as: .
[0077] By calculating in this way, the performance degradation problem caused by the conventional position interpolation method when the sliding window attention mechanism is extrapolated to a longer text can be solved.
[0078] Based on the above embodiment, the step of expanding the initial window size includes: The initial window sizes of the pre-filling phase and the decoding phase in the decoding process of the initial large language model are respectively expanded.
[0079] Specifically, the initial window size expansion of the sliding window attention mechanism can be divided into two steps: first, the initial window size of the prefill stage in the decoding process of the initial large language model is expanded. For example, its initial window size can be expanded to a preset length, and the results of the attention calculation are fully or partially cached to avoid information loss during subsequent word-by-word decoding. Then, the initial window size of the decoding stage in the decoding process is expanded, and the expanded window size should match the length of the information sequence cached in the prefill stage. When all the calculation results of the prefill stage are cached, the initial window size of the decoding stage will be expanded to the preset length. At this time, the sliding window attention mechanism is transformed into a global attention mechanism, and the long text processing capability of the initial large language model can be extrapolated to fully cover the preset length of reasoning.
[0080] On the basis of the above embodiment, the processing of the to-be-processed text based on the target large language model includes: Based on the target large language model, generating target position codes for the to-be-processed text segments in a fixed size according to the expanded position codes; Based on the target position code, the text to be processed is processed.
[0081] Specifically, in order to avoid the degradation of the processing performance of the target large language model on short texts due to the expansion operation, when generating the target position code, the text to be processed is segmented with a fixed size P, and multiple text segments such as [0, P], [P, 2P], ..., [L / / PP, L / / P+P] can be obtained. Among them, " / / " means integer division and downward evidence. The smaller P is, the finer the granularity of the segmentation is, and the smaller the processing performance loss for short texts is.
[0082] In summary, the large language model long text extrapolation method provided in the embodiment of the present invention greatly expands the processing capability of the large language model for long text with less resource consumption, thereby improving the efficiency and effect of long text processing.
[0083] like Figure 7 As shown, based on the above embodiment, an embodiment of the present invention provides a large language model long text extrapolation device, including: A text acquisition module 71 is used to acquire a text to be processed of a preset length; A size expansion module 72 is used for expanding the initial position encoding of the sliding window attention mechanism based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, and expanding the initial window size to obtain a target large language model if the preset length is greater than the sequence length of the training text of the initial large language model; The processing module 73 is used to process the text to be processed based on the target large language model.
[0084] On the basis of the above-mentioned embodiment, in the large language model long text extrapolation device provided in the embodiment of the present invention, the sliding window attention mechanism includes multiple attention heads; The size expansion module is specifically used for: Determine a test set, wherein the test set includes decoding results corresponding to a plurality of texts to be decoded; Based on the decoding result, determining the number of successful searches of each word in the to-be-decoded text by the multiple attention heads, selecting a search head from the multiple attention heads based on the number of successful searches, and determining a local head other than the search head from the multiple attention heads; The initial window size of the retrieval header is extended to the preset length, and the initial window size of the local header remains unchanged.
[0085] On the basis of the above-mentioned embodiment, in the large language model long text extrapolation device provided in the embodiment of the present invention, the cache information corresponding to the retrieval head and the local head are respectively stored in different cache spaces, and the retrieval head and the local head perform dot product attention calculation respectively.
[0086] On the basis of the above-mentioned embodiment, in the large language model long text extrapolation device provided in the embodiment of the present invention, the size expansion module is further specifically used for: For any attention head, based on the number of successful retrievals of the attention head and the sequence length and number of the decoding results in the test set, calculate the retrieval score of the attention head; Based on the retrieval score of any of the attention heads, determine whether any of the attention heads is the retrieval head.
[0087] On the basis of the above-mentioned embodiment, in the large language model long text extrapolation device provided in the embodiment of the present invention, the size expansion module is further specifically used for: Determining initial encoding parameters for each position in the initial position encoding; Calculating a target coding parameter based on a ratio of the preset length to the initial window size and the initial coding parameter; Based on the target coding parameters, the extended position coding is determined.
[0088] On the basis of the above-mentioned embodiment, in the large language model long text extrapolation device provided in the embodiment of the present invention, the size expansion module is further specifically used for: The initial window sizes of the pre-filling phase and the decoding phase in the decoding process of the initial large language model are respectively expanded.
[0089] On the basis of the above-mentioned embodiment, in the large language model long text extrapolation device provided in the embodiment of the present invention, the processing module is specifically used for: Based on the target large language model, generating target position codes for the to-be-processed text segments in a fixed size according to the expanded position codes; Based on the target position code, the text to be processed is processed.
[0090] Specifically, the functions of each module in the large language model long text extrapolation device provided in the embodiment of the present invention correspond one-to-one to the operation flow of each step in the above-mentioned method embodiment, and the effects achieved are also consistent. Please refer to the above-mentioned embodiment for details, and no further description will be given in the embodiment of the present invention.
[0091] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the large language model long text extrapolation method provided in the above embodiments.
[0092] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0093] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the large language model long text extrapolation method provided in the above embodiments.
[0094] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the large language model long text extrapolation method provided in the above embodiments.
[0095] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0096] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A long text extrapolation method for a large language model, characterized in that: include: Get the text to be processed of preset length; If the preset length is greater than the sequence length of the training text of the initial large language model, based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, the initial position encoding of the sliding window attention mechanism is expanded, and the initial window size is expanded to obtain a target large language model; The text to be processed is processed based on the target large language model.
2. The large language model long text extrapolation method according to claim 1, characterized in that: The sliding window attention mechanism includes multiple attention heads; The step of expanding the initial window size comprises: Determine a test set, wherein the test set includes decoding results corresponding to a plurality of texts to be decoded; Based on the decoding result, determining the number of successful searches of each word in the to-be-decoded text by the multiple attention heads, selecting a search head from the multiple attention heads based on the number of successful searches, and determining a local head other than the search head from the multiple attention heads; The initial window size of the retrieval header is extended to the preset length, and the initial window size of the local header remains unchanged.
3. The large language model long text extrapolation method according to claim 2, characterized in that: The cache information corresponding to the retrieval head and the local head are respectively stored in different cache spaces, and the retrieval head and the local head perform point product attention calculations respectively.
4. The large language model long text extrapolation method according to claim 2, characterized in that: The selecting a retrieval head from the plurality of attention heads based on the number of successful retrievals comprises: For any attention head, based on the number of successful retrievals of the attention head and the sequence length and number of the decoding results in the test set, calculate the retrieval score of the attention head; Based on the retrieval score of any of the attention heads, determine whether any of the attention heads is the retrieval head.
5. The large language model long text extrapolation method according to any one of claims 1 to 4, characterized in that: The initial position encoding of the sliding window attention mechanism is expanded based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, including: Determining initial encoding parameters for each position in the initial position encoding; Calculating a target coding parameter based on a ratio of the preset length to the initial window size and the initial coding parameter; Based on the target coding parameters, the extended position coding is determined.
6. The large language model long text extrapolation method according to any one of claims 1-4, characterized in that: The step of expanding the initial window size comprises: The initial window sizes of the pre-filling phase and the decoding phase in the decoding process of the initial large language model are respectively expanded.
7. The large language model long text extrapolation method according to any one of claims 1-4, characterized in that: The processing of the to-be-processed text based on the target large language model includes: Based on the target large language model, generating target position codes for the to-be-processed text segments in a fixed size according to the expanded position codes; Based on the target position code, the text to be processed is processed.
8. A large language model long text extrapolation device, characterized in that: include: A text acquisition module is used to acquire a text to be processed of a preset length; A size expansion module, configured to expand the initial position encoding of the sliding window attention mechanism based on the preset length and the initial window size of the sliding window attention mechanism of the initial large language model, and expand the initial window size to obtain a target large language model if the preset length is greater than the sequence length of the training text of the initial large language model; A processing module is used to process the text to be processed based on the target large language model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the large language model long text extrapolation method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the large language model long text extrapolation method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Data processing method and device, equipment and storage medium
CN120745871A