Long text processing method and device, equipment and medium
By integrating the state space model SSM and self-attention mechanism in the large model of Transformer architecture, the computational complexity can be effectively reduced and global semantic information can be captured when processing long text, solving the problem of semantic incompleteness in the existing technology during long text processing.
Patent Information
- Application Number
- CN202510031301.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
AI Technical Summary
When the existing Transformer architecture uses large models based on the Transformer architecture to process long text, the computational complexity of the self-attention mechanism increases square-level with the increase in sequence length, resulting in a sharp increase in resource demand and the inability to effectively capture the global semantic information of long text, resulting in incomplete or incoherent text modeling semantics.
A language processing model based on the fusion of state space model SSM and self-attention mechanism is adopted to perform word segmentation and prediction processing on long text, and the global attention of token is obtained through SSM, and local attention is obtained through self-attention mechanism. After the fusion is carried out to reduce the computational complexity and ensure semantic integrity.
While reducing the complexity of the model calculation, it can effectively capture the global semantic information of long text to ensure the semantic integrity and coherence of text modeling.
Smart Images

Figure CN119990136A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing, and in particular to a long text processing method, device, equipment and medium. Background Art
[0002] With the explosive growth of data and the significant improvement of computing power, natural language processing has entered the large model stage. Among the existing large model architectures, the Transformer architecture with self-attention mechanism as the core has demonstrated excellent performance in a variety of natural language processing (NLP) tasks since its introduction, relying on its powerful sequence modeling and parallel processing capabilities. It has now become the mainstream infrastructure for building large models.
[0003] Although the length of text sequences that can be processed by large models based on the Transformer architecture has increased compared to the Recurrent Neural Network (RNN) era, when faced with extremely long texts, the computational complexity of its self-attention mechanism increases quadratically with the increase in sequence length, and the required resource space also increases dramatically, which may even exceed the carrying capacity of hardware resources.
[0004] To address the above issues, existing technologies reduce the length or dimension of input text by text segmentation or text feature conversion to reduce the computational complexity of the self-attention mechanism. However, existing methods will cause large models to be unable to effectively capture the global semantic information of the input long text when predicting long texts, resulting in incomplete or incoherent text modeling semantics. Summary of the invention
[0005] The embodiments of the present application provide a long text processing method, apparatus, device and medium, so as to achieve the effect of reducing the computational complexity of the model while ensuring the semantic integrity and coherence of the model text modeling when modeling the long text.
[0006] In a first aspect, an embodiment of the present application provides a long text processing method, comprising:
[0007] Segment the long text to be processed to obtain a token sequence, wherein the token sequence includes multiple tokens;
[0008] The token sequence is predicted and processed using a pre-acquired language processing model to obtain a processed token sequence, wherein the processed token sequence includes the token sequence and at least one predicted token after the token sequence; the attention layer of the language processing model is obtained based on the fusion of a state space model SSM and a self-attention mechanism, wherein the SSM is used to obtain the global attention of the input token, and the self-attention mechanism is used to obtain the local attention of the input token.
[0009] In a possible implementation, the using a pre-acquired language processing model to perform prediction processing on the token sequence includes:
[0010] For the last token in the token sequence, obtaining the global attention of the token through the SSM part in the attention layer in the language processing model;
[0011] According to a preset window size, at least one token in the window where the token is located is obtained, and the self-attention mechanism in the attention layer is used for the at least one token to obtain local attention of the token;
[0012] Fusing the global attention and local attention of the token to obtain the final attention of the token;
[0013] Based on the final attention, the token sequence is predicted to obtain a predicted token, and the predicted token is placed at the last position of the token sequence to obtain a new token sequence. The above steps are repeated until the iteration termination condition is met to obtain the processed token sequence.
[0014] In a possible implementation, the iteration termination condition includes:
[0015] The predicted token obtained is the terminator;
[0016] or,
[0017] The number of characters in the new token sequence reaches the preset number.
[0018] In a possible implementation, before obtaining the global attention of the token, the method further includes:
[0019] Obtaining a vector corresponding to the token through the embedding layer of the language processing model;
[0020] Accordingly, obtaining the global attention of the token includes:
[0021] Based on the vector corresponding to the token, obtain the global attention of the token.
[0022] In a possible implementation, obtaining the global attention of the token based on the vector corresponding to the token includes:
[0023] Through the SSM in the language processing model, based on the vector corresponding to the token and the hidden state of the previous time step, the hidden state of the current time step is obtained, and based on the hidden state of the current time step, the global attention of the token is obtained.
[0024] In a possible implementation, fusing the global attention and the local attention of the token to obtain the final attention of the token includes:
[0025] Based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token.
[0026] In a possible implementation, the method further includes:
[0027] Based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token, including:
[0028] Based on the gating mechanism, a gating signal representing the weight ratio of global attention and local attention is obtained;
[0029] Based on the gating signal, the global attention and local attention of the token are weighted summed to obtain the final attention of the token.
[0030] In a second aspect, an embodiment of the present application provides a text processing device, including:
[0031] A first processing unit is used to segment the long text to be processed to obtain a token sequence, wherein the token sequence includes multiple tokens;
[0032] The second processing unit is used to use a pre-acquired language processing model to perform prediction processing on the token sequence to obtain a processed token sequence, wherein the processed token sequence includes the token sequence and at least one predicted token after the token sequence; the attention layer of the language processing model is obtained based on the fusion of a state space model SSM and a self-attention mechanism, wherein the SSM is used to obtain the global attention of the input token, and the self-attention mechanism is used to obtain the local attention of the input token.
[0033] In a possible implementation, the second processing unit includes:
[0034] A first acquisition module is used to acquire the global attention of the last token in the token sequence through the SSM part in the attention layer in the language processing model;
[0035] A second acquisition module is used to acquire at least one token in the window where the token is located according to a preset window size, and adopt the self-attention mechanism in the attention layer to the at least one token to acquire the local attention of the token;
[0036] A first processing module is used to fuse the global attention and local attention of the token to obtain the final attention of the token;
[0037] The second processing module is used to predict the token sequence based on the final attention to obtain a predicted token, put the predicted token into the last position of the token sequence to obtain a new token sequence, repeat the above steps until the iteration termination condition is met, and obtain the processed token sequence.
[0038] In a possible implementation manner, the iteration termination in the second processing module includes:
[0039] The predicted token obtained is the terminator;
[0040] or,
[0041] The number of characters in the new token sequence reaches the preset number.
[0042] In a possible implementation manner, the long text processing device further includes:
[0043] An acquisition unit, configured to acquire a vector corresponding to the token through an embedding layer of the language processing model;
[0044] Accordingly, the first processing module includes:
[0045] Based on the vector corresponding to the token, obtain the global attention of the token.
[0046] In a possible implementation manner, the first acquisition module is specifically configured to:
[0047] Through the SSM in the language processing model, based on the vector corresponding to the token and the hidden state of the previous time step, the hidden state of the current time step is obtained, and based on the hidden state of the current time step, the global attention of the token is obtained.
[0048] In a possible implementation manner, the first processing module is specifically configured to:
[0049] Based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token.
[0050] In a possible implementation manner, the first processing module is specifically configured to:
[0051] Based on the gating mechanism, a gating signal representing the weight ratio of global attention and local attention is obtained;
[0052] Based on the gating signal, the global attention and local attention of the token are weighted summed to obtain the final attention of the token.
[0053] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0054] The memory stores computer-executable instructions;
[0055] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible long text processing methods of the first aspect.
[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the first aspect above and / or various possible long text processing methods of the first aspect.
[0057] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the first aspect and / or various possible long text processing methods of the first aspect.
[0058] The long text processing method, apparatus, device and medium provided in the embodiments of the present application first segment the long text to be processed to obtain a token sequence, and then use a pre-acquired language processing model based on the fusion of a state space model SSM and a self-attention mechanism to predict the token sequence to obtain a processed token sequence, so as to achieve the purpose of obtaining the local attention of the input token by using the self-attention mechanism of a large model based on the Transformer architecture when processing the long text, so that the computational complexity of the model can be reduced; at the same time, the SSM is introduced into the large model based on the Transformer architecture, and the global attention of the input token is obtained by using the SSM, so that the model can capture the global semantic information of the long text, thereby ensuring the semantic integrity and coherence of the text modeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0060] Figure 1 A flowchart of a long text processing method provided in Example 1 of the present application;
[0061] Figure 2 A flowchart of a long text processing method provided in Example 2 of the present application;
[0062] Figure 3 A schematic diagram of the model architecture of a specific language processing model provided in Example 3 of the present application;
[0063] Figure 4 A schematic diagram of the principle of the long text processing method provided in Example 4 of the present application;
[0064] Figure 5 A schematic diagram of the structure of a long text processing device provided in Embodiment 5 of the present application;
[0065] Figure 6 A schematic diagram of the structure of a long text processing device provided in Example 6 of the present application;
[0066] Figure 7 A schematic diagram of the structure of the electronic device provided in this application.
[0067] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0068] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0069] First, the terms involved in this application are explained:
[0070] State Space Model: State Space Model (SSM) refers to a class of mathematical models widely used in time series analysis and signal processing. It captures the dynamic characteristics of input data by defining the evolution process of a hidden state and the relationship between the observed data and the hidden state.
[0071] Self-attention mechanism: The self-attention mechanism is a core component of the large model based on the Transformer architecture, which is used to capture the dependencies between positions in the input sequence.
[0072] Text modeling: Text modeling refers to the abstraction, representation and organization of text information such as semantics, grammar, structure, etc.
[0073] Token: In the field of artificial intelligence, token refers to "word unit", which is the smallest semantic unit that uses numbers to represent words in language models.
[0074] In order to better understand the technical solution provided by this application, the following is a detailed introduction to the background technology:
[0075] In the field of natural language processing, processing long texts has always been a complex and challenging task. In the traditional RNN era, when modeling long texts, it is difficult to use the RNN model to effectively process long-distance dependencies in long texts due to the problem of gradient vanishing or gradient exploding. Although the RNN model was later improved and a variety of RNN improved models (such as LSTM, GRU and other network structures) were proposed, the length of text that can be effectively processed is still limited to a shorter length (such as a few hundred characters). These improved models need to process in blocks when processing ultra-long texts, which limits their effect on understanding long texts; at the same time, because recurrent neural networks must process texts sequentially and cannot process texts in parallel, this also causes their processing speed to be relatively slow.
[0076] With the explosive growth of data and the significant improvement of computing power, natural language processing has entered the stage of large-scale pre-trained models (referred to as large models). Among them, the Transformer architecture with self-attention mechanism as the core has shown excellent performance in various NLP tasks since its introduction, relying on its powerful sequence modeling and parallel processing capabilities. It has now become the mainstream infrastructure for building large models. The length of text sequences that can be processed by the existing large models based on the Transformer architecture has been much larger than that of the RNN era, but with the increase of sequence length, the computational complexity of its self-attention mechanism also increases quadratically, and the required resource space also increases sharply, which may even exceed the carrying capacity of hardware resources.
[0077] In the prior art, the length or dimension of the text input into the large model based on the Transformer architecture is often reduced by processing the input text by methods such as text segmentation or text feature conversion, so as to alleviate the above-mentioned problems existing in the large model based on the Transformer architecture when processing long texts. However, this solution of the prior art cannot effectively capture the global semantic information of the input text sequence, and there may be problems with incomplete or incoherent semantics in the long text modeling of the model.
[0078] In response to the above technical problems, the inventors, while studying a method for processing long texts using a large model based on the Transformer architecture, found that a model obtained by fusing SSM with the model's own self-attention mechanism can use the SSM structure to capture global information when processing long texts. At the same time, the computational complexity can be reduced through the local self-attention mechanism.
[0079] Based on the above technical conception of the inventor, the long text processing method provided in the present application first obtains a token sequence by segmenting the long text to be processed, and then uses a pre-acquired language processing model based on the fusion of the state space model SSM and the self-attention mechanism to predict the token sequence to obtain a processed token sequence, so as to achieve the purpose of obtaining the local attention of the input token based on the self-attention mechanism when processing the long text, and obtaining the global attention of the input token based on the SSM structure, so as to achieve the effect of reducing the computational complexity of the model while ensuring the semantic integrity and coherence of the model text modeling.
[0080] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0081] Figure 1 A flowchart of a long text processing method provided in Example 1 of the present application is shown as follows: Figure 1 As shown, the method includes:
[0082] S101. Segment the long text to be processed to obtain a token sequence, wherein the token sequence includes multiple tokens.
[0083] In this step, in order to facilitate the extraction of text information, the long text to be processed needs to be preprocessed. Specifically, the continuous long text needs to be divided into multiple tokens according to certain rules and semantic units to obtain a token sequence corresponding to multiple tokens.
[0084] S102. Use the pre-acquired language processing model to perform prediction processing on the token sequence to obtain a processed token sequence, wherein the attention layer of the language processing model is obtained based on the fusion of the state space model SSM and the self-attention mechanism.
[0085] In this step, based on the multi-layer structure of the language processing model, the input token sequence is predicted and processed to obtain a processed token sequence, wherein the processed token sequence includes the token sequence and at least one predicted token after the token sequence.
[0086] Specifically, according to the SSM part of the self-attention layer of the language processing model, the global semantics of the token sequence input to the attention layer is captured to obtain the global attention of the last token in the token sequence;
[0087] A local token sequence including the last token is divided from the token sequence. According to the self-attention mechanism of the self-attention layer of the language processing model, the correlation between each token in the local token sequence and other tokens is calculated to obtain the local attention of the last token in the token sequence. Based on the global attention and local attention of the last token obtained, the model can predict the next token in the token sequence.
[0088] Based on this, the predicted next token is taken as the last bit of the token sequence to obtain a new token sequence. The language processing model self-attention layer continues to predict the next token of the new token sequence. The above operation is repeated until the iteration termination condition is met, and a processed token sequence including the token sequence and at least one predicted token after the token sequence is obtained.
[0089] In practical applications, in order to convert long text into a form that the model can process, after obtaining the token sequence, each token in the token sequence must be assigned a unique digital identifier (id), and a vocabulary table used to represent the mapping relationship between tokens and ids must be established. It should be understood that the token sequence input to the language processing model is actually the id corresponding to each token in the token sequence.
[0090] The long text processing method provided in this embodiment first performs word segmentation on the long text to be processed to obtain a token sequence, and then uses a pre-acquired language processing model based on the fusion of a state space model SSM and a self-attention mechanism to predict the token sequence to obtain a processed token sequence, so as to achieve the purpose of obtaining the local attention of the input token based on the self-attention mechanism when processing the long text, and at the same time, obtaining the global attention of the input token based on the SSM structure, so as to achieve the effect of reducing the computational complexity of the model while ensuring the semantic integrity and coherence of the model text modeling.
[0091] Figure 2 A flowchart of a long text processing method provided in Example 2 of the present application is shown as follows: Figure 2 As shown, the token sequence can be predicted in the following way:
[0092] S201. For the last token in the token sequence, obtain the global attention of the token through the SSM part in the attention layer in the language processing model.
[0093] In the specific implementation of this solution, the final attention is calculated for a token at each time step. When the next token in the token sequence is predicted for the first time, for each token in the token sequence, in the attention layer in the language processing model, starting from the first token in the token sequence, local attention and global attention are calculated for each token, and the final attention of each token is obtained based on the local attention and global attention of each token, until the final attention of the last token in the token sequence is calculated.
[0094] It is worth noting that after calculating the final attention of the last token in the token sequence, the token sequence can be predicted based on the final attention of the last token to obtain a predicted token;
[0095] Based on this, the predicted token is put into the last bit of the token sequence. After obtaining the new token sequence, the prediction of the next token of the new token sequence is continued. Specifically, in this prediction process, since the final attention of all tokens before the new last token in the new token sequence has been calculated, the local attention and global attention can be directly calculated for the new last token, and the final attention of the token can be obtained. Then, based on the final attention of the new last token, a predicted token is obtained. Repeat this step until the iteration termination condition is met to obtain the processed token sequence.
[0096] In this step, in order to predict the next token, it is necessary to obtain the global attention of the last token in the token sequence through the SSM part of the attention layer in the language processing model. The global attention of the last token in the token sequence is obtained based on the iterative calculation of the global attention of other tokens in the token sequence, which is the extraction of semantic information of all tokens in the token sequence.
[0097] In a specific implementation, before obtaining the global attention of the token, the vector corresponding to the token is obtained through the embedding layer of the language processing model. Accordingly, the method of obtaining the global attention of the token is specifically to obtain the global attention of the token based on the vector corresponding to the token.
[0098] In this implementation, the embedding layer of the text processing model obtains the vector corresponding to the token based on word embedding processing and position embedding processing; based on the vector corresponding to the token, the local attention and global attention of the token can be calculated.
[0099] Furthermore, based on the vector corresponding to the token, obtaining the global attention of the token may include:
[0100] Through the SSM in the language processing model, the hidden state of the current time step is obtained based on the vector corresponding to the token and the hidden state of the previous time step, and the global attention of the token is obtained based on the hidden state of the current time step.
[0101] Among them, one time step completes the calculation of the final attention of a token in the token sequence; the hidden state of the previous time step refers to the hidden state obtained when the global attention calculation is performed on the previous token of the currently processed token in the token sequence.
[0102] Specifically, the mathematical expression of SSM usually includes two equations:
[0103] (1) State update equation: describes how the hidden state evolves from the state of the previous time step to the current time step. This equation is usually related to the input data and the previous hidden state.
[0104]
[0105] In the above formula, is the hidden state at the current time step, is the hidden state of the previous time step, It is the vector corresponding to the current token.
[0106] (2) Observation equation: describes how to generate the observation output from the current hidden state. In the context of natural language processing, the observation output is usually the representation of the current token (i.e., the global attention of the current token).
[0107]
[0108] In the above formula, is the global attention of the current token, It is the mapping function from hidden state to observed output.
[0109] Based on this, based on the calculation of the global attention of the current token corresponding to the time step at each time step, the global attention of the last token in the token sequence can be obtained. The last token includes the last token in the original token sequence and the new last token in at least one subsequent new token sequence.
[0110] It should be understood that SSM uses linear calculation to process long text sequences. This step takes advantage of the low computational complexity of SSM and applies it to long text processing tasks, so that the computational burden of the model will not increase significantly due to the increase in sequence length. At the same time, it also enables the model to gradually accumulate and transmit global semantic information starting from the first token in the token sequence through SSM, so that the global attention of the last token can effectively capture the global dependencies in the long text to achieve the capture of global semantic information.
[0111] S202. According to a preset window size, obtain at least one token in the window where the token is located, and use the self-attention mechanism in the attention layer for the at least one token to obtain the local attention of the token.
[0112] In this step, in order to reduce the computational complexity of processing long texts based on the self-attention mechanism and realize the prediction of the next token, when performing attention calculation on the token sequence based on the self-attention mechanism, a local window mechanism is used to perform attention calculation on all tokens in the window where the last token in the token sequence is located.
[0113] Specifically, in each time step, the self-attention mechanism in the attention layer is used to obtain the local attention of the current token, which may include:
[0114] Step 1: According to the preset window size, obtain at least one token in the window where the current token is located.
[0115] In this step, all other tokens in the window where the current token is located are obtained according to the preset window size to calculate the local attention of the current token. The local attention of the current token represents the correlation information between the current token and other tokens in the window; it should be understood that the current token is located at the last position in the window; in the process of processing a long text, the window size is fixed, but the number of tokens obtained according to the window size is variable.
[0116] Exemplarily, the preset window size may be 1024. If the current token is the 1028th token in the token sequence, the tokens obtained are the tokens at positions 5 to 1028 in the token sequence. If the current token is the 28th token in the token sequence, the tokens obtained are the tokens at positions 1 to 28 in the token sequence.
[0117] The pre-set window size in this application can be determined according to actual conditions and is not limited by this application.
[0118] Step 2: Construct a matrix based on at least one token in the window where the current token is located and the vector corresponding to each token obtained in the embedding layer ,in, Is the number of tokens obtained, is the embedding dimension of each token.
[0119] Step 3: Based on the self-attention mechanism, three linear transformations are performed on the input matrix to obtain the query matrix, key matrix and value matrix corresponding to the matrix. The three matrices are recorded as , and . It can be specifically expressed by the following formula:
[0120]
[0121] In the above formula, , , is the learned parameter matrix.
[0122] Step 4: Calculate the dot product of Query and Key to get the attention score matrix, and then perform weighted summation on Value according to the attention score to get the output matrix, where the output matrix includes the local attention of each token; the local attention of each token integrates the correlation information between the token and other tokens. The output matrix can be expressed as follows:
[0123]
[0124] in,
[0125] is a scaling factor to prevent the dot product value from being too large.
[0126] The output matrix is a new sequence matrix with the same sequence length as the input matrix. According to the output matrix, the local attention of the current token can be obtained.
[0127] Based on this, based on the calculation of the local attention of the current token corresponding to the time step at each time step, the local attention of the last token in the token sequence can be obtained. The last token includes the last token in the original token sequence and the new last token in at least one subsequent new token sequence.
[0128] It should be understood that this step takes advantage of the self-attention mechanism to capture any dependencies between positions in the sequence without being affected by the position distance, and calculates the attention of the token. Considering that its computational complexity is , that is, as the length of the sequence increases, the amount of calculation increases quadratically. In order to reduce the computational complexity, a local window mechanism is introduced in this step, which only performs self-attention calculations on the tokens within the window to obtain the local attention of the token. This method greatly reduces the computational complexity while maintaining the efficient modeling capability of the self-attention mechanism, and limits the resource space usage to a certain range, making the model more suitable for processing long text sequences.
[0129] S203: Fuse the global attention and local attention of the token to obtain the final attention of the token.
[0130] In this step, in order to obtain the final attention of the token that captures both the global semantic information and the correlation information with other tokens in the window, the global attention and local attention of the token need to be fused. In order to predict the next token, the final attention of the last token in the token sequence needs to be obtained.
[0131] Optionally, the global attention and local attention of the token may be weighted and summed according to their respective pre-set weights to obtain the final attention of the token.
[0132] Preferably, the global attention and local attention of the token can be fused based on the gating mechanism to obtain the final attention of the token.
[0133] Specifically, in each time step, a gating mechanism is used to regulate the weight ratio of the global attention and local attention of the current token. Based on the determined weight ratio, the global attention and local attention of the current token obtained in the current time step are fused to obtain the final attention of the current token.
[0134] In a specific implementation, the following method can be used to obtain the final attention of the current token at each time step:
[0135] Step 1: Based on the gating mechanism, obtain a gating signal representing the weight ratio of global attention and local attention.
[0136] In this step, a small feedforward neural network is usually used to obtain the gating signal, which can be calculated using the following formula:
[0137]
[0138] In the above formula, represents the gating signal of the current time step, Is a nonlinear activation function (usually a sigmoid function); and are the weights and biases that need to be learned; Represents the global attention of the current token corresponding to the current time step; Represents the local attention of the current token corresponding to the current time step; Represents the concatenation operation of global attention and local attention.
[0139] Step 2: Based on the gating signal, the global attention and local attention of the token are weighted summed to obtain the final attention of the token.
[0140] In this step, the gate signal The value of is between 0 and 1, indicating the weight ratio of global attention and local attention. According to the weight ratio indicated by the gate signal, the global attention and local attention of the token are weighted and summed to obtain the final attention of the token. The final attention of the token can be calculated using the following formula:
[0141]
[0142] In the above formula, is the final attention of the current token corresponding to the current time step.
[0143] Based on this, based on the calculation of the final attention of the current token corresponding to the time step at each time step, the final attention of the last token in the token sequence can be obtained. The last token includes the last token in the original token sequence and the new last token in at least one subsequent new token sequence.
[0144] It should be understood that the gating mechanism changes dynamically during the entire long text processing process. The gating signal at each time step is are calculated based on the current global attention and local attention output. When it is close to 1, the model is more inclined to global information; When it is close to 0, the model relies more on local information. This step dynamically adjusts the weight ratio of global attention and local attention through the gating mechanism, so that the model can adaptively adjust the fusion mode of global information and local information when processing different parts of the input token sequence.
[0145] S204: Predict the token sequence based on the final attention to obtain a predicted token, put the predicted token into the last position of the token sequence to obtain a new token sequence, repeat the above steps until the iteration termination condition is met to obtain the processed token sequence.
[0146] In this step, in order to obtain the processed token sequence, it is necessary to perform iterative calculations for multiple time steps. Specifically, starting from the first token in the original token sequence, the final attention of each token is obtained through calculations for multiple time steps until the final attention of the last token in the original token sequence is obtained; then, based on the final attention of the last token in the original token sequence and the subsequent multi-layer structure of the model, the next token in the original token sequence is predicted to obtain a predicted token, and the predicted token is placed at the end of the original token sequence to obtain a new token sequence;
[0147] Since the final attentions of all tokens before the new last token in the new token sequence have been calculated during the iteration process, the calculation of the next time step can be directly carried out for the new last token, and the final attention of the last token in the new token sequence can be obtained. The next token of the new token sequence is predicted to obtain a predicted token and a new token sequence. The above steps are repeated for the new token sequence until the iteration termination condition is met to obtain the processed token sequence.
[0148] In a specific implementation, the iteration termination condition includes: the obtained predicted token is a terminator; or the number of characters in the new token sequence reaches a preset number.
[0149] Specifically, the terminator can be a preset character indicating the end of a sentence, such as a period, an exclamation mark, etc., which is not limited in this application. The preset number of characters can be determined according to actual conditions, which is not limited in this application.
[0150] The long text processing method provided in this embodiment obtains the global attention of the token through the SSM part in the attention layer in the language processing model for the last token in the token sequence; then obtains the local attention of the token according to the preset window size and the self-attention mechanism; then, the global attention and the local attention of the token are merged to obtain the final attention of the token; finally, the token sequence is predicted based on the final attention to obtain a predicted token, and the predicted token is placed in the last position of the token sequence to obtain a new token sequence, and the above steps are repeated until the iteration termination condition is met to obtain the means of the processed token sequence. This method organically combines the SSM with the self-attention mechanism to capture the global semantic information of the long text and effectively process the text processing effect of the local dependency, thereby achieving the technical effect of not only greatly improving the accuracy of long text processing but also effectively reducing the computational complexity. In addition, this embodiment also provides a dynamic weight adjustment method based on a gating mechanism. By calculating a gating signal at each time step, the model can adaptively adjust the weighted ratio between the output of the SSM and the self-attention mechanism, without excessive reliance on the output of the SSM or self-attention mechanism, thereby further improving the accuracy of long text processing.
[0151] Figure 3 This is a schematic diagram of the model structure of a specific language processing model provided in Example 3 of the present application, such as Figure 3 As shown, the architecture of the language processing model used in the specific long text processing method of this embodiment includes:
[0152] Embedding layer, normalization layer, attention layer, feedforward layer, residual connection layer and output layer. Among them, the attention layer includes SSM module, local window self-attention module and gating mechanism; the feedforward layer includes fully connected layer 1 and fully connected layer 2. Furthermore, the normalization layer, attention layer, feedforward layer and residual connection layer are in a multi-layer processing unit, which includes 12 sub-units, each of which includes a normalization layer, an attention layer, a feedforward layer and a residual connection layer.
[0153] Specifically, the functions of each functional layer of the language processing model are as follows:
[0154] Embedding layer 301: used to convert the input token sequence into an embedded representation (i.e., vector representation) of a fixed dimension. The embedding dimension is set to 768.
[0155] Multi-layer processing unit 302: For each sub-unit of the multi-layer processing unit, the following normalization layer, attention layer, feed-forward layer and residual connection layer are used to repeatedly process the token sequence.
[0156] Normalization layer 3021: Perform layer normalization on the output of the input embedding layer to stabilize the network training process. The calculation formula for layer normalization is:
[0157]
[0158] In the above formula, and Respectively represent the mean and standard deviation of the input token vector, are learnable parameters. Represents the vector of tokens input to the normalization layer; is a small constant used to prevent the denominator from being zero.
[0159] Attention layer 3022: used to obtain the final attention of the current token.
[0160] (1) SSM module in the attention layer: used to capture global attention, and the hidden state dimension is set to 128. The output of the SSM module represents the global information of the current token and all previous tokens.
[0161] (2) Local window self-attention module of the attention layer: used to capture the local attention of the current token within the window range. The window size is set to 1024 tokens.
[0162] (3) Gating mechanism in the attention layer: used to dynamically adjust the output weights of the SSM module and the local window self-attention module to obtain the final attention of the current token. The gating signal is generated by a feed-forward neural network consisting of a hidden layer with a hidden layer size of 1024 and a sigmoid function as the activation function.
[0163] Feedforward layer 3023: After the gating mechanism completes the weighting of the global attention and local attention of the last token in the token sequence, the final attention of the token is input to the feedforward layer. The feedforward layer is used to further integrate and transmit information of the input token sequence. It consists of two layers of fully connected networks, including:
[0164] (1) The first fully connected layer has an input dimension of 1024 (the same as the output dimension of the gating mechanism) and an output dimension of 4096. This higher dimension allows the model to capture more complex features when performing nonlinear transformations.
[0165] Optional, activation function: Use the ReLU activation function after the first fully connected layer to introduce non-linearity to the model and help capture complex patterns in the input data.
[0166] (2) The second fully connected layer: reduces the output dimension of the first layer from 4096 to 1024 to keep it consistent with the input of the subsequent processing units of the model.
[0167] Optionally, before the output of the feedforward layer, a Dropout mechanism is introduced to prevent overfitting and improve the generalization ability of the model. The Dropout ratio is 0.1.
[0168] Residual connection layer 3024: The output of the input embedding layer is added to the output of the feedforward layer to form a residual connection. The main purpose of the residual connection is to alleviate the gradient vanishing problem in deep networks and promote the effective propagation of information, thereby improving the effect of the model.
[0169] Output layer 303: used to generate a prediction for the next token of the token sequence, where the output layer is a linear layer with an output dimension of 1024.
[0170] It should be understood that, depending on the internal structure of the output layer, it can be used to process different downstream natural language processing tasks, such as text classification, machine translation, etc. This application does not limit the type of tasks processed by the output layer.
[0171] The values of the embedding dimension, input dimension, output dimension, hidden state dimension, creation size, Dropout ratio, hidden layer size, etc. in this embodiment are schematically given based on experience. In the specific application of the solution, the selection of these values is determined according to the actual situation, and this application does not impose specific restrictions on the selection of these values.
[0172] In specific applications, after initially building the initial model architecture of the above-mentioned language processing model, based on the acquired text prediction, the model is trained with "predicting the next token" as the target task, and the trained language processing model is obtained to implement the long text processing method in this solution.
[0173] The model architecture of the language processing model provided in this embodiment can implement the long text processing method in this solution after training. For the input token sequence, the model can obtain the local attention of the input token based on the self-attention mechanism, and obtain the global attention of the input token based on the SSM structure, so as to achieve the effect of reducing the computational complexity of the model while ensuring the semantic integrity and coherence of the model text modeling.
[0174] Furthermore, Figure 4 The schematic diagram of the principle of the long text processing method provided in the fourth embodiment of the present application is as follows: Figure 4As shown, the long text processing method provided by this embodiment first performs word segmentation on the long text to be processed to obtain a token sequence; then, the vector corresponding to the token is obtained through the embedding layer of the language processing model; then, based on the vector corresponding to the token, the SSM part is used to obtain the global attention of the token; based on the vector corresponding to the token, the self-attention mechanism is used to obtain the local attention of the token; then, based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token represented by the vector; finally, the token sequence is predicted based on the final attention of the last token in the token sequence to obtain a predicted token.
[0175] Figure 5 This is a schematic diagram of the structure of the long text processing device provided in Example 5 of the present application, such as Figure 5 As shown, the long text processing device 40 provided in this embodiment includes:
[0176] The first processing unit 401 is used to segment the long text to be processed to obtain a token sequence, wherein the token sequence includes multiple tokens;
[0177] The second processing unit 402 is used to use a pre-acquired language processing model to perform prediction processing on the token sequence to obtain a processed token sequence, wherein the processed token sequence includes the token sequence and at least one predicted token after the token sequence; the attention layer of the language processing model is obtained based on the fusion of the state space model SSM and the self-attention mechanism, the SSM is used to obtain the global attention of the input token, and the self-attention mechanism is used to obtain the local attention of the input token.
[0178] The long text processing device 40 provided in this embodiment can execute the long text processing method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.
[0179] Figure 6 This is a schematic diagram of the structure of the long text processing device provided in Example 6 of the present application, such as Figure 5 As shown, based on the above embodiment, the long text processing device 40 provided in this embodiment further includes:
[0180] The acquisition unit 403 is used to acquire the vector corresponding to the token through the embedding layer of the language processing model.
[0181] In a possible implementation, the second processing unit 402 includes:
[0182] A first acquisition module 4021 is used to acquire the global attention of the last token in the token sequence through the SSM part in the attention layer in the language processing model;
[0183] A second acquisition module 4022 is used to acquire at least one token in the window where the token is located according to a preset window size, and adopt the self-attention mechanism in the attention layer to the at least one token to acquire the local attention of the token;
[0184] A first processing module 4023 is used to fuse the global attention and the local attention of the token to obtain the final attention of the token;
[0185] The second processing module 4024 is used to perform prediction processing on the token sequence based on the final attention to obtain a predicted token, put the predicted token into the last position of the token sequence to obtain a new token sequence, repeat the above steps until the iteration termination condition is met, and obtain the processed token sequence.
[0186] In a possible implementation, the iteration termination in the second processing module 4024 includes:
[0187] The predicted token obtained is the terminator;
[0188] or,
[0189] The number of characters in the new token sequence reaches the preset number.
[0190] In a possible implementation manner, the first processing module 4023 is specifically configured to:
[0191] Based on the vector corresponding to the token, obtain the global attention of the token.
[0192] In a possible implementation manner, the first acquisition module 4021 is specifically configured to:
[0193] Through the SSM in the language processing model, based on the vector corresponding to the token and the hidden state of the previous time step, the hidden state of the current time step is obtained, and based on the hidden state of the current time step, the global attention of the token is obtained.
[0194] In a possible implementation manner, the first processing module 4023 is specifically configured to:
[0195] Based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token.
[0196] In a possible implementation manner, the first processing module 4023 is specifically configured to:
[0197] Based on the gating mechanism, a gating signal representing the weight ratio of global attention and local attention is obtained;
[0198] Based on the gating signal, the global attention and local attention of the token are weighted summed to obtain the final attention of the token.
[0199] The long text processing device 40 provided in this embodiment can execute the long text processing method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.
[0200] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.
[0201] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned long text processing method.
[0202] The specific implementation process of the processor 501 can be found in the above-mentioned long text processing method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0203] The present application also provides a computer program product, including a computer program, which implements the long text processing method when executed by a processor.
[0204] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above-mentioned long text processing method is implemented.
[0205] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A long text processing method, characterized in that: include: Segment the long text to be processed to obtain a token sequence, wherein the token sequence includes multiple tokens; The token sequence is predicted and processed using a pre-acquired language processing model to obtain a processed token sequence, wherein the processed token sequence includes the token sequence and at least one predicted token after the token sequence; the attention layer of the language processing model is obtained based on the fusion of a state space model SSM and a self-attention mechanism, wherein the SSM is used to obtain the global attention of the input token, and the self-attention mechanism is used to obtain the local attention of the input token.
2. The method according to claim 1, characterized in that The using a pre-acquired language processing model to perform prediction processing on the token sequence includes: For the last token in the token sequence, obtaining the global attention of the token through the SSM part in the attention layer in the language processing model; According to a preset window size, at least one token in the window where the token is located is obtained, and the self-attention mechanism in the attention layer is used for the at least one token to obtain local attention of the token; Fusing the global attention and local attention of the token to obtain the final attention of the token; Based on the final attention, the token sequence is predicted to obtain a predicted token, and the predicted token is placed at the last position of the token sequence to obtain a new token sequence. The above steps are repeated until the iteration termination condition is met to obtain the processed token sequence.
3. The method according to claim 2, characterized in that The iteration termination conditions include: The predicted token obtained is the terminator; or, The number of characters in the new token sequence reaches the preset number.
4. The method according to claim 2, characterized in that: Before obtaining the global attention of the token, the method further includes: Obtaining a vector corresponding to the token through the embedding layer of the language processing model; Accordingly, obtaining the global attention of the token includes: Based on the vector corresponding to the token, obtain the global attention of the token.
5. The method according to claim 3, characterized in that: The obtaining the global attention of the token based on the vector corresponding to the token includes: Through the SSM in the language processing model, based on the vector corresponding to the token and the hidden state of the previous time step, the hidden state of the current time step is obtained, and based on the hidden state of the current time step, the global attention of the token is obtained.
6. The method according to any one of claims 2 to 5, characterized in that: The global attention and local attention of the token are fused to obtain the final attention of the token, including: Based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token.
7. The method according to claim 6, characterized in that Based on the gating mechanism, the global attention and local attention of the token are fused to obtain the final attention of the token, including: Based on the gating mechanism, a gating signal representing the weight ratio of global attention and local attention is obtained; Based on the gating signal, the global attention and local attention of the token are weighted summed to obtain the final attention of the token.
8. A long text processing device, characterized in that: include: A first processing unit is used to segment the long text to be processed to obtain a token sequence, wherein the token sequence includes multiple tokens; The second processing unit is used to use a pre-acquired language processing model to perform prediction processing on the token sequence to obtain a processed token sequence, wherein the processed token sequence includes the token sequence and at least one predicted token after the token sequence; the attention layer of the language processing model is obtained based on the fusion of a state space model SSM and a self-attention mechanism, wherein the SSM is used to obtain the global attention of the input token, and the self-attention mechanism is used to obtain the local attention of the input token.
9. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the long text processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the long text processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Long text processing method and device based on large language model, equipment and storage medium
CN120705288A
Data processing method and device, electronic equipment and storage medium
CN122333392A
Data processing method and device, electronic equipment and storage medium
CN122333392B