Long text generation method and system based on layered sliding window mechanism
By building a Transformer model with a hierarchical sliding window mechanism, dynamically adjusting the attention range, solving the calculation complexity and memory usage problems of the traditional Transformer architecture in long-sequence text processing, and achieving efficient long-text processing.
Patent Information
- Application Number
- CN202510861607.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing long-sequence text, the traditional Transformer architecture cannot dynamically adapt to multi-scale dependencies in the text, resulting in high computational complexity and high memory usage.
A model containing N-layer Transformer blocks is constructed, where the window size of each layer increases layer by layer, combining layer normalization, causal self-attention and multi-layer perceptron, dynamically adjust the attention range through a layered sliding window mechanism, and optimize the model training using gradient cropping and cosine decay learning rate scheduling strategies.
It improves the computing efficiency of long text processing, reduces memory footprint, solves the computational complexity and memory footprint problems of the traditional Transformer architecture in long-sequence text processing, and enhances the model's expression and generalization capabilities.
Smart Images

Figure CN120354893A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of natural language processing and deep learning, and particularly to a long text generation method and system based on a hierarchical sliding window mechanism. Background Art
[0002] When dealing with long sequence texts, the traditional Transformer architecture usually relies on a fixed global attention range or a single context modeling strategy. This mode cannot dynamically adapt to the multi-scale characteristics of dependencies in the text, resulting in problems of high computational complexity and high memory occupancy when dealing with long-distance texts. When there is a complex narrative structure spanning multiple paragraphs in the sequence, the fixed-range attention mechanism of the traditional Transformer architecture either misses global associations due to too narrow a field of view or introduces a large amount of irrelevant calculations and memory overhead due to too wide a field of view. Summary of the Invention
[0003] The main purpose of this application is to provide a long text generation method and system based on a hierarchical sliding window mechanism, aiming to solve the problems of incomplete capture of long-distance dependencies and waste of computing resources caused by the fixed attention range of the traditional Transformer architecture.
[0004] To achieve the above purpose, this application provides a long text generation method based on a hierarchical sliding window mechanism. The method includes: constructing a Transformer model, where the Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks, and N is an integer greater than or equal to 2; training the Transformer model, and performing long text processing based on the trained Transformer model.
[0005] Optionally, the method further includes: determining the window size of each Transformer block using the following formula (1):
[0006] In the formula, represents the window size of the -th layer of Transformer blocks, represents a preset basic window size, represents the total number of layers of Transformer blocks in the Transformer model.
[0007] Optionally, each Transformer block includes layer normalization, causal self-attention, and a multi-layer perceptron.
[0008] Optionally, training the Transformer model includes: calculating the gradients of the Transformer model and performing gradient clipping, calculating and updating the parameters of the Transformer model based on the clipped gradients and a preset optimizer; dynamically adjusting the learning rate of the Transformer model based on a cosine decay learning rate scheduling strategy.
[0009] Optionally, calculating the gradients of the Transformer model and performing gradient clipping includes: performing forward propagation on the pre-trained long text data input to the Transformer model and calculating the gradients of the Transformer model; clipping the gradients based on a preset clipping threshold, such that the L2 norm of the clipped gradients is within a preset threshold range.
[0010] Optionally, calculating and updating the parameters of the Transformer model based on the clipped gradients and a preset optimizer includes: calculating the parameters of the Transformer model based on the clipped gradients; updating the parameters of the Transformer model based on a preset gradient accumulation strategy and a preset optimizer.
[0011] Optionally, updating the parameters of the Transformer model based on a preset gradient accumulation strategy and a preset optimizer includes: calculating the average gradient of a preset number of gradients and updating the parameters of the Transformer model based on the average gradient and a preset optimizer.
[0012] In addition, to achieve the above object, the present application further provides a long text generation device based on a hierarchical sliding window mechanism, including: a model construction module for constructing a Transformer model, the Transformer model including N Transformer blocks, wherein the window size used in the (i + 1)-th layer of the N Transformer blocks is greater than the window size used in the i-th layer of the N Transformer blocks; a model training module for training the Transformer model and performing long text processing based on the trained Transformer model.
[0013] A long text generation method and system based on a hierarchical sliding window mechanism proposed in this application constructs a Transformer model containing N layers of Transformer blocks, where the window size of each layer increases layer by layer, realizing dynamic adaptation to multi-scale dependency relationships in long texts. This not only improves the computational efficiency of the model when processing long-distance texts but also effectively reduces memory occupancy, solving the problems of high computational complexity and high memory occupancy rate faced by traditional Transformer architectures when processing long-sequence texts. Description of the Drawings
[0014] Figure 1 It is a flowchart of the long text generation method based on the hierarchical sliding window mechanism provided by an embodiment of this application; Figure 2 It is a block diagram of the structure of the long text generation system based on the hierarchical sliding window mechanism provided by an embodiment of this application; Figure 3 It is a schematic diagram of the structure of the long text generation device based on the hierarchical sliding window mechanism provided by an embodiment of this application.
[0015] The realization, functional features, and advantages of the purpose of this application will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0016] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0017] When processing long-sequence texts, traditional Transformer architectures usually rely on a fixed global attention range or a single context modeling strategy. This mode cannot dynamically adapt to the multi-scale characteristics of dependency relationships in texts (such as local detail dependencies and global structure dependencies), resulting in high computational complexity and high memory occupancy when dealing with long-distance dependencies. When there is a complex narrative structure spanning multiple paragraphs in the sequence, the fixed-range attention mechanism of the traditional Transformer architecture either misses global associations due to too narrow a field of view or introduces a large amount of irrelevant calculations and memory overhead due to too wide a field of view.
[0018] In addition, although existing sparsification or cyclic mechanisms attempt to alleviate the problem, they often face performance losses or scalability bottlenecks when capturing long-distance dependencies across different text scales.
[0019] For example, the Transformer architecture proposed by Vaswani et al. still faces significant computational and memory bottlenecks when dealing with long sequences. To alleviate this problem, Child et al. (2019) designed a sparse Transformer, reducing the computational complexity through a factored self-attention mechanism. However, this method may sacrifice model performance when capturing long-range dependencies. Dai et al. further proposed Transformer-XL, introducing a recurrence mechanism to maintain the context coherence of long sequences. Nevertheless, its computational efficiency is still restricted. Recently, Lin et al. and Li et al. respectively explored the application of sparse attention mechanisms in image processing and recommendation systems, but they did not specifically optimize for the multi-scale context modeling requirements in long text tasks.
[0020] To solve the above problems, this application provides a long text generation method and system based on a hierarchical sliding window mechanism. The following is a detailed introduction to the solution of this application.
[0021] Figure 1 For the flowchart of the long text generation method based on the hierarchical sliding window mechanism provided in the first embodiment of this application, refer to Figure 1 This long text generation method based on a hierarchical sliding window mechanism may include the following steps: S11. Construct a Transformer model. The Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks; S12. Train the Transformer model and perform long text processing based on the trained Transformer model.
[0022] It should be noted that when there is a complex narrative structure spanning multiple paragraphs in a long text sequence, due to the fixed-range attention mechanism of the traditional Transformer architecture, there may be a situation where global associations are missed due to too narrow a field of view, or a large amount of irrelevant calculations and memory overhead are introduced due to too wide a field of view.
[0023] Based on this, the hierarchical sliding window mechanism proposed in the embodiments of this application can use windows of different sizes in different layers of the Transformer model to capture multi-scale context information. Specifically, the embodiments of this application use smaller windows in the lower layers to capture detailed local context, while using larger windows in the higher layers to capture broader context information. The hierarchical sliding window mechanism proposed in the embodiments of this application is achieved by modifying the attention mechanism of the traditional model, ensuring that local and global information can be effectively integrated when processing long texts, while reducing the computational burden and memory usage.
[0024] In a specific implementation process, first, a Transformer model is constructed. The Transformer model includes N layers of Transformer blocks, and each Transformer block includes layer normalization, causal self-attention, and a multi-layer perceptron.
[0025] Specifically, taking any Transformer block as the target Transformer block, this embodiment will take the target Transformer block as an example for illustration. Layer normalization is applied at the start and end positions of the target Transformer block, and the following calculation steps are executed: First, calculate the mean of the input vector x: E[x] = mean(x); then calculate the variance of the input vector x: Var[x] = variance(x); secondly, perform the normalization operation: (x - E[x]) / sqrt(Var[x] + ε), where ε is a preset constant to prevent the denominator from being zero; finally, apply the learnable parameters: where, and are learnable parameters.
[0026] It can be understood that the input vector is the result of vectorizing the input long text data. Layer normalization is used to stabilize the input distribution and accelerate the convergence of the model. The steps of layer normalization can be expressed by the following formula (2): ; Furthermore, the traditional causal self-attention is improved by using the hierarchical sliding window mechanism proposed in this embodiment. First, perform a linear projection to calculate the query vector: Q = W_q 、the key vector: K = W_k and the value vector: V = W_v , where W_q, W_k, and W_v are learnable weight matrices; then, use the hierarchical sliding window mechanism in this example to determine the window size used by each layer of Transformer blocks, where the window size used by the (i + 1)-th layer of Transformer blocks is larger than the window size used by the i-th layer of Transformer blocks.
[0027] Specifically, taking the -th layer of Transformer blocks as an example, the window size of the -th layer of Transformer blocks can be determined by the following formula (1):
[0028] In the formula, represents the window size of the -th layer of Transformer blocks, denotes the preset base window size, represents the total number of layers of Transformer blocks in the Transformer model.
[0029] It should be noted that in this embodiment is determined by grid search on the validation set, ranges from [32, 128], and the implementer can set according to the specific implementation scenario.
[0030] Furthermore, calculate the attention dispersion within each window based on the window sizes used in each layer of Transformer blocks: , where is normalization, is the dimension of the key vector, T represents the target position, that is, any position within the window. Perform weighted summation over all positions within the window: O = A V, to obtain the output of the improved causal self-attention in this embodiment.
[0031] Furthermore, perform a linear transformation using a multi-layer perceptron: MLP(x) = W2 GELU(W1 + b1) + b2, where W1, W2, b1, b2 are learnable parameters, and GELU is the activation function, represents the output of the causal self-attention after layer normalization in the current Transformer block.
[0032] It should be noted that in this application, the window sizes of each layer are determined by presetting the base window size and the total number of Transformer blocks, enabling the model to automatically adjust the attention range according to feature representations at different levels, thereby more accurately capturing long-range dependencies in the text. At the same time, the Transformer block composed of layer normalization, causal self-attention, and multi-layer perceptron enhances the model's expressive ability and generalization ability.
[0033] Thus, a Transformer model based on a hierarchical sliding window mechanism is constructed, and then the constructed Transformer model needs to be trained.
[0034] In one embodiment, in step S12, training the Transformer model may specifically include: S121. Calculate the gradients of the Transformer model and perform gradient clipping, and calculate and update the parameters of the Transformer model based on the clipped gradients and a preset optimizer; S122. Dynamically adjust the learning rate of the Transformer model based on the cosine decay learning rate scheduling strategy.
[0035] It should be noted that during the training process, the embodiments of the present application adopt gradient clipping and cosine decay learning rate scheduling strategies, effectively avoiding the problems of gradient explosion and overfitting, and improving the stability and convergence speed of the model.
[0036] In the specific implementation process, first perform forward propagation on the pre-trained long text data input to the Transformer model, and calculate the gradient of the Transformer model.
[0037] Specifically, the forward propagation process of each Transformer block can be expressed as: x' = x + CausalSelfAttention(LayerNorm(x)); y = x' + MLP(LayerNorm(x')), where CausalSelfAttention(LayerNorm(x)) represents the output of causal self-attention after layer normalization, and MLP(LayerNorm(x') represents the result output by the multi-layer perceptron.
[0038] Furthermore, clip the gradient based on a preset clipping threshold: g_clipped = g min(1, threshold / ||g||), where g is the gradient, threshold is the preset clipping threshold, and ||g|| is the L2 norm of the gradient. It should be noted that the gradient clipping step set in this embodiment can be used to prevent gradient explosion and ensure the stability of training.
[0039] In one embodiment, in step S121, calculating and updating the parameters of the Transformer model based on the clipped gradient and a preset optimizer may specifically include: S1211. Calculate the parameters of the Transformer model based on the clipped gradient; S1212. Update the parameters of the Transformer model based on a preset gradient accumulation strategy and a preset optimizer.
[0040] In the specific implementation process, this embodiment first calculates the average gradient of a preset number of gradients, and updates the parameters of the Transformer model based on the average gradient and a preset optimizer.
[0041] Specifically, accumulate gradients of multiple small batches, that is, a preset number of gradients, to simulate a larger gradient size: g_accumulated = (1 / K , where K is the cumulative number of steps, i.e., the preset number, and g_i is the gradient of the i-th mini-batch.
[0042] Furthermore, in this embodiment, the AdamW optimizer is used as the preset optimizer for parameter update, and its update rule can be expressed as: ; ; ; ; ; where t is the time step, i.e., the current iteration number, and t starts counting from 1; represents the first moment estimate, i.e., the exponential moving average of the gradient; represents the second moment estimate, i.e., the exponential moving average of the squared gradient; represents the first moment estimate after bias correction; represents the second moment estimate after bias correction; represents the updated model parameters, i.e., the model parameters at the current iteration number; represents the model parameters before update, i.e., the model parameters at the previous iteration number; represents the gradient at the current iteration number; represents the square of each gradient at the current iteration number, which can be used to calculate the second moment; represents the first moment decay rate; represents the second moment decay rate; represents the learning rate; represents the weight decay coefficient; is a constant to prevent the denominator from being zero.
[0043] It should be noted that AdamW is a variant of Adam, which adds weight decay, thus better preventing the model from overfitting. In this embodiment, the AdamW optimizer is used as the preset optimizer, and the model parameters are updated by calculating the average value of the cumulative gradient, which can not only improve the training efficiency but also make the model more stable.
[0044] Furthermore, the cosine annealing learning rate scheduling strategy is used to dynamically adjust the learning rate of the Transformer model, and its expression can be as follows:
[0045]
[0046] where t is the current iteration number, is the warm-up iteration number, is the total iteration number, and are the maximum learning rate and the minimum learning rate, respectively.
[0047] It can be understood that the specific implementation of the cosine annealing learning rate scheduling strategy is as follows: First, set the initial learning rate, the minimum learning rate, and the total number of training epochs. Then, calculate the learning rate decay coefficient according to the current training epoch. Finally, multiply the initial learning rate by the learning rate decay coefficient to obtain the learning rate for the current epoch. Among them, the learning rate decay coefficient can be calculated through the cosine function.
[0048] It should be noted that in this embodiment, the cosine annealing learning rate scheduling strategy is adopted to dynamically adjust the learning rate of the Transformer model. The cosine annealing learning rate scheduling strategy is a commonly used learning rate adjustment method. It dynamically adjusts the learning rate according to the training epoch, making the learning rate larger in the initial stage of training. As the number of training epochs increases, the learning rate gradually decreases and finally stabilizes. Adopting this adjustment method in this embodiment can make the model converge quickly in the initial stage of training and be more stable in the later stage of training, thereby improving the performance of the model. In addition, in this embodiment, by presetting the gradient accumulation strategy and the preset optimizer to update the model parameters, the training efficiency and performance of the model can be further improved.
[0049] Therefore, by adopting the above embodiments, the long text generation method based on the hierarchical sliding window mechanism provided in this embodiment can effectively integrate local and global information when processing long texts, while reducing the computational burden and memory usage. In addition, by adopting the AdamW optimizer and the cosine annealing learning rate scheduling strategy, the training efficiency and performance of the model can be further improved. The long text generation method based on the hierarchical sliding window mechanism in the embodiments of this application can be used in the fields of natural language processing, text generation, machine translation, etc.
[0050] A long text generation method based on the hierarchical sliding window mechanism proposed in the embodiments of this application first constructs a Transformer model including N layers of Transformer blocks, where the window size of each layer increases layer by layer, realizing the dynamic adaptation to multi-scale dependency relationships in long texts, not only improving the computational efficiency of the model when processing long-distance texts, but also effectively reducing the memory occupancy, and solving the problems of high computational complexity and high memory occupancy rate faced by the traditional Transformer architecture when processing long sequence texts.
[0051] Based on the above embodiments, Figure 2 is a structural block diagram of a long text generation system based on the hierarchical sliding window mechanism according to an embodiment of this application. As Figure 2 shown, the long text generation system 200 based on the hierarchical sliding window mechanism may include: a model construction module 210 and a model training module 220, where The model construction module 210 is used to construct a Transformer model, which includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks; The model training module 220 is used to train the Transformer model and perform long text processing based on the trained Transformer model.
[0052] In an exemplary embodiment, the model construction module 210 can also be used to determine the window size of each Transformer block using the following formula (1):
[0053] In the formula, represents the window size of the -th layer of Transformer blocks, represents a preset basic window size, represents the total number of layers of Transformer blocks in the Transformer model.
[0054] In an exemplary embodiment, each Transformer block in the system includes layer normalization, causal self-attention, and a multi-layer perceptron.
[0055] In an exemplary embodiment, the model training module 220 can also be used to calculate the gradients of the Transformer model and perform gradient clipping, calculate and update the parameters of the Transformer model based on the clipped gradients and a preset optimizer; dynamically adjust the learning rate of the Transformer model based on the cosine decay learning rate scheduling strategy.
[0056] In an exemplary embodiment, the model training module 220 can also be used to perform forward propagation on the pre-trained long text data input to the Transformer model and calculate the gradients of the Transformer model; clip the gradients based on a preset clipping threshold, and the L2 norm of the clipped gradients is within a preset threshold range.
[0057] In an exemplary embodiment, the model training module 220 can also be used to calculate the parameters of the Transformer model based on the clipped gradients; update the parameters of the Transformer model based on a preset gradient accumulation strategy and a preset optimizer.
[0058] In an exemplary embodiment, the model training module 220 may further be configured to calculate the average gradient of a preset number of gradients, and update the parameters of the Transformer model based on the average gradient and a preset optimizer.
[0059] Those skilled in the art should understand that the division of each module in the embodiment is only a logical function division. In actual application, they can be fully or partially integrated into one or more actual carriers, and these modules can all be implemented in the form of software called by a processing module, or all be implemented in the form of hardware, or be implemented in the form of a combination of software and hardware. It should be noted that each module in a long text generation device based on a hierarchical sliding window mechanism in this embodiment corresponds one by one to each step in a long text generation method based on a hierarchical sliding window mechanism in the foregoing embodiment. Therefore, the specific implementation manner of this embodiment can refer to the implementation manner of the foregoing long text generation method based on a hierarchical sliding window mechanism, which will not be elaborated here.
[0060] Based on the above embodiments, Figure 3 FIG. is a schematic structural diagram of a long text generation device based on a hierarchical sliding window mechanism according to an embodiment of the present application. As Figure 3 shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340. The processor 310 may call logical instructions in the memory 330 to execute a long text generation method based on a hierarchical sliding window mechanism, and the method includes: constructing a Transformer model, where the Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks, and N is an integer greater than or equal to 2; training the Transformer model, and performing long text processing based on the trained Transformer model.
[0061] In addition, when the logical instructions in the above-mentioned memory 330 can be implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0062] On the basis of the above-mentioned embodiments, on the other hand, the present invention further provides a computer program product. This computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a long text generation method based on a hierarchical sliding window mechanism provided by the above-mentioned various methods. The method includes: constructing a Transformer model, where the Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks, and N is an integer greater than or equal to 2; training the Transformer model, and performing long text processing based on the trained Transformer model.
[0063] On the basis of the above-mentioned embodiments, on the one hand, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute a long text generation method based on a hierarchical sliding window mechanism provided by the above-mentioned various methods. The method includes: constructing a Transformer model, where the Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks, and N is an integer greater than or equal to 2; training the Transformer model, and performing long text processing based on the trained Transformer model.
[0064] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A long text generation method based on a hierarchical sliding window mechanism, characterized in that The method includes: Construct a Transformer model, where the Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks, and N is an integer greater than or equal to 2; Train the Transformer model and perform long text processing based on the trained Transformer model.
2. The method according to claim 1, wherein The method further includes: Use the following formula (1) to determine the window size of each of the Transformer blocks: In the formula, represents the window size of the th layer Transformer block, represents the preset base window size, represents the total number of layers of Transformer blocks in the Transformer model.
3. The method according to claim 1, characterized in that Each of the Transformer blocks includes layer normalization, causal self-attention, and a multi-layer perceptron.
4. The method according to claim 1, characterized in that, The training of the Transformer model includes: Calculate the gradient of the Transformer model and perform gradient clipping, and calculate and update the parameters of the Transformer model based on the clipped gradient and a preset optimizer; Dynamically adjust the learning rate of the Transformer model based on a cosine decay learning rate scheduling strategy.
5. The method according to claim 4, wherein The calculation of the gradient of the Transformer model and performing gradient clipping includes: Perform forward propagation on the pre-trained long text data input to the Transformer model and calculate the gradient of the Transformer model; Clip the gradient based on a preset clipping threshold, and the L2 norm of the clipped gradient is within a preset threshold range.
6. The method according to claim 5, characterized in that, The calculation and update of the parameters of the Transformer model based on the clipped gradient and a preset optimizer includes: Calculate the parameters of the Transformer model based on the clipped gradient; Update the parameters of the Transformer model based on a preset gradient accumulation strategy and a preset optimizer.
7. The method according to claim 6, characterized in that, The update of the parameters of the Transformer model based on a preset gradient accumulation strategy and a preset optimizer includes: Calculate the average gradient of a preset number of gradients and update the parameters of the Transformer model based on the average gradient and a preset optimizer.
8. A long text generation system based on a hierarchical sliding window mechanism, characterized in that, Includes: A model construction module for constructing a Transformer model, where the Transformer model includes N layers of Transformer blocks, and the window size used in the (i + 1)-th layer of the N layers of Transformer blocks is larger than the window size used in the i-th layer of the N layers of Transformer blocks; A model training module for training the Transformer model and performing long text processing based on the trained Transformer model.
9. The system according to claim 8, wherein The model construction module is further used to determine the window size of each of the Transformer blocks using the following formula (1): Wherein, represents the window size of the layer of Transformer blocks, represents the preset basic window size, represents the total number of layers of Transformer blocks in the Transformer model.
10. The system according to claim 8, wherein The model training module is further configured to calculate the gradients of the Transformer model and perform gradient clipping, calculate and update the parameters of the Transformer model based on the clipped gradients and a preset optimizer; dynamically adjust the learning rate of the Transformer model based on a cosine decay learning rate scheduling strategy.
Citation Information
Patent Citations
Long text processing method, related equipment and readable storage medium
CN112527992A