Large model optimization method and system based on sequence-dependent hierarchical propagation mechanism

By introducing a sequence-dependent hierarchical propagation mechanism and optimizing the Transformer structure, the problems of high computational complexity and large memory usage of the self-attention mechanism of the large language model are solved, and more efficient long-sequence data processing is achieved.

CN120066802AActive Publication Date: 2025-05-30SHANDONG UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510541168.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The high computational complexity and large memory usage of existing large language models limit their application and promotion in long-sequence data modeling scenarios.

Method used

A sequence dependency hierarchical propagation mechanism is introduced, the sequence is divided into multiple blocks, and the sequence dependencies are effectively learned within and between these blocks. By sharing Key and Value matrices and optimizing the Transformer stacking layer structure, the computational complexity and memory usage are reduced.

Benefits of technology

It significantly reduces the computational complexity and memory usage of the traditional self-attention mechanism, and improves the efficiency and performance of the model when processing long sequence data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066802A_ABST
    Figure CN120066802A_ABST
Patent Text Reader

Abstract

The invention provides a large model optimization method and system based on a sequence-dependent hierarchical propagation mechanism, and relates to the technical field of large language models.The large model comprises a representation learning module and a downstream task module, and the representation learning module comprises the specific steps that natural language processing is conducted on an input text to obtain a word vector sequence; the method comprises the following steps: uniformly partitioning a word vector sequence based on a sequence-dependent hierarchical propagation mechanism, and generating a Key matrix and a Value matrix by respectively using attention mechanisms between fragments and in the fragments; and carrying out attention mechanism calculation of Transform by utilizing a Key matrix and a Value matrix to obtain a final sequence representation, and taking the final sequence representation as the input of a downstream task module. According to the method, a sequence dependency hierarchy propagation mechanism is introduced, the sequence is divided into a plurality of blocks, and the sequence dependency relationship is effectively learned in the blocks and among the blocks, so that the calculation complexity of a traditional self-attention mechanism is remarkably reduced, and the memory occupation is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of large language models, and specifically relates to an optimization method and system for large models based on a sequence-dependent hierarchical propagation mechanism. Background Art

[0002] As a core breakthrough in the field of artificial intelligence, the Large Language Model (LLM) is profoundly reconstructing the knowledge processing paradigms in fields such as healthcare, finance, education, and scientific research. Through capabilities such as intelligent dialogue, knowledge reasoning, and content generation, the LLM has significantly improved industry efficiency, spawned innovative application scenarios, and simultaneously promoted the evolution of the human-machine collaboration model. The LLM generally includes a representation learning module and a downstream task module. The representation learning module is used to perform vector representation on the text input by the user, and the downstream task module takes the vector representation of the text as input and performs specific downstream tasks such as text classification, text generation, time series classification, and time series regression prediction. The core model architecture of the existing representation learning module is Transformer, and its Self-Attention Mechanism (SA) breaks through the length limitation of traditional sequence models. Through parallel computing and dynamic weight allocation, the model can accurately capture long-distance semantic associations. This technology provides the underlying support for the birth of epoch-making models such as GPT and BERT, and its high scalability has directly promoted the leapfrog development of model parameters from the hundreds of millions to the trillions, laying the technical foundation for the intelligent leap of large language models.

[0003] The Transformer model effectively captures long-distance dependencies through the SA mechanism and supports efficient parallel computing, thus overcoming the gradient vanishing problem of traditional recurrent neural networks when processing long sequences. However, the main defect of the Transformer model lies in the computational complexity of its self-attention mechanism being , where L is the sequence length, and its high memory occupancy results in too high a computational cost for the LLM, limiting its further application and promotion.

[0004] Existing improved models based on Transformer, such as Reformer, Flowformer, and Informer, although reducing the computational complexity of LLM to a certain extent and improving efficiency, still have obvious defects. These improved models usually optimize the self-attention mechanism by introducing complex auxiliary calculation processes or additional operations (such as locality-sensitive hashing, network flow calculation, probabilistic sparse attention, etc.). However, these improvements often bring additional computational overhead, larger model parameters, or lower training efficiency, limiting their universality and scalability in practical applications. In the scenario of long sequence data modeling, existing optimization methods still fail to fully solve the problems of high memory occupancy and high computational complexity of large models, and there is still room for further improvement. Summary of the Invention

[0005] To solve the above problems, this application proposes an optimization method and system for large models based on a sequence-dependent hierarchical propagation mechanism. By introducing the sequence-dependent hierarchical propagation mechanism, the sequence is divided into multiple blocks, and the sequence-dependent relationships are effectively learned within and between these blocks, thus significantly reducing the computational complexity of the traditional self-attention mechanism and effectively reducing memory occupancy.

[0006] According to some embodiments, this application adopts the following technical solutions: An optimization method for large models based on a sequence-dependent hierarchical propagation mechanism, the large model includes a representation learning module and a downstream task module, and the specific steps of the representation learning module are: Perform natural language processing on the input text to obtain a sequence of word vectors; Based on the sequence-dependent hierarchical propagation mechanism, evenly divide the sequence of word vectors into blocks, and use the attention mechanisms between and within segments to generate matrix and matrix; Use matrix and matrix to perform the attention mechanism calculation of Transformer to obtain the final sequence representation, which is used as the input of the downstream task module for specific downstream tasks.

[0007] According to some embodiments, this application adopts the following technical solutions: An optimization system for large models based on a sequence-dependent hierarchical propagation mechanism, the large model includes a representation learning module and a downstream task module, characterized in that the representation learning module includes: A language processing module configured to perform natural language processing on the input text to obtain a sequence of word vectors; Hierarchical Dependency Module, configured to: Based on the sequence dependency hierarchical propagation mechanism, evenly divide the word vector sequence into chunks, and respectively use the attention mechanisms between and within the segments to generate matrix and matrix; Sequence Representation Module, configured to: Utilize matrix and matrix to perform the attention mechanism calculation of the Transformer, obtain the final sequence representation, and use it as the input of the downstream task module for performing specific downstream tasks.

[0008] According to some embodiments, the present application adopts the following technical solutions: A computer program product, including a computer program, which when executed by a processor implements the described method for optimizing a large model based on the sequence dependency hierarchical propagation mechanism.

[0009] According to some embodiments, the present application adopts the following technical solutions: A non-transitory computer-readable storage medium, which is used to store computer instructions, and when the computer instructions are executed by a processor, the described method for optimizing a large model based on the sequence dependency hierarchical propagation mechanism is implemented.

[0010] According to some embodiments, the present application adopts the following technical solutions: An electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the described method for optimizing a large model based on the sequence dependency hierarchical propagation mechanism.

[0011] Compared with the prior art, the beneficial effects of the present application are: The present invention introduces a sequence dependency hierarchical propagation mechanism, divides the sequence into multiple chunks, and effectively learns the sequence dependency relationships within and between these chunks, thereby significantly reducing the computational complexity of the traditional self-attention mechanism and effectively reducing the memory occupancy. In addition, the present invention further reduces the model parameter scale and significantly improves the efficiency and performance of the model in processing long sequence data by sharing the Key and Value matrices and optimizing the Transfomer stacking layer structure.

[0012] The present invention utilizes the sequence dependency hierarchical propagation mechanism to ensure that in the generated matrix, the representation of each segment incorporates the dependency information of the entire sequence, reducing the information loss caused by the generation operation; in addition, the time complexities of the attention mechanisms between and within the segments are respectively and , the time complexity of the final sequence-dependent hierarchical propagation mechanism is , compared with the traditional Transformer based on the sequence-dependent hierarchical propagation mechanism, its computational complexity is significantly reduced, and the computational complexity of the attention mechanism is reduced from to , which means that the hierarchical propagation mechanism has advantages in terms of efficiency and resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application.

[0014] Figure 1 is a flowchart of the method for Embodiment 1.

[0015] Figure 2 is a detailed diagram of the method for Embodiment 1.

[0016] Figure 3 is a flowchart of the sequence-dependent hierarchical propagation mechanism for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The present application will be further described below in conjunction with the accompanying drawings and embodiments.

[0018] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0019] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0020] Embodiment 1 An embodiment of the present application, relying on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of General Artificial Intelligence in Shandong Province's institutions of higher learning, provides an optimization method for large models based on a sequence-dependent hierarchical propagation mechanism. The large model includes a representation learning module and a downstream task module. The specific steps of the representation learning module are as follows: Step S1: Perform natural language processing on the input text to obtain a sequence of word vectors; Step S2: Based on the sequence-dependency hierarchical propagation mechanism, evenly divide the word vector sequence, and use the inter-segment and intra-segment attention mechanisms respectively to generate matrix and matrix; Step S3: Utilize the matrix and matrix to perform the attention mechanism calculation of the Transformer, obtain the final sequence representation, and use it as the input of the downstream task module for specific downstream tasks.

[0021] As an embodiment, an optimization method for large models based on the sequence-dependency hierarchical propagation mechanism of the present application divides the sequence into multiple blocks by introducing the sequence-dependency hierarchical propagation mechanism, and effectively learns the sequence-dependency relationships within and between these blocks, thereby significantly reducing the computational complexity of the traditional self-attention mechanism and effectively reducing the memory occupancy; Figure 1 is the flowchart of the large model optimization method, Figure 2 is the detailed diagram of the large model optimization method. As shown in Figure 1 , 2 shown, the specific steps of the large model optimization method are as follows: Step 1: Obtain the word vector sequence through steps such as word segmentation, vocabulary construction, and word embedding.

[0022] Specifically, first, use the subword decomposition algorithm (Byte Pair Encoding, BPE) to perform word segmentation on the original text to improve the generalization ability of the model for low-frequency words and out-of-vocabulary words. Secondly, based on the word segmentation results, construct a vocabulary. The size of the vocabulary is usually set according to task requirements and computing resources to balance the representational ability and computational overhead. Then, in order to enable the model to understand the semantic information of words, perform word embedding on each subword in the vocabulary and map it to a high-dimensional continuous space, thereby providing rich semantic information for subsequent deep learning models. Finally, process the text into a sequence according to the vocabulary.

[0023] Step 2: Utilize the sequence-dependency hierarchical propagation mechanism to evenly divide the representation of the sequence, and then use the inter-segment and intra-segment attention mechanisms respectively to generate the Key matrix and Value matrix for calculating the attention scores of the Transformer layer. Specifically: 1. Perform high-dimensional embedding and positional encoding on the word vector sequence to obtain an embedding representation containing positional information.

[0024] For the word vector sequence, in this embodiment, it is represented by the matrix , where represents the length of the sequence, that is, the number of word vectors , Indicates the number of features of the word vector. To obtain rich sequence dependency information from sequence data, first perform embedding through a linear layer to embed the original data into a high-dimensional space; then apply positional encoding to the embedded data to preserve the order information of the sequence; finally, obtain the embedded representation containing positional information.

[0025] Specifically, given a sequence sample of word vectors , embed it into a new space and perform positional encoding to obtain the embedded representation , which is expressed by the formula:

[0026] where and are the parameters of the learnable embedding layer, represents the positional encoding operation, and its formal definition is as follows:

[0027]

[0028] where represents the position of the th feature vector in the sequence. , represent the odd and even positions of the feature respectively.

[0029] 2. Generate the Key matrix and the Value matrix based on the embedded representation containing positional information.

[0030] Using the sequence dependency hierarchical propagation mechanism, generate the compressed matrix and the matrix for subsequent attention calculation to reduce the complexity of the self-attention mechanism. As shown in Figure 3 , the specific steps are as follows: (1) Pad the embedded representation containing positional information to adapt its length to the calculation requirements; then evenly divide the padded sequence into blocks and perform mean pooling on each block to obtain the preliminary segment-level representation.

[0031] Specifically, perform the padding-block operation on the entire embedded representation , then perform mean pooling on each segment, and use the pooled vector as the representation of each sequence segment. Finally, obtain the preliminary segment-level representation , which is expressed by the formula:

[0032]

[0033]

[0034]

[0035] Among them, Chunk, Pad, mean are the sequence uniform slicing operation, padding operation, and average pooling operation respectively, indicating that the embedding representation is divided into segments, representing the i-th sequence segment representation, representing the floor operation.

[0036] (2) Use the inter-segment attention mechanism and intra-segment attention mechanism on the preliminary segment-level representations respectively to obtain the final inter-segment representation and the fine-grained intra-segment representation.

[0037] In order to make the generated and matrices fully reflect the sequence dependency representation of the entire sequence, this embodiment uses a hierarchical propagation mechanism to process the sequence dependency information outside and inside the segments. Among them, using the inter-segment attention mechanism, the inter-segment dependency relationship is obtained, and by calculating the attention distribution between different segments, the final segment-level representation is obtained, which is expressed by the formula:

[0038]

[0039] Among them, represents the preliminary segment-level representation, represents the softmax function, the vector x represents the input of the softmax function, and x i …x h represents the i-th to h-th elements of x.

[0040] Use the intra-segment attention mechanism to obtain the intra-segment dependency relationship, and by calculating the attention matrix within each segment, the fine-grained intra-segment representation is obtained, which is expressed by the formula:

[0041]

[0042] (3) Concatenate the inter-segment representation and the intra-segment representation to obtain the Key matrix and the Value matrix, which is expressed by the formula:

[0043]

[0044]

[0045] Among them, are the Value matrix and the Key matrix respectively, represents a single linear layer, are all the parameters of this linear layer, represents the tensor concatenation operation.

[0046] Step 3: In the Transformer, use the Key matrix and the Value matrix to perform low-complexity attention mechanism calculation, and share the Key matrix and the Value matrix between layers, and finally obtain the representation of the sequence.

[0047] Adopt the generated and matrices to perform multi-head self-attention mechanism calculation in the stacked layers of the Transformer, and share the Key matrix and the Value matrix in each stacked layer; in addition, use a linear layer to replace the original Position-wise Feed-Forward layer in the Transformer, further reducing the parameter scale; this method adopts the multi-head self-attention mechanism, each head independently calculates the Query matrix in the attention mechanism, and combines the Key matrix and the Value matrix generated in Step 2 to obtain the attention representation; the results of multiple attention heads are concatenated and embedded using a linear layer to generate the final context representation, that is, the final sequence representation. In this way, the long-range dependencies in the sequence can be efficiently modeled.

[0048] Specifically, in the first stacked layer of the Transformer, the multi-head self-attention mechanism is defined as follows (the other stacked layers are similar):

[0049]

[0050]

[0051] Among them, is the number of heads of the multi-head self-attention, i represents the i-th head, is the attention score of the i-th head, represents the softmax function, is the number of dimensions of the feature, is the input of the first stacked layer of the Transformer, , are the Value matrix and the Key matrix respectively, represents the multi-head self-attention, is the learnable parameter of the multi-head self-attention mechanism, represents the tensor concatenation operation, is the output of the multi-head self-attention.

[0052] The remaining content of the Transformer stack layer is as follows:

[0053]

[0054]

[0055]

[0056] Among them, and are model parameters, represents the layer normalization operation, where and respectively represent the mean and variance, and respectively represent the parameter vectors for scaling and translation, and is the minimum amount to prevent division-by-zero errors. As the input of the next stack layer, it continues to participate in the calculation. After stack layers, the final sequence representation is obtained.

[0057] This method reduces the scale of model parameters through the inter-layer sharing mechanism and parameter sharing strategy. Between multiple block layers of the Transformer, it shares the parameters of the compressed Key matrix and Value matrix to reduce the computational amount; it uses a linear layer to replace the feed-forward network layer in the traditional Transformer to further reduce the computational cost; these optimization measures improve the scalability of the Transformer in large-scale data processing.

[0058] Step Four: Input the sequence representation into the downstream task module to perform downstream tasks.

[0059] In the downstream task module, the input sequence representation is used to perform specific downstream tasks, such as text classification, text generation, time series classification, time series regression prediction, etc. The downstream task layer generally consists of several linear layers.

[0060] Embodiment 2 In an embodiment of the present application, an optimization system for a large model based on a sequence-dependent hierarchical propagation mechanism is provided. The large model includes a representation learning module and a downstream task module, and is characterized in that the representation learning module includes: A language processing module, configured to: perform natural language processing on the input text to obtain a sequence of word vectors; A hierarchical dependency module, configured to: based on a sequence dependency hierarchical propagation mechanism, evenly divide the sequence of word vectors into blocks, and respectively use the attention mechanisms between and within segments to generate matrix and matrix; A sequence representation module, configured to: utilize matrix and matrix to perform the attention mechanism calculation of Transformer to obtain the final sequence representation, which is used as the input of the downstream task module for performing specific downstream tasks.

[0061] Example 3 In an embodiment of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned large model optimization method based on a sequence dependency hierarchical propagation mechanism.

[0062] Example 4 In an embodiment of the present application, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the above-mentioned large model optimization method based on a sequence dependency hierarchical propagation mechanism.

[0063] Example 5 In an embodiment of the present application, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the above-mentioned large model optimization method based on a sequence dependency hierarchical propagation mechanism.

[0064] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0065] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks. Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for realizing the functions specified in one block or a plurality of blocks.

[0066] Although the specific embodiments of the present application have been described above in conjunction with the accompanying drawings, they are not intended to limit the protection scope of the present application. Those skilled in the art should understand that, based on the technical solutions of the present application, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present application.

Claims

1. A large model optimization method based on a sequence-dependent hierarchical propagation mechanism, wherein the large model includes a representation learning module and a downstream task module, characterized in that: The specific steps of the representation learning module are: Perform natural language processing on the input text to obtain a word vector sequence; Based on the sequence-dependent hierarchical propagation mechanism, the word vector sequence is evenly divided into blocks, and the attention mechanism between segments and within segments is used to generate Matrix and matrix; use Matrix and The matrix is ​​used to calculate the Transformer's attention mechanism to obtain the final sequence representation, which is used as the input of the downstream task module for specific downstream tasks.

2. A large model optimization method based on a sequence-dependent hierarchical propagation mechanism as claimed in claim 1, characterized in that: The natural language processing includes word segmentation, word list construction, and word embedding.

3. A large model optimization method based on a sequence-dependent hierarchical propagation mechanism as claimed in claim 1, characterized in that: The generation Matrix and Matrix, the specific steps are: Perform high-dimensional embedding and position encoding on the word vector sequence to obtain an embedding representation containing position information; Fill the embedded representation containing position information, divide the filled embedded representation into blocks evenly, and perform mean pooling on each block to obtain a preliminary fragment-level representation; The inter-fragment attention mechanism and the intra-fragment attention mechanism are used on the preliminary fragment-level representation to obtain the final inter-fragment representation and the fine-grained intra-fragment representation; The inter-fragment representation and the intra-fragment representation are concatenated to obtain the Key matrix and the Value matrix.

4. A large model optimization method based on a sequence-dependent hierarchical propagation mechanism as claimed in claim 3, characterized in that: The inter-segment attention mechanism is to obtain the dependencies between segments and obtain the final segment-level representation by calculating the attention distribution between different segments; The intra-fragment attention mechanism is to obtain intra-fragment dependencies and obtain fine-grained intra-fragment representations by calculating the attention matrix in each fragment.

5. A large model optimization method based on a sequence-dependent hierarchical propagation mechanism as claimed in claim 1, characterized in that: The attention mechanism calculation of the Transformer is to perform multi-head self-attention mechanism calculation in the Transformer stacking layer, the Key matrix and the Value matrix between multiple stacking layers, and use the linear layer to replace the feedforward network layer in the traditional Transformer.

6. A large model optimization method based on a sequence-dependent hierarchical propagation mechanism as claimed in claim 1, characterized in that: The downstream tasks include, but are not limited to, text classification, text generation, time series classification, and time series regression prediction.

7. A large model optimization system based on a sequence-dependent hierarchical propagation mechanism, wherein the large model includes a representation learning module and a downstream task module, characterized in that: The representation learning module includes: The language processing module is configured to: perform natural language processing on the input text to obtain a word vector sequence; The hierarchical dependency module is configured to evenly divide the word vector sequence into blocks based on the sequence dependency hierarchical propagation mechanism, and use the attention mechanism between segments and within segments to generate Matrix and matrix; The sequence representation module is configured to: Matrix and The matrix is ​​used to calculate the Transformer's attention mechanism to obtain the final sequence representation, which is used as the input of the downstream task module for specific downstream tasks.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements a large model optimization method based on a sequence-dependent hierarchical propagation mechanism as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, a large model optimization method based on a sequence-dependent hierarchical propagation mechanism as described in any one of claims 1 to 6 is implemented.

10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a large model optimization method based on a sequence-dependent hierarchical propagation mechanism as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Natural language processing method and device and storage medium

    CN111368536A

  • Statement text aspect level sentiment classification method and system

    CN113157919A

  • BERT-improved text semantic matching device, system and method and storage medium

    CN113239700A

  • Text classification method based on attention mechanism

    CN115017314A

  • Software system performance fault prediction method and system

    CN116069606A