Controllable attention method and system based on feature domain division and medium
By introducing feature domain division into attention calculation method and dividing it into local and global attention heads, problems such as high computing resource consumption and insufficient long-distance dependency capture in the existing technology are solved, and more efficient text generation and better controllability and quality are achieved.
Patent Information
- Application Number
- CN202510637497.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing attention calculation methods consume high computing resources when processing long texts, insufficient long-distance dependency capture, uneven information distribution, difficult multi-dimensional information fusion, and poor balance between controllability and generation quality.
The controllable attention method based on feature domain division is adopted. By dividing the multi-head attention into local attention heads and global attention heads, the detailed information in the feature domain and long-distance dependence in the sequence are respectively processed to achieve fine control of text features.
It significantly reduces the computational complexity, improves the controllability and quality of the generated text, improves the harmony and rationality of the generated content, and accelerates the convergence and reasoning of the model.
Smart Images

Figure CN120197509A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and relates to a method for calculating attention in natural language text processing and music processing of artificial intelligence. More specifically, it relates to a controllable attention method, system and medium based on feature domain division. Background Art
[0002] In the fields of natural language processing, text generation, and music generation, after years of development of modeling and controllable generation technologies, a large number of deep learning-based models and methods have emerged. Currently, the technology of AI (Artificial Intelligence) generating text mainly relies on a model called 'Transformer'. Like humans, it can pay attention to every detail of the entire text at the same time, but the disadvantage is that the computational cost is large. To solve this problem, subsequent improved models (such as Transformer-XL, Sparse Transformer, Transformer-LS) attempt to narrow the attention range, but important information may be missed.
[0003] First of all, the Transformer model (Transformer, 2017) adopts a global self-attention mechanism, which can capture the dependencies between any two positions in the sequence and has achieved breakthrough results in tasks such as machine translation and text summarization. However, the computational complexity of global self-attention is the square of the sequence length, which greatly increases the computational cost and memory occupancy when processing long texts and large-scale data, limiting the scalability and real-time generation ability of the model. Even after introducing improvement means such as position encoding, the modeling of long sequences still faces the problem of computational bottlenecks.
[0004] Subsequent long-distance dependence modeling methods based on the cyclic mechanism alleviate the high overhead problem of global attention calculation to a certain extent by introducing segment overlap and relative position encoding. Although this method has made progress in capturing long-term dependencies, it still relies on the basic framework of global attention and still faces high computational costs and difficulties in finely controlling the generated content when generating large-scale texts.
[0005] To address the above problems, by sparsifying the attention matrix design, selectively calculating the dependencies between some key positions, the computational complexity is significantly reduced. This method solves the efficiency problem in long-sequence text generation to a certain extent. However, in specific applications, the design of the sparse pattern often needs to make a trade-off between information integrity and computational efficiency, and it is easy to have insufficient capture of global context and less coherent generated content.
[0006] Meanwhile, the local attention model divides the input sequence into fixed or dynamic windows and performs attention calculations only within a local range, thus significantly reducing the computational cost. Although local attention has obvious advantages in efficiency compared to traditional global attention, its drawback is that it limits the model's ability to capture long-range dependencies, making it easy for the generated text to break in terms of long-distance information association, thereby affecting the overall semantic coherence and consistency.
[0007] In addition, axial attention has shown good results in image generation and language modeling. However, in language modeling tasks, text data inherently has more complex temporal and context dependencies. Simple axial decomposition may be difficult to comprehensively capture semantic information across different positions and levels in the text, and at the same time, problems such as the mismatch between local information and global semantics are likely to occur when introducing control conditions.
[0008] In addition, in recent years, controllable generation methods based on pre-trained models (such as GPT-2 / CTRL, etc.) often have two main problems in actual operation: one is that the integration of control signals and the internal knowledge of the language model is not tight enough, resulting in the generated content may lose the fluency of natural language while meeting specific control requirements; the other is that in the case of continuous expansion of the model scale, how to balance generation efficiency and controllability while ensuring generation quality remains an urgent problem to be solved.
[0009] Generally speaking, the current attention calculation methods have at least the following technical problems: (1) High computational resource consumption: Traditional Transformer models, due to the adoption of global self-attention mechanisms, have a computational complexity proportional to the square of the sequence length, resulting in significant computational resource challenges when dealing with long texts.
[0010] (2) Insufficient capture of long-range dependencies: Although Transformer-XL has solved the long-range dependency problem to a certain extent by introducing segment overlap and relative position encoding, its overall framework still relies on global attention calculation, making it difficult for the model to balance the capture of local details and global semantics when generating extremely long texts.
[0011] (3) Uneven information distribution: While reducing the computational cost, Sparse Transformer and local attention methods tend to miss some key long-range dependency information due to the limitation of the attention calculation range, thus affecting the coherence and semantic integrity of the generated text.
[0012] (4) Multidimensional Information Fusion Problem: When axial attention decomposes and calculates multidimensional data, although the calculation efficiency is improved, in language generation tasks, how to effectively fuse the temporal information, grammatical structure, and semantic content in the text to ensure that the generated results are both in line with local details and globally consistent remains a major challenge.
[0013] (5) Balance Problem between Controllability and Generation Quality: Current controllable generation methods based on control codes are prone to causing abrupt or unnatural generated content when achieving specific control objectives. Especially when dealing with abstract attributes such as emotion and style, it is difficult to reconcile the contradiction between the control signal and the inherent semantics of the model.
[0014] Therefore, there is an urgent need for a controllable attention scheme based on feature domain division that can accelerate the convergence and reasoning of the model in language modeling tasks, while improving the controllability and quality of the generated content, and enhancing the harmony and rationality of the generated content. Summary of the Invention
[0015] In view of the above problems, the purpose of the present invention is to provide a controllable attention method, system, and medium based on feature domain division to solve the technical problems of high computational resource consumption, insufficient capture of long-distance dependencies, unbalanced information distribution, difficult multidimensional information fusion, and poor balance between controllability and generation quality in existing attention calculation methods.
[0016] In the first aspect, an embodiment of the present application provides a controllable attention method based on feature domain division. The controllable attention method based on feature domain division includes: Pass the original data through a linear layer and position embedding processing to form a word vector representation, and obtain an attention matrix based on the word vector representation; Divide the preset multi-head attention to obtain local attention heads and global attention heads, and calculate the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix; Perform attention fusion on the local attention output and the global attention output to obtain an attention output result.
[0017] Optionally, passing the original data through a linear layer and position embedding processing to form a word vector representation, and obtaining an attention matrix based on the word vector representation includes: Obtain the events in the original data; Perform feature domain division corresponding to the attribute quantity of each event to obtain attribute information in at least two dimensions; Obtain attribute embedding matrices corresponding to each piece of attribute information; Obtain embedding vectors of each piece of attribute information according to the embedding matrix; Concatenate the embedded vectors to obtain a comprehensive representation of the event, and stack the comprehensive representations of each event to obtain a structured representation of the original data; Combine the structured representation with the pre-acquired positional encoding to obtain a hidden layer state, and obtain an attention matrix based on the hidden layer state.
[0018] Optionally, obtaining the attention matrix based on the hidden layer state includes: Perform a linear transformation on the hidden layer state based on a preset projection matrix to obtain a query matrix, a key matrix, and a value matrix; Divide the query matrix, the key matrix, and the value matrix into a preset number of sub-matrices respectively, and use the sub-matrices of the query matrix, the sub-matrices of the key matrix, and the sub-matrices of the value matrix as the attention matrix; where the preset number is the total number of heads h of the multi-head attention.
[0019] Optionally, dividing the preset multi-head attention to obtain local attention heads and global attention heads includes: Obtain the total number of heads h of the multi-head attention; where If the total number of heads h is even, use the attention heads with head numbers from 1 to h / 2 as global attention heads, and use the attention heads with head numbers from h / 2 + 1 to h as local attention heads; If the total number of heads h is odd, use the attention heads with head numbers from 1 to (h + 1) / 2 as global attention heads, and use the attention heads with head numbers from (h + 3) / 2 to h as local attention heads.
[0020] Optionally, calculating the local attention output of the local attention heads based on the attention matrix includes: Divide the local attention heads into feature domain heads of a specific number of categories according to a preset indicator function; the specific number is the number of attributes; Obtain a feature attention score matrix based on the attention matrix, and make the feature domain heads perform attention calculation for specific attributes corresponding to the feature domain heads according to the feature attention score matrix to obtain an attention output; where Concatenate and fuse the attention outputs of the specific number of categories to obtain the local attention output.
[0021] Optionally, calculating the global attention output of the global attention heads based on the attention matrix includes: Obtain a global attention score matrix through the attention matrix; Perform standard attention calculation through the global attention heads based on the global attention score matrix to obtain the global attention output.
[0022] Optionally, after obtaining the attention output result, it further includes: Obtaining an output projection matrix based on the attention output result; Performing a linear transformation on the output projection matrix to obtain hybrid head information; Performing layer normalization on the hybrid head information to obtain corrected data.
[0023] In a second aspect, an embodiment of the present application provides a controllable attention system based on feature domain division, implementing the controllable attention method based on feature domain division as described above. The system includes: An attention acquisition unit, configured to subject the original data to linear layer and position embedding processing to form a word vector representation, and obtain an attention matrix based on the word vector representation; A local-global segmentation unit, configured to divide the preset multi-head attention to obtain a local attention head and a global attention head, and calculate the local attention output of the local attention head and the global attention output of the global attention head based on the attention matrix; An attention fusion unit, configured to perform attention fusion on the local attention output and the global attention output to obtain an attention output result.
[0024] Optionally, the total number of heads of the multi-head attention is h; wherein, If h is an even number, the attention heads with head numbers from 1 to h / 2 are used as global attention heads, and the attention heads with head numbers from h / 2 + 1 to h are used as local attention heads; If h is an odd number, the attention heads with head numbers from 1 to (h + 1) / 2 are used as global attention heads, and the attention heads with head numbers from (h + 3) / 2 to h are used as local attention heads.
[0025] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any optional controllable attention method based on feature domain division in the first aspect.
[0026] As can be seen from the above technical solutions, for the controllable attention method, system, and computer-readable storage medium provided by the present application, after the original data is subjected to linear layer and position embedding processing to form a word vector representation, and an attention matrix is obtained based on the word vector representation, the preset multi-head attention is divided to obtain a local attention head and a global attention head, and the local attention output of the local attention head and the global attention output of the global attention head are calculated based on the attention matrix, and then the local attention output and the global attention output are subjected to attention fusion to obtain an attention output result. Compared with the prior art, the present application has the following beneficial effects: Taking music generation as an example, through experimental comparison, whether it is chord accuracy, voice part overlap degree, perplexity or value fluctuation, it is better than the prior art. It can make the generated music of higher quality, closer to real music, reduce uncertainty at the same time, and be more efficient with faster inference speed. Briefly speaking, by introducing feature domain division, fine control of text features is achieved, which improves the controllability and quality of the generated text while reducing computational complexity. Local attention helps to focus on the feature regions of the text, while global attention attempts to capture broader context information. The combination of the two enables the improvement of the controllability and quality of the generated content, as well as the harmony and rationality of the generated content while significantly reducing computational complexity. Brief Description of the Drawings
[0027] By referring to the following description of the specification in conjunction with the drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the drawings: Figure 1 is a flowchart of a controllable attention method based on feature domain division according to an embodiment of the present invention; Figure 2 is a schematic diagram of specific execution in a controllable attention method based on feature domain division according to an embodiment of the present invention; Figure 3 is a logic block diagram of a controllable attention system based on feature domain division according to an embodiment of the present invention. Detailed Embodiments
[0028] The attention calculation methods in the prior art have at least the following technical problems: (1) High consumption of computing resources; (2) Insufficient capture of long-distance dependencies; (3) Unbalanced information distribution; (4) Difficulties in multi-dimensional information fusion; (5) The balance problem between controllability and generation quality.
[0029] To address the above problems, the present invention provides a controllable attention method, system and computer-readable storage medium based on feature domain division. First, the original data is processed through a linear layer and positional embedding to form a word vector representation, and after obtaining an attention matrix based on the word vector representation, the preset multi-head attention is divided to obtain local attention heads and global attention heads, and the local attention output of the local attention heads and the global attention output of the global attention heads are calculated based on the attention matrix. Then, the local attention output and the global attention output are fused to obtain an attention output result. Compared with the prior art, the present application has the following beneficial effects: Taking music generation as an example, through experimental comparison, whether it is chord accuracy, voice overlapping degree, perplexity or value fluctuation, it is better than the existing technology, which can make the generated music of higher quality, closer to real music, reduce uncertainty at the same time, and be more efficient with faster inference speed. Briefly speaking, by introducing feature domain division, fine control of text features is achieved, which improves the controllability and quality of the generated text while reducing the computational complexity. Local attention helps to focus on the feature regions of the text, while global attention attempts to capture more extensive context information. The combination of the two enables the controllability and quality of the generated content to be improved, and the harmony and rationality of the generated content to be enhanced while significantly reducing the computational complexity.
[0030] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. The description of the following exemplary embodiments is actually only illustrative and in no way limits the invention and its application or use. Technologies and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies and devices should be regarded as part of the specification. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0031] Figure 1 The flowchart of a controllable attention method based on feature domain division provided by an embodiment of this application is as Figure 1 shown, and this method includes: S1: Subject the original data to linear layer and positional embedding processing to form a word vector representation, and obtain an attention matrix based on the word vector representation; S2: Divide the preset multi-head attention to obtain local attention heads and global attention heads, and calculate the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix; S3: Perform attention fusion on the local attention output and the global attention output to obtain an attention output result.
[0032] Specifically, step S1 is a process of subjecting the original data to linear layer and positional embedding processing to form a word vector representation, and obtaining an attention matrix based on the word vector representation.
[0033] In this embodiment, subjecting the original data to linear layer and positional embedding processing to form a word vector representation, and obtaining an attention matrix based on the word vector representation includes: S11: Obtain the events in the original data; S12: Perform feature domain partitioning corresponding to the attribute quantities of each event to obtain attribute information in at least two dimensions; S13: Obtain the attribute embedding matrices corresponding to each piece of attribute information; S14: Obtain the embedding vectors of each piece of attribute information according to the embedding matrices; S15: Concatenate the embedding vectors to obtain the comprehensive representation of the event, and stack the comprehensive representations of each event to obtain the structured representation of the original data; S16: Combine the structured representation and the pre-obtained position encoding to obtain the hidden layer state, and obtain the attention matrix based on the hidden layer state.
[0034] Obtaining the attention matrix based on the hidden layer state includes: S161: Perform a linear transformation on the hidden layer state based on a preset projection matrix to obtain a query matrix, a key matrix, and a value matrix; S162: Divide the query matrix, the key matrix, and the value matrix into a preset number of sub-matrices respectively, and use the sub-matrices of the query matrix, the sub-matrices of the key matrix, and the sub-matrices of the value matrix as the attention matrix; wherein, the preset number is the total number of heads h of the multi-head attention.
[0035] As Figure 2 shown, in a specific embodiment, assume that the original data (original language data, original music data, original text data, etc.) contains T events, and each event contains n attributes, corresponding to n feature domains.
[0036] Let the original representation of each event i be: ; wherein, the meanings of each attribute are as follows: : Information of the first attribute in the i th event; : Information of the second attribute in the i th event; : Information of the third attribute in the i-th event; : Information of the fourth attribute in the i-th event; … and other possible event attributes.
[0037] Taking computer music as an example, ui(1) ui(2) ui(3) ui(4) can represent the instrument, relative time, pitch, and duration attribute information of the i-th note event, and the order is not fixed.
[0038] Taking dance skeletal data as an example, ui(1) ui(2) ui(3) ui(4) can represent the torso, hand, leg, and head attribute information of the i-th action, and the order is not fixed.
[0039] For each attribute, an independent embedding matrix can be used to map discrete categorical information or continuous numerical values into a vector space of a fixed dimension. Denote the embedding matrices of each attribute as: ; where represents the vocabulary size of this attribute, represents the dimension of the corresponding embedding vector.
[0040] For event i , the embedding vectors obtained by looking up the table are respectively: ; Concatenate the embedding vectors of each attribute to obtain the comprehensive representation of the event: ; where the semicolon ";" represents the concatenation operation of vectors, is the dimension of the final vector.
[0041] Stack the comprehensive representations of all T events to obtain the structured representation input matrix with structure alignment: ; This structured representation (structured representation input matrix) X provides a unified vector representation for subsequent attention calculation.
[0042] In order to capture the sequential information between events, positional encoding also needs to be introduced. Let the positional encoding matrix be , then the final hidden layer state can be written as: ; The hidden layer state is the vector representation obtained by passing the input data through the linear layer and positional encoding, reflecting the preliminary feature extraction result of the data. The hidden layer state provides a complete representation containing data content and positional features for the subsequent attention mechanism.
[0043] where is thei Encoding vectors for each position, usually using sine and cosine functions: ,
[0044] Here k is the dimension index. This encoding method enables the model to distinguish different positions and learn relative position information.
[0045] Among them, in this specific embodiment, the specific steps for obtaining the attention matrix are as follows: For the hidden layer input , obtain the query ( Query ), key ( Key ), and value ( Value ) matrices through linear transformation: ,
[0046] Among them, the projection matrix , , is used to map to the same dimensional space. Usually, d is evenly divided into h heads, and the dimension of each head is .
[0047] Here, , , are learnable projection matrices used to map the hidden layer state to the query, key, and value spaces respectively. The sizes of these matrices are all d×d , ensuring that the dimensions of the output Q, K, and V matrices are consistent with the input, i.e., all .
[0048] In the multi-head attention mechanism, to enable the model to capture information from different perspectives, the dimension d is usually evenly divided into h heads. The dimension of each head is . Specifically: The Q, K, and V matrices are split into h sub-matrices, and the split sub-matrices are used as the subsequent attention matrices. The dimension of each sub-matrix is . Among them, this h is equal to the total number of heads h of the multi-head attention mentioned later.
[0049] Each head calculates the attention independently, and then the results are concatenated and fused to generate the final output.
[0050] Through the projection matrices , , for linear transformation, the main functions are as follows: Feature extraction: Different projection matrices extract features suitable for queries, keys, and values from such that Q, K, and V focus on different roles respectively.
[0051] Dimension alignment: Ensure that Q, K, and V have the same dimension d for subsequent calculations (such as dot product and weighted summation).
[0052] Flexibility: The projection matrix is learnable, and the model can adjust the mapping method according to the task requirements.
[0053] It should be noted that the Q, K, and V matrices each perform their own functions in the attention mechanism and jointly complete the selection and transmission of information. The following are their functions: Query (Q) matrix: Meaning: Q represents the attention requirement of the current token (or a certain position in the sequence) for other tokens.
[0054] Function: Each query vector (the i th row of Q) is like a "question" used to "ask" for information at other positions in the sequence. It determines which parts the current token hopes to focus on.
[0055] Key (K) matrix: Meaning: K represents the "key" to the information of each token in the sequence, used to match the query.
[0056] Function: The key vector (the j th row of K) calculates the similarity with the query vector (usually through dot product) and determines the degree of attention to the position j .
[0057] Value (V) matrix: Meaning: V contains the actual information content to be transmitted.
[0058] Function: The value vector (the j th row of V) is weighted and summed in the attention calculation to form the final output. The weights are determined by the similarity between Q and K.
[0059] Step S2 is a process of dividing the preset multi-head attention to obtain local attention heads and global attention heads, and calculating the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix.
[0060] In this embodiment, the preset multi-head attention is divided to obtain local attention heads and global attention heads, including: S21: Obtain the total number of heads h of the multi-head attention; where S21A: If the total number of heads h is even, the attention heads with head numbers from 1 to h / 2 are used as global attention heads, and the attention heads with head numbers from h / 2 + 1 to h are used as local attention heads; S21B: If the total number of heads h is odd, the attention heads with head numbers from 1 to (h + 1) / 2 are used as global attention heads, and the attention heads with head numbers from (h + 3) / 2 to h are used as local attention heads.
[0061] Calculating the local attention output of the local attention heads based on the attention matrix includes: S221: Divide the local attention heads into feature domain heads of a specific number of categories according to a preset indication function; the specific number is the attribute quantity; S222: Obtain a feature attention score matrix based on the attention matrix, and perform attention calculation for a specific attribute by the feature domain heads according to the feature attention score matrix to obtain an attention output, where; S223: Concatenate and fuse the attention outputs of the specific number of categories to obtain a local attention output.
[0062] Calculating the global attention output of the global attention heads based on the attention matrix includes: S231: Obtain a global attention score matrix through the attention matrix; S232: Perform standard attention calculation through the global attention heads based on the global attention score matrix to obtain a global attention output.
[0063] The core of the multi-head attention mechanism in Transformer is to capture information in different subspaces through multiple independent attention heads. To achieve the focused capture of local features and the fusion of global information, this embodiment divides the traditional multi-head attention into global and local types of heads, and makes the local attention calculated within the feature domain, so that the local attention helps to focus on the feature regions of the text, while the global attention attempts to capture more extensive context information, and introduces the step of feature domain division, thereby achieving fine control of text features. This improves the controllability and quality of the generated text while reducing the computational complexity.
[0064] In a specific embodiment, the h heads are divided into two groups: Global attention heads: Head numbers i=1,…, h / 2. The traditional global attention calculation method is retained, allowing information interaction between any two positions to capture global dependencies; Local attention head: head number i= h / 2 + 1,…, h. Each local attention head only focuses on the tokens within the same feature domain. For example, for the same music event, its different features (instrument, time, pitch, duration, etc.) form several "feature domains". The local attention heads only perform information interaction within the same feature domain.
[0065] The original intention of this specific embodiment is that the balance between local and global information is considered to contribute to the generation quality, so a scheme of evenly dividing the attention heads is adopted. Considering that h may be odd, when h is even, the local attention heads and the global attention heads each account for half; when h is odd, it is recommended to set the number of local attention heads to (rounded down), and the number of global attention heads to (rounded up). In addition, in certain specific tasks, the ratio between the two can be adjusted according to actual needs to optimize the balance between local detail capture and global dependency modeling. The indicator function of this specific embodiment is as follows. Let the representation of the input event have been sorted or chunked according to each feature domain. Define the indicator function: ; In the local attention head According to the indicator function divide out n types of feature domain heads: ;
[0066] Among them, the working principle of the indicator function is to select token 1 as a, and then loop from b to 10 and all match the same feature, so 1~10 is divided out; then when it comes to 11 and it is found that it does not belong to the same feature, 11 is selected as a and the loop continues backward for b. And so on.
[0067] "Feature domain" refers to a subset of features related to this part.
[0068] For the music generation task, the input data can be divided into 4 feature domains: pitch, duration, instrument, and start time, and the boundaries of each feature domain are preset.
[0069] Through this feature domain division, the model can focus on information interaction of specific attributes in the local attention heads, providing a more refined input basis for subsequent attention calculation.
[0070] For the i th head, the attention score matrix is defined as: ;
[0071] Subsequently, the attention output is: ;
[0072] Wherein: represents the query matrix of the i -th head, represents the key matrix, represents the value matrix, is the dimension for each head.
[0073] Here, softmax represents normalizing each row so that the sum of the weights in each row is 1.
[0074] Perform standard attention calculation on each type of feature domain head separated by the local attention head to obtain: ;
[0075] Subsequently, splice the feature domain outputs to obtain the local attention output: ;
[0076] In this embodiment, the global attention head directly calculates the standard attention as in steps S231 and S232, which will not be elaborated here.
[0077] After that, fuse the global and local attentions to obtain the attention output result: ;
[0078] Wherein, " Concat " represents the concatenation operation along the last dimension, and the dimension of the concatenated vector is d.
[0079] Thus, the complete attention output result is obtained.
[0080] As described above, the controllable attention method based on feature domain division provided by the present invention first obtains the word vector representation (word embedding, Word Embedding), that is, the original data passes through a linear layer (Linear Layer) and a position embedding (Position Embedding) to obtain a hidden state (Hidden State), which is the word vector representation of the original data; then performs projection (Projection), that is, passes the word vector representation through a weight matrix , , Project into query (Q), key (K), value (V) matrix; then divide the head and feature domain, that is, divide the h heads (literally translated from English Head, indicating the meaning of an operation unit) of Multi-Head Attention into two types of heads: h / 2 global attention heads, retaining the traditional global attention mechanism. h / 2 local attention heads, the word units of each head are divided and arranged according to the number of preset feature domains; step 4, fusion and output of attention information, the attention calculation results of different feature domains are aggregated into local attention, the local and global attention are fused to obtain controllable local attention based on feature domain, and then the attention results are output through linear layer and layer normalization. The local attention mechanism is similar to paying attention to the conversation of only some people in a group of people and ignoring the voices of others, which can improve processing efficiency. Global attention means that everyone in the whole group can hear the voices of others, helping to capture more extensive information. This effectively solves the technical problems of high attention calculation complexity, poor fine-grained control ability and unstable generation quality in the existing technology, helps to accelerate the convergence and reasoning of the model in language modeling tasks, while improving the controllability and quality of the generated content, and improving the harmony and rationality of the generated content.
[0081] In addition, in this embodiment, after obtaining the attention output result, it also includes: S41: Obtaining an output projection matrix through the attention output result; S42: performing a linear transformation on the output projection matrix to obtain mixing head information; S43: Perform layer normalization on the mixing head information to obtain correction data.
[0082] Specifically, the concatenated vectors need to pass through an output projection matrix Perform a linear transformation to mix the information from each head: ; Next, in order to stabilize the training process and correct the scale of each part's output, layer normalization is introduced. The formula for layer normalization is: ; Layer normalization is similar to "standardizing" the output of each neuron to make its numerical distribution more stable and avoid gradient explosion or disappearance during training.
[0083] Among them, the layer normalization operation is defined as: ;
[0084] Here: represents the representation at the \(i\)-th position in vector \(Z\), \(\mu\) and \(\sigma\) respectively represent the mean and standard deviation of, \(\gamma\) and \(\beta\) are learnable scaling and translation parameters, is a small constant to prevent division by zero, “⊙” represents element-wise multiplication.
[0085] To prove the progress of the controllable attention method based on feature domain division in this embodiment compared with the prior art, it is necessary to prove the effectiveness of the controllable local attention method (FCLA) based on feature domain division in language modeling and controllable generation tasks. Therefore, multiple groups of experiments are designed on the symbolic music (specifically computer music in MIDI format) generation task, covering the objective evaluation and comparative analysis of the experiments.
[0086] Although this embodiment is mainly based on language modeling theory, the music generation task can also be regarded as a sequence generation problem, and its technical principle is highly similar to natural language processing. Therefore, the MIDI music generation task is used as a verification platform in the experimental part to illustrate the applicability of this method in different types of sequence generation tasks. In subsequent work, data from natural language generation tasks can be further introduced to more comprehensively verify the universality of the method.
[0087] This experiment uses the Lakh MIDI dataset (LMD), https: / / colinraffel.com / projects / lmd as the benchmark dataset, and selects popular song segments (about 12,000 songs) in the LMD-matched subset.
[0088] To compare the performance of FCLA, we select the state-of-the-art (SOTA) model as the baseline: The large model of the controllable attention method based on feature domain division in this embodiment is based on the Transformer architecture and integrates feature domain division and local and global attention mechanisms. We conduct experiments on the self-developed model Pop-Diffuseq using the FCLA mechanism.
[0089] The model is trained on 4 NVIDIA 4090 GPUs, with a batch size of 64, using the Adam optimizer, a learning rate of 2e-4, and 50 training epochs.
[0090] I. The evaluation metrics are as follows: ① Music metric Chord Accuracy (CA): Chord accuracy measures the degree of match between the harmonic progression of the generated music (i.e., the sequence of chords) and that of the real music. Specifically, it calculates the proportion of the correctly identified chord sequence in the generated music to the total chord sequence.
[0091] Distribution Overlap: Distribution overlap is used to measure the similarity between the generated music and the real music in certain musical attributes: 1. Pitch Distribution (DP): DP measures the similarity between the frequencies of various pitches (such as C, D, E, etc., regardless of octave) in the generated music and those in the real music.
[0092] 2. Duration Distribution (DD): DD measures the similarity between the distribution of note durations (such as whole notes, half notes, etc.) in the generated music and that in the real music.
[0093] 3. Onset Time Distribution (DO): DO measures the similarity between the distribution of note onset time intervals (i.e., the time differences between notes) in the generated music and that in the real music.
[0094] DP, DD, and DO are respectively the average values of the overlapping areas of the pitch, duration, and onset time distributions.
[0095] Let the generated distribution be DG (generated distributions) and the real distribution be DT (truth distributions). The generated and real distributions of pitch are DPG and DPT. The generated and real distributions of duration are DDG and DDT. The generated and real distributions of onset time are DOG and DOT. Then: DP = E[min(DPG, DPT) / max(DPG, DPT)] DD = E[min(DDG, DDT) / max(DDG, DDT)] DO = E[min(DOG, DOT) / max(DOG, DOT)] Among them, the function E represents taking the mean value of all the test set data, the function min represents taking the minimum value, and the function max represents taking the maximum value.
[0096] The pitch distribution DP of a piece of music is the frequency of occurrence of each pitch category. There are 12 categories in total, without dividing into octave groups.
[0097] The duration distribution DD and the onset time distribution DO of a piece of music represent the average duration and the average onset time in the music respectively, both of which are scalar information.
[0098] Suppose a piece of real music has the following pitch data: The C note appears 50 times, the D note appears 30 times, and the E note appears 20 times (a total of 100 times).
[0099] Frequencies: C = 50%, D = 30%, E = 20%.
[0100] The pitch data of the generated music is: The C note appears 45 times, the D note appears 35 times, and the E note appears 20 times (a total of 100 times).
[0101] Frequencies: C = 45%, D = 35%, E = 20%.
[0102] Calculate DP: For the C note: min(50%, 45%) / max(50%, 45%) = 45% / 50% = 0.9 For the D note: min(30%, 35%) / max(30%, 35%) = 30% / 35% = 0.857 For the E note: min(20%, 20%) / max(20%, 20%) = 20% / 20% = 1.0 DP = average(0.9, 0.857, 1.0) ≈ 0.919 This value is close to 1, indicating that the pitch distribution of the generated music is very similar to that of the real music ② Generation quality metrics Perplexity (PPL) is a commonly used metric to measure the quality of sequence generation. Its calculation method is based on the cross-entropy loss. Specifically, for a given generated sequence and the corresponding real sequence, first calculate the cross-entropy loss for each token, and then take the exponential of the average cross-entropy loss of all tokens, that is: where N represents the total number of tokens, represents the cross-entropy loss of the i-th token. Taking the exponential of the output of the average cross-entropy loss here is to convert the loss into an intuitive value related to probability, so as to facilitate understanding the uncertainty of the model during generation. A lower PPL value means that the model fits the real data distribution better and the quality of the generated sequence is higher.
[0103] In the actual experiment, we calculated the token-level PPL every 20 training steps and evaluated the training process of the model by observing the stability of the PPL curve.
[0104] II. The experimental results are as follows: ① Performance comparison Table 1 shows the comparison of objective indicators between the method of the present invention (FCLA) and the baseline model. The experimental results show that: FCLA is basically superior to the baseline in terms of chord accuracy (CA = 0.64 ± 0.05), pitch (DP = 0.78 ± 0.1), duration (DD = 0.83 ± 0.1), and onset time (DO = 0.68 ± 0.1), and the generated results are closer to the real data distribution.
[0105]
[0106] Table 1 Comparison of objective indicators (mean ± standard deviation, the higher the value, the higher the generation quality); Perplexity (PPL): The PPL curve of FCLA is stable, converging faster and fluctuating less than DiffuSeq.
[0107] In the reference experiment, the PPL value of Pop-Diffuseq remained stable and low, while the PPL value of DiffuSeq fluctuated greatly, indicating that the former has an obvious advantage in generation quality.
[0108] Further experiments found that when the FCLA mechanism was introduced into DiffuSeq, its PPL value decreased significantly and became more stable. This shows that the FCLA mechanism can effectively improve the performance of DiffuSeq, making the generated music of higher quality and closer to real music, while reducing the uncertainty in the generation process. The perplexity (PPL) curve was obtained through experiments (the lower the value, the higher the generation quality). The perplexity curve obtained by the controllable attention method based on feature domain division in this embodiment greatly shows a significant improvement in generation quality.
[0109] ② Efficiency comparison · GPU memory occupancy: The peak video memory of FCLA is reduced by 26.9% compared to DiffuSeq, adapting to the generation requirements of long sequences.
[0110] · Inference speed: It only takes 37.86 seconds to generate 447 music segments with a single card, and the efficiency is increased by 18.89% compared to DiffuSeq. Therefore, the inference speed and inference efficiency are also greatly improved compared to the prior art.
[0111] It can be seen that the controllable attention method based on feature domain division provided in this embodiment enables the original data to be processed through a linear layer and positional embedding to form a word vector representation, and after obtaining an attention matrix based on the word vector representation, divides the preset multi-head attention to obtain local attention heads and global attention heads, and calculates the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix, and then fuses the local attention output and the global attention output to obtain an attention output result. Compared with the prior art, the present application has the following beneficial effects: Through experimental comparison, whether it is chord accuracy, voice overlapping degree, perplexity or value fluctuation, it is better than the prior art, which can make the generated music quality higher, closer to real music, reduce uncertainty at the same time, and has higher efficiency and faster inference speed. Briefly speaking, by introducing feature domain division, fine control of text features is achieved, which improves the controllability and quality of the generated text while reducing the computational complexity. Local attention helps to focus on the feature regions of the text, while global attention attempts to capture more extensive context information. The combination of the two enables the improvement of the controllability and quality of the generated content, as well as the harmony and rationality of the generated content while significantly reducing the computational complexity.
[0112] In summary, the controllable attention method based on feature domain division provided in this embodiment realizes fine control of text features by combining local and global attention and introducing feature domain division. This improves the controllability and quality of the generated text while reducing the computational complexity, enabling local attention to help focus on the feature regions of the text and global attention to attempt to capture more extensive context information. Specifically, the following beneficial effects can be achieved: (1) Reducing computational complexity: By localizing some attention heads, large-scale matrix operations are split into multiple small-scale matrix operations, and the overall complexity is significantly reduced compared to traditional Transformer. This improvement is particularly significant in long sequence generation tasks, which can greatly reduce computational and memory overhead.
[0113] The feature domain mentioned here refers to the region obtained by dividing the input data (such as text or music events) according to its attributes (such as grammar, semantics, emotion, etc.), so that the model can perform specialized processing for different features.
[0114] (2) Considering both local and global information: The present invention designs a hybrid attention scheme to organically combine local attention and global attention, ensuring the capture of local details while not missing global context information, thereby generating content that is both coherent and detailed.
[0115] By dividing the multi-head attention into local attention heads and global attention heads, the detailed information within the feature domain and the long-range dependencies in the sequence are processed separately.
[0116] (3) Improving the controllability of the generated text: By introducing the feature domain partitioning technique, the present invention can more precisely control various dimensions such as grammar, semantics, sentiment, and style in text generation, making the generated text not only realistic in natural language expression but also able to meet specific control requirements.
[0117] (4) Enhancing the scalability and application scope of the model: The method proposed by the present invention can be widely applied to various language modeling tasks such as machine translation, text summarization, and content creation, and has good promotion value and practical application prospects.
[0118] Next, a controllable attention system based on feature domain partitioning provided by the embodiments of the present application will be introduced. The controllable attention system based on feature domain partitioning described below can be mutually corresponded and referred to with the controllable attention method based on feature domain partitioning described above.
[0119] As Figure 3 shown, the present invention also provides a controllable attention system 100 based on feature domain partitioning, which can implement the controllable attention method based on feature domain partitioning as described above. The system includes: An attention acquisition unit 110, configured to subject the original data to linear layer and position embedding processing to form a word vector representation, and obtain an attention matrix based on the word vector representation; A local-global segmentation unit 120, configured to partition the preset multi-head attention to obtain local attention heads and global attention heads, and calculate the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix; An attention fusion unit 130, configured to perform attention fusion on the local attention output and the global attention output to obtain an attention output result.
[0120] Among them, the total number of heads of the multi-head attention is h; among them, if h is an even number, the attention heads with head numbers from 1 to h / 2 are used as global attention heads, and the attention heads with head numbers from h / 2 + 1 to h are used as local attention heads; if h is an odd number, the attention heads with head numbers from 1 to (h + 1) / 2 are used as global attention heads, and the attention heads with head numbers from (h + 3) / 2 to h are used as local attention heads.
[0121] The controllable attention system 100 based on feature domain division provided by the present invention is calculated in the same way as the aforementioned controllable attention method based on feature domain division. For a more specific implementation process, reference can be made to the specific embodiments of the above-mentioned controllable attention method based on feature domain division.
[0122] As described above, the controllable attention system 100 based on feature domain division provided by the present invention, like the above-mentioned controllable attention method based on feature domain division, can also solve the technical problems in the existing attention calculation methods, such as high computational resource consumption, insufficient capture of long-distance dependencies, unbalanced information distribution, difficulty in multi-dimensional information fusion, and poor balance between controllability and generation quality. By introducing feature domain division, fine control of text features is achieved, which reduces the computational complexity while improving the controllability and quality of the generated text. Local attention helps to focus on the feature regions of the text, while global attention attempts to capture more extensive context information. The combination of the two enables significant reduction of computational complexity while improving the controllability and quality of the generated content, and enhancing the harmony and rationality of the generated content.
[0123] An embodiment of the present invention also provides a computer-readable storage medium (not shown in the figure). The storage medium can be non-volatile or volatile. The storage medium stores a computer program, and when the computer program is executed by a processor, it realizes: Making the original data undergo linear layer and position embedding processing to form a word vector representation, and obtaining an attention matrix based on the word vector representation; Dividing the preset multi-head attention to obtain local attention heads and global attention heads, and calculating the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix; Performing attention fusion on the local attention output and the global attention output to obtain an attention output result.
[0124] Specifically, for the specific implementation method when the computer program is executed by the processor, reference can be made to the description of the relevant steps in the controllable attention method based on feature domain division in the embodiment, which will not be elaborated here.
[0125] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0126] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0128] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0129] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0130] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. The terms such as "second" are used to denote names and do not denote any particular order.
[0131] It should be noted that the embodiments in this specification are all described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for methods, devices, electronic devices and computer-readable storage media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The methods, devices, electronic devices and media described above are only illustrative. The units described as separation components may or may not be physically separated. The components indicated as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0132] The controllable attention method, system and storage medium based on feature domain division proposed according to the present invention are described above by way of example with reference to the accompanying drawings. However, those skilled in the art should understand that various improvements can be made to the above-mentioned controllable attention method, system and storage medium based on feature domain division proposed by the present invention without departing from the content of the present invention. Therefore, the protection scope of the present invention should be determined by the content of the appended claims.
Claims
1. A controllable attention method based on feature domain division, characterized in that: include: Processing the original data through a linear layer and position embedding to form a word vector representation, and obtaining an attention matrix based on the word vector representation; Divide the preset multi-head attention to obtain local attention heads and global attention heads, and calculate the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix; The local attention output and the global attention output are subjected to attention fusion to obtain an attention output result.
2. The controllable attention method based on feature domain division as claimed in claim 1, characterized in that: The original data is processed through a linear layer and position embedding to form a word vector representation, and an attention matrix is obtained based on the word vector representation, including: Obtaining events in the original data; Dividing each event into feature domains corresponding to the attribute quantity of the event to obtain attribute information of at least two dimensions; Obtaining an attribute embedding matrix corresponding to each attribute information; Obtain the embedding vector of each attribute information according to the embedding matrix; Concatenating the embedding vectors to obtain a comprehensive representation of the event, and stacking the comprehensive representations of the events to obtain a structured representation of the original data; The structured representation and the pre-acquired position encoding are combined to obtain a hidden layer state, and an attention matrix is obtained based on the hidden layer state.
3. The controllable attention method based on feature domain division as claimed in claim 2, characterized in that: Acquiring an attention matrix based on the hidden layer state includes: Performing a linear transformation on the hidden layer state based on a preset projection matrix to obtain a query matrix, a key matrix and a value matrix; The query matrix, key matrix and value matrix are respectively divided into a preset number of sub-matrices, and the sub-matrix of the query matrix, the sub-matrix of the key matrix and the sub-matrix of the value matrix are used as the attention matrix; wherein the preset number is the total number of heads h of the multi-head attention.
4. The controllable attention method based on feature domain division as claimed in claim 2, characterized in that: The preset multi-head attention is divided to obtain local attention heads and global attention heads, including: Get the total number of heads h of the multi-head attention; where, If the total number of heads h is an even number, the attention heads numbered 1 to h / 2 are regarded as global attention heads, and the attention heads numbered h / 2+1 to h are regarded as local attention heads; If the total number of heads h is an odd number, the attention heads numbered 1 to (h+1) / 2 are taken as global attention heads, and the attention heads numbered (h+3) / 2 to h are taken as local attention heads.
5. The controllable attention method based on feature domain division as claimed in claim 4, characterized in that: Calculating a local attention output of the local attention head based on the attention matrix includes: Dividing the local attention head into feature domain heads of a specific number of categories according to a preset indicator function; the specific number is the attribute quantity; A feature attention score matrix is obtained based on the attention matrix, and the feature domain head is made to perform an attention calculation of a specific attribute for the attribute corresponding to the feature domain head according to the feature attention score matrix to obtain an attention output; wherein, The attention outputs of the specific number of categories are concatenated and fused to obtain the local attention output.
6. The controllable attention method based on feature domain division as claimed in claim 5, characterized in that: Calculating a global attention output of the global attention head based on the attention matrix includes: Obtain a global attention score matrix through the attention matrix; A standard attention calculation is performed by the global attention head based on the global attention score matrix to obtain a global attention output.
7. The controllable attention method based on feature domain division as claimed in claim 6, characterized in that: After obtaining the attention output results, it also includes: Obtain an output projection matrix through the attention output result; Performing a linear transformation on the output projection matrix to obtain mixing head information; The mixing head information is layer normalized to obtain correction data.
8. A controllable attention system based on feature domain division, characterized in that: To implement the controllable attention method based on feature domain division as described in any one of claims 1 to 7, the system includes: An attention acquisition unit, used to process the original data through a linear layer and position embedding to form a word vector representation, and acquire an attention matrix based on the word vector representation; A local-global segmentation unit, configured to divide the preset multi-head attention to obtain local attention heads and global attention heads, and calculate the local attention output of the local attention heads and the global attention output of the global attention heads based on the attention matrix; An attention fusion unit is used to fuse the local attention output and the global attention output to obtain an attention output result.
9. The controllable attention system based on feature domain division as claimed in claim 8, characterized in that: The total number of heads of the multi-head attention is h; where If h is an even number, the attention heads numbered from 1 to h / 2 are regarded as global attention heads, and the attention heads numbered from h / 2+1 to h are regarded as local attention heads; If h is an odd number, the attention heads numbered from 1 to (h+1) / 2 are regarded as global attention heads, and the attention heads numbered from (h+3) / 2 to h are regarded as local attention heads.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the controllable attention method based on feature domain division described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Electroencephalogram cognitive load assessment method and system based on multi-feature-domain attention network
CN116584955A
Han machine translation system based on dynamic fusion attention model
CN118520886A
Text classification method based on pre-training language model fusion deep convolutional network
CN120011558A
System and method for enhancing machine learning model for audio / video understanding using gated multi-level attention and temporal adversarial training
US20220300740A1