A method and system for imaginary speech classification based on attention-guided tensor network

By adopting a classification method based on attention-guided tensor network in the BCI system, using data enhancement and multi-head attention mechanisms to extract EEG data features, the problem of low classification performance of imaginary speech is solved, and the decoding accuracy and practicality of the BCI system is improved.

CN116597824BActive Publication Date: 2025-05-16HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310580969.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2025-05-16
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

In the prior art, the multi-category classification performance of imaginary voice is low, making it difficult to effectively decode the user's intuitive intentions, which limits the practicality of the application of the BCI system.

Method used

Using a classification method based on attention-guided tensor network, through data augmentation and multi-head attention mechanisms, the characteristic information of EEG data in the time dimension is extracted, and a robust feature representation is constructed to improve classification performance.

Benefits of technology

The multi-category classification performance of imaginary voice is improved, allowing the BCI system to more accurately decode user intentions, and enhance the practicality of the system and the freedom of user operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597824B_ABST
    Figure CN116597824B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for imaginary speech classification based on an attention-guided tensor network. Obtain imaginary speech EEG data and its corresponding label; perform data enhancement on the imaginary speech EEG data to construct a training data set; construct an attention-guided tensor network, train using the data-enhanced training set in the data set, and test using the unenhanced test set in the data set; use the trained and verified attention-guided tensor network to implement EEG imaginary speech classification. The method of the present invention combines data enhancement with the tensor network technology guided by the classification identifier of the attention mechanism to achieve high-precision imaginary speech EEG classification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention belongs to the field of brain-computer interface, and relates to a method and system for classifying imagined speech based on an attention-guided tensor network. Specifically, the present invention is a method for classifying imagined speech based on data enhancement and a tensor network technology guided by classification identifiers of an attention mechanism, which performs data enhancement and feature extraction on imagined speech EEG data, thereby determining the category of imagined speech. Background Art

[0002] It is hoped that in the future, BCI will be able to decode people's intuitive imagination and output it into a real environment. Once the imagined words or dialogues are decoded by the BCI system, it can be used as a neural command to output the user's imagined words through speech synthesis, or to control robots and devices based on words. Therefore, the effectiveness and practicality of imagined speech decoding are important issues that cannot be ignored. In order to realize these types of BCI, research on extracting relevant features of imagined speech paradigms can improve the effectiveness of capturing brain activities related to speech. Recently, researchers have studied various methods, especially deep learning methods, which have developed with the development of natural language processing technology to accurately capture phoneme-level speech from brain signals.

[0003] Imagined speech can be a key paradigm for developing intuitive systems that are easy for users to operate. Recognizing the user's intuitive intentions and translating them into commands for the external world is one of the key functions of BCI. Using the imagined speech paradigm, BCI communication can be significantly improved because it can directly convey the user's intentions through imagined speech or words themselves, rather than through the spelling of individual letters. At the same time, the technology can apply this decoded result to control external devices. Imagined speech is an emerging paradigm that can transfer the user's intentions to external devices. Compared with traditional BCI paradigms such as MI, the imagined speech paradigm can provide crucial advantages. For example, increasing the number of classes in MI relies on the movement of body parts. When many classes are required, the movement of body parts may naturally overlap. On the contrary, the speech properties of different classes can allow more variations between classes without overlapping concepts. In addition, the decoded imagined speech can directly match the interaction between user intentions and device feedback in real environments. Ultimately, this feature of the imagined speech paradigm may help develop more practical BCI systems that provide users with a high degree of freedom. Therefore, BCI is more inclined to a technology that decodes human intuitive intentions. However, compared with conventional BCI paradigms such as MI or ERP, the multi-class classification performance of imagined speech is still at a relatively low level. Effective feature selection or classification methods for imagined speech may help improve the decoding performance. Improving the multi-class classification performance of imagined speech to the level of conventional BCI paradigms could enable simple communication or control of the external environment through internal speech. Summary of the invention

[0004] The purpose of the present invention is to address some deficiencies in the prior art. The multi-classification performance of imagined speech is still at a relatively low level. A method and system for imaginary speech classification based on an attention-guided tensor network is proposed. Data enhancement technology is used to introduce prior knowledge into the model, so that the model learns more robust features to solve problems such as the small number of samples in the existing data set. A multi-head attention mechanism is used to effectively extract feature information of the data in the time dimension. A tensor network is used to solve the problem of small sample sizes in the data set and improve the classification performance of the model.

[0005] In a first aspect, the present invention provides an imaginary speech classification method based on an attention-guided tensor network, comprising the following steps:

[0006] Step S1: Obtain the thought imagination speech EEG data and its corresponding label label;

[0007] Step S2: When the model is trained, data enhancement is performed on the above-mentioned thought imagination speech EEG data to construct a training data set;

[0008] Step S3: construct an attention-guided tensor network, train it using the data-augmented training set in the dataset, and test it using the unaugmented test set in the dataset;

[0009] Step S4: Use the trained and verified attention-guided tensor network to implement EEG imagined speech classification.

[0010] In a second aspect, the present invention provides an imagined speech classification system, comprising a trained and validated attention-guided tensor network.

[0011] The beneficial effects of the present invention are as follows:

[0012] The present invention utilizes a multi-head self-attention mechanism to simultaneously focus on information of different time steps and different channels in the EEG signal, thereby better capturing the temporal and spatial correlations in the EEG signal. These correlations are then used as weights to calculate the importance of each time step and channel. This can help the model automatically learn important features in the EEG signal, thereby improving the performance of the model. Different feature representations can also be learned through the multi-head self-attention mechanism, and these feature representations are combined to form a final representation. This can help the model better process different EEG signals, thereby improving the generalization ability of the model.

[0013] A potential problem of deep learning is the large number of parameters. Therefore, fitting requires a large number of samples, training the model takes a lot of time, and EEG samples are usually insufficient. The present invention introduces a method of converting the weight matrix of the fully connected layer into a tensor format in the tensor learning network, which greatly reduces the number of parameters while retaining the expressive power of the layer.

[0014] In summary, the present invention proposes an imaginary speech classification method based on an attention-guided tensor network, which combines data enhancement and the tensor network technology guided by the classification flag of the attention mechanism. The network includes a data acquisition module, a data enhancement module, a multi-head attention module and a classification module, which are respectively used for data enhancement, feature acquisition, and classification results. At the same time, the network adopts random position encoding, adds classification flags, and a tensor network to achieve high-precision imaginary speech EEG classification performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Flowchart of the method for imagining speech classification;

[0016] Figure 2 Diagram of the experimental paradigm process for the dataset used

[0017] Figure 3 Schematic diagram of feature extraction based on multi-head attention mechanism;

[0018] Figure 4 Schematic diagram of the classification module based on tensor network. DETAILED DESCRIPTION

[0019] The method of the present invention is described in detail below with reference to the accompanying drawings.

[0020] like Figure 1 As shown in FIG. 1 , a method for classifying imagined speech based on an attention-guided tensor network is shown in FIG. 1 . The specific steps are as follows:

[0021] A method for imaginary speech classification based on an attention-guided tensor network comprises the following steps:

[0022] Step S1: Obtain the thought imagination speech EEG data and its corresponding label label.

[0023] The data used in this experiment is a public dataset. The experiment recorded the EEG data of 15 subjects (S1-S15; aged 20-30 years old) as shown in the attached Figure 2As shown. During the experiment, the subjects sat in a comfortable chair in front of a 24-inch LCD monitor. The subjects were asked to imagine the silent pronunciation of a given word or phrase as if they were performing real speech, without moving any articulators or making sounds. The subjects were instructed not to perform any brain activity other than the given task. They were asked not to move or blink while imagining or receiving the cue. All imagination experiments were conducted using a black screen so that the subjects were not exposed to any stimulation to avoid any other factors that could affect brain activity. The auditory cue representing one of the five words / phrases was randomly presented for 2 seconds, followed by a cross mark for 0.8 seconds to 1.2 seconds. The researchers asked the subjects to imagine the given cue immediately after the cross mark disappeared from the screen. Each random cue was sequentially presented with 4 cross mark phases (0.8-1.2s) and an imagined speech phase (2s). After four phases of imagined speech, a 3s relaxation phase was allowed to clear the subject's mind for the next word / phrase. EEG data were recorded using a signal amplifier (BrainAmp, BrainProducts GmbH, Germany). Raw data were recorded using BrainVision (BrainProducts GmbH, Germany) and MATLAB 2019a (The MathWorks Inc. USA) using 64 EEG electrodes following the 10-20 international configuration. The ground and reference channels were placed on Fpz and FCz, respectively. The impedance of all electrodes between the sensor and the scalp skin was kept below 15k.

[0024] This experiment recorded the EEG of 5 types of imagined words / phrases. T is the time dimension, and its size is 795. C is the channel dimension, and its size is 64. That is, the original data size is 64*795.

[0025] Step S2: When the model is trained, data enhancement is performed on the above-mentioned thought imagination speech EEG data to construct a training data set;

[0026] The data enhancement uses the Mixup linear interpolation method, specifically:

[0027]

[0028]

[0029] Where (X i ,Y i ) and (X j ,Y j ) are two samples randomly selected from the training data, X i ,X jis the original data input, Y i ,Y j is the one-hot encoding of the corresponding category, λ∈[0,1];

[0030] Mixup is a data augmentation technique used to improve the performance of a model. It creates new training examples by randomly combining a pair of examples from different categories. This technique can make the model more robust, reduce overfitting, and improve generalization. Specifically, Mixup combines two input samples X i ,X j Generate a new sample by linear interpolation with a random ratio λ∈[0,1] At the same time, their labels Y i ,Y j Interpolate at the same ratio to generate a new label In this way, the model can learn more features and similarities between different categories, thereby improving the generalization ability. After data enhancement, the data size remains unchanged at 795*64.

[0031] This method introduces prior knowledge into the model: by introducing this prior knowledge, the enhanced data can enable the model to learn more robust features and improve the generalization of the deep learning model.

[0032] Step S3: construct an attention-guided tensor network, train it using the data-augmented training set in the dataset, and test it using the unaugmented test set in the dataset;

[0033] The attention-guided tensor network includes a feature extraction module of a cascaded multi-head attention mechanism and a classification module of a tensor learning network;

[0034] 1) If Figure 3 The feature extraction module of the cascade multi-head attention mechanism includes an embedding layer, a classification identification bit Class Token layer, a position encoding layer, a first LN regularization layer, a multi-head self-attention layer, a first residual connection layer, a second LN regularization layer, a feedforward network layer, a second residual connection layer, and a third LN regularization layer connected in series in sequence;

[0035] The embedding layer described in 1.1 upsamples the channel dimension of the 795*64 size EEG data through a fully connected layer to increase the data dimension and extract more fine-grained information to obtain 795*1024 size data;

[0036] 1.2 The Class Tokens layer uses random initialization to generate a vector of size 1*1024, which is concatenated into the data header of the embedding layer to realize statistical global feature information and reduce the interference of local feature information. At this time, the data size is 796*1024; using Class Tokens, it can be continuously updated as the network is trained, and it can encode the statistical characteristics of the entire imagined speech data. While aggregating the information on all other Tokens (global feature aggregation), since it is not based on the content of the data itself, it can avoid bias towards a specific Token in the data. Secondly, using fixed position encoding for the Class Tokens can effectively prevent the output from being interfered by the position encoding.

[0037] The position encoding layer described in 1.3 adopts a random position encoding method, specifically: generating a random number matrix with the same format as the input data, and adding the above random number matrix to the input data as the output of the position encoding layer; the position encoding layer adopts a random position encoding method, which can solve the problem that the model cannot capture the position relationship in the time dimension of the input sequence. The position encoding layer assigns a position code to each vector on the time dimension in each input sequence, and this position code will be added to the time dimension vector, so that the time dimension vector can contain information about its position in the input sequence. In this way, the neural network can better understand the order and relationship of the time dimension in the input sequence, thereby improving the performance of the model. Specifically, the position encoding layer outputs a random number matrix with the same format as its input data, and adds it to the input data as the input data of the multi-head self-attention layer.

[0038] 1.4 The first LN regularization layer normalizes the output data of the position encoding layer;

[0039] The multi-head self-attention mechanism layer described in 1.5 maps the output data of the LN regularization layer to different subspaces, and then performs a dot multiplication operation on all subspaces to calculate the attention vector; finally, the attention vectors calculated in all subspaces are concatenated and mapped to the original input space to obtain the final attention vector, so as to realize the feature correlation of the statistically imagined speech data in the time dimension;

[0040] The role of the multi-head self-attention mechanism layer is to enhance the model's ability to understand and express input data, and improve the model's accuracy and generalization ability. Specifically, the multi-head attention layer can divide the input data into multiple heads, and each head can focus on different parts of the input data to extract different feature information. These heads can be calculated in parallel, thereby speeding up the training of the model. Finally, the calculation results of multiple heads are merged to obtain the final output result. The expression of the multi-head self-attention layer is as follows (3):

[0041]

[0042] Where MultiHead(Q,K,V) represents the final output attention vector; Concat represents the concatenation operation; where head i represents the attention vector calculated in the i-th subspace;

[0043] i represents different subspaces. The query vector Q, key vector K and value vector V are obtained by passing the output data of the first LN regularization layer through the fully connected layer as the input of the multi-head self-attention module. i Q is the mapping matrix of Q in different subspaces, W i K is the mapping matrix of K in different subspaces, W i V is the mapping matrix of V in different subspaces, W O By W in all subspaces i V Spliced ​​together;

[0044] The attention vector on a single subspace is calculated as follows: first, the query vector Q is dot-multiplied with the key vector K, and then divided by the square root of the dimension of the key vector K. The score matrix of the query vector Q is obtained, and the result is finally passed to the Softmax function, which is used to normalize the weight matrix, and then multiplied by the value vector V to obtain an attention vector of the subspace, which is expressed as follows (4):

[0045]

[0046] The parameter matrices of Q, K, and V have dimension d q , d k , and d v Both are 128, the number of attention heads is 8, and d model is 1024;

[0047] Through linear transformation, the query vector Q is transformed from d model The dimension is mapped to d q *head, transfer the key vector K from d model The dimension is mapped to d k *head, convert the value vector V from d model The dimension is mapped to d v *head;

[0048] Implicitly increasing the number of attention heads without reducing the hidden dimension assigned to each attention head can effectively extract global features and improve classification accuracy.

[0049] 1.6 The first residual connection layer performs a residual connection on the output of the multi-head self-attention mechanism layer to improve the network's ability to represent imagined speech data and effectively solve the problems of gradient disappearance and gradient explosion;

[0050] 1.7 The second LN regularization layer normalizes the output data of the first residual connection layer;

[0051] 1.8 The feed-forward network layer (FFN) is composed of two layers of feed-forward neural networks. The first layer of feed-forward network converts the output of the second LN regularization layer from d model Dimension mapping is 4*d model dimension, the activation function is GELU function, and the second layer of feedforward neural network is 4*d model The dimension is mapped back to d model Dimension, no activation function is used;

[0052] The expression of each layer of the feedforward network is as follows (5):

[0053]

[0054] Where W1 and W2 are randomly initialized weight vectors, b1 and b2 are randomly initialized biases; x represents the output of the second LN regularization layer;

[0055] 1.9 The second residual connection layer performs a residual connection on the output of the feedforward network layer to enhance the network's ability to represent imagined speech data;

[0056] The third LN regularization layer described in 1.10 normalizes the output data of the second residual connection layer;

[0057] 2) The classification module of the tensor learning network obtains the data with ClassTokens in the output data of the feature extraction module of the cascaded multi-head attention mechanism, and performs prediction and classification on it;

[0058] like Figure 4 The classification module of the tensor learning network includes a tensor network, an activation layer, and a fully connected layer connected in series in sequence;

[0059] The tensor network performs tensor quantization processing on the input data of size 1*1024 to enable the network to extract linear relationship features in high-dimensional imaginary speech data, specifically:

[0060] A linear transformation is performed on an N-dimensional input vector, so its mathematical expression is shown in formula (6):

[0061] y1=Wx1+b (6)

[0062] in is the weight matrix, For input data, is bias;

[0063] The element y(i) in y is expressed as shown in formula (7):

[0064]

[0065] According to the tensor learning idea, y, W, x, and b are all converted into tensor representations, recorded as y, W, x, and b; specifically:

[0066] First, x∈R N*S Converted to a 5-dimensional tensor Denoted as x(j1,...,j5), where N*S=S1*S2*S3*S4*S5, that is, the input Class Tokens vector 1*1024 is converted into a five-dimensional tensor of size 4*4*4*4*4;

[0067] Through the bijective function F(i) = (f1(i), f2(i), f3(i), f4(i), f5(i)) = (i1, i2, i3, i4, i5), the vectors y and b are linked to the five-dimensional tensor representations of y(i1, i2, i3, i4, i5) and b(i1, i2, i3, i4, i5) through the index i as shown in formula (8):

[0068] y(F(i))=y(i1,i2,i3,i4,i5)=y(i)

[0069]

[0070] where y,b∈R M , d=5, i∈1,2,...,M; y(i),b(i) are the elements in y,b, and y(i1,i2,i3,i4,i5) is also a five-dimensional tensor of size 4*4*4*4*4;

[0071] The same is true for the weight matrix See formula (9):

[0072] F(i)=(f1(i),f2(i),f3(i),f4(i),f5(i))=(i1,i2,i3,i4,i5)

[0073] G(j)=(g1(j),g2(j),g3(j),g4(j),g5(j))=(j1,j2,j3,j4,j5) (9)

[0074] The weight matrix W can be associated with its corresponding tensor W and converted into the tensor column format TensorTrain Format (TT-format) as shown in formula (10):

[0075]

[0076] Where each g[i k ,j k ] In the case of the same k, it is represented as i k *j k *r k-1 *r k , k∈1,2,...,5, where r0=r5=1, where r k-1 *r k It is called tensor column rank TT-rank, and TT-rank is [1,8,8,8,8,1];

[0077] Finally, equation (6) can be converted into the tensor form shown in equation (11):

[0078]

[0079] The activation layer uses a RELU activation function to pass the output data of the tensor network through the activation layer into the full connection to obtain the classification result;

[0080] The corresponding classification label is output by the tensor network classification module, and the loss function is calculated by comparing it with the real label. The present invention adopts the cross entropy loss function, and the specific formula is as follows:

[0081]

[0082] Where M1 is the number of trials, N1 is the number of categories, represents the true label of the m1th trial, It represents the predicted probability of the m1th trial of category n1. When combined with model training, it is recorded as criterion, and the loss calculation method is as follows:

[0083] loss = λ*criterion(pred,Y i )+(1-λ)criterion(pred,Y j ) (13)

[0084] in

[0085] When the present invention is used specifically, Adam with fast convergence speed is used as the optimizer, the initial learning rate is set to 8e-5, and the batch size is 8.

[0086] Step S4: Use the trained and verified attention-guided tensor network to implement EEG imagined speech classification.

[0087] Table (1) Accuracy of imagined speech classification by different subjects using the above methods

[0088]

[0089]

Claims

1. A method for imaginary speech classification based on attention-guided tensor networks, characterized in that The method comprises the following steps: Step S1: Obtain the thought imagination speech EEG data and its corresponding label label; Step S2: Perform data enhancement on the above-mentioned thought imagination speech EEG data to construct a training data set; Step S3: construct an attention-guided tensor network, train it using the data-augmented training set in the dataset, and test it using the unaugmented test set in the dataset; The attention-guided tensor network includes a feature extraction module of a cascaded multi-head attention mechanism and a classification module of a tensor learning network; 1) The feature extraction module of the cascaded multi-head attention mechanism includes an embedding layer, a classification identification bit ClassToken layer, a position encoding layer, a first LN regularization layer, a multi-head self-attention layer, a first residual connection layer, a second LN regularization layer, a feedforward network layer, a second residual connection layer, and a third LN regularization layer connected in series in sequence; The multi-head self-attention mechanism layer maps the output data of the LN regularization layer to different subspaces, and then performs a dot multiplication operation on all subspaces to calculate an attention vector; finally, the attention vectors calculated in all subspaces are concatenated and mapped to the original input space to obtain a final attention vector, so as to realize the feature correlation of the statistically imagined speech data in the time dimension; 2) The classification module of the tensor learning network obtains the data with Class Tokens in the output data of the feature extraction module of the cascaded multi-head attention mechanism, and performs prediction and classification on it; The classification module of the tensor learning network includes a tensor network, an activation layer, and a fully connected layer connected in series in sequence; The tensor network performs tensorization processing on the input data to enable the network to extract linear relationship features in the high-dimensional imagined speech data; Step S4: Use the trained and verified attention-guided tensor network to implement EEG imagined speech classification.

2. The method according to claim 1, characterized in that The data enhancement in step S2 is specifically as follows: Where (X i ,Y i ) and (X j ,Y j ) are two samples randomly selected from the training data, X i ,X j is the original data input, Y i ,Y j is the one-hot encoding of the corresponding category, λ∈[0,1].

3. The method according to claim 1, characterized in that In step S3, the embedding layer in the feature extraction module of the cascaded multi-head attention mechanism upsamples the channel dimension of the EEG data through a fully connected layer to increase the data dimension and extract more fine-grained information to obtain data of 795*1024 size; The Class Tokens layer generates a vector of size 1*1024 by random initialization, and splices it into the data header of the embedding layer to realize statistical global feature information and reduce the interference of local feature information. At this time, the data size is 796*1024; The position coding layer adopts a random position coding method, specifically: generating a random number matrix with the same format as the input data, and adding the random number matrix to the input data as the output of the position coding layer; The first LN regularization layer normalizes the output data of the position encoding layer.

4. The method according to claim 1, characterized in that In step S3, the expression of the multi-head self-attention layer is as follows (3): Where MultiHead(Q,K,V) represents the final output attention vector; Concat represents the concatenation operation; wherehead i represents the attention vector calculated in the i-th subspace; i represents different subspaces. The query vector Q, key vector K and value vector V are obtained by passing the output data of the first LN regularization layer through the fully connected layer as the input of the multi-head self-attention module. i Q is the mapping matrix of Q in different subspaces, W i K is the mapping matrix of K in different subspaces, W i V is the mapping matrix of V in different subspaces, W O By W in all subspaces i V Spliced ​​together; The attention vector on a single subspace is calculated as follows: first, the query vector Q is dot-multiplied with the key vector K, and then divided by the square root of the dimension of the key vector K. The score matrix of the query vector Q is obtained, and the result is finally passed to the Softmax function, which is used to normalize the weight matrix, and then multiplied by the value vector V to obtain an attention vector of the subspace, which is expressed as follows (4): The parameter matrices of Q, K, and V have dimension d q , d k , and d v Both are 128, the number of attention heads is 8, and d model is 1024; Through linear transformation, the query vector Q is transformed from d model The dimension is mapped to d q *head, transfer the key vector K from d model The dimension is mapped to d k *head, convert the value vector V from d model The dimension is mapped to d v *head.

5. The method according to claim 4, characterized in that In step S3, the first residual connection layer performs a residual connection on the output of the multi-head self-attention mechanism layer to improve the network's ability to represent imagined speech data; The second LN regularization layer normalizes the output data of the first residual connection layer; The feedforward network layer consists of two layers of feedforward neural networks. The first layer of feedforward network converts the output of the second LN regularization layer from d model Dimension mapping is 4*d model dimension, the activation function is GELU function, and the second layer of feedforward neural network is 4*d model The dimension is mapped back to d model Dimension, no activation function is used; The second residual connection layer performs a residual connection on the output of the feedforward network layer to improve the network's ability to represent imagined speech data; The third LN regularization layer normalizes the output data of the second residual connection layer.

6. The method according to claim 5, characterized in that In step S3, the expression of each layer of the feedforward network is as follows (5): Where W1 and W2 are randomly initialized weight vectors, b1 and b2 are randomly initialized biases; x represents the output of the second LN regularization layer.

7. The method according to claim 1, characterized in that In step S3, the tensor network is specifically: An N-dimensional input vector is linearly transformed, so its mathematical expression is shown in formula (6): y1=Wx1+b(6) in is the weight matrix, For input data, is bias; The element y(i) in y1 is expressed as shown in formula (7): According to the tensor learning idea, y, W, x, and b are all converted into tensor representations, denoted as y, W, x, and b; Specifically: first, Converted to a 5-dimensional tensor Denoted as x(j1,...,j5), where N*S=S1*S2*S3*S4*S5, that is, the input Class Tokens vector 1*1024 is converted into a five-dimensional tensor of size 4*4*4*4*4; Through the bijective function F(i) = (f1(i), f2(i), f3(i), f4(i), f5(i)) = (i1, i2, i3, i4, i5), the vectors y and b are linked to the five-dimensional tensor representations of y(i1, i2, i3, i4, i5) and b(i1, i2, i3, i4, i5) through the index i as shown in formula (8): y(F(i))=y(i1,i2,i3,i4,i5)=y(i) b(F(i))=b(i1,i2,i3,i4,i5)=b(i)(8) where y,b∈R M , d=5, i∈1,2,...,M; y(i),b(i) are the elements in y,b, and y(i1,i2,i3,i4,i5) is also a five-dimensional tensor of size 4*4*4*4*4; The same is true for the weight matrix d = 5, See formula (9): F(i)=(f1(i),f2(i),f3(i),f4(i),f5(i))=(i1,i2,i3,i4,i5) G(j)=(g1(j),g2(j),g3(j),g4(j),g5(j))=(j1,j2,j3,j4,j5)(9) The weight matrix W is associated with its corresponding tensor W and converted into the tensor column format TensorTrainFormat (TT-format) as shown in formula (10): Where each g[i k ,j k ] In the case of the same k, it is represented as i k *j k *r k-1 *r k , k∈1,2,...,5, where r0=r5=1, where r k-1 *r k It is called tensor column rank TT-rank, and TT-rank is [1,8,8,8,8,1]; Finally, equation (6) is converted into the tensor form shown in equation (11):

8. The method according to claim 1, characterized in that The activation layer uses the RELU activation function.

9. The method according to claim 1, characterized in that The loss function of the attention-guided tensor network adopts the cross entropy loss function, and the specific formula is as follows: Where M1 is the number of trials, N1 is the number of categories, represents the true label of the m1th trial, represents the predicted probability of the m1th trial of category n1; when combined with model training, the cross entropy loss function is recorded as criterion, and the loss calculation method is as follows: loss = λ*criterion(pred,Y i )+(1-λ)criterion(pred,Y j ) (13) in 10. A classification system for implementing the method according to any one of claims 1 to 9, characterized in that Includes training and validating good attention-guided tensor networks.

Citation Information

Patent Citations

  • Multi-head attention model compression method for image classification

    CN115713109A

  • Multi-modal sentiment analysis method fusing multiple features and attention mechanism

    CN116028846A