Text classification method and device of multi-channel convolutional capsule network based on multi-head attention network
By introducing multi-head attention network and multi-channel convolutional capsule network into the short text classification model, the problem of missing position information and keyword characteristics in text embedding is solved, deep features are extracted and useful information is retained, and the accuracy of text classification is improved.
Patent Information
- Application Number
- CN202510101067.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
AI Technical Summary
The existing short text classification model ignores word position information, the differences in importance of the same word in different texts, and the importance of text keyword characteristics when using text embedding. In addition, the one-dimensional convolutional neural network cannot extract deep features, and the pooling layer will delete useful information.
A multi-channel convolutional capsule network based on multi-head attention network is adopted, and location information and keyword features are introduced through text feature embedding modules, and deep features are extracted through multi-head attention mechanisms and capsule networks, replacing the traditional pooling layer to retain information.
It effectively solves the problem of missing position information and keyword characteristics in text embedding, extracts deep features of the text and retains useful information, and improves the accuracy of text classification.
Smart Images

Figure CN120086378A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text classification, and particularly to a text classification method and device based on a multi-channel convolutional capsule network with multi-head attention network. Background Art
[0002] The main objective of short text classification is to classify short text segments into predefined categories. This task has wide applications in many downstream tasks, such as sentiment analysis, spam filtering, topic classification, etc. Handling the short text classification task requires extracting key features from limited text information and automatically identifying the category to which the text belongs through model learning.
[0003] Common short text classification methods include machine learning-based methods and deep learning methods. Traditional machine learning methods usually use hand-designed features, such as bag-of-words model, TF-IDF, etc., and then train through a classifier. While deep learning methods focus more on end-to-end learning, automatically extracting abstract features in the text through neural networks to achieve more accurate classification. However, due to the limited text length in short text classification, problems such as data sparsity, ambiguity, and insufficient context occur, and the model needs to use some pre-trained language models, feature extraction and other processing means to understand context and semantic information.
[0004] Text feature embedding realizes a more effective representation of text information by mapping text data into a low-dimensional dense vector space. Compared with traditional bag-of-words models and hand-designed features, text feature embedding has stronger expressive power and can capture semantic relationships and context information between words. The core idea of text feature embedding is to map each word or subsequence in the text into a real number vector, so that similar semantic contents are closer in the vector space, thereby providing richer semantic information for machine learning models.
[0005] In recent years, with the rise of deep learning, neural network-based text embedding methods, such as WordEmbeddings and BERT, etc., have become mainstream. At the same time, with the continuous increase in the complexity of tasks and the diversity of application scenarios, text feature embedding faces some challenges and problems. First, the text feature requirements for different tasks are different, and traditional static embedding models such as Word2Vec are difficult to fully meet. Therefore, researchers have gradually turned to dynamic embedding models, such as BERT, which can better capture context information through context awareness and pre-training mechanisms. However, the complexity of dynamic embedding models leads to huge demands for computing and storage resources, restricting their applications in resource-constrained environments.
[0006] To achieve better classification performance, after the text features are embedded in vector form, machine learning or neural network methods are needed to learn the latent features. In traditional machine learning text classification methods, a one-dimensional convolutional neural network is commonly used to extract features along the text sequence direction. When using a one-dimensional convolutional neural network for feature mapping, explicit feature acquisition is avoided, and learning can be implicitly performed from the training data. However, the learning ability of the one-dimensional convolutional neural network is limited, and it is difficult to obtain deep features through mapping. Moreover, when the traditional convolutional layer performs the max-pooling operation, only the maximum value of each pooling size is taken, inevitably resulting in information loss and further reducing the classification performance. Summary of the Invention
[0007] The purpose of this application is to propose a text classification method and device based on a multi-channel convolutional capsule network with a multi-head attention network, so as to solve the problems that existing short text classification models often ignore the position information of words, the importance difference of the same word in different texts, and the importance of keyword features of texts in the classification task. At the same time, the one-dimensional convolutional neural network cannot extract deep features of texts, and the pooling layer of the convolutional neural network will delete useful information during feature extraction.
[0008] In the first aspect, the present invention provides a text classification method based on a multi-channel convolutional capsule network with a multi-head attention network, including the following steps:
[0009] Obtain the text to be classified;
[0010] Construct and train a text classification model to obtain a trained text classification model. The text classification model includes a text feature embedding module, a multi-channel convolutional capsule network based on multi-head attention, and a classification module. The multi-channel convolutional capsule network based on multi-head attention includes a multi-head attention layer and a fused multi-channel convolutional layer connected in sequence. The fused multi-channel convolutional layer is constructed by using a capsule network instead of the pooling layer on the basis of a multi-channel convolutional neural network;
[0011] The text to be classified is input into the trained text classification model. In the text feature embedding module, the text to be classified generates an embedding vector by combining the initial embedding vector generated by the pre-trained Word2vec module with the calculated TF-IDF-TDF value and the word position embedding, and the final embedding vector is obtained through filtering and screening. The keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are parallel to form the text information vector output by the text feature embedding module. The text information vector is input into the multi-channel convolutional capsule network based on multi-head attention for deep feature extraction of the text, and the deep features are input into the classification module to obtain the text classification result.
[0012] Preferably, the pre-trained Word2vec module adopts the CBOW model.
[0013] Preferably, the calculation process of the TF-IDF-TDF value includes:
[0014] Calculate the frequency tf of word i appearing in document j i,j :
[0015]
[0016] where n ij represents the number of times word t i appears in document d j , ∑ k n kj represents the total number of times all words appear in document d j , and k represents any one of all words;
[0017] Calculate the frequency idf of word i in the entire corpus i :
[0018]
[0019] where |D| represents the total number of documents in the corpus, and |j:t i ∈d j | represents the total number of documents j containing word t i ;
[0020] Calculate the arithmetic document frequency tdf i,j :
[0021]
[0022] where |T j | represents the total number of words in document d j , M represents the total number of all documents, represents the proportion of the number of word t j in document d i ;
[0023] The calculation formula for the TF-IDF-TDF value of word t i is as follows:
[0024] tit = tf i,j ·idf i ·tdf i,j .
[0025] Preferably, the text to be classified generates an embedding vector by combining the initial embedding vector generated by a pre-trained Word2vec module with the calculated TF-IDF-TDF value and the positional embedding of the word, and obtains the final embedding vector through filtering and screening. The keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are parallelized into the text information vector output by the text feature embedding module, specifically including:
[0026] The input text generates an initial embedding vector ST by a pre-trained Word2vec p ={st 1 , st 2 ,..., st q}, where p is the number of initial embedding vectors, q is the number of words corresponding to each initial embedding vector, and st represents the elements in the initial embedding vector;
[0027] Calculate the positional embedding of the words in each initial embedding vector according to the positional encoding method of the Transformer, as shown in the following formula:
[0028]
[0029] where pe represents the positional embedding, pos represents the positional information of the word, d model represents the dimension of the positional vector, the dimension of the positional vector is equal to the dimension of the initial embedding vector, 2e represents the even dimension, and 2e + 1 represents the odd dimension;
[0030] The embedding vector is represented as NT p ={st 1 ·tit 1 +pe 1 , st 2 ·tit 2 +pe 2 ,..., st q ·tit q +pe q};
[0031] Filter the embedding vector according to the TF-IDF-TDF value of the corresponding word in the embedding vector, as shown in the following formula:
[0032]
[0033] V p ={v 1 , v 2 ,..., v q};
[0034] where nt q =stq ·tit q +pe q ,ξ 1 represents the first threshold, V p represents the final embedding vector, v q represents an element in the final embedding vector;
[0035] Feature extraction is performed according to the TF-IDF-TDF value of the corresponding word in the embedding vector to obtain keyword features, as shown in the following formula:
[0036]
[0037] key p ={k 1 ,k 2 ,...,k q};
[0038] Among them, ξ 2 represents the second threshold, key p represents the keyword feature, k q represents an element in the keyword feature;
[0039] The text information vector is expressed as represents a parallel operation, and the parallel operation includes a splicing operation.
[0040] Preferably, the text information vector is input into the multi-head attention layer, and the multi-head attention mechanism is used to obtain the weighted attention scores of each word in the text, and the attention features are generated after aggregation.
[0041] Preferably, the fused multi-channel convolutional layer includes a multi-channel convolutional layer and a capsule network connected in sequence. The attention features are input into the multi-channel convolutional layer to obtain multi-channel text features, and the multi-channel text features are input into the capsule network to obtain deep features. The classification module includes a fully connected neural network, and the last layer in the fully connected neural network is a softmax function layer.
[0042] In a second aspect, the present invention provides a text classification device based on a multi-channel convolutional capsule network of a multi-head attention network, including:
[0043] A text acquisition module configured to acquire the text to be classified;
[0044] A model construction module, configured to construct and train a text classification model to obtain a trained text classification model. The text classification model includes a text feature embedding module, a multi-channel convolutional capsule network based on multi-head attention, and a classification module. The multi-channel convolutional capsule network based on multi-head attention includes a multi-head attention layer and a fused multi-channel convolutional layer connected in sequence. The fused multi-channel convolutional layer is constructed by using a capsule network to replace the pooling layer on the basis of a multi-channel convolutional neural network;
[0045] A prediction module, configured to input the text to be classified into the trained text classification model. In the text feature embedding module, the text to be classified generates an embedding vector by combining the initial embedding vector generated by the pre-trained Word2vec module with the calculated TF-IDF-TDF value and the position embedding of the word, and obtains the final embedding vector through filtering and screening. The keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are parallelized to be the text information vector output by the text feature embedding module. The text information vector is input into the multi-channel convolutional capsule network based on multi-head attention for deep feature extraction of the text, and the deep features are input into the classification module to obtain the text classification result.
[0046] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0047] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0048] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] (1) The text classification method of the multi-channel convolutional capsule network based on the multi-head attention network proposed by the present invention solves the problems of lack of position information in text embedding, ignoring the differences in the importance of words for different documents, and keyword features in short text classification problems by establishing a text feature embedding module. In the text feature embedding module, position embeddings based on trigonometric functions of absolute positions used in transformers are added to introduce position information into the text embedding; on the basis of traditional TF-IDF, term document frequency (TDF) is proposed to improve the differences in the importance of words for different documents and incorporated into the text feature embedding.
[0051] (2) The text classification method of the multi-channel convolutional capsule network based on the multi-head attention network proposed by the present invention extracts deep features of the text through a multi-channel convolutional layer based on the multi-head attention mechanism and fuses the capsule network, solving the problems of few features extracted by one-dimensional convolution and useful features deleted by the convolutional pooling layer. Brief Description of the Drawings
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a schematic flowchart of the text classification method of the multi-channel convolutional capsule network based on the multi-head attention network for the embodiments of the present application;
[0054] Figure 2 It is a flowchart block diagram of the text classification method of the multi-channel convolutional capsule network based on the multi-head attention network for the embodiments of the present application;
[0055] Figure 3 It is a schematic diagram of the text classification device of the multi-channel convolutional capsule network based on the multi-head attention network for the embodiments of the present application;
[0056] Figure 4 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present invention. Detailed Embodiments
[0057] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following will further describe the present invention in detail with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0058] Figure 1 A text classification method of a multi-channel convolutional capsule network based on a multi-head attention network provided by an embodiment of the present application is shown, including the following steps:
[0059] S1. Obtain the text to be classified.
[0060] Specifically, a text data set is constructed based on a medical customer service live conversation corpus. The data in the text data set is cleaned and preprocessed (denoising, handling missing values, etc.), and some unbalanced index data is sampled to ensure data diversity. A vocabulary is constructed for the text data set using a tokenizer such as BPE for subsequent text feature embedding work such as word2vec. The principle of the BPE tokenizer is to first construct a vocabulary based on all Chinese characters / English words appearing in the text, and then combine adjacent words with the highest probability into subwords each time until the vocabulary size reaches a predetermined value. This text data set is used to train a text classification model to obtain a trained text classification model. Then, the text to be classified is obtained and input into the trained text classification model.
[0061] S2. Construct and train a text classification model to obtain a trained text classification model. The text classification model includes a text feature embedding module, a multi-channel convolutional capsule network based on multi-head attention, and a classification module. The multi-channel convolutional capsule network based on multi-head attention includes a multi-head attention layer and a fused multi-channel convolutional layer connected in sequence. The fused multi-channel convolutional layer is constructed by using a capsule network to replace the pooling layer on the basis of a multi-channel convolutional neural network.
[0062] Specifically, referring to Figure 2 , the text feature embedding module includes a pre-trained Word2vec module, a TF-IDF-TDF value calculation module, a position encoding module, a vector generation module, a vector filtering module, and a keyword feature extraction module. First, the pre-trained Word2vec module is used to represent the text, and then the TF-IDF-TDF value is calculated by the TF-IDF-TDF value calculation module for feature screening and keyword extraction. Then, the position embedding of the word is calculated by the position encoding module, and the position encoding module adopts the absolute position encoding method based on trigonometric functions of Transformer. Finally, the embedded vectors are filtered and output in parallel with the keyword features.
[0063] Furthermore, the text information vector output by the text feature embedding module is fed into the multi-channel convolutional capsule network based on multi-head attention for deep text feature extraction. Then, the deep features are input into the classification module to obtain the text classification result.
[0064] S3. The text to be classified is input into the trained text classification model. In the text feature embedding module, the text to be classified combines the initial embedding vectors generated by the pre-trained Word2vec module with the calculated TF-IDF-TDF values and the position embeddings of the words to generate embedding vectors, and the final embedding vectors are obtained through filtering and screening. The keyword features in the embedding vectors are screened out by the TF-IDF-TDF values, and the final embedding vectors and the keyword features are parallelized as the text information vectors output by the text feature embedding module. The text information vectors are input into the multi-channel convolutional capsule network based on multi-head attention for deep feature extraction of the text, and the deep features are input into the classification module to obtain the text classification result.
[0065] In a specific embodiment, the pre-trained Word2vec module adopts the CBOW model.
[0066] Specifically, Word2vec can convert words into computable and structured vectors. Its two training methods are the continuous bag of words (CBOW) model and the continuous skip-gram model. In the embodiments of this application, the CBOW model is used to calculate the probability of the central word w i-n , w i-n+1 , … w i-1 , w i+1 , …, w i+n-1 , w i+n appearing according to the n consecutive words w i before and after the central word:
[0067] p(w f |w f-r , w f-r+1 , … w f-1 , w f+1 , …, w f+r-1 , w f+r ).
[0068] Among them, the r consecutive words w f-r , w f-r+1 , …, w f-1 , w f+1 , …, w f+r-1 , w i+r can be abbreviated as w window . To maximize the value of this likelihood function is equivalent to minimizing the value of the following loss function:
[0069]
[0070] where F represents the length of the sentence. Substituting the specific probability formula gives the following formula:
[0071]
[0072] The above formula is the loss function of the CBOW model, where h r represents the average vector of the context word w window . The calculation of this average vector is obtained by multiplying a trainable vector of the vocabulary length × a given hyperparameter size by the one-hot vector matrix of w window and then summing them up; corpus represents the vocabulary constructed by the BPE tokenizer in the construction of the text dataset.
[0073] In a specific embodiment, the calculation process of the TF-IDF-TDF value includes:
[0074] Calculating the frequency tf i,j of the word i in the document j:
[0075]
[0076] where n ij represents the number of times the word t i appears in the document d j , ∑ k n kj represents the total number of times all words appear in the document d j , and k represents any one of all words;
[0077] Calculating the inverse document frequency idf i of the word i in the entire corpus:
[0078]
[0079] where |D| represents the total number of documents in the corpus, and |j:t i ∈d j | represents the total number of documents j that contain the word t i ;
[0080] Calculating the term document frequency tdf i,j :
[0081]
[0082] where |T j | represents the total number of words in the document d j , M represents the number of all documents, represents the proportion of the number of the word t j in the document d i ;
[0083] The calculation formula of the TF-IDF-TDF value of the word t i is as follows:
[0084] tit = tfi,j ·idf i ·tdf i,j 。
[0085] Specifically, since it is very labor-intensive to label data with supervised methods, the embodiments of this application extract keyword features while performing feature screening based on TF-IDF-TDF values. Term Frequency-Inverse Document Frequency (TF-IDF) is a commonly used text mining weighting technique, whose purpose is to evaluate the importance of a word to the documents in a corpus by the frequency of a specific word in a document and its frequency in the corpus. Since the importance of the same word is different in different documents and TF-IDF cannot be simply used to represent the importance of a word, the embodiments of this application propose the Term Document Frequency (TDF) to represent the importance of the same word in different documents. There is a +1 in the denominator of the calculation formula of the term document frequency to avoid the situation where the word i does not appear in the entire text, that is from occurring.
[0086] In a specific embodiment, the text to be classified generates an initial embedding vector through a pre-trained Word2vec module, combines it with the calculated TF-IDF-TDF value and the positional embedding of the word to generate an embedding vector, and obtains the final embedding vector through filtering and screening. The keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are parallel to form the text information vector output by the text feature embedding module, specifically including:
[0087] The input text generates an initial embedding vector ST through a pre-trained Word2vec p ={st 1 , st 2 ,..., st q}, where p is the number of initial embedding vectors, q is the number of words corresponding to each initial embedding vector, and st represents the elements in the initial embedding vector;
[0088] Calculate the positional embedding of the words in each initial embedding vector according to the positional encoding method of Transformer, as shown in the following formula:
[0089]
[0090] where pe represents the positional embedding, pos represents the positional information of the word, d model represents the dimension of the positional vector, the dimension of the positional vector is equal to the dimension of the initial embedding vector, 2e represents the even dimension, and 2e + 1 represents the odd dimension;
[0091] The embedding vector is represented as NT p ={st 1·tit 1 +pe 1 ,st 2 ·tit 2 +pe 2 ,...,st q ·tit q +pe q};
[0092] Filter the embedding vectors according to the TF-IDF-TDF values of the corresponding words in the embedding vectors, as shown in the following formula:
[0093]
[0094] V p ={v 1 ,v 2 ,...,v q};
[0095] Among them, nt q =st q ·tit q +pe q , ξ 1 represents the first threshold, V p represents the final embedding vector, v q represents the elements in the final embedding vector;
[0096] Extract features according to the TF-IDF-TDF values of the corresponding words in the embedding vectors to obtain keyword features, as shown in the following formula:
[0097]
[0098] key p ={k 1 ,k 2 ,...,k q};
[0099] Among them, ξ 2 represents the second threshold, key p represents the keyword features, k q represents the elements in the keyword features;
[0100] The text information vector is expressed as denotes a parallel operation, and the parallel operation includes a splicing operation.
[0101] Specifically, the magnitude of the TF-IDF-TDF value reflects the importance of a word to the text. Since there are usually some words in a document that do not reflect the document category, this does not contribute to the text classification result and may even be misleading. Therefore, removing words with smaller TF-IDF-TDF values can improve the classification accuracy. Thus, the elements in the embedding vector are filtered according to the TF-IDF-TDF values to obtain the final embedding vector V p Moreover, the larger the TF-IDF-TDF value, the more important the word is, which is the keyword feature of the text. The model can classify the text more accurately based on the keyword information. Therefore, extracting words with larger TF-IDF-TDF values can improve the classification accuracy. In the embodiments of this application, the keyword features in the embedding vector are screened according to the TF-IDF-TDF values. Finally, the final embedding vector and the keyword features are concatenated into a text information vector as the output of the final text feature embedding module. Here, concatenation means the concat operation in tensor operations. Specifically, the final embedding vector and the keyword features are vectors with shapes (batch size, vector length, model dimension 1) and (batch size, vector length, model dimension 2) respectively. Under the torch deep learning framework, the concat(dim=-1) operation is performed to merge them into a vector of (batch size, vector length, model dimension 1 + model dimension 2).
[0102] In a specific embodiment, the text information vector is input into a multi-head attention layer, and the multi-head attention mechanism is used to obtain the weighted attention scores of each word in the text, and after aggregation, attention features are generated.
[0103] Specifically, to enhance the expressive power of the model, the embodiments of this application use the multi-head attention mechanism through the multi-head attention layer to obtain the weighted attention scores of each word in the text. The multi-head attention mechanism introduces multiple attention heads on the basis of the attention mechanism. Each head learns different attention weights, allowing the model to simultaneously focus on different parts in different representation spaces. The basic attention mechanism combines each element in the input sequence with other elements with weights to obtain a weighted sum for generating the representation after attention aggregation. The multi-head attention enables the model to learn multiple different attention aggregations by introducing multiple independent attention heads, which helps the model capture the information of the input sequence more comprehensively. The calculation formula of the multi-head attention mechanism is as follows:
[0104]
[0105] Wherein, It is to prevent the attention (Q, K, V) from having an overly large scale factor. Q, K, and V represent Query, Key, and Value respectively. This formula is used to calculate the attention correlation between each token in the sentence, and they are respectively obtained by calculating with the input tensor and the parameter-trainable linear layer (which can be represented as torch.functional.nn.Linear(seq_len, d_model / / heads)).
[0106] They are respectively calculated, and the input tensor is the text information vector.
[0107] In this formula, QK T The calculation result is a tensor with both length and width equal to the sequence length, which represents the scores of the attention coefficients between each token. After this score goes through scaling (dividing by ), and then multiplying by V after passing through the softmax activation function, the obtained is the preliminary output tensor that simultaneously contains the attention scores and the weights of each token itself. After this tensor goes through subsequent operations such as residual connection, layer normalization, and feed-forward layer, it is the output of the multi-head attention layer. Assuming the number of 'heads' is h, then each 'head' is calculated as follows:
[0108]
[0109] Among them, respectively represent the linear transformation parameter matrices of each 'head'. The output of the multi-head attention layer is obtained by concatenating each head a together and performing a linear transformation, as shown in the following formula:
[0110] MultiHead = Concate(head 1 , head 2 ,..., head h )W O ;
[0111] Among them, W O represents the learnable parameter.
[0112] In a specific embodiment, the fused multi-channel convolutional layer includes a multi-channel convolutional layer and a capsule network connected in sequence. The attention features are input into the multi-channel convolutional layer to obtain multi-channel text features, and the multi-channel text features are input into the capsule network to obtain deep features. The classification module includes a fully connected neural network, and the last layer in the fully connected neural network is the softmax function layer.
[0113] Specifically, due to the complexity of the text classification task, the text data contains many deep features. The multi-channel convolutional neural network can capture the complex structures and correlations in the input text data more comprehensively than a single-channel one, thereby improving the network's representation ability. However, the pooling layer of the convolutional neural network will delete some useful information. Therefore, the embodiments of the present application use a capsule network to replace the multi-channel convolutional layer of the pooling layer to form a fused multi-channel convolutional layer for extracting multi-channel text features.
[0114] The multi-channel convolutional layer considers the information of multiple channels simultaneously during the convolutional operation. For each channel, there is a corresponding convolutional kernel, and these convolutional kernels are respectively used to extract features from each channel. Then, the convolutional results of each channel are added element-wise to form the final output. This can better retain the feature information in the multi-channel input data. Specifically, assume there are two input channels corresponding to two convolutional kernels. During the convolutional operation, the convolutional kernel of each channel performs a convolutional operation with the input of the corresponding channel to obtain two independent feature maps. Then, these two feature maps are added element-wise to form the final output. The calculation formula for the multi-channel text features output by the multi-channel convolutional layer is as follows:
[0115]
[0116] where, Z i,j,k represents the output value at the j-th row and k-th column of the i-th channel, that is, the multi-channel text feature, V l,j+m,k+n is the value at the (j + m)-th row and (k + n)-th column of the l-th channel in the attention feature, X i,l,m,n represents the input value at the m-th row and n-th column of the l-th channel of the i-th convolutional kernel, represents the element-wise multiplication operation, and relu() is the relu activation function.
[0117] In traditional CNNs, the information passed between layers is through weight matrices, while the capsule network introduces the concept of capsules, attempting to better model the spatial relationships and hierarchical structures between features. A capsule is a combination containing multiple neurons, and these neurons jointly describe a specific feature or concept. These capsules interact with each other through dynamic routing to transmit information in the hierarchical structure. Dynamic routing dynamically adjusts the weights between capsules so that relevant information can be transmitted more effectively. Therefore, the capsule network can learn the relevant information between the local and the entire text.
[0118] The capsules in the capsule network contain rich information such as spatial location, and adjacent nodes have strong correlations, which can retain the underlying details in the original data. These features only conform to the relationships and natural order between contexts in text data and can be well used for text classification tasks. The dynamic routing algorithm between capsules replaces the max pooling algorithm in traditional CNNs, avoiding information loss caused by pooling operations. Pass the multi-channel text features Z i,j,k as the input through the capsule network. If the number of layers of the multi-channel convolutional layer is n, then the number of layers of the capsule network is n - 1. According to the change in the size of the input text, the number of layers n of the multi-channel convolutional layer is an adjustable hyperparameter. The calculation formula of the capsule network is as follows:
[0119] u j′|i′ =W i′j′ u i′ ;
[0120]
[0121] The process of dynamically routing to update the weights is as follows:
[0122] s j′ =∑ i′ cout i′j′ u j′|i′ ;
[0123]
[0124] b i′j′ ←b i′j′ +v j′ u j′|i′ ;
[0125] Among them, u i′ represents the output of the capsules in the previous layer, u j′|i′ is the predicted vector transformed from u i′ , and W i′j′ is the transformation matrix. s j′ is the total input of the high-level capsules, and v j′ is the total output of the capsule network.
[0126] Input the deep features output by the capsule network into the classification module. The classification module uses a fully connected neural network, and the final text classification result is obtained through the last softmax function layer in the fully connected neural network. The text classification result is the probability corresponding to each text category, and the text category corresponding to the maximum probability is taken as the final classification result, as shown in the following formula:
[0127] label=argmax(softmax(V out ));
[0128] Among them, label represents the final classification result, and V out represents the input of the last layer of the fully connected neural network.
[0129] The above steps S1 - S3 do not necessarily represent the order between steps, but are step symbols. The order between steps can be adjusted.
[0130] For further reference Figure 3 , as an implementation of the methods shown in the above figures, an embodiment of a text classification device based on a multi - head attention network and a multi - channel convolutional capsule network is provided in the present application. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0131] An embodiment of the present application provides a text classification device based on a multi - head attention network and a multi - channel convolutional capsule network, including:
[0132] A text acquisition module 1, configured to acquire the text to be classified;
[0133] A model construction module 2, configured to construct and train a text classification model to obtain a trained text classification model. The text classification model includes a text feature embedding module, a multi - head attention - based multi - channel convolutional capsule network, and a classification module. The multi - head attention - based multi - channel convolutional capsule network includes a multi - head attention layer and a fused multi - channel convolutional layer connected in sequence. The fused multi - channel convolutional layer is constructed by using a capsule network to replace the pooling layer on the basis of a multi - channel convolutional neural network;
[0134] A prediction module 3, configured to input the text to be classified into the trained text classification model. In the text feature embedding module, the text to be classified generates an embedding vector by combining the initial embedding vector generated by a pre - trained Word2vec module with the calculated TF - IDF - TDF value and the word position embedding, and obtains a final embedding vector through filtering and screening. The keyword features in the embedding vector are screened out by the TF - IDF - TDF value, and the final embedding vector and the keyword features are parallelized to be the text information vector output by the text feature embedding module. The text information vector is input into the multi - head attention - based multi - channel convolutional capsule network for deep feature extraction of the text, and the deep features are input into the classification module to obtain the text classification result.
[0135] Figure 4 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. As shown in Figure 4As shown in the figure, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer-executable instructions; the processor 401 is used to execute the computer-executable instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0136] Optionally, the memory 402 can be either independent or integrated with the processor 401.
[0137] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401.
[0138] The embodiment of the present invention also provides a computer storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the above method is implemented.
[0139] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor, the above method is implemented.
[0140] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0141] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0142] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module exists physically alone, or two or more modules are integrated in one unit. The units formed by the above modules can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0143] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods of the various embodiments of the present application.
[0144] It should be understood that the above-mentioned processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.
[0145] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0146] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.
[0147] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0148] An exemplary storage medium is coupled to a processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master device.
[0149] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text classification method based on a multi-channel convolutional capsule network with multi-head attention network, characterized in that: The following steps are involved: Get the text to be classified; Constructing and training a text classification model to obtain a trained text classification model, wherein the text classification model includes a text feature embedding module, a multi-channel convolutional capsule network based on multi-head attention, and a classification module, wherein the multi-channel convolutional capsule network based on multi-head attention includes a multi-head attention layer and a fused multi-channel convolutional layer connected in sequence, wherein the fused multi-channel convolutional layer is constructed by using a capsule network instead of a pooling layer on the basis of a multi-channel convolutional neural network; The text to be classified is input into the trained text classification model. In the text feature embedding module, the text to be classified is embedded by combining the initial embedding vector generated by the pre-trained Word2vec module with the calculated TF-IDF-TDF value and the position embedding of the word to generate an embedding vector, and a final embedding vector is obtained through filtering and screening. The keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are parallelized to form a text information vector output by the text feature embedding module; the text information vector is input into the multi-channel convolutional capsule network based on multi-head attention to extract deep features of the text, and the deep features are input into the classification module to obtain a text classification result.
2. The text classification method based on a multi-channel convolutional capsule network with multi-head attention network according to claim 1, characterized in that: The pre-trained Word2vec module adopts the CBOW model.
3. The text classification method based on a multi-channel convolutional capsule network with multi-head attention network according to claim 1, characterized in that: The calculation process of the TF-IDF-TDF value includes: Calculate the frequency tf of word i in document j i,j : Among them, n ij Represents word t i In the document d j The number of times it appears in k n kj Indicates all words in document d j The number of times it appears in, k represents any one of all words; Calculate the frequency idf of word i in the entire corpus i : Where |D| represents the total number of documents in the corpus, |j:t i ∈d j | indicates that the word t is included i The total number of documents j; Calculate term document frequency tdf i,j : Among them, |T j | represents the total number of words in document dj, M represents the number of all documents, Represents document d j Chinese word t i The proportion of quantity; Word t i The TF-IDF-TDF value is calculated as follows: tit=tfi , j·idf i ·tdf i,j 。 4. The text classification method of the multi-channel convolutional capsule network based on the multi-head attention network according to claim 3 is characterized in that: The text to be classified is generated by combining the initial embedding vector generated by the pre-trained Word2vec module with the calculated TF-IDF-TDF value and the word position embedding to generate an embedding vector, and the final embedding vector is obtained by filtering and screening, and the keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are combined to form the text information vector output by the text feature embedding module, which specifically includes: The input text generates the initial embedding vector ST from the pre-trained Word2vec p = {st1, st2, ..., St q }, where p is the number of initial embedding vectors, q is the number of words corresponding to each initial embedding vector, and st represents an element in the initial embedding vector; The position embedding of words in each initial embedding vector is calculated according to the Transformer's position encoding method, as shown in the following formula: Among them, pe represents position embedding, pos represents the position information of the word, and d model represents the dimension of the position vector, which is equal to the dimension of the initial embedding vector, 2e represents an even dimension, and 2e+1 represents an odd dimension; The embedding vector is denoted as NT p ={st1·tit1+pe1, st2·tit2+pe2,..., st q ·tit q +pe q }; The embedding vector is filtered according to the TF-IDF-TDF value of the corresponding word in the embedding vector, as shown in the following formula: V p ={v1,v2,...,v q }; Among them, nt q =st q ·tit q +pe q , ξ1 represents the first threshold, V p represents the final embedding vector, v q Represents the elements in the final embedding vector; Feature extraction is performed based on the TF-IDF-TDF value of the corresponding word in the embedding vector to obtain keyword features, as shown in the following formula: key p ={k1,k2,...,k q }; Among them, ξ2 represents the second threshold, key p represents the keyword feature, k q Represents the elements in the keyword feature; The text information vector is represented as Represents a parallel operation, wherein the parallel operation includes a concatenation operation.
5. The text classification method based on a multi-channel convolutional capsule network with multi-head attention network according to claim 1, characterized in that: The text information vector is input into the multi-head attention layer, and a multi-head attention mechanism is used to obtain a weighted attention score for each word in the text, and after aggregation, an attention feature is generated.
6. The text classification method based on a multi-channel convolutional capsule network with multi-head attention network according to claim 5, characterized in that: The fused multi-channel convolution layer includes a multi-channel convolution layer and a capsule network connected in sequence, the attention feature is input into the multi-channel convolution layer to obtain a multi-channel text feature, the multi-channel text feature is input into the capsule network to obtain a deep feature, the classification module includes a fully connected neural network, and the last layer in the fully connected neural network is a softmax function layer.
7. A text classification device based on a multi-channel convolutional capsule network with multi-head attention network, characterized in that: include: A text acquisition module, configured to acquire text to be classified; A model building module is configured to build and train a text classification model to obtain a trained text classification model, wherein the text classification model includes a text feature embedding module, a multi-channel convolutional capsule network based on multi-head attention, and a classification module, wherein the multi-channel convolutional capsule network based on multi-head attention includes a multi-head attention layer and a fused multi-channel convolutional layer connected in sequence, wherein the fused multi-channel convolutional layer is constructed by using a capsule network instead of a pooling layer on the basis of a multi-channel convolutional neural network; The prediction module is configured to input the trained text classification model into the text to be classified. In the text feature embedding module, the text to be classified is embedded by combining the initial embedding vector generated by the pre-trained Word2vec module with the calculated TF-IDF-TDF value and the position embedding of the word to generate an embedding vector, and the final embedding vector is obtained through filtering and screening. The keyword features in the embedding vector are screened out by the TF-IDF-TDF value, and the final embedding vector and the keyword features are parallelized as the text information vector output by the text feature embedding module; the text information vector is input into the multi-channel convolutional capsule network based on multi-head attention to extract deep features of the text, and the deep features are input into the classification module to obtain a text classification result.
8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.