Text Classification Method Based on Dual-Channel Semantic Enhancement and Convolutional Neural Network

By introducing dual-channel convolution and semantic enhancement mechanisms into the text classification model, the existing model's insufficient generalization ability in different languages ​​and text lengths and the sparsity of short text semantic features are solved, and stronger semantic mining and classification performance are achieved.

CN119513318BActive Publication Date: 2025-07-01CHONGQING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411663418.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-07-01
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

When existing deep learning models deal with different languages ​​and different text lengths, the generalization ability and general applicability of the model need to be further strengthened, and the semantic feature sparsity and finiteness in short texts lead to poor classification results.

Method used

A text classification method based on dual-channel semantic enhancement and convolutional neural network is proposed. Dual-channel convolution feature extraction is performed through Conv1D and AtrousConv1D, and combined with weighted average attention and semantic enhancement modules, the model's semantic mining ability and long-distance dependency exploration ability are improved.

Benefits of technology

It is implemented for Chinese and English text classification tasks, with good classification performance and generalization capabilities, especially in both long text and short text categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119513318B_ABST
    Figure CN119513318B_ABST
Patent Text Reader

Abstract

The present invention proposes a text classification method based on dual-channel semantic enhancement and convolutional neural network, which includes the following steps: S1, first preprocess the text and perform word vector embedding to obtain a text matrix X; S2, respectively use Conv1D and AtrousConv1D on the generated text matrix X for dual-channel convolutional feature extraction to obtain the original semantic information C and the global text information A; S3, use weighted average attention on the generated original semantic information C and global text information A to generate attention scores C score and A score , and at the same time perform semantic enhancement on the text matrix X to obtain y k . Finally, splice the high-dimensional convolutional feature maps, and then map the spliced feature maps to the probability distribution of the labels through the Linear fully connected layer and the Softmax layer. The present invention can be applied to Chinese and English text classification tasks, and has good classification performance and generalization ability in both long texts and short texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text classification, and in particular to a text classification method based on dual-channel semantic enhancement and convolutional neural network. Background Art

[0002] With the popularization of social media, a large amount of text data has been generated on the Internet. Text classification, as a key area of natural language processing, plays an important role in tasks such as spam filtering, news classification, and sentiment analysis. However, traditional machine learning models require a large amount of manual cost for text feature annotation when dealing with text classification, especially when dealing with short texts and long texts. Short texts have poor classification effects due to limited context information, while long texts are difficult to capture long-distance dependency relationships. The emergence of deep learning models provides a new way to solve these problems, mainly through two stages of pre-trained word vector representation and end-to-end text classification, effectively overcoming the dilemma faced by manual annotation of text information. Among them, word vector embedding methods such as Word2Vec, GloVe, and FastText are widely used, and deep learning models such as TextCNN and TextRNN perform well in sentence semantic mining and long-distance dependency information capture. Nevertheless, these models still have certain limitations, such as the problem of gradient disappearance or gradient explosion faced by TextRNN.

[0003] Although deep learning models have achieved certain results in text classification tasks, they still face many challenges. First of all, when existing deep learning models handle different languages (such as Chinese and English) and different text lengths (such as short texts and long texts), the generalization ability and general applicability of the models need to be further strengthened. Secondly, although some research works capture deep semantic information by introducing modules such as attention mechanisms, neighborhood information, and feature reconstruction, as well as adopting integration strategies, the text mining ability of these methods is still poor in practical applications. For example, in response to the sparsity of semantic features and the limited semantic information of short texts, Ai W, Wei Y, Shao H et al. developed a graph attention network EMGAN based on an edge enhancement module (see the paper "Edge-enhanced minimum-margin graph attention network for short text classification" in 《Expert Systems with Applications》). The heterogeneous information graph constructed in this model can effectively represent the complex relationships of text features and enhance semantic information. However, such methods have high complexity and lack good scalability for text mining tasks in different environments to a certain extent. Therefore, how to construct a robust text classification model that can be applied to Chinese and English contexts and different text lengths is the key issue in the current research on text classification tasks. Summary of the Invention

[0004] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a text classification method based on dual-channel semantic enhancement and convolutional neural network.

[0005] To achieve the above object of the present invention, the present invention provides a text classification method based on dual-channel semantic enhancement and convolutional neural network, including the following steps:

[0006] S1. First, preprocess the text to divide the text into text units; then perform word vector embedding to convert words or phrases into vector representations; thus, obtain a text matrix X.

[0007] S2. Respectively use Conv1D (one-dimensional convolution) and AtrousConv1D (one-dimensional dilated convolution) on the generated text matrix X for dual-channel convolutional feature extraction to obtain original semantic information C and global text information A.

[0008] S3. Use weighted average attention on the generated original semantic information C and global text information A to generate attention scores C score and A score , and at the same time perform semantic enhancement on the text matrix X to obtain y k, finally, the high-dimensional convolutional feature maps are concatenated, and then the concatenated feature maps are mapped to the probability distribution of the labels through the Linear fully connected layer and the Softmax layer.

[0009] Preferably, the text in step S1 is a Chinese text or an English text. At this time, the word vector embedding is as follows:

[0010] When it is a Chinese text, each Chinese character is represented by a pre-trained word vector of a Chinese corpus;

[0011] When it is an English text, each word is represented by a pre-trained word vector of an English corpus.

[0012] Preferably, for the word vector embedding, converting the text unit into a vector representation includes: the text S passes through the corresponding index of the vocabulary V to form the text embedding vector X:

[0013] X = index(S) * V 1)

[0014] Among them, index(S) represents the process of generating X by the text through the indexed pre-trained corpus;

[0015] V represents the vocabulary.

[0016] Preferably, Conv1D (one-dimensional convolution) is used for feature extraction, and obtaining the original semantic information C includes the following steps:

[0017] For the i-th word representation vector of the text S being x i , under the one-dimensional convolution kernel w with a convolution kernel size of z for the text embedding vector X, the sliding convolution of X is performed through Equation (2) to generate the feature h i :

[0018] [x i :x i+z-1 = [x i ,x i+1 ,...,x i+z-1 2)

[0019] h i = σ(w1 · [x i :x i+z-1 + b1) 3)

[0020] Among them, [x i :x i+z-1 represents the sliding window range of the convolution kernel;

[0021] x i represents the word vector of the i-th word of a certain sentence;

[0022] x i+z-1The word vector representing the (i+z-1)-th word of a certain sentence;

[0023] w1 represents the feature weight, which is an optimizable parameter matrix;

[0024] h i Represents the convolutional feature information after the i-th sliding convolution of the convolutional kernel, which is generated by the dot product operation of the optimizable parameter w1 and the sentence vector plus the bias b1;

[0025] σ is a non-linear activation function, specifically LeakyReLU; this is because there is a situation where neurons are discarded in the activation function ReLU. Therefore, it is more appropriate to choose the activation function LeakyReLU when facing the problem of dead neurons in the network.

[0026] The convolutional feature h after convolution i Is mapped and represented as a new convolutional feature H:

[0027] H = [h1, h2,..., h l-z+1 5)

[0028] Thus, the extracted convolutional feature M is obtained:

[0029] M = [H1, H2,..., H m 6)

[0030] Among them, H1 represents the convolutional feature generated by the first convolutional channel;

[0031] H m Represents the convolutional feature generated by the m-th convolutional channel;

[0032] m is the total number of convolutional channels;

[0033] In order to reduce the loss of semantic information in the features extracted by the convolutional layer, we perform an average pooling operation on the convolutional feature map M to obtain the original semantic information of the one-dimensional convolutional channel. The pooling output C of the convolutional feature map M is expressed as:

[0034] C = avgpool(M)7)

[0035] Among them, C represents the original semantic information;

[0036] avgpool() represents average pooling, which is used for Conv1D to capture the original semantic information of the text.

[0037] Preferably, the convolutional kernel size z ≤ 4.

[0038] Preferably, the convolutional kernel size is set to z = {2, 3, 4}.

[0039] Preferably, AtrousConv1D (one-dimensional dilated convolution) is used for feature extraction to obtain the global text information A:

[0040] A = maxpool{LeakyReLU(Z)}9)

[0041] Among them, LeakyReLU() represents the non-linear activation function; this is because there is a situation where neurons are discarded in the activation function ReLU. Therefore, it is more appropriate to choose the activation function LeakyReLU when facing the problem of dead neurons in the network.

[0042] maxpool() represents the max pooling operation;

[0043] Z = AtrConv1D(X)

[0044] Among them, Z = [z1, z2,..., z K , after dilated convolution with a filter w of length K k generates the global output z i :

[0045]

[0046] Among them, x[i + p·k] represents the word vector after dilation of AtrousConv1D;

[0047] p represents the stride for dilated convolution of the input vector x[i + p·k], p ∈ P, P = {p1, p2, p3}, and p1, p2, p3 represent three semantic receptive field dilation rates.

[0048] The advantage of using AtrousConv1D is to integrate the global semantic acquisition of one-dimensional convolution and at the same time be able to skip the sequence of p·k to obtain long-distance semantic dependencies. Through the feature map M after AtroussConv1D and LeakyReLU activation a perform the maxpool operation to obtain the maximum feature map A of sentence-level category preferences.

[0049] Preferably, step S3 includes the following steps:

[0050] S3-1, use weighted average attention on the generated original semantic information C and global text information A to generate attention scores C score and A score :

[0051]

[0052] Among them, C score and A scoreWeighted attention scores obtained for Conv1D and AtrousConv1D convolutions respectively;

[0053] softmax(v c T u c ) i is the matrix product probability mapping of v c T and u c ;

[0054] softmax(v a T u a ) i is the matrix product probability mapping of v a T and u a ;

[0055] T is the transpose symbol;

[0056] C i is the i-th convolutional feature map of C;

[0057] A i is the i-th convolutional feature map of A;

[0058] l is the sentence length;

[0059] z is the convolutional kernel size;

[0060] v c and v a are the feature vectors H and H a respectively, where H is the mapped feature after the text matrix X passes through the Conv1D convolution; H a is the mapped feature after the text matrix X passes through the AtrousConv1D convolution;

[0061] u c and u a are the high-dimensional feature scores after linear transformation, as shown in Equation (12):

[0062]

[0063] where M, M a are the convolutional features passing through the one-dimensional convolutional channel and the convolutional features passing through the AtrousConv1D convolutional channel respectively;

[0064] tanh is the activation function for calculating the attention score;

[0065] w1 and w2 represent the feature weights of the Conv1D and AtrousConv1D convolutional channels respectively;

[0066] b1 and b2 respectively represent the biases of the Conv1D and AtrousConv1D convolution channels;

[0067] S3-2, using the semantic enhancement module, the y generated after activating and enhancing the text matrix X k , is expressed as:

[0068] y k = σ1(conv(X; w k , b k ))13)

[0069] where σ1 represents the ReLU activation function;

[0070] y k represents the enhanced feature;

[0071] conv() represents the convolution operation, which is a one-dimensional convolution;

[0072] w k represents the feature weights of the convolution channel;

[0073] b k represents the bias of the convolution channel;

[0074] k represents the optimizable feature weights of the convolution channel (Conv1D or AtrousConv1D);

[0075] S3-3, dividing y k after maxpool1D to obtain the enhanced features y k c and y k a ; then concatenating y k c and y k a with the weighted attention features extracted by two-channel convolution:

[0076]

[0077] where C and A are the original semantic information and global text information obtained by Conv1D and AtrousConv1D convolutions respectively;

[0078] C score and A score are the weighted attention scores obtained by Conv1D and AtrousConv1D convolutions respectively;

[0079] y k c represents the high-dimensional volume feature map C·C scoreSemantic enhancement features generated by splicing;

[0080] y k a Represents the high-dimensional convolutional feature map A·A score Semantic enhancement features generated by splicing;

[0081] Represents the splicing operation;

[0082] S3-4, map the spliced feature map H cla Through the Linear fully connected layer and the Sofrmax layer to the probability distribution of the label.

[0083] Preferably, the cross-entropy loss function is used to measure the difference between the predicted probability distribution and the true label, and the cross-entropy loss function is expressed as;

[0084]

[0085] Where Represents the loss for classification;

[0086] n is the number of test samples;

[0087] y is the true label;

[0088] Is the predicted label.

[0089] In summary, due to the above technical solutions, the beneficial effects of the present invention are:

[0090] 1) A text classification method applicable to Chinese and English text classification tasks is proposed, and it has good classification performance and generalization ability in both long texts and short texts.

[0091] 2) A dual-channel text convolutional neural network structure with the synergistic effect of semantic enhancement and weighted attention is constructed, which has stronger semantic mining ability and longer distance dependence exploration ability in sentence-level text classification tasks.

[0092] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Brief Description of the Drawings

[0093] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein:

[0094] Figure 1 Is a schematic diagram of the framework of the DcSeCNN model of the present invention.

[0095] Figure 2It is a schematic diagram of one-dimensional convolution and AtrousConv1D convolution of the present invention.

[0096] Figure 3 It is a schematic diagram of the comparison of all dataset models (Acc).

[0097] Figure 4 It is a schematic diagram of the influence of the dropout size in DcSeCNN (Acc).

[0098] Figure 5 It is a schematic diagram of the influence of the λ size in DcSeCNN (Acc).

[0099] Figure 6 It is a schematic diagram of the influence of the padsize on short texts in three comparison models.

[0100] Figure 7 It is a schematic diagram of the influence of the padsize on long texts in three comparison models. Detailed implementation manners

[0101] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the drawings below are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0102] The present invention aims to construct a lightweight text classification model with high robustness, called DcSeCNN; this model uses TextCNN for feature enhancement to cope with the performance challenges in complex text mining tasks. With the explosive growth of social media data, tasks in fields such as news annotation and information retrieval have put forward higher requirements for the accuracy of short text classification technology. However, in variable and complex situations, deep learning models based on deep semantic mining often fail to achieve ideal performance. To solve this problem, this invention patent proposes an innovative solution, that is, applying effective semantic enhancement and attention mechanism variants to TextCNN to develop a new lightweight and robust text classification model. Through this method, the model can more effectively mine the deep semantic information in the text and improve the classification accuracy under complex conditions.

[0103] The present invention not only utilizes the advantages of TextCNN in feature extraction, but also enhances the robustness and accuracy of the model in complex text classification tasks by introducing semantic enhancement and attention mechanism variants, providing strong support for the efficient processing of social media data.

[0104] The framework diagram of the DcSeCNN model for text classification tasks of the present invention is as Figure 1As shown. The framework mainly consists of three modules: the preprocessing stage for Chinese and English respectively, and dual-channel convolution, semantic enhancement, and weighted attention. This simple and lightweight DcSeCNN structure can perform effective semantic mining and knowledge capture in various text environments (different text scales, different text lengths, Chinese and English). Specifically: First, perform text preprocessing and word vector embedding on the text (Chinese or English). For Chinese text, each Chinese character is represented using pre-trained word vectors from the SougouNews corpus, and for English, each word is represented using pre-trained word vectors from the GloVe corpus. The generated text matrix X is respectively subjected to dual-channel convolution feature extraction using Conv1D and AtrousConv1D, and the generated C and A are used to generate attention scores C score and A score . At the same time, semantic enhancement is performed on X to obtain y k . Finally, the high-dimensional convolutional feature maps are concatenated and generated through the Linear fully connected layer and Sofrmax

[0105] 1 Text preprocessing stage

[0106] Text preprocessing is the process of converting sentences in a dataset into a real-valued matrix through pre-trained word vectors. The input sentence can be represented as S = {s1, s2,..., s l}, S ∈ S; where s i represents each word or Chinese character in the sentence, l is the length of a sentence, and can also be regarded as l tokens of this text representation. Here we use GloVe and pre-trained embeddings from the SougouNews corpus, and the dimension d of each word embedding vector is 300. Finally, the embeddings of all sentences in the dataset can form a vocabulary v represents the size of the vocabulary, and d represents the embedding dimension 300. All the word vectors of the vocabulary V constructed from a dataset can be understood as the word vectors corresponding to each Chinese character or word in this dataset. Each word s i is converted into a vector x after word vector embedding i . At this time, the text S is represented as a matrix x l n represents the word vector corresponding to the last word or Chinese character of the nth sentence, with a size of 1*300; its mapping relationship is shown in Equation (1).

[0107] X = index(S) * V 1)

[0108] where n represents the total number of sentences;

[0109] index(S) represents the process of generating X from the original text by indexing the pre-trained corpus;

[0110] V represents the vocabulary;

[0111] The corresponding index of the text S through the vocabulary V constitutes the text embedding vector X recognizable by the computer.

[0112] The text embedding vector X is divided into a training set X train , a validation set X dev and a test set X test , where X train and X dev are used as the input of the model, and the model after training and hyperparameter optimization classifies and predicts X test . To simplify the complexity of the model principle, we use X to represent the sentence-level embedding vector participating in training in the following introduction.

[0113] 2 Dual-channel convolutional module

[0114] The matrix X output in the previous stage is simultaneously fed forward into the dual channels of the classical one-dimensional convolution and the Atrous convolution to obtain the feature sequence. The one-dimensional convolution can obtain the local original semantic information through the sentence sequence, while the Atrous convolution can expand to obtain the global text information at the sentence level, increasing the probability of avoiding the overall semantic noise. The convolution processes of the two channels are as Figure 2 shown.

[0115] 2.1 One-dimensional convolution channel

[0116] For the i-th word representation vector of the original text S (i.e., text S) as x i , there is an input text vector (i.e., text embedding vector X) under the one-dimensional convolution kernel with a convolution kernel size of z , and the sliding convolution of X can be performed through Equation (2) to generate a higher-quality feature h i .

[0117] [x i :x i+z-1 = [x i ,x i+1 ,...,x i+z-1 2)

[0118] h i = σ(w1 · [x i :x i+z-1 + b1)3)

[0119] where [x i :x i+z-1 represents the sliding window range of the convolution kernel;

[0120] x i represents the word vector of the i-th word in a certain sentence;

[0121] x i+z-1 represents the word vector of the (i + z - 1)-th word of a certain sentence;

[0122] h i represents the convolutional feature information after the i-th sliding convolution of the convolutional kernel, which is generated by the dot product operation of the one-dimensional convolutional kernel and the sentence vector plus the bias b1.

[0123] σ uses LeakyReLU, a variant of ReLU, as the non-linear activation function.

[0124] The initial values of w1 and b1 are system pseudo-random numbers and are gradually optimized during the training process.

[0125] At this time, the word sliding window of the convolutional kernel for the sentence is:

[0126] {h1, h2,..., h l-z+1} = {[x1:x z , [x2:x z+1 ,..., [x l-z+1 :x l}4)

[0127] It means that each convolution operation of the convolutional kernel generates an h i .

[0128] The feature map after this convolution can be represented as the new convolutional feature H.

[0129] H = [h1, h2,..., h l-z+1 5)

[0130] It is found that in text classification, if the convolutional kernel is too large, it will not only increase the number of parameters, but also may incorporate the information of noise words in the phrase representation, resulting in much worse classification results. And when the size z of the convolutional kernel is ≤ 4, it is most beneficial to the feature extraction of text information. Therefore, in the paper, the size of the convolutional kernel is set to z = {2, 3, 4} to increase the diversity of feature extraction. The extracted convolutional features can be represented as M.

[0131] M = [H1, H2,..., H m 6)

[0132] where H1 represents the convolutional features generated by the first convolutional channel;

[0133] H m represents the convolutional features generated by the m-th convolutional channel;

[0134] m is the total number of convolutional channels.

[0135] To reduce the loss of semantic information in the features extracted by the convolutional layer, we use average pooling on the convolutional feature map M to obtain the original semantic information of the one-dimensional convolutional channels. The pooling output C of the convolutional feature map M = {C1 avg , C2 avg ,..., C m avg} is shown in Equation (7).

[0136] C i avg = avgpool{H i}

[0137] C = avgpool(M) 7)

[0138] C i avg represents the high-dimensional convolutional feature map generated after the i-th H passes through the average pooling operation;

[0139] Taking C extracted from the one-dimensional convolutional channels as the input of the weighted average attention module, the attention score C score is generated. After the dot product operation, it is used as one of the input features of the final fully connected layer and is used for the classification prediction of the final probability distribution through the fully connected layer.

[0140] 2.2 Atrous Convolutional Channels

[0141] The Atrous convolutional neural network was initially used in image semantic segmentation to expand the receptive field and thus obtain the global semantics of the image. In the text classification task, when using a relatively large one-dimensional convolutional kernel (the sliding window width z), the maximized semantic information cannot be obtained. Therefore, the Atrous convolutional channel with an extended receptive field p provides a new path for the extraction of sentence-level global semantics. For the short texts or longer texts processed by the present invention, we set the expansion rate corresponding to the convolutional kernel w to P = {p1, p2, p3}, where p1, p2, and p3 represent the expansion rates of three semantic receptive fields, and p1 = p2 = p3 = 1. When extended to sentence convolution, the Atrous convolution will also perform semantic extraction with jumps in the one-dimensional text sequence.

[0142] Specifically, the principle of the Atrous convolution is similar to that of the one-dimensional convolution. The difference is that the process of performing Atrous on the input sequence will extract the global convolutional feature map. Combining Equations (2) - (6), the Atrous one-dimensional convolutional module can be summarized as shown in Equation (9).

[0143]

[0144] A = maxpool{LeakyReLU(Z)} 9)

[0145] LeakyReLU() represents a non-linear activation function;

[0146] maxpool represents the max pooling operation; trConv1D uses max pooling to obtain the most important semantic information of the text.

[0147] Among them, all variables in Equation (8) have the same meaning as one-dimensional convolution. w2 and b2 are optimizable parameters for dilated convolution. The difference is that the three-stage high-dimensional features obtained by expanding the Atrous convolution kernel w with different convolution kernel sizes are respectively superscripted with a to facilitate differentiation from traditional convolution. The input vector after operating on X by its Atrous convolution in Equation (9) in the filter w of length K k The dilated convolution generates the global output z i is defined as in Equation (10).

[0148]

[0149] where p represents the stride for dilated convolution on the input vector x[i + p·k], and i represents the i-th word. Each layer of convolution can take values in P. The dilation rate of the standard Atrous convolution is 1, Figure 2 (b) is a schematic diagram, that is, Z = AtrConv1D(X). The advantage of using Atrous convolution is to integrate the global semantic acquisition of one-dimensional convolution, and at the same time be able to skip the sequence of p·k to obtain long-range semantic dependency relationships.

[0150] The feature map M after Atrous convolution and LeakyReLU activation a performs the maxpool operation to obtain the maximum feature map A of sentence-level category preferences.

[0151] 3 Semantic Enhancement Module

[0152] The attention mechanism has a good effect on obtaining key category words and context semantic information of sentence-level text. Therefore, we perform feature splicing with the high-dimensional feature map generated by weighted average attention after the semantic enhancement module. The combination with the semantic enhancement module improves the model's semantic perception ability of sentences. Weighted attention is mainly used to assign new weights to different parts of the input sequence to highlight the information of more important parts in the sentence, and input C and A respectively to generate different weighted attention scores, as shown in Equation (11).

[0153]

[0154] where C score and A score are the weighted attention scores obtained by one-dimensional convolution and Atrous convolution respectively, and their scales are respectively: and

[0155] softmax(v c T u c ) i is the matrix product probability mapping of v c T and u c ;

[0156] softmax(v a T u a ) i is the matrix product probability mapping of v a T and u a ;

[0157] · T is the transpose symbol;

[0158] C i is the i-th convolutional feature map of C;

[0159] A i is the i-th convolutional feature map of A;

[0160] l is the sentence length;

[0161] z is the convolutional kernel size;

[0162] v c and v a are the feature vectors H and H a , in order to distinguish the operations of different modules, different characters are used to represent them. u c and u a are the high-dimensional feature scores after linear transformation as shown in Equation (12).

[0163]

[0164] where M, M a are the convolutional features passing through the one-dimensional convolutional channel and the convolutional features passing through the Atrous convolutional channel respectively;

[0165] tanh calculates the attention score with the activation function;

[0166] w1, w2, b1, b2 represent the feature weights and biases of the one-dimensional convolutional and Atrous convolutional channels, which are optimizable parameters.

[0167] Adopt the semantic enhancement module to perform convolution operation on the original invention patent vector X and then activate to generate y k , expressed as Equation (13):

[0168] y k = σ1(conv(X; w k , b k ))13)

[0169] where σ1 represents the ReLU activation function,

[0170] k represents the convolutional kernel;

[0171] Then y k is partitioned after passing through maxpool1D to obtain the enhanced features y k c and y k a ,

[0172] Finally, y k c and y k a are respectively concatenated with the dual-channel convolutional features after weighted attention processing, as shown in Equation (14).

[0173]

[0174] where C and A are the original semantic information and global text information obtained by one-dimensional convolution and Atrous convolution respectively;

[0175] C score and A score are the weighted attention scores obtained by one-dimensional convolution and Atrous convolution respectively;

[0176] y k c represents the semantic enhanced feature concatenated with the high-dimensional convolutional feature map C·C score ;

[0177] y k a represents the semantic enhanced feature concatenated with the high-dimensional convolutional feature map A·A score ;

[0178] represents the concatenation operation, which concatenates the feature maps obtained by weighted average attention through one-dimensional convolution and Atrous convolution and the semantic enhanced features, and finally maps them to the probability distribution of the labels through the fully connected softmax layer for classification prediction. In this paper, the cross-entropy loss function is used:

[0179]

[0180] where represents the loss for classification;

[0181] n is the number of test samples;

[0182] y is the ground-truth label;

[0183] is the predicted label.

[0184] The cross-entropy loss function can effectively represent the difference between the distribution of sample estimates and the distribution of actual labels.

[0185]

[0186]

[0187] 4 Experimental verification

[0188] To verify the text mining performance of the DcSeCNN model, six different real-world text datasets and four classic evaluation metrics were used to conduct experimental comparisons on the proposed DcSeCNN method and other competing models. The server configuration information for the experiment was as follows: Intel(R) Xeon(R) Gold 6226R CPU @ 3.90GHz (32 cores and 64 threads), 128-GB memory, and 2 GPUs of the NVIDIA RTX A6000 model.

[0189] 4.1 Datasets and evaluation metrics

[0190] 4.1.1 Dataset introduction

[0191] The performance verification experiment of the model used five English datasets including MR, R8, R52, TREC, and IMDBR, and the THUCNews Chinese dataset, with 3 long texts and 3 short texts each. This facilitated the training and testing of text classification models of different languages and lengths using the same model, and verified the wide applicability of the proposed DcSeCNN in text mining. The specific information of the text datasets is shown in Table 1.

[0192] Table 1 Information of experimental datasets

[0193] Dataset #Text #Training Set #Validation Set / #Test Set #Average Length #Category MR 10662 6396 2133 20.35 2 R8 7674 4604 1534 103.36 8 R52 9100 5460 1820 110.56 52 TREC 5952 3570 1191 9.81 6 IMDBR 50000 30000 10000 240.56 2 THUCNews 200000 120000 40000 16.10 10

[0194] The dataset was divided into training set: validation set: test set = 3:1:1. To address the problem of inconsistent text lengths of different scales, we truncated or padded short texts and long texts to a certain length. GloVe embedding vectors and SougouNews embedding vectors were used for the above six datasets (GloVe for English texts and SougouNews for Chinese texts), and the maximum embedding vector dimension of 300 was selected.

[0195] MR1 (Movie Review) is a movie review text dataset, including two tags, pos and neg, collected by Pang and Lee for sentiment analysis experiments

[44] . We selected 6,396 samples as the training set, and 2,133 samples each for the validation set and the test set. This dataset is used for binary text classification, with short texts having an average sentence length of 20.35.

[0196] R8 and R52 2 (Reuters-21578) is one of the text datasets widely used by researchers in the TC field, collected from the Reuters financial news agency service in 1987. For R8, we selected 4,604 samples as the training set, and 1,534 samples each for the validation set and the test set. It has 8 categories and is a long text with an average sentence length of 103.36. For R52, our training set, validation set, and test set samples are 5,460, 1,820, and 1,820 respectively, with 52 categories and an average sentence length of 110.56.

[0197] TREC 3 (Text REtrieval Conference) is a classic dataset for information retrieval and text classification, collected and maintained by the National Institute of Standards and Technology (NIST) of the United States. We selected 3,570 samples as the training set, and 1,191 samples each for the validation set and the test set. It is a short text dataset with an average sentence length of 9.81 and 6 categories.

[0198] IMDBR 4 is a movie review dataset for polarity sentiment analysis. We selected 30,000 samples as the training set, and 10,000 samples each for the validation set and the test set. This dataset has 2 categories, positive sentiment and negative sentiment, and is a long text with an average length of 240.56 words.

[0199] THUCNews 5 is a Chinese news annotation dataset filtered and screened from Sina News. The data downloaded for this invention patent is a dataset of 200,000 articles containing 10 news types. The training set samples are 120,000, and the validation set and the test set are both 40,000. It is a Chinese short text with a sentence length of 16.10.

[0200] Dataset website:

[0201] 1.https: / / www.cs.cornell.edu / people / pabo / movie-review-data /

[0202] 2.https: / / huggingface.co / datasets / reuters21578

[0203] 3.https: / / trec.nist.gov / data.html

[0204] 4.http: / / ai.stanford.edu / ~amaas / data / sentiment /

[0205] 5.http: / / thuctc.thunlp.org / sendMessage

[0206] 4.1.2 Introduction to Evaluation Metrics

[0207] Four classic evaluation metrics, namely Accuracy, Precision, Recall, and F1-Score, are used to measure the text classification performance of the comparison models. Their definitions are given by Equations (16) - (19). To simplify the description of the performance metrics in the experimental analysis section, we use Acc, Pre, Rec, and F1 to replace the above evaluation metrics.

[0208]

[0209] Among them, TP, TN, FP, and FN represent the statistics of four cases: the prediction and the label are both 1, the prediction and the label are both 0, the prediction and the label are 0 and 1 respectively, and the prediction and the label are 1 and 0 respectively. Finally, each metric and the confusion matrix are obtained. To ensure the reproducibility of the model and reduce the randomness of the experiment, the model runs independently 5 times. The model is trained using the training set, then the validation set is used for hyperparameter optimization, and finally the test set is used for performance testing and statistics of different comparison models.

[0210] 4.2 Baseline Methods

[0211] We selected 7 models in two aspects, namely classic deep learning models and variants of deep learning models, as the baseline methods for the experiment. To ensure a fair comparison, each comparison model was reproduced, and a reasonable algorithm performance comparison was carried out according to our experimental settings. The information of each baseline method is as follows:

[0212] TextCNN was the first to apply CNN to the text classification task. It is an efficient text classification model that needs to consider text order and long-distance semantic dependencies, and has also been used by relevant personnel for comparison and optimization research due to its accelerated performance.

[0213] TextRNN is a text model that uses RNN to capture the sequential dependencies of sentences. The capture of temporal information effectively overcomes the long-distance dependencies existing in TextCNN and can, to a certain extent, overcome the problem of gradient explosion.

[0214] FastText is a word embedding n-gram feature model developed by the Facebook AI Research laboratory. This model has advantages over traditional text representation methods such as TF-IDF and NBOW and is superior to similar pre-training tools like Word2Vec and GloVe in terms of representation ability and computational speed.

[0215] Transformers abandons the traditional feature extraction method. The proposal of the Attention mechanism enables excellent performance in multiple fields and tasks of AI, greatly promoting the development of the NLP field and representing a major leap in the field of artificial intelligence.

[0216] TextRCNN combines TextCNN and TextRNN. The combination of RNN and CNN can effectively improve the computational performance of the model and also possess advanced text mining capabilities for long texts.

[0217] TextRNN_Att is a text recurrent neural network with an attention mechanism. Introducing the attention mechanism is beneficial for the model to focus on the important feature content of the input sequence and adopts a survival-of-the-fittest strategy to improve the computational performance and interpretability of the model.

[0218] DPCNN was proposed by Johnson et al.

[46] and is a variant TextCNN architecture mainly used for text data mining tasks. It has excellent performance in text classification tasks through the aggregation of information in the pyramid convolutional layer and a deeper network structure.

[0219] 4.3 Model Execution Details

[0220] In the text classification task, first, a text preprocessing stage including word segmentation, padding, and word embedding vectors is carried out. For English, word segmentation is performed using space characters, and for Chinese, Jieba library is used to divide a paragraph into each Chinese character. Then, each sample is filled using the short-fill-long-truncate strategy to obtain a standardized data sample. Finally, GloVe and SougouNews are used for word vector representation of Chinese and English respectively. Second, the word vectors are input into the DcSeCNN model for classical convolution and Atrous convolution in a dual-channel manner. Third, the dual-channel high-dimensional text features extracted are connected and used as the input of the semantic enhancement module for feature enhancement. In this process, weighted average attention is adopted to adjust the global text semantics. Finally, under the condition of setting fixed parameters in the system, the self-adaptive enhancement of the learning rate is carried out during the optimization process by introducing a learning rate decay coefficient.

[0221] Specifically, the operating environment of the model is set as follows: the number of epochs is 20, the learning rate is 0.0001, the size of the traditional convolutional kernel is (2, 3, 4), the dilation rate of the Atrous convolutional kernel is (1, 1, 1), the number of one-dimensional filters is 256, and the semantic enhancement channel is 512. Each model runs independently 5 times, and the average value is taken as the final statistical evaluation index. During this process, when there is no improvement in the model optimization iteration after 1000 times, the program is terminated. In addition, according to the scale and sentence length of different datasets in Table 1, the BatchSize and PadSize we set are as follows: 1) For the four datasets of MR, R8, R52, and TREC, the BatchSize is 64, and for the two datasets of IMDBR and THUCNews, the BatchSize is 128; 2) For the three datasets of MR, TREC, and THUCNews, the PadSize is 32, and for the three datasets of R8, R52, and IMDBR, the PadSize is 128.

[0222] Table 2 Parameter Settings

[0223] Parameter Default Value Range dropout 0.5 0.3,0.4,0.5,0.6,0.7 λ 0.95 0.91,0.93,0.95,0.97,0.99 pad 32 / 128 16 / 32 / 64,64 / 128 / 256

[0224] In the model experiment and performance analysis, we will conduct the following experiments: 1) Study the system parameter settings and conduct experiments on all comparison models to verify the effectiveness of DcSeCNN; 2) Explore the two models of TextCNN and DcSeCNN to conduct experiments on the four datasets of MR, R8, IMDBR, and THUCNews to determine the best dropout and learning rate decay coefficient λ for setting the optimal hyperparameter combination. The above four datasets cover most text classification situations such as Chinese and English, different scales, and different lengths of texts, and are used for parameter establishment and model portability; 3) Verify the influence of the short-padding and long-truncation strategy on dataset processing through different pad parameter performances; 4) Use ablation experiments to verify the contribution of each module in DcSeCNN to the proposed model. Table 2 shows the information on the settings of the three parameters of dropout, λ, and pad.

[0225] 4.4 Experimental Results and Analysis

[0226] 4.4.1 Analysis of Comparative Experiment Results

[0227] To verify the performance of the DcSeCNN model in the TC task, systematic control experiments were conducted on our method and 7 classic baseline models on 6 datasets. The experimental results are shown in Table 3 (where the data marked in bold is the first in performance, and the data marked with an underline is the second in performance). Among the 18 statistical data including three indicators of Pre, Rec, and F1, 72% of the statistical indicators of the DcSeCNN model ranked first, and 28% ranked second, which fully verified that our model has very good results in text data mining tasks under various conditions of different scales and lengths.

[0228] In Table 3, the performance of the models is compared and analyzed from the perspective of data scale. For the two larger-scale datasets of IMDBR and THUCNews, the models with better performance in each indicator are TextRCNN, TextRNN_Att, and our model. Specifically: DcSeCNN improved by 1.22%, 1.44%, and 1.59% respectively in the three indicators of IMDBR, and at the same time, the Pre of DcSeCNN in THUCNews was as high as 90.71%, and the other two indicators ranked second. It can be seen that the performance of the variants of the classic text classification model is generally better than that of the basic model in such datasets. In the classification performance of the other four smaller-scale datasets, the performance of other models is generally lower than that of our model.

[0229] Table 3 Pre, Rec, and F1 results of different models on 6 text datasets

[0230]

[0231]

[0232] From the perspective of text length, the performance of the models is compared and analyzed. For the three short texts of MR, TREC, and THUCNews, the DcSeCNN performance indicators ranked first and second each accounted for half. It shows that for short text datasets, due to the lack of semantic information, the model's ability to mine text information is poor, but there are certain improvements in the indicators of the variants of the deep model. In the remaining three long text datasets, DcSeCNN ranked first in all indicators and was better than all comparison models, and the F1 indicators were 94.07%, 88.42%, and 70.34 respectively. This verifies that our method has stronger text classification ability under the condition of sufficient semantic information. In addition, the comparison of the models that statistically calculate the Acc indicator is as Figure 3 shown.

[0233] As Figure 3As shown in the figure, the DcSeCNN model demonstrates excellent classification capabilities in terms of Acc. Specifically, the improvement effect in MR is relatively obvious. In R8 and R52, the Acc of TextCNN, DPCNN, and DcSeCNN is higher than that of other comparison models. In TREC and THUCNews, all models tend to reach the same classification level, but our method has better Acc. In IMDBR, the performance of Transformer and TextRCNN is poor, and DcSeCNN also has a higher Acc in other models. Through the above model performance comparison and visualization analysis, it can be seen that DcSeCNN has greater advantages in text mining capabilities under various conditions, which greatly verifies the effectiveness and generalization ability of DcSeCNN in text mining tasks.

[0234] 4.4.2 Analysis of the Influence of Hyperparameters

[0235] (1) Influence of the dropout size in DcSeCNN

[0236] DcSeCNN uses different dropout values to perform a certain ratio of screening on the dual-channel convolutional feature extraction, effectively ensuring the simplicity of the network structure and the effectiveness of feature representation learning. The proposed method is trained with different dropout settings, and the settings are shown in Table 2. Then, the influence of this hyperparameter on the model classification performance is analyzed through result visualization, and the dropout value corresponding to the best classification performance of different text datasets is obtained.

[0237] As Figure 4 shown, when the dropout size is 0.5, DcSeCNN achieves better Acc on the four datasets of MR, R8, TREC, and THUCNews. On the R52 dataset, when the parameter size is 0.7, the classification effect is much worse than other values, and an excessive dropout will cause the network structure to be unable to fit the high-dimensional features with more categories. At the same time, the performance on other values is also unstable, and the best classification performance of R52 is achieved when the dropout size is 0.4. This means that it is not appropriate to set the dropout too large or too small for classification tasks with more categories. On the larger-scale IMDBR dataset, as the dropout increases, the classification effect gradually decreases, and there is a large performance difference when crossing from 0.6 to 0.7. This shows that for the classification task of large-scale datasets, setting the dropout of DcSeCNN between 0.3 - 0.5 can better learn and represent high-dimensional features for more samples. Therefore, we believe that an excessive dropout will lead to poor data fitting of the network, and a smaller dropout will increase the burden on computing resources, and setting it to 0.5 is appropriate.

[0238] (1) Influence of the Size of λ in DcSeCNN

[0239] For the DcSeCNN model with a changed network structure, whether the learning rate is a fixed parameter or has a certain self - adaptability is crucial for the classification performance of the model. An overly large fixed learning rate will cause early convergence in the model optimization process, while an overly small learning rate will lead to a slow gradient convergence speed. Therefore, introducing a learning rate with an adaptive decay mechanism is of great significance for improving the performance of DcSeCNN and expanding downstream tasks. The visualization of the results obtained from the parameter ranges given in Table 2 is as Figure 5 shown.

[0240] From Figure 5 it can be seen that except for the IMDBR dataset, the sensitivities of the other 5 datasets to the learning rate decay coefficient are in the same distribution trend, showing a performance effect that is high in the middle and low on both sides. Among them, on the MR, R52, and YHUCNews datasets, λ = 0.95 has the best performance, while on the R8 and TREC datasets, the classification performance is lower than that when λ = 0.97. On the IMDBR dataset, within a certain range, the classification effect of the model is negatively correlated with λ, which indicates that our learning rate setting is too small on large - scale datasets, but the consistency with the results obtained by dropout verifies the rationality of our experimental work. Generally speaking, when the adaptive change rate of λ is too fast, it is not conducive to the semantic mining and representation of texts, and when it is too slow, it will lead to a performance decline. It is appropriate to set its adaptive decay of the learning rate to 0.95 with a moderate self - adaptability.

[0241] 4.4.3 Analysis of the Influence of Padding Scale

[0242] To explore the influence of the padsize short - filling and long - truncating strategy on the classification performance of the DcSeCNN model in different languages (Chinese and English) and texts of different lengths. In this paper, the average lengths corresponding to each dataset given in Table 1 are processed, and the padsize sizes of short texts and long texts are 32 and 128 respectively. When performing long truncation on them, it is set to 16 and 64, and when performing short filling, it is set to 64 and 256. Through system simulation experiments, the Acc comparisons of each model are shown in Table 4.

[0243] Table 4 Influence of the Size of padsize in DcSeCNN (Acc)

[0244]

[0245] For short texts such as MR, TREC, and THUCNews, our model still has better classification performance on MR after padding and truncating. When truncating on TREC, it is consistent with TextRCNN, ranking second by default, 0.02% lower than TextRCNN in terms of Acc, and about 0.8% lower than DPCNN and TextRCNN when truncating. In long texts, DcSeCNN has the best text classification performance under all conditions.

[0246] In addition, TextCNN and its variant DPCNN are selected in this paper for comparison with our DcSeCNN, and the classification performances of the three models on the same dataset are visualized. The visualizations of different padsizes for short texts and long texts are as Figure 6 and Figure 7 shown.

[0247] As Figure 6 shown, the comparison of different padsizes of MR can clearly show the effectiveness of our solution. On TREC, the overall classification ability is the same, both about 97%. In the Chinese text THUCNews, compared with the comparison models, DcSeCNN ranks second when the padsize is 32 and 64, and is at the same performance level as the best one. As Figure 7 shown, the performance of each model on IMDBR improves as the padsize increases, and our solution has better effects. There is also a slight improvement in performance on R8 and R52 as the padsize increases. At the same time, it will be found that the overall performance of DPCNN is lower than that of the basic TextCNN and our model. This means that our default settings of padsize as 32 and 128 are appropriate both in terms of improving computing performance and balancing computing resources. Based on the above performance analysis, it shows that the DcSeCNN model can perform text mining tasks under various conditions and has good classification performance.

[0248] 4.4.4 Ablation Experiment

[0249] The main contributions of this invention patent include three modules: dual-channel convolution, semantic enhancement, and weighted average attention. Dual-channel convolution is to overcome the problem of insufficient semantic information mining of the traditional convolutional kernel sliding window, and capture long-distance semantic connections with the new Atrous convolution. The semantic enhancement and attention modules can enhance strongly correlated semantics and denoise weakly correlated semantics during the extraction process. Specifically, we separately remove the learning rate decay factor λ, Atrous convolution channels (Dc), semantic enhancement (Se), and weighted average attention mechanism (Att) in DcSeCNN. The Acc and F1 metrics of the experimental results are statistically shown in Table 5 under the default parameter settings in Table 1.

[0250] Table 5 Ablation Experiment

[0251]

[0252] Table 5 shows that removing any one module will cause varying degrees of decline in the performance of the model. This indicates that each module or parameter tuning scheme introduced by the model is effective. However, from the underlined data in the experimental results, it can be seen that for MR, IMDBR, and THUCNews, the Att module has the least impact on them, while the Se module has the greatest impact. For the two datasets of R8 and R52 in the same series, λ has the least impact on them, and TextCNN ranks second, indicating that introducing a single module alone cannot have a profound impact on them, but the synergistic effect of multiple modules improves their classification performance to a certain extent. In addition, due to the strong normality of the data, the attenuation coefficient has the least impact on it, while semantic enhancement has the greatest contribution to the model on TREC. Generally speaking, each module plays a positive role to varying degrees in the optimization of the model, and through the synergistic effect, it has reliable performance index improvement and good applicability for the DcSeCNN model in handling text classification tasks under various conditions.

[0253] Thus, the advanced text classification performance of DcSeCNN is demonstrated on 6 text classification datasets (3 short texts, 3 long texts, 4 small-scale, 2 large-scale, 5 English, and 1 Chinese). Among them, there are obvious performance improvements on 4 datasets such as MR, R8, R52, and TREC, and the best performances on 2 datasets such as IMDBR and THUCNews tend to be at the same level.

[0254] To sum up, this invention patent proposes a text classification model based on DcSeCNN. First, the combination of Atrous convolution with long-distance semantic mining ability and traditional one-dimensional convolution is introduced, which can effectively capture the local and global semantic information of sentence-level text. To improve the quality of the convolutional feature map, we use a weighted average attention with an attention mechanism to enhance the features of the convolutional feature map. Finally, semantic enhancement operations are performed on the initial input text vector, and two different original semantics and the enhanced features are respectively fused to obtain a high-dimensional feature vector closer to the semantics of the original sentence. In addition, the model training process with an adaptive attenuation learning rate λ and drouput is dynamically tuned to obtain the DcSeCNN model. A large number of experiments and data analyses are carried out on this model, and the results prove that this method has excellent text classification performance and good generality under different conditions (different languages, lengths, and scales).

[0255] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A text classification method based on dual-channel semantic enhancement and convolutional neural network, characterized in that: The following steps are involved: S1, first preprocess the text and divide the text into text units; then perform word vector embedding to convert words or phrases into vector representations; thus, the text matrix X is obtained; S2, the generated text matrix X is subjected to dual-channel convolution feature extraction using Conv1D and AtrousConv1D to obtain the original semantic information C and global text information A; Using Conv1D for feature extraction, obtaining the original semantic information C includes the following steps: For the i-th word in text S, the vector representation is x i , the text embedding vector X is subjected to sliding convolution on X under a one-dimensional convolution kernel w with a convolution kernel size of z, and the feature h is generated by equation (2) i : [x i :x i+z-1 ]=[x i ,x i+1 ,...,x i+z-1 ]2) h i =σ(w1·[x i :x i+z-1 ]+b1)3) where [x i :x i+z-1 ] represents the sliding window range of the convolution kernel; x i Represents the word vector of the i-th word in a sentence; x i+z-1 Represents the word vector of the i+z-1th word in a sentence; w1 represents the feature weight; b1 represents bias; h i Represents the convolution feature information after the convolution kernel has undergone sliding convolution for the i-th time; σ is a nonlinear activation function, specifically LeakyReLU; The convolution feature h i The mapping is represented as a new convolutional feature H: H=[h1,h2,...,h l-z+1 ]5) Thus, the extracted convolution feature M is obtained: M=[H1,H2,...,H m ]6) Where H1 represents the convolution feature generated by the first convolution channel; H m Represents the convolution feature generated by the mth convolution channel; m is the total number of convolution channels; The average pooling operation is used on the convolution feature map M to obtain the original semantic information of the one-dimensional convolution channel. The pooling output C of the convolution feature map M is expressed as: C = avgpool(M)7) Among them, C represents the original semantic information; avgpool() means average pooling; AtrousConv1D is used for feature extraction to obtain global text information A: A=maxpool{LeakyReLU(Z)}9) Among them, LeakyReLU() represents a nonlinear activation function; maxpool() represents the maximum pooling operation; Z = AtrConv1D(X) Where Z = [z1,z2,...,z K ], in the filter w of length K k After dilation convolution, the global output z is generated i : Where x[i+p·k] represents the word vector after AtrousConv1D expansion; p represents the stride of the dilated convolution on the input vector x[i+p·k], p∈P, P={p1,p2,p3}, p1, p2, p3 represent the three semantic receptive field expansion rates; S3, the weighted average attention is used to generate the attention score C for the generated original semantic information C and the global text information A score and A score , and semantic enhancement is performed on the text matrix X to obtain y k Finally, the high-dimensional convolutional feature maps are concatenated, and then the concatenated feature maps are mapped to the probability distribution of the label through the Linear fully connected layer and the Sofrmax layer.

2. A text classification method based on dual-channel semantic enhancement and convolutional neural network according to claim 1, characterized in that: The text in step S1 is Chinese text or English text. At this time, the word vector embedding is: When the text is Chinese, the pre-trained word vector of the Chinese corpus is used to represent each Chinese character; When the text is in English, the pre-trained word vector of the English corpus is used to represent each word.

3. The text classification method based on dual-channel semantic enhancement and convolutional neural network according to claim 1, characterized in that: The word vector embedding is performed to convert the text unit into a vector representation; including: The text S is converted into a text embedding vector X through the corresponding index of the vocabulary V: X=index(S)*V 1) Among them, index(S) represents the process of generating X by indexing the pre-training corpus; V stands for vocabulary.

4. The text classification method based on dual-channel semantic enhancement and convolutional neural network according to claim 1, characterized in that: The convolution kernel size z≤4.

5. A text classification method based on dual-channel semantic enhancement and convolutional neural network according to claim 4, characterized in that: The convolution kernel size is set to z = {2, 3, 4}.

6. The text classification method based on dual-channel semantic enhancement and convolutional neural network according to claim 1, characterized in that: Step S3 includes the following steps: S3-1, use weighted average attention to generate attention score C for the generated original semantic information C and global text information A score and A score : Among them C score and A score The weighted attention scores obtained for Conv1D and AtrousConv1D convolutions respectively; softmax(v c T u c ) i v c T with u c The matrix product probability mapping of ; softmax(v a T u a ) i v a T with u a The matrix product probability mapping of ; T is the transpose symbol; C i is the i-th convolution feature map of C; A i is the i-th convolution feature map of A; l is the sentence length; z is the convolution kernel size; v c and v a The eigenvectors H and H are a , H is the mapping feature of the text matrix X after Conv1D convolution; H a It is the mapping feature of the text matrix X after AtrousConv1D convolution; u c and u a Then it is the high-dimensional feature score after linear change, as shown in formula (12): Among them, M a They are the convolution features after the one-dimensional convolution channel and the convolution features after the AtrousConv1D convolution channel respectively; Tanh is the activation function to calculate the attention score; w1 and w2 represent the feature weights of the convolution channels of Conv1D and AtrousConv1D respectively; b1 and b2 represent the bias of the convolution channels of Conv1D and AtrousConv1D respectively; S3-2, using the semantic enhancement module, semantically enhances the text matrix X and activates the generated y k , expressed as: y k =σ1(conv(X;w k ,b k ))13) Where σ1 represents the ReLU activation function; y k Indicates enhanced features; conv() represents the convolution operation; w k Represents the feature weight of the convolution channel; b k Represents the bias of the convolution channel; S3-3, y k After maxpool1D, the enhanced feature y is obtained by division. k c and k a ; Then y k c and k a Concatenate with the weighted attention features extracted by dual-channel convolution: Where C and A are the original semantic information and global text information obtained by the convolution of Conv1D and AtrousConv1D respectively; C score and A score The weighted attention scores obtained by Conv1D and AtrousConv1D convolutions respectively; y k c Representation and high-dimensional volume feature map C·C score Semantically enhanced features generated by splicing; y k a Representation and high-dimensional volume feature map A·A score Semantically enhanced features generated by splicing; Represents a splicing operation; S3-4, the concatenated feature map H cla It is mapped to the probability distribution of labels through the Linear fully connected layer and the Sofrmax layer.

7. The text classification method based on dual-channel semantic enhancement and convolutional neural network according to claim 1, characterized in that: The cross entropy loss function is used to measure the difference between the predicted probability distribution and the actual label. The cross entropy loss function is expressed as; in represents the loss used for classification; n is the number of test samples; y is the true value label; is the predicted label.

Citation Information

Patent Citations

  • Convolutional neural network matching text recognition method based on attention enhancement mechanism

    CN110298037A

  • Semantic reconstruction video description method based on time sequence Gaussian mixture cavity convolution

    CN113420179A