Text classification method based on pre-training language model fusion deep convolutional network

CN120011558APending Publication Date: 2025-05-16JIANGSU HONGCHENG DATA INTELLIGENCE RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510092996.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional text classification methods are difficult to extract global semantics and local features at the same time, resulting in insufficient accuracy and efficiency when processing long texts.

Method used

The fusion deep convolutional network method based on pre-trained language model is adopted to deeply mine global semantics through multi-head attention networks, and local features are extracted in combination with convolutional neural networks with multi-layer residual structures, and finally feature fusion is performed through the learnable dimension transformation matrix.

Benefits of technology

It significantly improves the accuracy and efficiency of text classification, can effectively capture global semantics and local features, and enhances the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011558A_ABST
    Figure CN120011558A_ABST
Patent Text Reader

Abstract

The invention discloses a text classification method based on a pre-training language model fused with a deep convolutional network, and the method comprises the following steps: firstly, carrying out the feature extraction of an input text through a RoBERTa model; then, the features output by the RoBERTa are input into the multi-head attention network, and deep semantic information is obtained; thirdly, inputting the output of multi-head attention into a convolutional neural network of a multi-layer residual structure, and obtaining deep convolutional network features; and finally, introducing the output of multi-head attention into a learnable dimension transformation matrix, generating a low-dimensional representation suitable for a classification task by optimizing dimension features, fusing the low-dimensional representation with the output of the multi-layer convolutional neural network after global maximum pooling operation, and unifying semantic representations of global and local features. And the fused features generate a final classification result through a full connection layer. According to the method, through fusion of the multi-head attention mechanism and the deep convolutional network, global semantics and local features can be fully extracted, so that the efficiency and the effect of a classification task are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and natural language processing, and relates to a text classification method based on a pre-trained language model fused with a deep convolutional network. Background Art

[0002] With the rapid development of big data, artificial intelligence and the Internet industry, a large amount of text data has been accumulated in various industries and fields. Text classification research aims to automatically extract and classify valuable information from unstructured text data, and is widely used in information retrieval, sentiment analysis, spam filtering, public opinion monitoring and other fields. For example, in social media analysis, text classification can help companies analyze users' emotional attitudes; in news recommendation, text classification can quickly match user interests based on news content and provide personalized recommendations. However, most traditional text classification methods rely on a large amount of labeled data, which poses a great challenge to applications in certain scenarios.

[0003] The research on text classification technology mainly focuses on the following methods: methods based on convolutional networks, methods based on recurrent neural networks, methods based on graph convolutional networks, and methods based on attention mechanism networks.

[0004] 1) Convolutional network-based methods: Convolutional neural networks (CNNs) extract n-gram features of text through local receptive fields, which are suitable for capturing local information in text, especially when processing short texts. CNNs can automatically learn features, reduce reliance on manual feature engineering, and have good generalization capabilities.

[0005] 2) Recurrent Neural Network-based Methods: Recurrent Neural Networks (RNNs) are good at modeling temporal dependencies in text, and are particularly suitable for processing long texts and complex contextual dependencies. However, RNNs are slow to train and are prone to gradient vanishing or exploding problems when processing long texts.

[0006] 3) Methods based on graph convolutional networks: Graph convolutional networks (GCNs) process entity relationships and dependency relationships in texts through graph structures, and are suitable for relational modeling of complex texts. GCNs can capture higher-order structural features, but their computational complexity is high, and there are performance bottlenecks when processing large-scale texts.

[0007] 4) Methods based on attention mechanism networks: The attention mechanism dynamically adjusts the attention weights of each part, allowing the model to focus on key text areas. The self-attention mechanism can capture global information and is particularly suitable for processing long texts and complex dependencies, solving the difficulties of traditional RNN in processing long texts.

[0008] Although the text classification method based on pre-trained language models has achieved remarkable results, it also has some limitations. Although pre-trained language models can capture rich contextual information, in some tasks, they may not fully understand the deep language features and have difficulty handling long sentence dependencies, especially in long texts where some local grammar and cross-sentence relationships may be ignored. In order to make up for this shortcoming and further explore the deep language features in the text, a text classification method based on a pre-trained model and a fused deep convolutional network is proposed. By taking advantage of the feature extraction of deep convolutional networks, the processing capability of long texts can be effectively improved. Summary of the invention

[0009] Purpose of the invention: Based on existing research, the present invention proposes a text classification method based on a pre-trained language model fused with a deep convolutional network, which fully combines the global semantic extraction ability of the pre-trained language model and the local feature capture ability of the convolution operation to improve the accuracy and robustness of the model in text classification tasks, and at the same time optimizes the training efficiency and performance of the deep network through residual connection and normalization mechanism.

[0010] Technical solution: To achieve the above-mentioned invention object, the present invention proposes a text classification method based on a pre-trained language model fused with a deep convolutional network, comprising the following steps:

[0011] (1) Feature representation extraction: extract features from the input text using the RoBERTa model;

[0012] (2) Deep semantic capture: The features output by RoBERTa are input into the multi-head attention network to obtain deep semantic information.

[0013] (3) Local feature enhancement: The output of the multi-head attention is input into a convolutional neural network with a multi-layer residual structure to obtain deep convolutional network features.

[0014] (4) Feature fusion and classification prediction: The output of the multi-head attention network is fused with the output of the convolutional neural network in the final stage. The fused feature representation is used to generate the final classification result logits through the fully connected layer.

[0015] Furthermore, in step (2), the multi-head attention network decomposes the feature representation extracted by RoBERTa into multiple independent subspaces, each subspace captures deep semantic information through an independent attention head, and performs weighted aggregation to generate a unified output representation.

[0016] Furthermore, in step (3), the residual convolutional neural network enhances local features through 3×3 convolution operations, and combines LayerNorm and GeLU activation functions for normalization and nonlinear transformation, thereby improving the expression ability of local context information. At the same time, the residual connection mechanism is used to superimpose the input and output features of the convolution module to enhance the feature representation of the deep convolutional network.

[0017] Furthermore, in step (4), the output of the multi-head attention is introduced into the learnable dimensional transformation method, and a low-dimensional representation suitable for the classification task is generated by optimizing the dimensional features. The low-dimensional representation is fused with the output of the multi-layer convolutional neural network after the global pooling operation, thereby unifying the semantic representation of global and local features, and generating the classification result through the fully connected layer.

[0018] Beneficial effects: The present invention can effectively solve the problem that traditional text classification methods are difficult to extract global semantics and local features at the same time, give full play to the deep semantic capture ability of the multi-head attention mechanism and the local pattern extraction advantages of the convolution operation, and optimize the feature expression in the final fusion stage by introducing a learnable dimensional transformation matrix, which significantly improves the accuracy and efficiency of text classification. First, the multi-head attention network can deeply explore the global semantics of the feature vectors extracted by RoBERTa, capture the dependency of features in different subspaces, and enhance global consistency; secondly, the convolutional neural network with a multi-layer residual structure effectively extracts local context information, while improving the stability of training and the recognition ability of features through normalization and nonlinear transformation; thirdly, the learnable dimensional transformation matrix acts independently on the multi-head attention output, realizes the optimized expression of global features in the fusion stage, and forms an effective complement with local features. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is the overall framework diagram of the method of the present invention; DETAILED DESCRIPTION

[0020] The present invention is further explained below in conjunction with the accompanying drawings. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0021] This paper proposes a text classification method based on a pre-trained language model fused with a deep convolutional network, aiming to fully combine the global semantic extraction capability of the pre-trained language model and the local feature capture capability of the convolution operation, solve the problem that traditional text classification methods are difficult to extract global semantics and local features at the same time, and significantly improve the accuracy and robustness of the model in text classification tasks. Figure 1As shown in the figure, the complete process of the present invention includes four stages: feature representation extraction, deep semantic capture, local feature enhancement, and feature fusion. The specific implementation is described as follows:

[0022] The feature representation extraction stage corresponds to step (1) of the technical solution. The specific implementation is as follows: Given the input text X = {x 1 , x 2 ,..., x n}, where X is a text containing n tokens. The RoBERTa model encodes each token through a multi-layer Transformer structure to obtain the context representation of each token:

[0023] RoBERTa(X) = RoBERTa(x 1 , x 2 ,..., x n ) = {v 1 , v 2 ,..., v n}

[0024] Among them, v i represents the representation of the i-th token, with a dimension of d, that is, v i ∈R d .

[0025] The feature RoBERTa(X) extracted by the RoBERTa model is the global semantic representation of the entire input text, reflecting the syntactic and semantic information of the context.

[0026] The deep semantic capture stage corresponds to step (2) of the technical solution. The specific implementation is as follows: The feature RoBERTa(X) = {v 1 , v 2 ,..., v n} obtained from the RoBERTa model is input into the multi-head attention network for processing. The multi-head attention mechanism extracts the deep semantic information of the text by calculating the weighted sum of different attention heads. The specific formula is:

[0027]

[0028] Among them, Q, K, and V are the query, key, and value matrices respectively, and d k is the dimension of the key vector. For the multi-head attention mechanism, its calculation process is:

[0029]

[0030] Among them, h represents the number of attention heads, and Attenttion i (X) represents the output of the i-th attention head.

[0031] The output of the multi-head attention mechanism, Attenttion(X), is a set of deep semantic representations that contain global context information and the relative relationship of each token.

[0032] The local feature enhancement stage corresponds to step (3) of the technical solution. The specific implementation method is: the feature Attenttion (X) obtained from the multi-head attention network layer is input into the deep convolutional network for processing. This implementation scheme designs a two-layer convolutional neural network to process the input feature map. Each layer of convolution operation is followed by LayerNorm normalization and GELU activation function. This combination enhances the nonlinear expression ability of the convolution layer and reduces the gradient disappearance problem. For each layer of convolution operation, a 3×3 convolution kernel with a step size of 1 is used to extract local features. Therefore, the output obtained after the lth layer of convolution is Y (l) , the formula is:

[0033] Y (l) =Conv2d(Attention(X) (l) , W (l) )+b (l)

[0034] Among them, W (l) is the weight matrix of the lth layer of convolution, b (l) is the bias term, and Conv2d represents the convolution operation.

[0035] In order to ensure that the convolution is only performed in the valid area, masked convolution is used. In the convolution operation of each layer, some invalid or filled positions are dynamically shielded by masking. By changing the feature values ​​of the corresponding positions to zero, it is ensured that the convolution operation only focuses on the valid features in the input. Therefore, the processing method of masked convolution is expressed as:

[0036] Attention(X) (l) =Attention(X) (l) ⊙Mask

[0037] Among them, ⊙ represents the element-wise product operation, mask is the mask matrix, when the value is 1, it represents the valid feature area, when the value is 0, it represents the invalid feature area. In the mask convolution layer, the mask is applied to the input feature map Attenttion(X) and then convolution is performed to obtain the local feature Y (l) :

[0038] Y (l) =Conv2d(Attention(X) (l) ⊙mask,W (l) )+b (l)

[0039] After each convolution operation, LayerNorm is used to normalize the feature map to help speed up training and improve stability. The role of LayerNorm is to normalize the mean and variance of all channels in each feature map. (l) , the output after the LayerNorm layer is the local feature Z (l) , and its calculation formula is:

[0040]

[0041] Among them, μ (l) is the mean of the feature map of layer l, σ (l) is the standard deviation, γ and β are learnable scaling and bias parameters, and ε is a small constant to prevent division by zero.

[0042] After the normalization operation, the GELU activation function is used to increase the nonlinear expression ability and prevent the model from linear degradation. The calculation formula of GELU is:

[0043]

[0044] In order to further improve the training efficiency and accelerate the convergence, a residual connection is added between the convolution layer and the normalization layer, and the input feature is added to the output T(l) after convolution, layer normalization and GELU activation function to obtain the output feature CNN(X) = Attenttion(X) (l) +T (l) ,This residual connection helps the gradient to be better transferred in the network, prevents the gradient disappearance problem, and improves the learning ability of the model.

[0045] The feature fusion stage corresponds to step (4) of the technical solution. The specific implementation method is: by initializing a shape W dim ∈R batch×output The learnable training weight matrix is ​​initialized by the following formula:

[0046]

[0047] Among them, batch represents the feature dimension of the input, and output represents the feature dimension of the output.

[0048] Through a learnable linear transformation W dim Perform dimensionality reduction on Attention(X) to generate a low-dimensional representation Attention low (X), the formula is:

[0049]

[0050] Through this transformation, the low-dimensional representation of Attention low (X) provides a suitable feature representation for the subsequent classification task. Then, the low-dimensional representation Attention low (X) is fused with the output of the convolutional neural network CNN(X). The fusion operation aggregates local features through global maximum pooling, and then combines global and local features through a concatenation operation to form the final feature representation Final(X).

[0051] Final(X)=Concat(MaxPool(CNN(X)),Attention low (X)

[0052] Among them, MaxPool represents the maximum pooling operation, and the final feature representation Final(X) is obtained through the Concat splicing operation.

[0053] Then, the fused feature Final(X) is input into the fully connected layer to generate the logits of the classification results.

[0054] Logits(X)=W f Final(X)+b f

[0055] Among them, W f is the weight matrix of the fully connected layer, b f is the bias term, and Logits(X) is the score for each category.

[0056] Finally, the logits are converted into probability distributions of various categories through the softmax activation function, and the category with the maximum probability is selected as the prediction result. The calculation formula is as follows:

[0057]

[0058] Among them, p(y k |X) is category y k The predicted probability, C is the total number of categories, Logits k (X) is category y k The corresponding logit value, if the category with the highest probability is y max , then the final classification label of the text is y max .

[0059] During the training process, in order to optimize the performance of the model, the cross entropy loss function is used To measure the difference between the prediction and the true label, the calculation formula is:

[0060]

[0061] Among them, y k is the 0 or 1 encoding of the true label, (p(y k |X) is the predicted category probability, and the optimization goal is to minimize the loss function And use the back propagation algorithm to update the model parameters W f , W, b f and b, etc., to train the model.

[0062] The present invention uses the accuracy of the model to measure the correctness of the model prediction. Specifically, if the data set contains N samples, and the model performs classification prediction on each sample, then the accuracy is the ratio of the number of samples correctly predicted by the model to the total number of samples N. The specific formula is:

[0063]

[0064] Among them, θ(predicted label i =True label i ) is an indicator function, which takes the value of 1 when the model predicts the same label as the true label for the i-th sample, and 0 otherwise.

[0065] This paper proposes a text classification method based on a pre-trained language model fused with a deep convolutional network. In order to test the effectiveness of the method, it is evaluated on two English text classification datasets MR and R8, and compared with other text classification methods.

[0066] The method of the present invention adopts a set of general hyperparameter configurations, the details of which are shown in Table 1. At the same time, some specific parameters are customized according to the characteristics of each data set, and the specific information of these parameters is shown in Table 2. The basic version of the RoBERTa model, RoBERTa-base, is used in model inference, which contains 12 Transformer encoder layers. In order to reduce the risk of overfitting, weight decay is introduced in the model training process, and its coefficient is set to 1e-5. We also adopt a linear growth warmup strategy, and its growth rate is set to 0.1. In the selection of optimization algorithm, AdamW is used. For the learning rate, it is also set to 1e-5 to ensure the stable update of the model parameters. In addition, the hidden vector dimension of the model is set to 768 to maintain the full expression of information. And according to different data sets, different batch sizes, dropout rates, training epochs, and sentence maximum lengths are set.

[0067] Table 1 General parameter setting information of the model

[0068]

[0069] Table 2 Hyperparameter settings for text classification dataset

[0070]

[0071] The comparative experimental results of different text classification algorithms on the MR and R8 datasets are shown in Table 3. From the experimental results, it can be seen that the text classification method based on the pre-trained language model fused with the deep convolutional network proposed in the present invention has achieved the best accuracy under different hyperparameter settings of the MR and R8 datasets, so the comparative experimental results prove the effectiveness of the algorithm of the present invention in text classification and recognition tasks.

[0072] Table 3 Experimental results of MR and R8 datasets (accuracy values ​​reported)

[0073]

[0074] The present invention achieved the highest accuracy on both MR and R8 datasets, which were 89.52% and 98.41% respectively. The present invention deeply mines the global semantics of the feature vector extracted by RoBERTa through a multi-head attention network, captures the dependency of features in different subspaces, and enhances global consistency; then effectively extracts local context information through a convolutional neural network with a multi-layer residual structure, and improves the stability of training and the recognition ability of features through normalization and nonlinear transformation; finally, a learnable dimensional transformation matrix is ​​used to act independently on the multi-head attention output, and the global features are optimized in the fusion stage, and effectively complemented with the local features; therefore, it is proved that the present invention can significantly improve the accuracy and efficiency of text classification.

[0075] The innovation of the present invention lies in the introduction of a deep convolutional neural network and the use of a learnable dimensional transformation matrix to efficiently complement global features and local features in the fusion stage. Experiments have shown that this method effectively improves the accuracy of text classification.

[0076] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A text classification method based on a pre-trained language model fused with a deep convolutional network, characterized in that: The method comprises the following steps: (1) Feature representation extraction: extract features from the input text using the RoBERTa model; (2) Deep semantic capture: Input the features output by RoBERTa into the multi-head attention network to obtain deep semantic information; (3) Local feature enhancement: The output of the multi-head attention is input into a convolutional neural network with a multi-layer residual structure to obtain deep convolutional network features; (4) Feature fusion and classification prediction: The output of the multi-head attention network is fused with the output of the convolutional neural network in the final stage. The fused feature representation is used to generate the final classification result logits through the fully connected layer.

2. The method according to claim 1, characterized in that In step (2), the multi-head attention network decomposes the feature representation extracted by RoBERTa into multiple independent subspaces, each subspace captures deep semantic information through an independent attention head, and performs weighted aggregation to generate a unified output representation.

3. The method according to claim 1, characterized in that In the step (3), the convolutional neural network with a multi-layer residual structure enhances local features through a 3×3 convolution operation, and combines LayerNorm and GeLU activation functions for normalization and nonlinear transformation, thereby improving the expression ability of local context information.

4. The method according to claim 1, characterized in that: In the step (4), the output of the multi-head attention is introduced into a learnable dimensional transformation matrix, and a low-dimensional representation suitable for the classification task is generated by optimizing the dimensional features. The low-dimensional representation is then fused with the output of the multi-layer convolutional neural network after a global maximum pooling operation to unify the semantic representation of global and local features.

Citation Information

Cited By

  • Controllable attention method and system based on feature domain division and medium

    CN120197509A

  • Prompt text detection method and device, program product and storage medium

    CN120849622A