Multi-channel feature collaborative text classification method and system
By using multi-channel feature collaboration module and tag-aware data enhancement method in text classification, the problem of insufficient comprehension capabilities of complex texts in the prior art is solved, and more efficient feature extraction and classification performance improvement is achieved.
Patent Information
- Application Number
- CN202510004928.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-02
AI Technical Summary
When existing text classification algorithms deal with complex text, it is difficult to fully explore the potential relationship between text and labels, and are not comprehensive enough when considering multi-angle information of features, resulting in insufficient model understanding of texts with different complexity.
A multi-channel feature collaborative text classification method is proposed. By using a multi-channel feature collaborative module, including a RoBERTa encoder and a RoFormer encoder, the feature vectors with different position information focus are extracted, and fusion and optimization are performed through the feature processing module. At the same time, the label-aware data enhancement method and rotational position coding technology are used to enhance the ability of feature extraction.
Through the multi-channel feature collaboration module and tag-aware data enhancement method, the location information and tag information of the text can be captured more comprehensively, and the model's discrimination ability and classification performance of complex semantics can be improved. The threshold tuning strategy module further optimizes the generalization ability of the model and significantly improves the accuracy and stability of text classification.
Smart Images

Figure CN119938914A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the field of text classification technology, and in particular relates to a multi-channel feature collaborative text classification method and system. Background Art
[0002] Text classification is a core task in the field of natural language processing and is widely used in many scenarios such as sentiment analysis, comment classification, and question-answering systems. However, the diversity and complexity of text data make it an urgent problem to improve the accuracy of text classification. The structural characteristics of complex texts make it inefficient and time-consuming to extract effective information from them, and traditional methods have limitations in dealing with text complexity.
[0003] Traditional deep learning text classification methods mainly include CNN, RNN, attention mechanism and models improved based on these methods. Kim et al. applied convolutional neural network (CNN) to text classification tasks, using convolution kernels of different sizes to extract key information in the text and better capture local correlation. Zhang et al. proposed a character-level convolutional neural network Char-CNN. Experiments show that when the training set is large enough, the convolutional network can achieve good results even without considering the meaning and grammatical structure of words. In the application of RNN, MIKOLOV et al. improved the RNN model and enhanced the model's ability to process sequence data. However, both traditional CNN and RNN models have certain shortcomings: RNN may have gradient vanishing or gradient exploding problems when processing text, and CNN is limited by long-term dependency problems. To this end, Lai et al. proposed an RCNN model that combines RNN and CNN, using RNN to capture contextual information and using maximum pooling to obtain potential semantic information in the text. The core idea is to combine the advantages of both and solve their respective limitations. Xiao et al. proposed a neural network architecture that uses convolutional layers and recurrent layers to effectively encode character inputs. This model achieves performance comparable to other models with fewer parameters.
[0004] The introduction of the attention mechanism enables neural networks to pay different attention to different sentences when training texts, achieving more reasonable natural language modeling. This method not only has stronger parallel capabilities, but also significantly improves training efficiency, so more and more neural networks are beginning to integrate the attention mechanism. The Att-BLSTM model proposed by ZHOU et al. combines the attention mechanism with the bidirectional LSTM to effectively capture the semantic information in the sentence and convert it into useful feature representations. The BERT model is completely based on the Transformer architecture. It not only has excellent generalization but also amazing performance. Therefore, many excellent models have begun to improve based on BERT. Lan et al. proposed a lightweight pre-training model ALBERT (ALite BERT), which successfully reduced the number of model parameters and improved performance by adopting technologies such as factorized embedding parameters, cross-layer parameter sharing, and sentence order prediction loss, which helps to apply natural language processing technology to more scenarios.
[0005] Chen et al. proposed a new semi-supervised text classification algorithm MixText, the key of which is to use a data augmentation method called TMix to generate new samples by interpolating in latent space. This method performs well in few-sample scenarios, surpassing pre-trained models and other semi-supervised methods. Meng et al. proposed a method LOTClass that uses only labels to train text classification models without any labeled documents as training data, achieving performance comparable to or even better than strong semi-supervised and supervised models. Lin et al. proposed a method that combines linear classifiers with advanced methods to apply to text features, and verified its effectiveness through experiments, in response to the fact that training and deploying advanced methods directly on large-scale pre-trained models is not always satisfactory. Kim et al. proposed a simple framework DRAFT that can classify arbitrary topics. The framework uses dense retrievers to build customized datasets and fine-tune the classifier. With extremely low number of parameters, DRAFT shows good performance in few-sample topic classification tasks.
[0006] In view of the above analysis, the technical problems that need to be solved urgently in the existing technology are: some existing text classification algorithms focus on fine-tuning the model or adjusting the classifier. However, not only is the potential relationship between text and labels not fully explored, but the multi-angle information of features is also not comprehensive enough. In terms of supervised training, most of the work still has room for further improvement in the tuning part to enhance the model's ability to understand texts of different complexities. Summary of the invention
[0007] In view of the problems existing in the prior art, the present invention provides a multi-channel feature collaborative text classification method and system
[0008] The present invention is implemented as follows: a multi-channel feature collaborative text classification method, comprising:
[0009] S1, in the encoding stage, the target text information is first input into the multi-channel feature collaboration module, and feature vectors with different position information emphases are extracted through two independent encoders;
[0010] S2, these features are processed by the Feature Processing Module (FPM) to obtain output features;
[0011] S3, before classifying the text, the output features need to be screened and judged by the threshold tuning strategy module.
[0012] Furthermore, in text classification tasks, text position information is crucial to the model's understanding and classification performance. The order of words and sentences in a text not only carries semantics, but also reflects the dependencies between contexts. Accurately capturing and utilizing this position information can help the model understand the structure and meaning of the text more deeply and improve its ability to discern complex semantics.
[0013] Based on the importance of text position information, a multi-channel feature collaboration module is designed. The multi-channel feature collaboration module consists of two main components: a feature encoder module and a feature processing module. The feature encoder module integrates two independent encoders, the RoBERTa encoder and the RoFormer encoder. In the encoding stage, in order to obtain richer text features, the two encoders receive different input texts and extract feature representations with specific focuses.
[0014] Furthermore, we use the label-aware data augmentation method to insert all labels before the sentence of the RoBERTa encoder to form a new sequence token_1. Using this sequence for feature extraction, we can obtain the potential connection between labels and sentences and obtain the text feature f based on label augmentation. b , as shown in Formula 5-6:
[0015] token_1=[[CLS],c1,c2,…,c k ,x1,…,x n ] (5)
[0016] f b =RoBERTa(token_1) (6)
[0017] Furthermore, for the RoFormer encoder, its unique rotation position encoding method increases the learning of relative position information between sequence elements. Since the RoBEREa encoder has already incorporated label information, the RoFormer encoder is used only to extract features of text sentences to obtain features containing relative position information. r As shown in formula 7-8:
[0018] token_2=[x1,x2,…,x n ] (7)
[0019] f r =RoFormer(token_2) (8)
[0020] Take the feature f r With f b Perform the splicing feature fusion method to form the output feature f c , as shown in Formula 9.
[0021] f c =cat(f r ,f b ) (9)
[0022] Furthermore, considering the dimensional disaster that may be caused by too long features, the concatenated feature vector is subjected to three-dimensionality reduction processing through the feature processing module. This module includes a linear layer, an activation function RELU layer, a BN layer, and a Dropout layer, as shown in formula 10-14.
[0023] Splicing feature f c First, after the first linear layer, its dimension parameter is reduced from 3328 to 2024, and the output feature f is obtained. l1 After passing through the activation function RELU and BN layer, it enters the Dropout layer. After many experiments, the parameter is set to 0.45.
[0024] f l1 =BatchNormld(RELU(Linear(f c ))) (10)
[0025] f l1 = Dropout(f l1 ,p=0.45) (11)
[0026] The dimension parameter of the second linear layer is reduced from 2024 to 1024, and the parameter of the Dropout layer is set to 0.3 after multiple experiments.
[0027] f l2 =BatchNormld(RELU(Linear(fl1 ))) (12)
[0028] f l2 = Dropout(f l2 ,p=0.3) (13)
[0029] After the last feature processing, the feature f l2 The dimension is reduced to 768, and the output feature f suitable for subsequent text classification tasks can be obtained.
[0030] f=Linear(f l2 ,output dim ) (14)
[0031] Furthermore, the threshold tuning strategy makes judgments based on upper and lower thresholds, and dynamically adjusts the definition of semantically ambiguous text and semantically simple text. Through the threshold screening function, different decisions are made based on the differences between the two types of text to improve the model's understanding ability. This processing method is more in line with human thinking habits in form, and the effectiveness of this method has been verified through experiments.
[0032] The semantic fuzzy text screening strategy is shown in Formula 15, where λ represents the fuzzy category screening threshold, i.e., the upper threshold of the threshold tuning strategy module. Assuming that the kth text is a semantic fuzzy text, its loss function is expressed as L yk , the average training loss is L avg For semantically ambiguous text that satisfies the formula, the model will use a re-learning method to deepen the model's understanding of this type of text.
[0033] L yk ≥λL avg (15)
[0034] Similarly, the semantically simple text category screening threshold μ is the lower threshold of the threshold tuning module. In order to avoid the model fitting being too biased towards semantically simple texts, the screened semantically simple texts will be excluded from this training. The semantically simple text screening strategy is shown in Formula 16:
[0035] L y k≤μL avg (16)
[0036] When the upper threshold λ of the threshold screening function is 1.10 and the lower threshold μ is 0.93, the model achieves the best accuracy improvement effect on the SST-2 dataset.
[0037] Another object of the present invention is to provide a multi-channel feature collaborative text classification system for implementing the multi-channel feature collaborative text classification method, comprising:
[0038] The two core components of the model are the multi-channel feature coordination module MCFSM (Multi-Channel Feature Synergy Module) and the threshold tuning policy module TTPM (Threshold Tuning Policy Module).
[0039] The multi-channel feature collaboration module consists of two main components: a feature encoder module and a feature processing module. The feature encoder module integrates two independent encoders - the RoBERTa encoder and the RoFormer encoder. In the encoding stage, in order to obtain richer text features, the two encoders receive different input texts and extract feature representations with specific focuses.
[0040] The threshold tuning strategy module makes judgments based on upper and lower thresholds and dynamically adjusts the definition of semantically ambiguous text and semantically simple text.
[0041] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the multi-channel feature collaborative text classification method.
[0042] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the multi-channel feature collaborative text classification method.
[0043] Another object of the present invention is to provide an information data processing terminal, which includes the multi-channel feature collaborative text classification system.
[0044] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0045] First, in view of the limitations of text information utilization in traditional text classification tasks and the lack of consideration of texts with different semantic complexity during model training, the present invention proposes a multi-channel feature collaborative text classification model FSTTS based on a threshold tuning strategy. Before extracting features from the input text, the model focuses on the relationship between label information and text information, and establishes multiple channels to mine text information. At the same time, based on the threshold tuning strategy, semantically ambiguous texts that are difficult for the model to understand and semantically simple texts that are easy to understand are processed separately, further improving the generalization ability and classification performance of the model. Experimental results on four public data sets have shown that the FSTTS model has good generalization and effectiveness.
[0046] Second, the patent of this invention proposes an efficient and accurate real-time monitoring method for online public opinion by making full use of the massive real-time text data in social platforms. This method is centered on text mining and natural language processing technology, and realizes dynamic capture and trend analysis of online public opinion by analyzing multi-dimensional information such as keywords, emotional tendencies, and user behavior patterns in social platforms. Unlike traditional public opinion monitoring systems, the present invention does not require additional hardware support and can run efficiently based on existing computing resources, greatly reducing installation and maintenance costs. This lightweight design is not only suitable for the public opinion management needs of governments and enterprises, but also convenient for families and individuals to use in daily life.
[0047] The present invention is particularly aimed at the social background that the current online public opinion has penetrated into the daily life of families, and designs a set of real-time public opinion monitoring solutions applicable to various life scenarios. Through the efficient data collection and processing module, users can quickly obtain the latest dynamics of online public opinion, including the propagation path of hot topics, the diffusion range of negative emotions, and early warning information of potential public opinion crises. At the same time, the present invention is also combined with sentiment analysis algorithms to accurately identify emotional fluctuations in public opinion content, thereby providing users with a more intuitive public opinion status display and response suggestions.
[0048] At the technical implementation level, the present invention ensures the stability and real-time performance of the system in a high-concurrency data environment through large-scale data processing optimization algorithms and adaptive model adjustment mechanisms. In addition, the system supports flexible customization functions, and users can set public opinion keywords, monitoring ranges, and warning thresholds according to their own needs, thereby achieving accurate monitoring of specific topics or events. This highly intelligent public opinion analysis method greatly reduces the user's dependence on technical background knowledge, allowing non-professional users to easily master and operate it.
[0049] The practical application value of the present invention is reflected in its purification of the network environment and maintenance of a healthy public opinion ecology. By grasping the trends of online public opinion in real time, users can quickly respond to public opinion crises, convey positive voices in a timely manner, and prevent the spread of false information. Especially in the era of social media where information is spread quickly and widely, the present invention provides strong technical support for building a harmonious network environment, helps to improve the efficiency of public opinion management, and promotes the healthy development of cyberspace. In summary, the present invention meets the needs of the current era and has important social significance and economic value. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of the structure of the RoBERTa model provided by an embodiment of the present invention;
[0051] Figure 2 is a schematic diagram of implementing the rotational position encoding provided by an embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the FSTTS model structure provided by an embodiment of the present invention;
[0053] Figure 4 It is a FSTTS training flow chart provided by an embodiment of the present invention;
[0054] Figure 5 This is a schematic diagram of an example of binary classification of semantically ambiguous text provided by an embodiment of the present invention;
[0055] Figure 6 is a flow chart of a multi-channel feature collaborative text classification method provided by an embodiment of the present invention;
[0056] Figure 7 It is a structural diagram of a multi-channel feature collaborative text classification system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] In the encoding stage, this method extracts features from the input target text information through a multi-channel feature collaboration module. This module contains two independent encoders, each of which generates feature vectors based on different position information. For example, one encoder may focus more on the global semantic information of the text, while the other encoder focuses more on the fine-grained features of the local context. Through collaborative encoding, the two encoders perform in-depth analysis of the text content from different dimensions, ensuring that the extracted feature information is comprehensive and rich, laying the foundation for subsequent feature processing and classification.
[0059] After completing the multi-channel feature extraction, these features will be passed to the feature processing module for further processing. The main function of this module is to fuse and optimize the feature vectors from the two channels, unifying the dimensions and semantic information of different features through technical methods such as feature alignment, weighting or dimensionality reduction. This fusion processing can eliminate feature redundancy and retain the key information of the text, thereby generating high-quality output features and providing accurate semantic support for subsequent classification tasks.
[0060] Before classifying the text, the output features need to be screened and judged by the threshold tuning strategy module. This module further optimizes the features based on the preset threshold conditions, screens out the features that contribute significantly to the classification task, and filters out noise or redundant information. This process enhances the adaptability to the classification target by dynamically adjusting the activation weights of the features, thereby improving the classification accuracy and stability of the model.
[0061] Finally, the features after multi-channel feature extraction, feature processing and threshold tuning are input into the classifier for classification decision. Due to the collaborative optimization of the previous modules, this method can comprehensively consider the semantic information of the text from both global and local dimensions, significantly improving the accuracy and robustness of the classification results. This multi-channel collaborative method is particularly suitable for complex text classification tasks and can efficiently deal with semantic diversity and uncertainty in multi-category text classification scenarios.
[0062] The BERT model integrates the MLM (Masked Language Model) task and the bidirectional Transformer encoder structure, and is trained on a large-scale unlabeled corpus to obtain deep feature representations of the input text. While using the attention mechanism to globally understand contextual information, the model uses a fixed random masking method to prevent overfitting. In addition, the NSP (Next Sentence Prediction) task helps improve the model's ability to adapt to different downstream tasks and provide rich semantic expressions for various downstream tasks.
[0063] The RoBERTa model is an improved pre-training model based on the BERT model. Compared with the BERT model, RoBERTa introduces a dynamic mask mode in the data preprocessing stage, and generates a dynamic mask for each input sequence. This improvement enables the model to gradually adapt to different masking strategies and learn different language representations under the continuous input of a large amount of data. In the pre-training stage, RoBERTa uses the BPE (Byte Pair Encoding) method, expands the vocabulary size, and does not pre-process the input text. Based on the above improvements, RoBERTa has better robustness and accuracy in capturing deep language details.
[0064] The RoFormer model is a language model based on the Transformer architecture, which uses the Rotary Position Embedding (RoPE) method. This method uses the rotation matrix to encode the absolute position of the word embedding and embeds the relative position information into the self-attention mechanism. The core idea is to encode the relative position by multiplying the word embedding representation (q and k in the Transformer) with the rotation matrix, realizing the fusion of relative position encoding and linear self-attention, providing flexibility in sequence length, and making the dependency between characters decrease as the relative position increases. The implementation of RoPE is as follows: Figure 2 shown.
[0065] In order to achieve relative position encoding, RoPE adds absolute position information to q and k through formula (1).
[0066]
[0067] where q m is the mth element of q, k n Similarly, in order to include relative position information, the present invention uses a function g, which represents the m With k n The inner product between the generated vectors represents that this function only takes the word embeddings q and k and their relative positions mn as input variables, as shown in formula (2).
[0068] q m T k n = <f q (q,m),f k (k,n)>=g(q,k,mn) (2)
[0069] The function f(x) containing the absolute position information is converted into a matrix form using a complex formula, as shown in formula (3). Since matrix multiplication usually represents a spatial transformation operation, it can be seen that formula (3) represents a rotation operation on q, so it is called a rotation position encoding method.
[0070]
[0071] The expression generalized to high dimensions can be simplified as formula (4).
[0072]
[0073] The RoPE method adopted by RoFormer has been theoretically analyzed to show that relative positions can be naturally represented by vector products in the self-attention mechanism. This method uses the rotation angle between vectors to represent the relative position relationship between features, achieving the effect of relative position encoding by absolute position encoding, thereby enhancing the model's ability to understand semantically ambiguous texts. This method has the advantages of flexibility, attenuation, and scalability.
[0074] This paper proposes a text classification model FSTTS that combines threshold tuning strategy and multi-channel feature collaboration. Figure 3 The two core components of the model are the Multi-Channel Feature Synergy Module (MCFSM) and the Threshold Tuning Policy Module (TTPM).
[0075] In the encoding stage, the target text information is first input into the multi-channel feature collaboration module, and feature vectors with different position information focuses are extracted through two independent encoders. Subsequently, these features are processed by the feature processing module (FPM) to obtain output features. Before classifying the text, the output features must also be screened and judged by the threshold tuning strategy module.
[0076] In text classification tasks, text position information is crucial to the model's understanding and classification performance. The order of words and sentences in a text not only carries semantics, but also reflects the dependencies between contexts. Accurately capturing and utilizing this position information can help the model understand the structure and meaning of the text more deeply and improve its ability to discern complex semantics.
[0077] Based on the importance of text position information, the present invention designs a multi-channel feature collaboration module. The multi-channel feature collaboration module consists of two main components: a feature encoder module and a feature processing module. Figure 4 As shown in Figure 2, the feature encoder module integrates two independent encoders, the RoBERTa encoder and the RoFormer encoder. In the encoding stage, in order to obtain richer text features, the two encoders receive different input texts and extract feature representations with specific focuses.
[0078] There is usually an implicit association between the labels in the dataset and the text. For example, in the sentence "george purchased a small interest in which baseball team?" in the TREC dataset, the subject "george" is a person's name, so the text should be attributed to the label "human". If the label information can be associated with the text information, it may effectively enhance the model's understanding ability. Inspired by this, the present invention introduces a label-aware data enhancement method for the RoBERTa encoder.
[0079] The input of the RoBERTa encoder usually consists of tags and sentences for classification, but such input obviously ignores the potential impact of tags on classification. In order to solve the problem of lack of tag information, the present invention adopts a tag-aware data enhancement method to insert all tags before the sentence of the RoBERTa encoder to form a new sequence token_1. Using this sequence for feature extraction, the potential connection between tags and sentences can be obtained, and the text feature f based on tag enhancement can be obtained. b , as shown in Formula 5-6:
[0080] token_1=[[CLS],c1,c2,…,c k,x1,…,x n ] (5)
[0081] f b =RoBERTa(token_1) (6)
[0082] For the RoFormer encoder, its unique rotation position encoding method increases the learning of relative position information between sequence elements. Since the RoBEREa encoder has already incorporated label information, the RoFormer encoder is used only to extract features of text sentences to obtain features containing relative position information. r As shown in formula 7-8:
[0083] token_2=[x1,x2,…,x n ] (7)
[0084] f r =RoFormer(token_2) (8)
[0085] Zhang Haifeng et al. proposed to extract domain common features and characteristic features based on the dual BERT model fused with Fpnet, and purify them in combination with feature projection to improve the model classification effect. However, this method is sensitive to the selection of hyperparameters and has poor interpretability. Therefore, in order to ensure that the final output features retain as much information as possible, the present invention adopts the method of converting feature f r With f b Perform the splicing feature fusion method to form the output feature f c , as shown in Formula 9.
[0086] f c =cat(f r ,f b ) (9)
[0087] Considering the dimensional disaster that may be caused by overly long features, the present invention performs three-dimensionality reduction processing on the spliced feature vectors through a feature processing module. The module includes a linear layer, an activation function RELU layer, a BN layer, and a Dropout layer, as shown in formulas 10-14.
[0088] Splicing feature f c First, after the first linear layer, its dimension parameter is reduced from 3328 to 2024, and the output feature f is obtained. l1 After passing through the activation function RELU and BN layer, it enters the Dropout layer. After many experiments, the parameter is set to 0.45.
[0089] f l1 =BatchNormld(RELU(Linear(f c ))) (10)
[0090] f l1 = Dropout(f l1 ,p=0.45) (11)
[0091] The dimension parameter of the second linear layer is reduced from 2024 to 1024, and the parameter of the Dropout layer is set to 0.3 after multiple experiments.
[0092] f l2 =BatchNormld(RELU(Linear(f l1 ))) (12)
[0093] f l2 = Dropout(f l2 ,p=0.3) (13)
[0094] After the last feature processing, the feature f l2 The dimension is reduced to 768, and the output feature f suitable for subsequent text classification tasks can be obtained.
[0095] f=Linear(f l2 ,output dim ) (14)
[0096] In the field of text classification, the final performance of the model depends on the effectiveness of the training process. In order to obtain good classification results, many existing models will optimize parameters through continuous iteration until a pre-set performance threshold is reached. However, this traditional training mechanism fails to fully consider that texts with a large difference in complexity from regular texts during the training process have a certain impact on the model's understanding ability. In fact, distinguishing between semantically ambiguous texts and semantically simple texts can effectively improve the overall understanding ability of the model.
[0097] like Figure 5 As shown in the figure, for a text binary classification task, the traditional training iteration method may cause the model to lack depth in understanding ambiguous semantic texts with unclear meanings or prone to ambiguity, making the model insensitive to the differences between features.
[0098] From the perspective of loss function, semantically ambiguous text has a larger loss during training, while semantically simple text has a smaller loss. Based on this difference, the present invention designs and implements a threshold tuning strategy. This strategy makes judgments based on upper and lower thresholds and dynamically adjusts the definition of semantically ambiguous text and semantically simple text. Figure 5 As shown in the figure, through the threshold screening function, different decisions are taken according to the differences between the two types of text to improve the model's understanding ability. This processing method is more in line with human thinking habits in form, and the effectiveness of this method has been verified through experiments.
[0099] The semantic fuzzy text screening strategy is shown in Formula 15, where λ represents the fuzzy category screening threshold, i.e., the upper threshold of the threshold tuning strategy module. Assuming that the kth text is a semantic fuzzy text, its loss function is expressed as L yk For semantically ambiguous texts that satisfy the formula, the model will use a re-learning method to deepen the model's understanding of this type of text.
[0100] L yk ≥λL avg (15)
[0101] Similarly, the semantically simple text category screening threshold μ is the lower threshold of the threshold tuning module. In order to avoid the model fitting being too biased towards semantically simple texts, the screened semantically simple texts will be excluded from this training. The semantically simple text screening strategy is shown in Formula 16:
[0102] L yk ≤μL avg (16)
[0103] In the actual training process, in order to screen out the most appropriate upper threshold λ and lower threshold μ, the present invention conducted multiple experiments on the SST-2 dataset. Multiple rounds of experiments have shown that when the upper threshold λ of the threshold screening function is 1.10 and the lower threshold μ is 0.93, the model has the best accuracy improvement effect on the SST-2 dataset. The experimental results are shown in Section 3.4 of the present invention.
[0104] In order to verify the effectiveness of the model of the present invention, the present invention selects 4 benchmark text classification data sets for experiments.
[0105] The SST-2 dataset is a binary classification dataset for sentiment analysis, labeled as positive or negative sentiment.
[0106] The SUBJ dataset is a review dataset with sentences labeled as subjective or objective.
[0107] TREC dataset is a question-answering dataset divided into 6 categories for question type classification
[0108] The CR dataset is a sentiment analysis dataset consisting of customer reviews, which is used to judge the positive or negative sentiment of product reviews.
[0109] Table 1 summarizes the statistics of these datasets.
[0110] Table 1 Dataset statistics
[0111]
[0112] The environments of the present invention and the ablation experiment are both based on Python-3.10 and PyTorch2.1.0, the CPU is AMD EPYC7T83, the graphics card is NVIDIA GeForceRTX4090-24G, and the CUDA version is 12.2, see Table 2.
[0113] Table 2 Experimental environment
[0114]
[0115] The present invention uses accuracy (Acc) to evaluate the text classification results of the present invention, and the specific calculation is shown in Formula 17:
[0116]
[0117] Among them, TP (True Positive) means that the positive sample in text classification is predicted to be positive, FP (False Posi tive) means that the negative sample in text classification is predicted to be positive, TN (True Negative) means that the negative sample in text classification is predicted to be negative, and FN (False Negative) means that the negative sample in text classification is predicted to be positive.
[0118] 3.2 Experimental Results and Analysis
[0119] In order to verify the superiority of the model of the present invention in the text classification task, the present invention has conducted multiple comparative experiments with BERT, RoBERTa, EFL, LM-CPPF and other models, and the experimental results are shown in Table 3. Among them, the best results in each data set are marked in bold. As can be seen from Table 3, the model of the present invention after optimization of semantics and features is comprehensively superior to the benchmark model DualCL (using RoBERTa encoder), and achieved accuracies of 94.29%, 97.50%, 98.20% and 94.58% on the SST-2, SUBJ, TREC and CR data sets, respectively. Among them, the improvement is the largest on the SST-2 data set, with the accuracy increased by 1.54%.
[0120] In addition, from the overall experimental results of the four data sets, it can be seen that the FSTTS model has the best overall performance. Although the performance on the SST-2 data set is slightly lower than that of the EFL model, the accuracy of the FSTT S model exceeds that of the EFL model on the other three data sets, with an increase of 0.40% and 2.08% respectively. The analysis found that the EFL model used a larger pre-trained encoder RoBERTa-large, and in the text classification task, the use of a larger pre-trained model has a significant impact on performance improvement.
[0121] Table 3 Experimental results of each model on different datasets
[0122]
[0123]
[0124] Summarizing the performance of the FSTTS model on the four datasets of SST-2, SUBJ, TREC, and CR, it can be inferred that the labels in the SST-2 dataset are more closely related to the text information, and the distribution of its samples in the feature space is closer to the decision boundary than other datasets. Therefore, after the FSTTS model is optimized for features and semantics, the accuracy on this dataset is more significantly improved, and the results of the ablation experiment also verify this. The excellent performance on the SUBJ, TREC, and CR datasets may be due to the fact that the threshold tuning strategy effectively distinguishes semantic ambiguity from simple text, enhancing the generalization ability of the model.
[0125] In order to verify the effectiveness of the multi-channel feature collaboration module and the threshold tuning strategy, the present invention conducted an ablation experiment. As shown in Table 4, the ablation experiment was conducted on four public datasets in three parts, and the best result of each dataset is shown in bold.
[0126] Table 4 Ablation experiment
[0127]
[0128] The ablation experiment design in the first part uses only the multi-channel feature collaboration module, aiming to explore whether there are certain differences in text understanding between the two independent encoders that capture different position information of the text. From the perspective of experimental results, the FSTTS w / o TTP model with the addition of the multi-channel feature collaboration module performs better than the baseline model DualCL on the four datasets. This not only proves that there is a close connection between labels and text, and that fused labels can improve the overall understanding ability of the model, but also shows that the fusion of position information plays an important role in improving the classification effect of the model. Therefore, the addition of this module significantly enhances the model's understanding of text.
[0129] The second part of the experiment adjusts the model training process and increases the model's processing of semantically ambiguous text and semantically simple text, which is more in line with human thinking habits from the perspective of training. From the perspective of experimental results, the classification performance of the FSTTS w / o MCFS model using the threshold tuning strategy has improved to a certain extent on the four data sets. From the extent of the improvement, it is inferred that the sentences in the SST-2 data set are mostly distributed near the classification decision line in the feature space. The experiment proves the effectiveness of this strategy.
[0130] The experiments in the third part prove that the classification ability of the model can be further improved by using the multi-channel feature collaboration module and the threshold tuning module at the same time. The experimental results show that the two modules can not only cooperate with each other in improving the classification effect of the model, but also have a certain generalization ability.
[0131] The threshold tuning strategy module uses a specific threshold parameter design, and its parameter value determines the model's tolerance for differences in different texts. Since different datasets have different text labels, text quality, and text length, the choice of threshold parameters may also vary.
[0132] Table 5 Accuracy of different upper and lower thresholds on the SST-2 dataset
[0133]
[0134] In order to explore the difference in the impact of different threshold parameter selections in the threshold tuning strategy module on the model in the text classification task, a threshold selection analysis experiment was designed on the SST-2 dataset. The experimental results are shown in Table 5. It can be seen from the figure that when different threshold ranges are adopted, the accuracy of the FSTTS model on the SST-2 dataset is significantly different. When the upper threshold is selected as 1.10, its mean accuracy is numerically better than the other upper thresholds. When the lower threshold is selected as 0.93, its performance in mean accuracy is also better than the other lower thresholds. Comprehensive comparative analysis found that on the SST-2 dataset, when the upper and lower thresholds are 1.10 and 0.93, the performance of the FSTTS model reaches the best.
[0135] like Figure 6 As shown, the multi-channel feature collaborative text classification method includes:
[0136] S1, in the encoding stage, the target text information is first input into the multi-channel feature collaboration module, and feature vectors with different position information emphases are extracted through two independent encoders;
[0137] S2, these features are processed by the feature processing module to obtain output features;
[0138] S3, before classifying the text, the output features need to be screened and judged by the threshold tuning strategy module.
[0139] like Figure 7 As shown in the figure, the multi-channel feature collaborative text classification system includes:
[0140] The two core components of the model are the multi-channel feature coordination module MCFSM and the threshold tuning strategy module TTPM;
[0141] The multi-channel feature collaboration module consists of two main components: the feature encoder module and the feature processing module. The feature encoder module integrates two independent encoders, the RoBERTa encoder and the RoFormer encoder. In the encoding stage, in order to obtain richer text features, the two encoders receive different input texts and extract feature representations with specific focuses.
[0142] The threshold tuning strategy module makes judgments based on upper and lower thresholds and dynamically adjusts the definition of semantically ambiguous text and semantically simple text.
[0143] The multi-channel feature collaborative text classification system optimizes the text classification task by combining the multi-channel feature collaborative module (MCFSM) with the threshold tuning strategy module (TTPM). The core of the system is to comprehensively utilize the multi-channel feature representation capability and dynamic threshold tuning strategy to improve the accuracy and generalization ability of text classification, especially showing higher robustness when processing semantically complex and ambiguous texts.
[0144] The feature encoder module in MCFSM contains two independent encoders - RoBERTa encoder and RoFormer encoder.
[0145] The RoBERTa encoder aims to efficiently capture the deep semantic features of text and focuses on modeling global semantic information.
[0146] The RoFormer encoder introduces rotational position encoding, focusing on capturing local features and positional relationships.
[0147] In the encoding stage, the two encoders receive different input texts respectively and extract feature representations with specific focuses, thus forming rich multi-channel text features.
[0148] The feature processing module fuses and optimizes the feature representations output by the RoBERTa and RoFormer encoders.
[0149] -Through the feature alignment mechanism, the scale inconsistency problem between multi-channel features is solved.
[0150] -Use the attention mechanism to give more important features higher weights to ensure the capture of key semantic information.
[0151] The goal of the feature processing module is to generate multi-channel collaborative feature representations to provide high-quality input for subsequent classification.
[0152] TTPM uses dynamic upper and lower thresholds to evaluate the semantic complexity of text.
[0153] Semantically ambiguous text: When the feature complexity is higher than the upper threshold, an enhanced classification strategy is adopted to improve the classification accuracy by increasing the weight distribution calculation of the feature processing module.
[0154] Semantically simple text: When the feature complexity is below the lower threshold, the feature processing flow is simplified to improve computational efficiency.
[0155] This dynamic tuning strategy enables the system to work efficiently in different semantic environments.
[0156] When processing text, the system first extracts diversified features through the multi-channel feature collaboration module, and then optimizes the features in combination with the threshold tuning strategy module. Finally, the optimized features are input into the classifier to complete the accurate classification of the text. This process is also suitable for multiple types of text, especially in scenarios with semantic uncertainty, which significantly improves the classification effect.
[0157] The system can be deployed on different types of computing devices, including:
[0158] A computer device that cooperatively executes the stored multi-channel feature collaborative text classification method through a memory and a processor;
[0159] Computer-readable storage media for loading and execution on different computing platforms;
[0160] Information data processing terminal, suitable for intelligent terminals and server environments, supporting large-scale text classification tasks.
[0161] Through the collaborative design of hardware and algorithms, the system has good scalability and a wide range of practical application scenarios. The present invention has been extensively experimentally verified on a variety of text classification datasets, and the two-category datasets SST-2, SUBJ, CR and the six-category dataset TERC are selected as typical test datasets. These datasets cover different text classification scenarios, such as sentiment classification (SST-2), subjectivity analysis (SUBJ), product review classification (CR) and news classification (TERC). By using these datasets, the classification performance of the present invention has been comprehensively evaluated in diverse tasks.
[0162] In the binary classification task, the classification accuracy of the present invention on the SST-2, SUBJ and CR datasets reached 96.90%, 97.20% and 94.02% respectively. On the six-classification task TERC dataset, the classification accuracy of the present invention reached 98.10%. The experimental results show that the present invention has excellent performance in different types of text classification tasks, especially in complex multi-classification tasks such as TERC. The high accuracy of the present invention further verifies its robustness and wide applicability.
[0163] By comparing the experimental results with the existing optimal model, the present invention shows better classification performance overall. Compared with other models, the present invention not only achieves higher classification accuracy on a single data set, but also maintains stability and consistency in classification performance on multiple data sets. This advantage stems from the innovative design of the present invention in feature extraction, model optimization and classification strategy, which ensures more accurate capture of text features and more effective decision-making capabilities in classification tasks.
[0164] The experimental results fully verify the superior performance of the present invention in the field of text classification. It not only reaches the leading level of current technology in binary classification tasks, but also shows significant accuracy improvement in multi-classification tasks. The present invention provides a more efficient and reliable technical solution for solving text classification tasks, which can be widely used in scenarios such as sentiment analysis, news classification, product evaluation, etc., and has important value for practical industrial and research applications.
[0165] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0166] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A multi-channel feature collaborative text classification method, characterized in that: include: S1. Feature extraction: In the encoding stage, the target text information is input into the multi-channel feature collaboration module, and feature vectors with different position information focuses are extracted through two independent encoders. The first encoder extracts global semantic features, and the second encoder extracts local context features. S2, feature processing: input the feature vector into the feature processing module, fuse and optimize the feature vectors from the two channels, and generate a unified output feature through feature alignment, weighting or dimensionality reduction method; S3, threshold tuning: input the output features into the threshold tuning strategy module, screen effective features based on preset threshold conditions, filter noise or redundant information, and obtain optimized features; S4, classification decision: input the optimized features into the classifier, classify the target text, and generate classification results.
2. The multi-channel feature collaborative text classification method according to claim 1, characterized in that: In text classification tasks, text position information is crucial to the model's understanding and classification performance; the order of words and sentences in a text not only carries semantics, but also reflects the dependencies between contexts; accurately capturing and utilizing this position information can help the model understand the structure and meaning of the text more deeply and improve its ability to discern complex semantics; Based on the importance of text location information, a multi-channel feature collaboration module is designed; The multi-channel feature collaboration module consists of two main components: a feature encoder module and a feature processing module. The feature encoder module integrates two independent encoders - the RoBERTa encoder and the RoFormer encoder. In the encoding stage, in order to obtain richer text features, the two encoders receive different input texts and extract feature representations with specific focuses.
3. The multi-channel feature collaborative text classification method according to claim 1, characterized in that: Using the label-aware data enhancement method, all labels are inserted before the sentence of the RoBERTa encoder to form a new sequence token_1; using this sequence for feature extraction, the potential connection between the label and the sentence can be obtained to obtain the text feature f based on label enhancement. b , as shown in Formula 5-6: token_1=[[CLS],c1,c2,…,c k ,x1,…,x n ] (5) f b =RoBERTa(token_1) (6)。 4. The multi-channel feature collaborative text classification method according to claim 1, characterized in that: For the RoFormer encoder, its unique rotation position encoding method increases the learning of relative position information between sequence elements; since the RoBEREa encoder has already incorporated label information, the RoFormer encoder is used only to extract features of text sentences to obtain features containing relative position information. r ; As shown in formula 7-8: token_2=[x1,x2,…,x n ] (7) f r =RoFormer(token_2) (8) Take the feature f r With f b Perform the splicing feature fusion method to form the output feature f c , as shown in formula 9; f c =cat(f r ,f b ) (9)。 5. The multi-channel feature collaborative text classification method according to claim 1, characterized in that: Considering the possibility of dimensional disaster caused by overly long features, the concatenated feature vector is subjected to three-dimensionality reduction processing through the feature processing module; this module includes a linear layer, an activation function RELU layer, a BN layer, and a Dropout layer, as shown in formula 10-14; Splicing feature f c First, after the first linear layer, its dimension parameter is reduced from 3328 to 2024, and the output feature f is obtained. l1 After passing through the activation function RELU and BN layer, it enters the Dropout layer. After multiple experiments, the parameter is set to 0.45; f l1 =BatchNormld(RELU(Linear(f c ))) (10) f l1 =Dropout(f l1 ,p=0.45) (11) The dimension parameter of the second linear layer is reduced from 2024 to 1024, and the parameter of the Dropout layer is set to 0.3 after multiple experiments; f l2 =BatchNormld(RELU(Linear(f l1 ))) (12) f l2 =Dropout(f l2 ,p=0.3) (13) After the last feature processing, the feature f l2 The dimension is reduced to 768, and the output feature f suitable for subsequent text classification tasks can be obtained; f=Linear(f l2 ,output dim ) (14)。 6. The multi-channel feature collaborative text classification method according to claim 1, characterized in that: The threshold tuning strategy makes judgments based on upper and lower thresholds, and dynamically adjusts the definition of semantically ambiguous text and semantically simple text. Through the threshold screening function, different decisions are made based on the differences between the two types of text to improve the model's understanding ability. This processing method is more in line with human thinking habits in form, and the effectiveness of this method has been verified through experiments. The semantic fuzzy text screening strategy is shown in Formula 15, where λ represents the fuzzy category screening threshold, that is, the upper threshold of the threshold tuning strategy module; assuming that the kth text is a semantic fuzzy text, its loss function is expressed as L yk ,For semantically ambiguous texts that satisfy the formula, the model will adopt a re-learning method to deepen the model’s understanding of this type of text; L yk ≥λL avg (15) Similarly, the semantically simple text category screening threshold μ is the lower threshold of the threshold tuning module. In order to avoid the model fitting being too biased towards semantically simple texts, the screened semantically simple texts will be excluded from this training. The semantically simple text screening strategy is shown in Formula 16: L yk ≤μL avg (16) When the upper threshold λ of the threshold screening function is 1.10 and the lower threshold μ is 0.93, the model achieves the best accuracy improvement effect on the SST-2 dataset.
7. A multi-channel feature collaborative text classification system that implements the multi-channel feature collaborative text classification method according to any one of claims 1 to 6, characterized in that: include: The multi-channel feature collaboration module consists of two main components: the feature encoder module and the feature processing module. The feature encoder module integrates two independent encoders, the RoBERTa encoder and the RoFormer encoder. In the encoding stage, in order to obtain richer text features, the two encoders receive different input texts and extract feature representations with specific focuses. The threshold tuning strategy module makes judgments based on upper and lower thresholds and dynamically adjusts the definition of semantically ambiguous text and semantically simple text.
8. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the multi-channel feature collaborative text classification method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the multi-channel feature collaborative text classification method as described in any one of claims 1 to 6.
10. An information data processing terminal, comprising the multi-channel feature collaborative text classification system according to claim 7.
Citation Information
Patent Citations
BERT improved model-based text sentiment analysis method
CN114781392A
RoBERTa-BiLSTM-CRF voice dialogue text named entity recognition system fused with attention mechanism
CN117010387A
Text classification method and system based on multi-granularity text features
CN117668235A