Text classification method, system, equipment and medium

By introducing a multi-headed attention layer and a parallel attention layer into the convolutional neural network, the problem of context information loss in complex text classification is solved, and the classification accuracy and robustness are significantly improved.

CN120011569APending Publication Date: 2025-05-16SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178344.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When traditional convolutional neural networks process complex text, pooling operations lead to loss of context-related features, making it difficult to capture long-distance dependency semantic associations, resulting in low classification accuracy.

Method used

Multiple cascaded feature extraction modules are adopted, each module includes a convolutional layer and a feature aggregation layer. The feature aggregation layer is a multi-headed attention layer or a parallel attention layer. The semantic relationships inside the text are learned from multiple angles through the parallel attention mechanism.

Benefits of technology

It effectively retains more context information, enhances the modeling ability of long-distance dependency features, and improves the accuracy and robustness of text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011569A_ABST
    Figure CN120011569A_ABST
Patent Text Reader

Abstract

The invention relates to a text classification method, system and device and a medium. The method comprises the steps of obtaining a to-be-classified target text; inputting the target text into a feature extraction module group of a text classification model to generate text features; wherein the feature extraction module group comprises a plurality of cascaded feature extraction modules, each feature extraction module comprises at least one convolution layer and a feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and at least one feature aggregation layer is a parallel attention layer; and inputting the text features into a full connection layer of a text classification model to obtain a category to which the target text belongs. According to the invention, the text classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a text classification method, system, device and medium. Background Art

[0002] Text classification refers to the classification of text into predefined categories based on its content. It is widely used in various fields such as sentiment analysis, news classification, and spam detection. Convolutional neural networks are widely used in text classification tasks because they have good generalization ability in extracting local features. Traditional convolutional neural networks usually extract features by a convolution layer, then reduce the dimension of features through a pooling layer, and finally pass the reduced features into a deep neural network for learning and classification. However, when processing complex text information, the pooling operation will inevitably lose some of the contextual association features, and it is difficult to retain the semantic associations of long-distance dependencies, resulting in the model's final poor accuracy in complex text classification, which limits the application of the model in high-precision text analysis tasks. Therefore, it is necessary to provide a text classification method, system, device, and medium. Summary of the invention

[0003] In view of the above-mentioned shortcomings of the prior art, an object of the present invention is to provide a text classification method, system, device and medium, which improve the problem of poor accuracy in the prior art when performing complex text classification.

[0004] To achieve the above-mentioned purpose and other related purposes, the present invention provides a text classification method, comprising: obtaining a target text to be classified; inputting the target text into a feature extraction module group of a text classification model to generate text features; wherein the feature extraction module group includes a plurality of cascaded feature extraction modules, each feature extraction module includes at least one convolutional layer and a feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and there is at least one feature aggregation layer that is a parallel attention layer; inputting the text features generated by the last feature extraction module into the fully connected layer of the text classification model to obtain the category to which the target text belongs.

[0005] In one embodiment of the present invention, the step of obtaining the target text to be classified includes: obtaining an initial target text; and performing data cleaning on the initial target text to obtain a final target text.

[0006] In one embodiment of the present invention, the target text is input into the feature extraction module group of the text classification model to generate text features, including: performing word segmentation processing on the target text to obtain multiple text words, and arranging the multiple text words in sequence to obtain a text word sequence; performing word embedding processing on each text word in the text word sequence to obtain a word embedding word sequence; inputting the word embedding sequence into the feature extraction module group of the text classification model, performing feature processing through each feature extraction module, and finally generating text features.

[0007] In one embodiment of the present invention, when the feature aggregation layer is a multi-head attention layer, the feature processing is performed through each feature extraction module to finally generate text features, including: for each feature extraction module: the word embedding vocabulary sequence or intermediate text feature is input into the convolution layer of the corresponding feature extraction module for feature extraction to obtain local semantic features; the local semantic features are input into the feature aggregation layer of the corresponding feature extraction module, and based on the multi-head parallel attention mechanism, the deep features of the corresponding local semantic features are extracted through each attention head, and the extracted deep features are weighted and fused to generate the intermediate text features corresponding to the feature extraction module; the intermediate text features generated by the last feature extraction module are used as the final text features.

[0008] In one embodiment of the present invention, for each attention head, the process of extracting deep features from local semantic features includes: multiplying the local semantic features with a preset key weight matrix, value weight matrix and query weight matrix respectively to generate corresponding key matrix, value matrix and query matrix; calculating the similarity between the key matrix and the query matrix to generate an attention weight matrix; weighting the value matrix by the attention weight matrix to obtain a weighted feature matrix; residually connecting the local semantic features and the weighted feature matrix to extract the deep features of the local semantic features.

[0009] In one embodiment of the present invention, the local semantic features are multiplied with preset key weight matrix, value weight matrix and query weight matrix respectively to generate corresponding key matrix, value matrix and query matrix, including: evenly dividing the local semantic features into a preset number of sub-semantic features; for each sub-semantic feature: multiplying the sub-semantic feature with its preset key weight matrix, value weight matrix and query weight matrix respectively to generate corresponding key matrix, value matrix and query matrix; splicing the key matrices, value matrices and query matrices of all sub-semantic features according to corresponding feature dimensions to generate key matrix, value matrix and query matrix corresponding to the local semantic feature.

[0010] In one embodiment of the present invention, the attention weight matrix is ​​obtained based on a sparse attention mechanism.

[0011] In one embodiment of the present invention, a text classification system is also provided, the system comprising: a text acquisition module, used to acquire a target text to be classified; a text feature extraction module, used to input the target text into a feature extraction module group of a text classification model to generate text features; wherein the feature extraction module group comprises a plurality of cascaded feature extraction modules, each feature extraction module comprises at least one convolutional layer and a feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and there is at least one feature aggregation layer which is a parallel attention layer; a text classification module, used to input the text features generated by the last feature extraction module into the fully connected layer of the text classification model to obtain the category to which the target text belongs.

[0012] In one embodiment of the present invention, an electronic device is also provided, comprising: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements any of the text classification methods described above.

[0013] In one embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer executes any of the above-mentioned text classification methods.

[0014] As described above, a text classification method, system, device and medium of the present invention have the following beneficial effects: combining convolutional layers and different types of feature aggregation layers, effectively improving the text classification model's ability to extract text features. Among them, at least one feature aggregation layer adopts a parallel attention mechanism, so that the model can learn the semantic relationship within the text from multiple angles, and enhance the modeling ability of long-distance dependent features. Compared with the traditional method that only relies on pooling operations, the present invention can retain more contextual information and avoid the loss of key features due to dimensionality reduction operations. In addition, by inputting the deep text features generated by the last feature extraction module into the fully connected layer, the model can more accurately perform category discrimination, thereby improving the accuracy and robustness of classification. The present invention is applicable to a variety of text classification tasks, especially when processing complex texts or long texts, it can significantly improve the classification effect of the model, while also taking into account computational efficiency and information integrity. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A flowchart of a text classification method provided by an embodiment of the present invention;

[0016] Figure 2 Shown is a structural block diagram of a text classification system provided by an embodiment of the present invention;

[0017] Figure 3 Shown is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0019] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0020] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0021] The inventors found that in the current text classification task, the traditional convolutional neural network extracts text features through the convolution layer and uses the pooling layer for dimensionality reduction, but this method is difficult to capture long-distance semantic dependencies, resulting in low classification accuracy, especially in complex text classification scenarios, where the accuracy and recall rates are often less than 60%. Although the pooling operation can reduce the amount of calculation, it will lose some contextual information, making it difficult for the model to effectively learn the global features of the text. Based on this, in order to improve the classification effect of the convolutional neural network, a large amount of labeled data is usually required as a training set, which not only increases the data cost, but also puts higher requirements on computing resources.

[0022] Based on the above problems, the present invention provides a text classification method, which combines convolutional layers and different types of feature aggregation layers (including pooling layers and multi-head attention layers) to effectively improve the text classification model's ability to extract text features. Among them, at least one feature aggregation layer adopts a parallel attention mechanism, so that the model can learn the semantic relationship within the text from multiple angles and enhance the modeling ability of long-distance dependent features. Compared with the traditional method that only relies on pooling operations, the present invention can retain more contextual information and avoid the loss of key features due to dimensionality reduction operations. In addition, by inputting the deep text features generated by the last feature extraction module into the fully connected layer, the model can more accurately perform category discrimination, thereby improving the accuracy and robustness of classification. The present invention is applicable to a variety of text classification tasks, especially when processing complex texts or long texts, it can significantly improve the classification effect of the model, while also taking into account computational efficiency and information integrity.

[0023] See also Figure 1 ,The text classification method includes the following steps:

[0024] S1. Obtain the target text to be classified.

[0025] In one embodiment of the present invention, the step of obtaining the target text to be classified includes:

[0026] Get the initial target text;

[0027] Perform data cleaning on the initial target text to obtain the final target text.

[0028] The target text in the present invention is a long text sequence with complex semantic relationships. When a certain piece of text information needs to be analyzed to determine its type, the initial target text needs to be obtained first. The initial target text refers to the original text data without any preprocessing, and its acquisition method includes but is not limited to user input, database, web crawling or news articles. Considering that the original text may contain special symbols, meaningless characters or other noises that affect subsequent analysis, it is necessary to clean the initial target text to remove null values, invalid characters, etc. to obtain the final target text.

[0029] S2. Input the target text into a feature extraction module group of a text classification model to generate text features; wherein the feature extraction module group includes a plurality of cascaded feature extraction modules, each feature extraction module includes at least one convolutional layer and one feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and there is at least one feature aggregation layer that is a parallel attention layer.

[0030] In the present invention, considering that the pooling operation often loses some key information in the process of feature aggregation, especially when dealing with complex semantic relationships, its ability to capture long-distance dependencies is weak, resulting in limited classification accuracy of the model. Therefore, the present invention replaces at least one pooling layer in the text classification model with a multi-head attention layer to enhance the model's ability to analyze global features of the text. Specifically, the target text is input into a feature extraction module group of the text classification model, which is composed of a plurality of cascaded feature extraction modules, each of which is used to extract the intermediate text features of the text, wherein the text classification model can be any deep learning model capable of realizing text classification, including but not limited to CNN, RNN, LSTM, BERT, RoBERTa, ALBERT and other models based on Transformer structure.

[0031] In one embodiment of the present invention, the step of inputting the target text into a feature extraction module group of a text classification model to generate text features includes:

[0032] Performing word segmentation processing on the target text to obtain a plurality of text words, and arranging the plurality of text words in sequence to obtain a text word sequence;

[0033] Performing word embedding processing on each text word in the text word sequence to obtain a word embedding word sequence;

[0034] The word embedding sequence is input into the feature extraction module group of the text classification model, and feature processing is performed through each feature extraction module to finally generate text features.

[0035] The target text is segmented to divide the original natural language into multiple independent vocabulary units for subsequent model analysis and processing. After the word segmentation, the multiple text words obtained by the segmentation are sorted according to their positions in the original target text to obtain a text vocabulary sequence. For each text word in the text vocabulary sequence, word embedding processing is performed to map the text vocabulary to a vector space of a preset dimension, thereby converting the entire target text into a word vector sequence to obtain a word embedding vocabulary sequence. Among them, the word embedding processing method includes but is not limited to processing through pre-trained models such as BERT and Word2Vec. The word embedding vocabulary sequence is input into the trained text classification model, and deep semantic features are extracted layer by layer through multiple cascaded feature extraction modules to finally generate text features.

[0036] In one embodiment of the present invention, when the feature aggregation layer is a multi-head attention layer, the feature processing is performed step by step through each feature extraction module to finally generate text features, including:

[0037] For each feature extraction module:

[0038] Input the word embedding sequence or intermediate text features into the convolution layer of the corresponding feature extraction module for feature extraction to obtain local semantic features;

[0039] The local semantic features are input into the feature aggregation layer of the corresponding feature extraction module. Based on the multi-head parallel attention mechanism, the deep features of the corresponding local semantic features are extracted through each attention head, and the extracted deep features are weighted and fused to generate the intermediate text features corresponding to the feature extraction module.

[0040] The intermediate text features generated by the last feature extraction module are used as the final text features.

[0041] Since the text classification model includes multiple cascaded feature extraction modules, the intermediate text features output by the previous feature extraction module are used as the input of the next feature extraction module, and the input of the first feature extraction module is a word embedding sequence. For each feature extraction module: Since the feature extraction module first extracts features through several convolutional layers to extract local semantic features at the phrase level, so that the model can understand the association between each word. After all the convolutional layers of the feature extraction module have been extracted, the local semantic features extracted by the last convolutional layer are input to the feature aggregation layer for processing to integrate the features. When the feature aggregation layer is a multi-head attention layer, the feature aggregation layer includes multiple parallel attention heads, each of which independently processes the input local semantic features to generate corresponding deep features. Since different attention heads pay attention to different feature patterns, for example, some attention heads pay attention to long-distance dependencies, and some attention heads pay attention to the relationship between short-distance words. Through this method of parallel and independent processing of multiple attention heads, a more comprehensive and rich text feature representation can be obtained. Finally, the deep features generated by all attention heads are weighted and fused to obtain the intermediate text features generated by the current feature extraction module. The intermediate text features of the last feature extraction module are used as the final text features and input into the subsequent fully connected layer analysis.

[0042] It is understandable that different text classification models include different numbers of feature extraction modules, and the number of convolutional layers included in each feature extraction module also varies. The specific number can be adaptively set by technical personnel in this field based on task requirements and is not limited here.

[0043] In one embodiment of the present invention, for each attention head, the process of extracting deep features from local semantic features includes:

[0044] Multiplying the local semantic features with a preset key weight matrix, a value weight matrix, and a query weight matrix respectively to generate corresponding key matrices, value matrices, and query matrices;

[0045] Calculating the similarity between the key matrix and the query matrix to generate an attention weight matrix;

[0046] Weighting the value matrix by the attention weight matrix to obtain a weighted feature matrix;

[0047] The local semantic features are connected to the weighted feature matrix residual to extract the deep features of the local semantic features.

[0048] Since the text classification model is pre-trained, the relevant parameters of all query weight matrices, key weight matrices and value weight matrices involved therein have been determined. For each attention head, the input local semantic features are multiplied by the preset query weight matrix, key weight matrix and value weight matrix respectively to obtain the corresponding query matrix, key matrix and value matrix. By calculating the dot product of the query matrix and the key matrix, the attention weight matrix is ​​obtained. Furthermore, in order to avoid gradient instability caused by excessive values, the present invention also divides the dot product of the query matrix and the key matrix by the square root of the feature dimension, and performs Softmax normalization on the calculation results to obtain the corresponding attention weight matrix. Among them, the attention weight matrix is ​​used to characterize the degree to which different text contents are paid attention to. In one embodiment of the present invention, the attention weight matrix is ​​obtained based on a sparse attention mechanism. Compared with the traditional fully connected attention, the computational complexity and memory usage are significantly reduced, and it is particularly suitable for long text processing. The attention weight matrix is ​​multiplied by the value matrix of the attention head to obtain a weighted feature matrix. In this way, text information with high attention takes up a higher proportion in the final deep features, and text information with low attention takes up a smaller proportion, thereby reflecting the dependency between the various text information while retaining local information. Furthermore, in order to ensure the stability of deep feature extraction, the present invention also adds the original local semantic features and the weighted feature matrix so that the model can retain the original information and the new attention enhancement information at the same time to avoid excessive deformation of the features. The generated features are normalized by layer normalization to obtain the deep features generated by the feature extraction module.

[0049] In one embodiment of the present invention, the local semantic features are multiplied by a preset key weight matrix, a value weight matrix, and a query weight matrix to generate corresponding key matrices, value matrices, and query matrices, including:

[0050] Evenly dividing the local semantic feature into a preset number of sub-semantic features;

[0051] For each sub-semantic feature: multiply the sub-semantic feature with its preset key weight matrix, value weight matrix and query weight matrix respectively to generate the corresponding key matrix, value matrix and query matrix;

[0052] The key matrices, value matrices and query matrices of all sub-semantic features are concatenated according to the corresponding feature dimensions to generate the key matrix, value matrix and query matrix corresponding to the local semantic feature.

[0053] Considering that the amount of data calculated by each attention head is large, resulting in a slow reasoning speed of the model, in order to improve this situation, the present invention divides the local semantic features into several sub-semantic features in the multi-head attention mechanism by means of grouped linear transformation, and independently calculates the query matrix, key matrix and value matrix of each sub-semantic feature to improve the calculation rate. Specifically, for each attention head in each feature extraction module: the input local semantic features are evenly divided into several sub-semantic features, and each sub-semantic feature independently calculates the corresponding query matrix Q, key matrix K and value matrix V. Specifically, each sub-semantic feature is multiplied by the query weight matrix, key weight matrix and value weight matrix preset by the sub-semantic feature to obtain the corresponding query matrix, key matrix and value matrix. Since each sub-semantic feature performs the above calculation in a subspace of lower dimension, it can not only effectively reduce the computational complexity, but also enable the model to capture semantic information at different levels in different subspaces. The query matrix Q, key matrix K and value matrix V corresponding to all sub-semantic features are spliced ​​according to the feature dimension to restore the complete query matrix, key matrix and value matrix. Through the above process, since each subspace can learn different types of attention information, it can effectively improve the model's ability to learn and analyze complex text features, greatly reduce the amount of calculation, and improve the processing rate.

[0054] S3. Input the text features generated by the last feature extraction module into the fully connected layer of the text classification model to obtain the category to which the target text belongs.

[0055] The text features generated by the last feature extraction module of the text classification model are input into the fully connected layer, and the text features are mapped to the preset category space to obtain the probability of the target text in different categories, and several categories with the highest probability or probability higher than the preset threshold are selected as the category of the target text. For example, the target text is "Yesterday, the men's 100m final of the World Track and Field Championships was held in the United States. The contestants sprinted at an amazing speed, and the Jamaican contestant finally won the championship in 9.85 seconds." It needs to be classified to determine which category it belongs to, including sports, finance, entertainment, and technology. It is input into the aforementioned text classification model, and it can be obtained that the probability of it belonging to the sports type is the highest, so the category of this target text is the sports type.

[0056] See also Figure 2, the text classification system 100 includes: a text acquisition module 110, a text feature extraction module 120 and a text classification module 130. The above-mentioned text acquisition module 110 is used to obtain the target text to be classified. The text feature extraction module 120 is used to input the target text into the feature extraction module group of the text classification model to generate text features; wherein the feature extraction module group includes a plurality of cascaded feature extraction modules, each feature extraction module includes at least one convolution layer and a feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and there is at least one feature aggregation layer that is a parallel attention layer. The text classification module 130 is used to input the text features generated by the last feature extraction module into the fully connected layer of the text classification model to obtain the category to which the target text belongs.

[0057] For the specific definition of the text classification system, please refer to the definition of the text classification method above, which will not be repeated here. Each module in the above text classification system can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in or independent of the processor in the computer device in hardware format, or stored in the memory of the computer device in software format, so that the processor can call the corresponding operations of each of the above modules.

[0058] It should be noted that, in order to highlight the innovative part of the present invention, the present embodiment does not introduce modules that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other modules in the present embodiment.

[0059] See also Figure 3 The electronic device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a text classification program.

[0060] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (for example: SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Further, the memory 12 can also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 12 can not only be used to store application software and various types of data installed in the electronic device 1, such as text classification codes, etc., but also can be used to temporarily store data that has been output or is to be output.

[0061] In some embodiments, the processor 13 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 13 is the control core (Control Unit) of the electronic device 1, and uses various interfaces and lines to connect various components of the entire electronic device 1, and executes or executes programs or modules (such as text classification programs, etc.) stored in the memory 12, and calls data stored in the memory 12 to execute various functions of the electronic device 1 and process data.

[0062] The processor 13 executes the operating system and various installed application programs of the electronic device 1. The processor 13 executes the application programs to implement the steps in the above text classification method.

[0063] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a text acquisition module 110, a text feature extraction module 120, and a text classification module 130.

[0064] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium, and the computer-readable storage medium can be non-volatile or volatile. The above-mentioned software function module is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to perform part of the functions of the text classification method described in each embodiment of the present application.

[0065] In summary, the present invention discloses a text classification method, system, device and medium, which combine convolutional layers and different types of feature aggregation layers (including pooling layers and multi-head attention layers) to effectively improve the text classification model's ability to extract text features. Among them, at least one feature aggregation layer adopts a parallel attention mechanism, so that the model can learn the semantic relationship within the text from multiple angles and enhance the modeling ability of long-distance dependent features. Compared with the traditional method that only relies on pooling operations, the present invention can retain more contextual information and avoid the loss of key features due to dimensionality reduction operations. In addition, by inputting the deep text features generated by the last feature extraction module into the fully connected layer, the model can more accurately perform category discrimination, thereby improving the accuracy and robustness of classification. The present invention is applicable to a variety of text classification tasks, especially when processing complex texts or long texts, it can significantly improve the classification effect of the model, while also taking into account computational efficiency and information integrity. Compared with traditional convolutional neural networks, the present invention replaces at least one pooling layer after the convolutional layer with a multi-head attention layer. Under the premise of ensuring the lightweight of the model, the feature extraction method is optimized to improve the accuracy and recall of text classification, so that it has better classification performance in complex semantic tasks. Experiments have shown that under the same data, compared with traditional convolutional neural networks, the overall accuracy of the classification results obtained by this method of the present invention is improved by 27%. In addition, the present invention adopts the attention mechanism to replace part of the pooling operation, while improving the classification accuracy and recall rate, it also reduces the dependence on large-scale labeled data, thereby reducing training costs and improving training efficiency. Therefore, the present invention effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.

[0066] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A text classification method, characterized in that: The method comprises: Get the target text to be classified; Input the target text into a feature extraction module group of a text classification model to generate text features; wherein the feature extraction module group includes a plurality of cascaded feature extraction modules, each feature extraction module includes at least one convolutional layer and a feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and at least one feature aggregation layer is a parallel attention layer; The text features are input into the fully connected layer of the text classification model to obtain the category to which the target text belongs.

2. The text classification method according to claim 1, characterized in that: The step of obtaining the target text to be classified includes: Get the initial target text; Perform data cleaning on the initial target text to obtain the final target text.

3. The text classification method according to claim 1, characterized in that: The step of inputting the target text into a feature extraction module group of a text classification model to generate text features comprises: Performing word segmentation processing on the target text to obtain a plurality of text words, and arranging the plurality of text words in sequence to obtain a text word sequence; Performing word embedding processing on each text word in the text word sequence to obtain a word embedding word sequence; The word embedding sequence is input into the feature extraction module group of the text classification model, and feature processing is performed through each feature extraction module to finally generate text features.

4. The text classification method according to claim 3, characterized in that: When the feature aggregation layer is a multi-head attention layer, feature processing is performed through each feature extraction module to finally generate text features, including: For each feature extraction module: Inputting the word embedding vocabulary sequence or intermediate text features into the convolution layer of the corresponding feature extraction module for feature extraction to obtain local semantic features; The local semantic features are input into the feature aggregation layer of the corresponding feature extraction module. Based on the multi-head parallel attention mechanism, the deep features of the corresponding local semantic features are extracted through each attention head, and the extracted deep features are weighted and fused to generate the intermediate text features corresponding to the feature extraction module. The intermediate text features generated by the last feature extraction module are used as the final text features.

5. The text classification method according to claim 4, characterized in that: For each attention head, the process of extracting deep features from local semantic features includes: Multiplying the local semantic features with a preset key weight matrix, a value weight matrix, and a query weight matrix respectively to generate corresponding key matrices, value matrices, and query matrices; Calculating the similarity between the key matrix and the query matrix to generate an attention weight matrix; Weighting the value matrix by the attention weight matrix to obtain a weighted feature matrix; The local semantic features are connected to the weighted feature matrix residual to extract the deep features of the local semantic features.

6. The text classification method according to claim 5, characterized in that: The method of multiplying the local semantic features with a preset key weight matrix, a value weight matrix and a query weight matrix to generate corresponding key matrices, value matrices and query matrices includes: Evenly dividing the local semantic feature into a preset number of sub-semantic features; For each sub-semantic feature: multiply the sub-semantic feature with its preset key weight matrix, value weight matrix and query weight matrix respectively to generate the corresponding key matrix, value matrix and query matrix; The key matrices, value matrices and query matrices of all sub-semantic features are concatenated according to the corresponding feature dimensions to generate the key matrix, value matrix and query matrix corresponding to the local semantic feature.

7. The text classification method according to claim 5, characterized in that: The attention weight matrix is ​​obtained based on the sparse attention mechanism.

8. A text classification system, characterized in that: The system comprises: A text acquisition module is used to acquire the target text to be classified; A text feature extraction module, used for inputting the target text into a feature extraction module group of a text classification model to generate text features; wherein the feature extraction module group includes a plurality of cascaded feature extraction modules, each feature extraction module includes at least one convolutional layer and one feature aggregation layer, the feature aggregation layer is a pooling layer or a multi-head attention layer, and at least one feature aggregation layer is a parallel attention layer; The text classification module is used to input the text features into the fully connected layer of the text classification model to obtain the category to which the target text belongs.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enables the electronic device to implement the text classification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the text classification method according to any one of claims 1 to 7.