Text classification method and device, nonvolatile storage medium and electronic equipment

By using two gated networks in the multitasking processing model to control the output feature weights of the shared expert network and the specific expert network respectively, the problem of weight imbalance in the multitasking classification model is solved, and the accuracy of classification results is improved.

CN120045716APending Publication Date: 2025-05-27CHINA TELECOM CORP LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510192494.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the existing multi-task classification model, the weight imbalance caused by the use of a gated network to simultaneously control the output feature weights of the shared expert network and the specific expert network, which in turn caused the model to overfit or underfit on certain tasks, affecting the accuracy of the final classification results.

Method used

The multitasking model is adopted to control the output feature weights of the shared expert network and the specific expert network through two gated networks to avoid weight imbalance. The model includes multiple feature extraction layers, each feature extraction layer containing a shared expert network and multiple task-specific expert networks.

Benefits of technology

By avoiding weight imbalance, overfitting or underfitting of multi-task processing models is effectively avoided, and the accuracy of the final classification results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045716A_ABST
    Figure CN120045716A_ABST
Patent Text Reader

Abstract

The invention discloses a text classification method and device, a nonvolatile storage medium and electronic equipment. The method comprises the following steps: determining a text feature vector of a to-be-processed text and task types of a plurality of classification tasks; a multi-task processing model is adopted to process the text feature vectors to obtain task feature vectors corresponding to the multiple classification tasks, the multi-task processing model comprises multiple feature extraction layers, and each feature extraction layer is provided with a first gating network used for controlling output feature weights of the shared expert network and a second gating network used for controlling output feature weights of the shared expert network; and a second gating network for controlling output feature weights of the task specific expert network; and processing the task feature vector corresponding to the classification task through a classifier corresponding to the classification task to obtain a classification result of the classification task. According to the method and the device, the technical problem that the final classification result is low in accuracy due to the fact that one gating network is adopted to control the output feature weights of the shared expert network and the specific expert network at the same time in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular, to a text classification method, apparatus, non-volatile storage medium, and electronic device. Background Art

[0002] Currently, multi-task classification models in related technologies usually use a gating network to simultaneously control the output feature weights of a shared expert network and a specific expert network, which leads to the problem of weight imbalance, resulting in overfitting or underfitting of the multi-task classification model in certain types of tasks, affecting the accuracy of the final classification result.

[0003] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present application provide a text classification method, apparatus, non-volatile storage medium, and electronic device, so as to at least solve the technical problem of low accuracy of the final classification result caused by using a single gating network to simultaneously control the output feature weights of a shared expert network and a specific expert network in related technologies.

[0005] According to one aspect of the embodiments of the present application, a text classification method is provided, including: determining a text feature vector of a text to be processed, and task types of multiple classification tasks, where different classification tasks have different task types; processing the text feature vector by using a multi-task processing model to obtain task feature vectors respectively corresponding to the multiple classification tasks, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple specific task expert networks, a first gating network for controlling the output feature weight of the shared expert network, and a second gating network corresponding to each specific task expert network for controlling the output feature weight of the specific task expert network, and each classification task corresponds to a specific task expert network; determining classifiers corresponding to the respective classification tasks, and processing the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain classification results of the classification tasks.

[0006] Optionally, a multitask processing model is adopted to process the text feature vector to obtain task feature vectors corresponding to respective classification tasks in multiple classification tasks, including: inputting the input feature vector into a shared expert network and respective specific task expert networks in a target feature extraction layer. Wherein, when the target feature extraction layer is the first feature extraction layer in the multitask processing model, the input feature vector is the text feature vector; when the target feature extraction layer is not the first feature extraction layer, the input feature vector is the output feature vector of the previous feature extraction layer adjacent to the target feature extraction layer; adjusting the weights of each element in the first output feature vector of the shared expert network through a first gating network, and adjusting the weights of each element in the second output feature vector of the corresponding specific task expert network through a second gating network; when the target feature extraction layer is not the last feature extraction layer in the classification task processing model, concatenating the first output feature vector with adjusted weights and the second output feature vector with adjusted weights to obtain the output feature vector of the target feature extraction layer; when the target feature extraction layer is the last feature extraction layer in the classification task processing model, determining the task feature vectors corresponding to respective classification tasks according to the first output feature vector with adjusted weights and the second output feature vector with adjusted weights.

[0007] Optionally, each classification task corresponds to a first output feature vector with adjusted weights; determining the task feature vectors corresponding to respective classification tasks according to the first output feature vector with adjusted weights and the second output feature vector with adjusted weights includes: concatenating the first output feature vector with adjusted weights corresponding to the classification task and the second output feature vector with adjusted weights to obtain the task feature vector corresponding to the classification task.

[0008] Optionally, adjusting the weights of each element in the first output feature vector of the shared expert network through the first gating network includes: when the target feature extraction layer is the last feature extraction layer, adjusting the weights of each element in the first output feature vector through the first gating networks corresponding to respective classification tasks to obtain the first output feature vectors with adjusted weights corresponding to respective classification tasks; when the target feature extraction layer is not the last feature extraction layer, adjusting the weights of each element in the first output feature vector through the first gating networks corresponding to respective classification tasks to obtain the first output feature vectors with adjusted weights corresponding to respective classification tasks, and adjusting the weights of each element in the first output feature vector through the first gating network corresponding to the common features of all classification tasks to obtain the first output feature vector with adjusted weights corresponding to the common features of all classification tasks.

[0009] Optionally, determining the text feature vector of the text to be processed includes: removing hypertext markup tags, punctuation marks, and characters of a preset character type from the text to be processed; processing the text to be processed after removal through a bidirectional transducer encoder model to obtain a text feature vector.

[0010] Optionally, both the first gating network and the second gating network include a linear neural network and a Softmax output layer.

[0011] Optionally, the text classification method further includes: determining the preset task types corresponding to the respective specific task expert networks in the multi-task processing model; determining a target specific task expert network based on the preset task type corresponding to the specific task expert network and the task types of multiple classification tasks, where the target specific task expert network is the specific task expert network whose corresponding preset task type is among the task types of multiple classification tasks; determining that the target specific task expert network is the specific task expert network participating in the model training process of the classification model and the multi-task processing model.

[0012] According to another aspect of the embodiments of the present application, there is also provided a text classification device, including: a first processing module, configured to determine the text feature vector of the text to be processed and the task types of multiple classification tasks, where the task types of different classification tasks are different; a second processing module, configured to process the text feature vector by using a multi-task processing model to obtain task feature vectors respectively corresponding to multiple classification tasks, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple specific task expert networks, as well as a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each specific task expert network for controlling the output feature weights of the specific task expert network, and each classification task corresponds to a specific task expert network; a third processing module, configured to determine the classifiers corresponding to the respective classification tasks, and process the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain the classification results of the classification tasks, where the activation functions in different classifiers are different.

[0013] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium storing a program, where when the program runs, it controls the device where the non-volatile storage medium is located to execute the text classification method.

[0014] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run the program stored therein, and when the program runs, it executes the text classification method.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program which, when executed by a processor, implements the text classification method.

[0016] In the embodiments of the present application, the text feature vector of the text to be processed and the task types of multiple classification tasks are determined, where the task types of different classification tasks are different; a multi-task processing model is used to process the text feature vector to obtain task feature vectors respectively corresponding to multiple classification tasks. The multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple specific task expert networks, a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each specific task expert network for controlling the output feature weights of the specific task expert network. Each classification task corresponds to a specific task expert network; the classifiers corresponding to each classification task are determined, and the task feature vectors corresponding to the classification tasks are processed by the classifiers corresponding to the classification tasks to obtain the classification results of the classification tasks. By using the first gating network and the second gating network to control the output feature weights of the shared expert network and the specific expert network respectively, the purpose of avoiding weight imbalance is achieved, thereby realizing the technical effect of avoiding overfitting or underfitting of the multi-task processing model, and further solving the technical problem of the low accuracy of the final classification result caused by using a single gating network to control the output feature weights of the shared expert network and the specific expert network in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0018] Figure 1 is a schematic structural diagram of a computer terminal (portable terminal) provided according to an embodiment of the present application;

[0019] Figure 2 is a schematic flowchart of a text classification method provided according to an embodiment of the present application;

[0020] Figure 3 is a schematic structural diagram of a text classification model provided according to an embodiment of the present application;

[0021] Figure 4 is a schematic flowchart of a multi-task text classification process provided according to an embodiment of the present application;

[0022] Figure 5 is a schematic structural diagram of a multi-task classification device provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the scope of protection of this application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained as follows:

[0026] BERT (Bidirectional Encoder Representations from Transformers): BERT is a pre-trained language model based on the Transformer architecture, aiming to capture deep bidirectional representations of language through pre-training on large-scale text data.

[0027] PLE (Progressive Layered Extraction): PLE is a multi-task learning model, and its core idea is to separately process shared components and task-specific components, and introduce a progressive routing mechanism to extract and separate deeper semantic knowledge.

[0028] Text classification is a core technology in the field of natural language processing, aiming to automatically assign given text data to one or more predefined categories. This technology is widely used in many scenarios such as information retrieval, sentiment analysis, topic annotation, spam filtering, news classification, etc., and is of great significance for improving information processing efficiency and optimizing user experience.

[0029] In the related art, although some text classification methods are provided, they mainly focus on dealing with a single text classification task. However, in practical applications, it is often necessary to handle multiple tasks simultaneously, that is, to use the same set of data to complete multiple different text classification tasks. In this case, if a traditional single-task text classification model is adopted, the model needs to be run multiple times to complete all tasks, which will significantly increase the time cost. In contrast, adopting a multi-task classification model can handle these tasks more efficiently, thus saving costs.

[0030] In the field of multi-task learning, existing technologies mainly build models based on shared representations, shared parameters, and shared structures. However, these technologies have the following problems when dealing with complex problems in the multi-task field:

[0031] Interference between tasks. Multi-task learning aims to improve the generalization ability of the model by simultaneously learning multiple related tasks, but there may be interference problems between different tasks. For example, the features of some tasks may have a negative impact on the learning of other tasks, resulting in a decline in model performance. This interference is particularly obvious when the multi-task gradient directions are inconsistent, the task convergence speeds are inconsistent, and the magnitude differences of task losses are large, which may lead to oscillations in model parameters and negative transfer between tasks.

[0032] Determination of task weights. In multi-task learning, the importance of different tasks may be different, and how to determine the weights of tasks is a key issue. Incorrect weight allocation may lead to overfitting or underfitting of the model on some tasks. Although researchers have proposed some solutions, such as task weight learning mechanisms, how to automatically and accurately learn the weights of each task is still a challenge.

[0033] Insufficient utilization of information. Although multi-task learning methods based on parameter sharing can measure the correlation between related learning tasks, the utilization of multi-task learning information is still insufficient. This may cause the model to fail to fully capture the potential connections between tasks, thus limiting the improvement of model performance.

[0034] To solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0035] According to the embodiments of the present application, a method embodiment of a text classification method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0036] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The following shows a hardware block diagram of a computer terminal (or mobile device) for implementing a text classification method. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b,..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.

[0037] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this document. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0038] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the text classification method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned text classification method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0039] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0040] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0041] Under the above operating environment, an embodiment of the present application provides a text classification method, as Figure 2 shown, the method includes the following steps:

[0042] Step S202, determining a text feature vector of the text to be processed and task types of multiple classification tasks, where the task types of different classification tasks are different;

[0043] In the technical solution provided in step S202, the step of determining the text feature vector of the text to be processed includes: clearing hypertext markup tags, punctuation marks, and characters of a preset character type in the text to be processed; processing the cleared text to be processed through a bidirectional transducer encoder model to obtain a text feature vector.

[0044] In some embodiments of the present application, for the text to be processed or the data set of the text classification task, data cleaning processing can be performed first, including removing noise, word segmentation, removing stop words, etc. Then, clarify the classification objectives of each task, such as binary classification, multi-classification, or multi-label classification, etc.

[0045] As an alternative implementation, preprocessing operations such as data cleaning on the text to be processed or the text classification task data set include removing HTML tags and removing special characters and punctuation marks, such as #, @, $, %, &, *,?,!,.,,, etc. If the text contains HTML content, it needs to be deleted. Each piece of data obtained after preprocessing can be regarded as a text sequence, which is represented as T = [w 1 , w 2 , … w i . Where w i is the i-th word in the text sequence.

[0046] In some embodiments of the present application, after obtaining the text sequence, the BERT model can be used as the base model to obtain the text feature vector, so as to extract the representation information in the text sequence. BERT is a pre-trained deep bidirectional representation model that has obtained rich semantic and syntactic knowledge through unsupervised learning on a large amount of text data. Selecting BERT as the base model can utilize its powerful text representation ability to provide strong support for subsequent multi-task learning.

[0047] Optionally, obtaining the text feature vector using BERT includes the following steps:

[0048] First step, each word w in the text sequence T i is converted into an input form that can be processed by the BERT model, including converting the word into the corresponding word ID, and adding special tokens (such as [CLS] and [SEP]) to identify the start and end of the sentence, and finally the word position embedding, indicating the position of the word in the sentence;

[0049] Second step, the above input is passed to the Transformer layer of the BERT model, and after being processed by multiple layers of self-attention mechanisms and feed-forward neural networks, the embedding representation of each word in the context is obtained;

[0050] Third step, after obtaining the embedding representations of all words through the Transformer layer, the hidden state corresponding to the first token ([CLS] token) is extracted from these representations. This state is a high-dimensional vector (768-dimensional), which contains the aggregated information of the entire text sequence and is used to represent the semantic information of each sentence in the text sequence.

[0051] Step S204, using a multi-task processing model to process the text feature vector to obtain task feature vectors corresponding to multiple classification tasks respectively, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple task-specific expert networks, as well as a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each task-specific expert network for controlling the output feature weights of the task-specific expert network, and each classification task corresponds to a task-specific expert network;

[0052] In some embodiments of the present application, since a gating neural network is used in the PLE model to simultaneously control the output feature weights of the shared expert network and the task-specific expert networks, it is easy to have the problem of weight imbalance, that is, some feature weights are too high while some feature weights are too low. This results in the final output result being largely determined only by the part of the features with too high feature weights, which in turn leads to overfitting or underfitting of the model and affects the accuracy of the final classification result.

[0053] In the method provided in step S204, to solve the problem of weight imbalance, different gating networks are adopted in this application to adjust the weights of the shared expert network and the task-specific expert network respectively. Among them, both the first gating network and the second gating network include a linear neural network and a Softmax output layer.

[0054] In some embodiments of this application, the structure of the text classification model is as Figure 3 shown, including a BERT model, a multi-task classification model, and classifiers corresponding to each classification task. It can be seen from Figure 3 that the multi-task classification model includes multiple feature extraction layers, and each feature extraction layer contains a shared expert network (Experts Shared), and task-specific expert networks (Expert A and Expert B) corresponding to each classification task. The output of each expert network will be adjusted by the corresponding gating neural network, and then the outputs of each expert network after weight adjustment will be vector concatenated as the input of the next feature extraction layer. If it is the last feature extraction layer, the outputs of each expert network after weight adjustment will not be all concatenated together, but the output features of the shared expert network and the corresponding task-specific expert network after weight adjustment will be concatenated according to the task type of the classification task, and the concatenation result will be used as the input of the classifier corresponding to this task type.

[0055] It should be noted that Figure 3 for the convenience of description, there are only two types of classification tasks, A and B, and there are only two layers of feature extraction layers. In actual application, the number of types of classification tasks and the number of layers of feature extraction layers can be adjusted according to actual needs.

[0056] In addition, for Figure 3 any layer of feature extraction layer in, the output combination of the expert networks in this layer can be expressed as:

[0057]

[0058] In the above formula, S k,j (x) represents the output matrix of the expert network used by the kth task in the jth layer.

[0059] represents the transpose of the output vector of the ith expert network in the jth layer for task k.

[0060] The transpose of the output vector of the ith expert network of the shared expert network in the jth layer.

[0061] The superscript T in this formula represents the transpose operation, which is used to transpose the output vector of the expert network into a row vector for subsequent matrix combination and calculation. By transposing the output vectors of each expert network, they can be vertically stacked to form a matrix. Each row of this matrix corresponds to the output of an expert network, and the number of columns is equal to the dimension of the expert network output vector.

[0062] Figure 3 The superscript N of each gating neural network Gate in [description] indicates that the gating neural network adjusts the feature weights of the outputs of each expert network in the Nth layer feature extraction layer. The subscript A indicates that the output features of the specific task processing model corresponding to task A are weighted. The subscript B indicates that the output features of the specific task processing model corresponding to task B are weighted. The subscript S indicates that the output features of the shared expert network are weighted. When S has no subscript, it means that the gating neural network adjusts the output features based on the common features of all tasks. When S has a subscript, it means that the output features of the shared neural network are adjusted according to the task type corresponding to the subscript. For example, when the subscript of S is A, it means that the output features are adjusted according to the task features of task A type, and when the subscript is B, it means that the output features are adjusted according to the task features of task B type.

[0063] As an alternative implementation, as Figure 3 shown, the steps of using a multi-task processing model to process text feature vectors to obtain task feature vectors corresponding to each classification task in multiple classification tasks include: inputting the input feature vectors into the shared expert network and each specific task expert network in the target feature extraction layer respectively. Among them, when the target feature extraction layer is the first layer feature extraction layer in the multi-task processing model, the input feature vector is the text feature vector. When the target feature extraction layer is not the first layer feature extraction layer, the input feature vector is the output feature vector of the previous layer feature extraction layer adjacent to the target feature extraction layer; adjusting the weights of each element in the first output feature vector of the shared expert network through the first gating network, and adjusting the weights of each element in the second output feature vector of the corresponding specific task expert network through the second gating network; when the target feature extraction layer is not the last layer feature extraction layer in the classification task processing model, concatenating the weight-adjusted first output feature vector and the weight-adjusted second output feature vector to obtain the output feature vector of the target feature extraction layer; when the target feature extraction layer is the last layer feature extraction layer in the classification task processing model, determining the task feature vectors corresponding to each classification task based on the weight-adjusted first output feature vector and the weight-adjusted second output feature vector.

[0064] In some embodiments of the present application, each classification task corresponds to a first output feature vector with adjusted weights; the step of determining the task feature vector corresponding to each classification task based on the first output feature vector with adjusted weights and the second output feature vector with adjusted weights includes: concatenating the first output feature vector with adjusted weights corresponding to the classification task and the second output feature vector with adjusted weights to obtain the task feature vector corresponding to the classification task.

[0065] As an alternative implementation, as Figure 3 shown, the step of adjusting the weights of each element in the first output feature vector of the shared expert network through the first gating network includes: when the target feature extraction layer is the last feature extraction layer, adjusting the weights of each element in the first output feature vector through the first gating network corresponding to each classification task to obtain the first output feature vector with adjusted weights corresponding to each classification task; when the target feature extraction layer is not the last feature extraction layer, adjusting the weights of each element in the first output feature vector through the first gating network corresponding to each classification task to obtain the first output feature vector with adjusted weights corresponding to each classification task, and adjusting the weights of each element in the first output feature vector through the first gating network corresponding to the common features of all classification tasks to obtain the first output feature vector with adjusted weights corresponding to the common features of all classification tasks.

[0066] In some embodiments of the present application, as Figure 3 shown, when using a multi-task processing model for multi-task learning, the improved PLE model will gradually extract the information representations formed by the expert network. The output features of the shared expert network and the specific task expert network in each extraction network will be vector concatenated after weight adjustment and used as the input of the shared expert network and the specific task expert network in the next layer.

[0067] The calculation formula of the PLE model before improvement is as follows:

[0068] g k,j (x) = w k,j (g k,j-1 (x))S k,j (x)

[0069] where w k,j (g k,j-1 (x)) is a weight vector that determines the weights of the outputs of each expert network when combined. These weights are calculated through the gating network, and the input of the gating network is the output vector g k,j-1 (x) of the gating network in the previous layer (i.e., the j - 1 layer).

[0070] S k,j$(x)$ is the output matrix of the expert network used by the $k$-th task at the $j$-th layer. These expert networks include shared expert networks (i.e., expert networks used by all tasks) and task-specific expert networks (i.e., expert networks used only by task $k$). The output of each expert network is a vector, and these vectors are vertically stacked in the matrix $S$ k,j (x).

[0071] $g$ k,j (x) is the output vector of the gating network, which is calculated by multiplying the weight vector $w$ k,j (g k,j-1 (x)) with the output matrix $S$ k,j (x). This output vector contains both information from the shared expert network and the task-specific expert network, which are dynamically combined through the gating network to form the final representation of task $k$ at the $j$-th layer.

[0072] For the improved PLE model, the combination of the specific expert network outputs is:

[0073]

[0074] Denotes the output matrix of the specific expert network used by task $k$ at the $j$-th layer.

[0075] The combination of the shared expert network outputs is:

[0076]

[0077] Denotes the output matrix of the shared expert network used by all tasks at the $j$-th layer.

[0078] Formula for the specific expert gating network:

[0079] For task $k$ at the $j$-th layer, there is the following formula:

[0080]

[0081] Where, The weight matrix of the specific expert gating network for task $k$ at the $j$-th layer, $g$ k,j-1 (x) represents the concatenation of the unique gating network output vector and the shared gating network output vector of task $k$ at the $(j - 1)$-th layer (or the initial input if $j = 1$).

[0082] Denotes the weight vector of the specific expert gating network for task $k$ at the $j$-th layer.

[0083] Output of the specific expert gating network:

[0084]

[0085] Shared expert gating network calculation formula:

[0086] For all tasks at the j-th layer, the following formula holds:

[0087]

[0088] Where, represents the shared expert gating network weight matrix for all tasks at the j-th layer, and g j-1 (x) represents the concatenation of the output vectors of the shared expert gating network of the previous layer for all tasks (or the initial input if j = 1).

[0089] represents the shared expert gating network weight vector for all tasks at the j-th layer.

[0090] Shared expert gating network output:

[0091]

[0092] In each feature extraction layer, the input of the specific expert network is composed of the concatenation of the output of the specific expert gating network of the corresponding task in the previous layer and the output of the shared expert gating network. The input of the shared expert network is composed of the concatenation of the output vectors of the shared expert gating network of all tasks in the previous layer.

[0093] In the last layer, only the outputs of the specific task expert gating network and the shared expert gating network need to be concatenated to obtain the final vector for classification operations, and output it to the classifier corresponding to the task.

[0094] After being processed by multiple feature extraction layers, the task feature vectors corresponding to each classification task of the final output can be obtained and classified.

[0095] In some embodiments of the present application, the text classification method further includes: determining the preset task types corresponding to the respective specific task expert networks in the multi-task processing model; determining the target specific task expert network according to the preset task types corresponding to the specific task expert networks and the task types of the multiple classification tasks, where the target specific task expert network is the specific task expert network whose corresponding preset task type is among the task types of the multiple classification tasks; determining that the target specific task expert network is the specific task expert network participating in the model training process of the classification model and the multi-task processing model.

[0096] As an alternative implementation, pruning processing or parameter freezing processing can be directly performed on the non-target specific task expert networks.

[0097] Step S206: Determine the classifier corresponding to each classification task, and process the task feature vector corresponding to the classification task through the classifier corresponding to the classification task to obtain the classification result of the classification task.

[0098] In the technical solution provided in step S206, according to the specific task nature, multiple classifiers with different activation functions can be set. If it is a binary classification or multi-classification task, the activation function of the classifier is set to Softmax. If it is a multi-label classification task, the activation function is set to Sigmoid. The calculation formulas of the two activation functions are as follows:

[0099]

[0100] The Softmax function can perform normalization operations on the data and convert it into values between (0, 1). These values can be regarded as probability distributions and used as the target prediction values for binary classification or multi-classification. Among them, n represents the total number of categories, and i represents the index of the current category. The final output of the Softmax function represents the probability that the sample belongs to category i.

[0101] The Sigmoid function can calculate the output values [x 1 , x 2 , … x n of each node, and then obtain the probabilities of each category. These probabilities do not affect each other and are only related to the value of x i . Therefore, the sum of these probabilities may be greater than 1 or less than 1.

[0102] After being processed by the classifier, the classification result of each task on its corresponding category can be obtained.

[0103] As an alternative implementation, adversarial training can also be introduced to further improve the performance of the model. That is, during the model training process, adversarial samples are introduced, that is, new samples that can mislead the model are generated by making minor modifications to the original text. By training the model to correctly classify these adversarial samples, its robustness and generalization ability can be improved.

[0104] In some embodiments of the present application, a multi-task text classification process as shown in Figure 4 is also provided, including the following steps:

[0105] Step S402: Preprocess the text to be classified;

[0106] Step S404: Extract the text features of the preprocessed text to be classified;

[0107] Step S406: Process the text features using a multi-task classification model to obtain task features corresponding to each classification task;

[0108] Step S408: Process the corresponding task features using each classifier to obtain classification results for each classification task.

[0109] By determining the text feature vector of the text to be processed and the task types of multiple classification tasks, where the task types of different classification tasks are different; using a multi-task processing model to process the text feature vector to obtain task feature vectors corresponding to multiple classification tasks respectively, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple task-specific expert networks, as well as a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each task-specific expert network for controlling the output feature weights of the task-specific expert network, and each classification task corresponds to a task-specific expert network; determining the classifiers corresponding to each classification task, and processing the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain the classification results of the classification tasks, by using the first gating network and the second gating network to control the output feature weights of the shared expert network and the task-specific expert network respectively, the purpose of avoiding weight imbalance is achieved, thereby realizing the technical effect of avoiding overfitting or underfitting of the multi-task processing model, and further solving the technical problem of the low accuracy of the final classification result caused by using a single gating network to control the output feature weights of the shared expert network and the task-specific expert network in the related technology.

[0110] In addition, compared with the PLE model in the related technology, the gating network in the PLE model in the related technology needs to simultaneously control the feature adoption weights extracted by the task-specific expert network and the shared expert network. In the actual training process, the situation of weight imbalance is very likely to occur. Improving this single gating network to use two gating networks to control the weights of the features extracted by the task-specific expert network and the shared expert network respectively can effectively handle the situation of weight imbalance, and at the same time helps the model to extract multiple features according to different tasks.

[0111] In this application, the BERT model is used for initialization and fine-tuned in combination with domain-specific data to achieve domain adaptation. Then, in combination with a multi-task learning model, deep extraction of text features is performed. In this way, rich prior knowledge is provided for the pre-trained model, reducing the time and resource consumption of training from scratch. Through domain adaptation, the model can better understand the text features of a specific domain and improve the classification performance in that domain. Compared with directly training on a small dataset, this method significantly improves the starting point and final performance of the model. At the same time, in combination with the feature extraction ability of the multi-task learning model, the fused model can achieve performance improvement on multiple tasks simultaneously.

[0112] An embodiment of this application provides a text classification device. Figure 5 It is a schematic structural diagram of the device. As can be seen from Figure 5 it, the device includes: a first processing module 50, configured to determine the text feature vector of the text to be processed and the task types of multiple classification tasks, where the task types of different classification tasks are different; a second processing module 52, configured to process the text feature vector using a multi-task processing model to obtain task feature vectors corresponding to the multiple classification tasks respectively, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple specific-task expert networks, as well as a first gating network for controlling the output feature weights of the shared expert network and a second gating network corresponding to each specific-task expert network for controlling the output feature weights of the specific-task expert network, and each classification task corresponds to a specific-task expert network; a third processing module 54, configured to determine the classifiers corresponding to the respective classification tasks and process the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain the classification results of the classification tasks, where the activation functions in different classifiers are different.

[0113] In some embodiments of this application, the steps for the first processing module 50 to determine the text feature vector of the text to be processed include: clearing the hypertext markup tags, punctuation marks, and characters of a preset character type in the text to be processed; processing the cleared text to be processed through a bidirectional transducer encoder model to obtain the text feature vector.

[0114] In some embodiments of this application, both the first gating network and the second gating network include a linear neural network and a Softmax output layer.

[0115] In some embodiments of the present application, the step of the second processing module 52 using a multi-task processing model to process the text feature vector to obtain task feature vectors corresponding to each classification task in multiple classification tasks includes: inputting the input feature vector into the shared expert network and each specific task expert network in the target feature extraction layer respectively, where, when the target feature extraction layer is the first feature extraction layer in the multi-task processing model, the input feature vector is the text feature vector, and when the target feature extraction layer is not the first feature extraction layer, the input feature vector is the output feature vector of the previous feature extraction layer adjacent to the target feature extraction layer; adjusting the weights of each element in the first output feature vector of the shared expert network through the first gating network, and adjusting the weights of each element in the second output feature vector of the corresponding specific task expert network through the second gating network; when the target feature extraction layer is not the last feature extraction layer in the classification task processing model, concatenating the first output feature vector with adjusted weights and the second output feature vector with adjusted weights to obtain the output feature vector of the target feature extraction layer; when the target feature extraction layer is the last feature extraction layer in the classification task processing model, determining the task feature vectors corresponding to each classification task according to the first output feature vector with adjusted weights and the second output feature vector with adjusted weights.

[0116] In some embodiments of the present application, each classification task corresponds to a first output feature vector with adjusted weights; the step of the second processing module 52 determining the task feature vectors corresponding to each classification task according to the first output feature vector with adjusted weights and the second output feature vector with adjusted weights includes: concatenating the first output feature vector with adjusted weights corresponding to the classification task and the second output feature vector with adjusted weights to obtain the task feature vector corresponding to the classification task.

[0117] In some embodiments of the present application, the step of the second processing module 52 adjusting the weights of each element in the first output feature vector of the shared expert network through the first gating network includes: when the target feature extraction layer is the last feature extraction layer, adjusting the weights of each element in the first output feature vector through the first gating network corresponding to each classification task respectively to obtain the first output feature vector with adjusted weights corresponding to each classification task; when the target feature extraction layer is not the last feature extraction layer, adjusting the weights of each element in the first output feature vector through the first gating network corresponding to each classification task respectively to obtain the first output feature vector with adjusted weights corresponding to each classification task, and adjusting the weights of each element in the first output feature vector through the first gating network corresponding to the common features of all classification tasks to obtain the first output feature vector with adjusted weights corresponding to the common features of all classification tasks.

[0118] In some embodiments of the present application, the text classification device is further configured to: determine the preset task types corresponding to the respective specific task expert networks in the multi-task processing model; determine the target specific task expert network according to the preset task types corresponding to the specific task expert networks and the task types of the multiple classification tasks, where the target specific task expert network is the specific task expert network whose corresponding preset task type is among the task types of the multiple classification tasks; determine that the target specific task expert network is the specific task expert network participating in the model training process of the classification model and the multi-task processing model.

[0119] It should be noted that each module in the above text classification device may be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it may be presented in the following forms, but not limited thereto: the manifestation of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.

[0120] According to an embodiment of the present application, there is also provided a non-volatile storage medium storing a program, where, when the program runs, it controls the device where the non-volatile storage medium is located to execute the following text classification method: determine the text feature vector of the text to be processed, and the task types of multiple classification tasks, where the task types of different classification tasks are different; process the text feature vector using a multi-task processing model to obtain task feature vectors corresponding to the multiple classification tasks respectively, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple specific task expert networks, as well as a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each specific task expert network for controlling the output feature weights of the specific task expert network, and each classification task corresponds to a specific task expert network; determine the classifiers corresponding to the respective classification tasks, and process the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain the classification results of the classification tasks.

[0121] According to an embodiment of the present application, an electronic device is further provided, including: a memory and a processor, where the processor is configured to run a program stored in the memory. When the program runs, the following text classification method is executed: determining a text feature vector of the text to be processed and task types of multiple classification tasks, where the task types of different classification tasks are different; processing the text feature vector by using a multi-task processing model to obtain task feature vectors respectively corresponding to the multiple classification tasks, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple task-specific expert networks, a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each task-specific expert network for controlling the output feature weights of the task-specific expert network, and each classification task corresponds to a task-specific expert network; determining classifiers corresponding to the respective classification tasks, and processing the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain classification results of the classification tasks.

[0122] According to an embodiment of the present application, a computer program product is further provided, including a computer program, which when executed by a processor, implements the following text classification method: determining a text feature vector of the text to be processed and task types of multiple classification tasks, where the task types of different classification tasks are different; processing the text feature vector by using a multi-task processing model to obtain task feature vectors respectively corresponding to the multiple classification tasks, where the multi-task processing model includes multiple feature extraction layers, each feature extraction layer includes a shared expert network and multiple task-specific expert networks, a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each task-specific expert network for controlling the output feature weights of the task-specific expert network, and each classification task corresponds to a task-specific expert network; determining classifiers corresponding to the respective classification tasks, and processing the task feature vectors corresponding to the classification tasks through the classifiers corresponding to the classification tasks to obtain classification results of the classification tasks.

[0123] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0124] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0125] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0127] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.

[0128] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A text classification method, characterized in that: include: Determining a text feature vector of a text to be processed and task types of a plurality of classification tasks, wherein the task types of different classification tasks are different; Processing the text feature vector using a multi-task processing model to obtain task feature vectors corresponding to the multiple classification tasks, wherein the multi-task processing model includes multiple feature extraction layers, each of the feature extraction layers includes a shared expert network and multiple specific task expert networks, and a first gating network for controlling the output feature weights of the shared expert network, and a second gating network corresponding to each specific task expert network for controlling the output feature weights of the specific task expert network, and each of the classification tasks corresponds to one specific task expert network; The classifier corresponding to each of the classification tasks is determined, and the task feature vector corresponding to the classification task is processed by the classifier corresponding to the classification task to obtain a classification result of the classification task.

2. The text classification method according to claim 1, characterized in that: The text feature vector is processed using a multi-task processing model to obtain a task feature vector corresponding to each of the classification tasks in the multiple classification tasks, including: Inputting the input feature vector into the shared expert network and each of the specific task expert networks in the target feature extraction layer respectively, wherein, when the target feature extraction layer is the first feature extraction layer in the multi-task processing model, the input feature vector is the text feature vector, and when the target feature extraction layer is not the first feature extraction layer, the input feature vector is the output feature vector of the previous feature extraction layer adjacent to the target feature extraction layer; Adjusting the weight of each element in the first output feature vector of the shared expert network through the first gating network, and adjusting the weight of each element in the second output feature vector of the corresponding task-specific expert network through the second gating network; In a case where the target feature extraction layer is not the last feature extraction layer in the classification task processing model, concatenating the first output feature vector after weight adjustment and the second output feature vector after weight adjustment to obtain an output feature vector of the target feature extraction layer; In the case where the target feature extraction layer is the last feature extraction layer in the classification task processing model, the task feature vectors corresponding to each of the classification tasks are determined according to the first output feature vector after weight adjustment and the second output feature vector after weight adjustment.

3. The text classification method according to claim 2, characterized in that: Each of the classification tasks corresponds to a weight-adjusted first output feature vector; determining the task feature vectors corresponding to each of the classification tasks according to the weight-adjusted first output feature vector and the weight-adjusted second output feature vector comprises: The first output feature vector after weight adjustment and the second output feature vector after weight adjustment corresponding to the classification task are concatenated to obtain the task feature vector corresponding to the classification task.

4. The text classification method according to claim 2, characterized in that: Adjusting the weights of each element in the first output feature vector of the shared expert network by the first gating network includes: In the case where the target feature extraction layer is the last feature extraction layer, weights of each element in the first output feature vector are adjusted respectively through the first gating network corresponding to each of the classification tasks to obtain the first output feature vector after weight adjustment corresponding to each of the classification tasks; In the case that the target feature extraction layer is not the last feature extraction layer, the weights of each element in the first output feature vector are adjusted respectively through the first gating network corresponding to each of the classification tasks to obtain the first output feature vector after the weight adjustment corresponding to each of the classification tasks, and the weights of each element in the first output feature vector are adjusted through the first gating network corresponding to the common features of all the classification tasks to obtain the first output feature vector after the weight adjustment corresponding to the common features of all the classification tasks.

5. The text classification method according to claim 1, characterized in that: Determining the text feature vector of the text to be processed includes: Clearing hypertext markup tags, punctuation marks and characters of a preset character type from the text to be processed; The cleaned text to be processed is processed by a bidirectional transformer encoder model to obtain the text feature vector.

6. The text classification method according to claim 1, characterized in that: The first gating network and the second gating network both include a linear neural network and a Softmax output layer.

7. The text classification method according to claim 1, characterized in that: The text classification method also includes: Determining a preset task type corresponding to each of the specific task expert networks in the multi-task processing model; Determine a target specific task expert network according to the preset task type corresponding to the specific task expert network and the task types of the multiple classified tasks, wherein the target specific task expert network is a specific task expert network corresponding to the preset task type among the task types of the multiple classified tasks; The target specific task expert network is determined to be a specific task expert network that participates in a model training process of a classification model and the multi-task processing model.

8. A text classification device, characterized in that: include: A first processing module, used for determining a text feature vector of a text to be processed and a task type of a plurality of classification tasks, wherein the task types of different classification tasks are different; A second processing module is used to process the text feature vector using a multi-task processing model to obtain task feature vectors corresponding to the multiple classification tasks, wherein the multi-task processing model includes multiple feature extraction layers, each of the feature extraction layers includes a shared expert network and multiple specific task expert networks, and a first gating network for controlling the output feature weight of the shared expert network, and a second gating network corresponding to each specific task expert network for controlling the output feature weight of the specific task expert network, and each of the classification tasks corresponds to one specific task expert network; The third processing module is used to determine the classifier corresponding to each of the classification tasks, and process the task feature vector corresponding to the classification task through the classifier corresponding to the classification task to obtain the classification result of the classification task, wherein the activation functions in different classifiers are different.

9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the text classification method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the text classification method according to any one of claims 1 to 7 when running.

11. A computer program product, characterized in that The invention comprises a computer program, which implements the text classification method according to any one of claims 1 to 7 when being executed by a processor.

Citation Information

Cited By

  • Data processing method

    CN120234286A

  • Text classification method, and deep learning model training method and device

    CN121144908A

  • Text classification method, deep learning model training method and device

    CN121144908B