Text classification method and device based on multiple tags, electronic equipment and storage medium

Through text encoding and hierarchical label processing of the multi-label text classification model, the problem of fine-grained label semantic boundaries in multi-label classification is solved, and the accuracy of text classification is improved.

CN120257030APending Publication Date: 2025-07-04SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311868883.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has the problem of poor classification of fine-grained labels in multi-label classification, resulting in low classification accuracy.

Method used

Using a multi-label text classification model, text encoding, feature extraction and classification are performed through the combination of text encoding network, hierarchical label processing subnet and classification subnet, and the tag information of the hierarchical label tree is used for auxiliary classification to determine the target hierarchical label.

Benefits of technology

It improves the accuracy of text multi-label classification, ensures the accuracy of label depth and path, and solves the problem of unclear semantic boundaries of fine-grained labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257030A_ABST
    Figure CN120257030A_ABST
Patent Text Reader

Abstract

The invention provides a multi-label-based text classification method and device, electronic equipment and a storage medium, and belongs to the field of financial science and technology. The method comprises the steps of obtaining a target description text; performing coding processing on the target description text based on a text coding network of a multi-label text classification model to obtain target text semantic features; classifying the target text semantic features and the root tag features of the hierarchical tag tree based on a classification sub-network to obtain original hierarchical tags, and extracting hierarchical tag features of the original hierarchical tags based on a hierarchical tag processing sub-network; classifying the target text semantic features and the hierarchical tag features based on a classification sub-network to obtain prediction hierarchical tags; and if it is determined that the label depth of the predicted hierarchy label reaches a preset label depth, determining a target hierarchy label of the target description text based on a first path from the predicted hierarchy label to the root label feature in the hierarchy label tree. According to the method, the multi-label classification accuracy of the text can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a multi-label based text classification method, apparatus, electronic device and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, in data mining, text classification is one of the core issues in various business fields. For example, in financial scenarios such as item recommendation, financial business handling, and logistics business consultation, it is often necessary to assign multiple labels to various types of descriptive texts to refine the category to which each descriptive text belongs. For example, for a to-be-processed logistics task, it is often necessary to accurately determine the belonging department and personnel of the logistics task based on the task description text of the logistics task, so as to improve the processing efficiency of the task.

[0003] However, when the number of labels is large, common classification methods such as ten-classification and twenty-classification often cannot well distinguish the relationships between label categories, and there is a problem that the fine-grained semantic boundaries of the labels are unclear, which will lead to low accuracy of multi-label classification of texts. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a multi-label based text classification method, apparatus, electronic device and storage medium, aiming to improve the accuracy of multi-label classification of texts.

[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a multi-label based text classification method, the method comprising:

[0006] Obtain a target description text;

[0007] Perform encoding processing on the target description text based on the text encoding network of a preset multi-label text classification model to obtain the target text semantic feature of the target description text, wherein the multi-label text classification model further comprises a hierarchical label processing sub-network and a classification sub-network;

[0008] Perform classification processing on the target text semantic feature and the root label feature of a preset hierarchical label tree based on the classification sub-network to obtain the original hierarchical label of the target description text, the hierarchical label tree comprising a plurality of hierarchical labels;

[0009] Perform feature extraction on the original hierarchical label based on the hierarchical label processing sub-network to obtain a hierarchical label feature;

[0010] Perform classification processing on the target text semantic feature and the hierarchical label feature based on the classification sub-network to obtain the predicted hierarchical label of the target description text;

[0011] If it is determined that the label depth of the predicted hierarchical label reaches the preset label depth, determine a first path from the predicted hierarchical label to the root label feature in the hierarchical label tree;

[0012] Based on the hierarchical labels on the first path, determine the target hierarchical label of the target description text.

[0013] In some embodiments, the text encoding network includes a position encoding layer, a fully convolutional sub-network, and an attention sub-network. The text encoding network based on the preset multi-label text classification model encodes the target description text to obtain the target text semantic feature of the target description text, including:

[0014] Perform sentence splitting on the target description text to obtain a plurality of target description sentences;

[0015] For each target description sentence, perform position encoding on the target description sentence based on the position encoding layer to obtain a position encoding vector of the target description sentence, where the position encoding vector is used to indicate the positions of each word in the target description sentence;

[0016] Based on the position encoding vector of each target description sentence, perform feature extraction on the plurality of target description sentences through the fully convolutional sub-network to obtain target description embedding features;

[0017] Perform context feature extraction on the target description embedding features based on the attention sub-network to obtain the target text semantic feature.

[0018] In some embodiments, the hierarchical label processing sub-network includes an embedding layer and a self-attention layer. The feature extraction of the original hierarchical label based on the hierarchical label processing sub-network to obtain a hierarchical label feature includes:

[0019] Perform vectorization processing on the original hierarchical label based on the embedding layer to obtain a hierarchical label embedding vector;

[0020] Perform self-attention calculation on the hierarchical label embedding vector based on the self-attention layer to obtain the hierarchical label feature.

[0021] In some embodiments, the classification sub-network includes an attention layer, a feed-forward layer, a fully connected layer, and an activation function. The classification processing of the target text semantic feature and the hierarchical label feature based on the classification sub-network to obtain the predicted hierarchical label of the target description text includes:

[0022] Perform linear projection on the target text semantic feature to obtain a key vector and a value vector, and perform linear projection on the hierarchical label feature to obtain a query vector;

[0023] Performing attention calculation on the key vector, the value vector, and the query vector based on the attention layer to obtain a fused text feature;

[0024] Performing feature dimensionality transformation on the fused text feature based on the feed-forward layer to obtain a dimensionality-transformed text feature;

[0025] Performing a non-linear transformation on the dimensionality-transformed text feature based on the fully-connected layer to obtain a text label feature;

[0026] Performing label prediction on the text label feature based on the activation function and the hierarchical label tree to obtain the predicted hierarchical label.

[0027] In some embodiments, the performing label prediction on the text label feature based on the activation function and the hierarchical label tree to obtain the predicted hierarchical label includes:

[0028] Performing label classification scoring on the text label feature based on the activation function and the hierarchical label tree to obtain classification scoring data of the target description text belonging to each of the hierarchical labels;

[0029] Determining the predicted hierarchical label from among the multiple hierarchical labels based on the classification scoring data.

[0030] To achieve the above object, a second aspect of the embodiments of the present application proposes a method for training a multi-label text classification model, the method including:

[0031] Obtaining a sample description text and a sample hierarchical label of the sample description text;

[0032] Performing encoding processing on the sample description text based on a text encoding network of a preset multi-label text classification model to obtain a sample text semantic feature of the sample description text, wherein the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network;

[0033] Performing classification processing on the sample text semantic feature and a root label feature of a preset hierarchical label tree based on the classification sub-network to obtain a first sample hierarchical label of the sample description text, the hierarchical label tree including multiple hierarchical labels;

[0034] Extraction step: Performing feature extraction on the first sample hierarchical label based on the hierarchical label processing sub-network to obtain a sample hierarchical label feature;

[0035] Classification step: Performing classification processing on the sample text semantic feature and the sample hierarchical label feature based on the classification sub-network to obtain a second sample hierarchical label of the sample description text;

[0036] Replace the first sample hierarchical label with the second sample hierarchical label, and repeatedly execute the extraction step to the classification step until the label depth of the second sample hierarchical label reaches the preset label depth, and determine a first sample path from the second sample hierarchical label to the root label feature in the hierarchical label tree;

[0037] Based on the hierarchical labels on the first sample path, determine the sample prediction hierarchical label of the sample description text;

[0038] Based on the comparison between the sample prediction hierarchical label and the sample hierarchical label, adjust the model parameters of the multi-label text classification model to train the multi-label text classification model.

[0039] In some embodiments, the adjusting the model parameters of the multi-label text classification model based on the comparison between the sample prediction hierarchical label and the sample hierarchical label includes:

[0040] Calculate a loss based on the sample prediction hierarchical label and the sample hierarchical label to obtain a target loss function;

[0041] Adjust the model parameters of the multi-label text classification model based on the target loss function.

[0042] To achieve the above object, a third aspect of the embodiments of the present application proposes a multi-label based text classification device, the device includes:

[0043] A text acquisition module, configured to acquire a target description text;

[0044] An encoding module, configured to perform encoding processing on the target description text based on a text encoding network of a preset multi-label text classification model to obtain target text semantic features of the target description text, wherein the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network;

[0045] A first classification module, configured to perform classification processing on the target text semantic features and root label features of a preset hierarchical label tree based on the classification sub-network to obtain an original hierarchical label of the target description text, the hierarchical label tree including multiple hierarchical labels;

[0046] A feature extraction module, configured to perform feature extraction on the original hierarchical label based on the hierarchical label processing sub-network to obtain hierarchical label features;

[0047] A second classification module, configured to perform classification processing on the target text semantic features and the hierarchical label features based on the classification sub-network to obtain a predicted hierarchical label of the target description text;

[0048] A first determination module, configured to determine a first path from the predicted hierarchical label to the root label feature in the hierarchical label tree if it is determined that the label depth of the predicted hierarchical label reaches a preset label depth.

[0049] A second determination module, configured to determine a target hierarchical label of the target description text based on the hierarchical labels on the first path.

[0050] To achieve the above object, a fourth aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect or the method described in the second aspect above is implemented.

[0051] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect or the method described in the second aspect above is implemented.

[0052] The multi-label based text classification method, the training method of the multi-label text classification model, the multi-label based text classification device, the electronic device and the storage medium proposed in the present application obtain a target description text; encode the target description text through a text encoding network of a preset multi-label text classification model to obtain a target text semantic feature of the target description text, which can improve the feature quality of the target text semantic feature. Further, based on a classification sub-network, classify the target text semantic feature and the root label feature of a preset hierarchical label tree to obtain an original hierarchical label of the target description text, and the label information of the root label of the hierarchical label tree can be used to assist classification, improving the determination accuracy of the original hierarchical label. Further, based on a hierarchical label processing sub-network, extract features from the original hierarchical label to obtain a hierarchical label feature; based on the classification sub-network, classify the target text semantic feature and the hierarchical label feature to obtain a predicted hierarchical label of the target description text, and the label information of the original hierarchical label of the hierarchical label tree can be used to assist classification, improving the determination accuracy of the predicted hierarchical label. Further, if it is determined that the label depth of the predicted hierarchical label reaches the preset label depth, determine a first path from the predicted hierarchical label to the root label feature in the hierarchical label tree, and the timing of determining the first path can be accurately obtained according to the relationship between the label depth and the preset label depth, improving the accuracy and rationality of path determination. Finally, based on the hierarchical labels on the first path, determine the target hierarchical label of the target description text, which can improve the multi-label classification accuracy of the text. Description of the Drawings

[0053] Figure 1It is a flowchart of a multi-label based text classification method provided by an embodiment of the present application;

[0054] Figure 2 It is Figure 1 a flowchart of step S102 in

[0055] Figure 3 It is Figure 1 a flowchart of step S104 in

[0056] Figure 4 It is Figure 1 a flowchart of step S105 in

[0057] Figure 5 It is Figure 4 a flowchart of step S405 in

[0058] Figure 6A - Figure 6B It is a schematic diagram of the implementation process of a multi-label based text classification method provided by an embodiment of the present application;

[0059] Figure 7 It is a flowchart of a training method for a multi-label text classification model provided by an embodiment of the present application;

[0060] Figure 8 It is Figure 7 a flowchart of step S708 in

[0061] Figure 9 It is a schematic diagram of the structure of a multi-label based text classification device provided by an embodiment of the present application;

[0062] Figure 10 It is a schematic diagram of the structure of a training device for a multi-label text classification model provided by an embodiment of the present application;

[0063] Figure 11 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0065] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart in the flowchart. Terms such as "first" and "second" in the description and claims and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0067] First, some nouns involved in this application are analyzed as follows:

[0068] Artificial intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0069] Natural language processing (NLP): NLP uses a computer to process, understand, and apply human languages (such as Chinese, English, etc.). NLP belongs to a branch of artificial intelligence and is an interdisciplinary subject of computer science and linguistics, and is often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is often used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intention recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing, etc.

[0070] With the rapid development of artificial intelligence technology, in data mining, the text classification problem is one of the core problems in various business fields. For example, in financial scenarios such as item recommendation, financial business handling, and logistics business consultation, it is often necessary to assign multiple labels to various types of descriptive texts to refine the categories to which the descriptive texts belong. For example, for a logistics task to be processed, it is often necessary to accurately determine the department and personnel to which the logistics task belongs based on the task description text of the logistics task, so as to improve the processing efficiency of the task.

[0071] However, when the number of labels is large, common classification methods such as ten-class classification and twenty-class classification often cannot well distinguish the relationships between label categories. There is a problem of unclear fine-grained semantic boundaries of labels, which will lead to low accuracy of multi-label classification of text.

[0072] Based on this, the embodiments of the present application provide a multi-label based text classification method, a training method for a multi-label text classification model, a multi-label based text classification device, an electronic device and a storage medium, aiming to improve the accuracy of multi-label classification of text.

[0073] The multi-label based text classification method, the training method for the multi-label text classification model, the multi-label based text classification device, the electronic device and the storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the multi-label based text classification method in the embodiments of the present application is described.

[0074] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0075] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0076] The multi-label based text classification method and the training method for the multi-label text classification model provided by the embodiments of the present application relate to the field of fintech technology. The multi-label based text classification method and the training method for the multi-label text classification model provided by the embodiments of the present application can be applied to terminals, can also be applied to the server side, or can be software running on the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the multi-label based text classification method and the training method for the multi-label text classification model, etc., but is not limited to the above forms.

[0077] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0078] It should be noted that in each specific implementation manner of this application, when it comes to performing relevant processing based on data related to the identity or characteristics of an object, such as object information, object behavior data, object historical data, and object location information, the permission or consent of the object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the personal information of an object, the object's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the object's separate permission or separate consent, the necessary object-related data for the normal operation of the embodiments of this application will be obtained.

[0079] Figure 1 is an optional flowchart of the speech recognition method provided by the embodiments of this application, Figure 1 The method in may include but is not limited to steps S101 to S107.

[0080] Step S101, obtain the target description text;

[0081] Step S102, perform encoding processing on the target description text based on the text encoding network of a preset multi-label text classification model to obtain the target text semantic features of the target description text;

[0082] Step S103, perform classification processing on the target text semantic features and the root label features of a preset hierarchical label tree to obtain the original hierarchical labels of the target description text;

[0083] Step S104, perform feature extraction on the original hierarchical labels based on the hierarchical label processing sub-network to obtain hierarchical label features;

[0084] Step S105: Classify the target text semantic features and hierarchical label features based on the classification sub-network to obtain the predicted hierarchical label of the target description text.

[0085] Step S106: If it is determined that the label depth of the predicted hierarchical label reaches the preset label depth, determine the first path from the predicted hierarchical label to the root label feature in the hierarchical label tree.

[0086] Step S107: Determine the target hierarchical label of the target description text based on the hierarchical labels on the first path.

[0087] Steps S101 to S107 illustrated in the embodiments of the present application can improve the feature quality of the target text semantic features by obtaining the target description text and encoding the target description text based on the text encoding network of the preset multi-label text classification model to obtain the target text semantic features of the target description text. Further, by classifying the target text semantic features and the root label features of the preset hierarchical label tree based on the classification sub-network to obtain the original hierarchical label of the target description text, the label information of the root label of the hierarchical label tree can be utilized to assist classification and improve the determination accuracy of the original hierarchical label. Further, by extracting features from the original hierarchical label based on the hierarchical label processing sub-network to obtain hierarchical label features, and classifying the target text semantic features and the hierarchical label features based on the classification sub-network to obtain the predicted hierarchical label of the target description text, the label information of the original hierarchical label of the hierarchical label tree can be utilized to assist classification and improve the determination accuracy of the predicted hierarchical label. Further, if it is determined that the label depth of the predicted hierarchical label reaches the preset label depth, determining the first path from the predicted hierarchical label to the root label feature in the hierarchical label tree can accurately obtain the timing for determining the first path according to the relationship between the label depth and the preset label depth, improving the accuracy and rationality of path determination. Finally, determining the target hierarchical label of the target description text based on the hierarchical labels on the first path can improve the multi-label classification accuracy of the text.

[0088] In step S101 of some embodiments, the target description text is obtained.

[0089] The target description text refers to the text data for describing an object in each business field. For example, in the insurance field, the target description text refers to the product description text of a certain insurance product, and based on the product description text, the hierarchical label to which the insurance product belongs can be determined. For another example, in the enterprise management field, the target description text refers to the training empowerment text for enterprise employees, which is used to provide guidance for the regular training and new employee training of enterprise employees. For still another example, in the logistics field, the target description text refers to the text describing the transportation and transfer operation process of a certain type of item.

[0090] In the specific implementation of this embodiment, the target description text is uploaded to the server by the relevant object. Based on this, the server will receive the target description text.

[0091] In another embodiment, the target description text is stored in the cloud or the database by the relevant object. When it is necessary to determine the hierarchical label of the target description text, under the condition of authorized permission, the target description text is obtained by downloading from the cloud or by making a data call from the database.

[0092] In the embodiment of the present application, the multi-label text classification method depends on the multi-label text classification model for implementation. Specifically, the multi-label text classification model in the embodiment of the present application is based on the encoder-decoder network structure of the transformer. The multi-label text classification model includes a text encoding network, a hierarchical label processing sub-network, and a classification sub-network. Among them, the text encoding network is used to vectorize the description text to extract the deep semantic information in the description text; the hierarchical label processing sub-network is used to extract the features of a certain hierarchical label obtained in a certain recursion round; the classification sub-network is used to perform text classification according to the deep semantic information output by the text encoding network and the label features of the hierarchical label obtained in the previous recursion round, and output the hierarchical label of the current recursion round when the hierarchical label of the recursion round is the next-level label of the hierarchical label obtained in the previous recursion round. For the preset hierarchical label tree, the hierarchical label sub-network and the classification sub-network are used to recursively identify from the first layer to the Nth layer in sequence, and the output of the previous layer is used as the recognition assistance for the next layer, so as to obtain the hierarchical labels of the description text at multiple levels.

[0093] To save space, the training process of the multi-label text classification model in the embodiment of the present application will be described in detail below and will not be elaborated here.

[0094] Please refer to Figure 2 , in some embodiments, the text encoding network includes a position encoding layer, a fully convolutional sub-network, and an attention sub-network. Step S102 may include but is not limited to steps S201 to S204:

[0095] Step S201, perform clause splitting on the target description text to obtain a plurality of target description sentences;

[0096] Step S202, for each target description sentence, perform position encoding on the target description sentence based on the position encoding layer to obtain the position encoding vector of the target description sentence;

[0097] Step S203, based on the position encoding vector of each target description sentence, perform feature extraction on the plurality of target description sentences through the fully convolutional sub-network to obtain the target description embedding feature;

[0098] Step S204: Extract context features from the target description embedding features based on the attention sub-network to obtain the target text semantic features.

[0099] The following provides a detailed description of steps S201 to S204.

[0100] In step S201 of some embodiments, the position encoding vector is used to indicate the positions of the respective words in the target description sentence. Specifically, when performing sentence splitting on the target description text, the target description text can be segmented into multiple complete sentences according to predetermined punctuation marks, thereby obtaining multiple description sentences. Among them, the predetermined punctuation marks may include, but are not limited to, full stops, exclamation marks, question marks, and so on.

[0101] In step S202 of some embodiments, for each target description sentence, perform absolute position encoding or relative position encoding on the target description sentence based on the position encoding layer to obtain a position encoding vector that can indicate the position information of the respective words in the target description sentence.

[0102] In step S203 of some embodiments, first, based on the position encoding vector of each target description sentence, determine the position information of each description sentence in the entire target description text in sequence. Then, through the fully convolutional sub-network, in the position order of the target description sentence indicated by the position encoding vector, use the convolutional layer, deconvolutional layer, and pooling layer of the fully convolutional sub-network to extract features from each target description sentence in sequence, first capture the local semantic features of the target description sentence, and then map the captured local semantic features into global semantic features. Further, in the position order of each target description sentence in the target description text, splice the global semantic features of the target description sentences in sequence to obtain the target description embedding features.

[0103] In step S204 of some embodiments, the attention sub-network includes a self-attention mechanism layer, a normalization layer, a feed-forward layer, and a normalization layer connected in sequence. Specifically, first, a linear transformation is performed on the target description embedding feature to obtain a key vector, a value vector, and a query vector corresponding to the target description embedding feature. Then, the self-attention mechanism layer performs self-attention calculation on the key vector, the value vector, and the query vector to obtain a self-attention calculation result. Further, the self-attention calculation result and the target description embedding feature are feature concatenated to obtain a feature concatenation result, and the feature concatenation result is input into the normalization layer for normalization processing to obtain a normalized feature. Further, the normalized feature is input into the feed-forward layer, and the feed-forward layer performs a linear transformation on the normalized feature to obtain a linear transformation result. Then, the linear transformation result and the normalized feature are concatenated, and the concatenation result is input into the normalization layer, and the output of the normalization layer is used as the target text semantic feature. Among them, the target text semantic feature is used to represent the deep semantic feature information of the target description text.

[0104] Through the above steps S201 to S204, it is possible to use the position encoding method to identify the positions of each target description sentence in the target description text in the text, and also use the fully convolutional sub-network and the attention sub-network to vectorize the target description text and extract the deep semantic information of the target description text, improving the accuracy and efficiency of text feature extraction, thereby improving the feature quality of the target text semantic feature.

[0105] In step S103 of some embodiments, based on the classification sub-network, the target text semantic feature and the root label feature of the preset hierarchical label tree are classified to obtain the original hierarchical label of the target description text.

[0106] In the embodiments of the present application, the preset hierarchical label tree is a multi-way tree extending from top to bottom, and the hierarchical label tree includes multiple hierarchical labels. The highest level of the hierarchical label tree is the root label root of the hierarchical label tree, and the root label root has fixed label information. The next level of the root label is the first level of the hierarchical label tree, which includes multiple first-level labels; the next level of the first level of the hierarchical label tree is the second level; in the second level, there are sub-labels belonging to each first-level label, that is, second-level labels. The next level of the second level of the hierarchical label tree is the third level, and in the third level, there are sub-labels belonging to each second-level label, that is, third-level labels; and so on, until all labels are stored in the same multi-way tree according to the hierarchical relationship to form a hierarchical label tree.

[0107] It should be noted that the hierarchical label tree in the embodiments of the present application is customized according to the business scenario and has the characteristics of flexibility and variability.

[0108] For example, in the scenario of training and empowering enterprise employees in the logistics service field, the first-level label includes training and empowerment. At the next level of this first-level label, there are second-level labels such as new employee training, regular training, and cloud guidance. At the next level of the second-level label "new employee training", there are third-level labels such as pre-job training and intensive training. At the next level of the second-level label "regular training", there are third-level labels such as hot topic training and weak point training.

[0109] It should be noted that the root label feature is obtained by extracting the features of the root label of the hierarchical label. Its specific implementation process is similar to the feature extraction process of step S104, and its specific implementation process refers to the following specific description of step S104.

[0110] When this embodiment is specifically implemented, the classification sub-network includes an attention layer, a feed-forward layer, a fully connected layer, and an activation function. Specifically, first, perform a linear projection on the semantic features of the target text to obtain a key vector and a value vector, and perform a linear projection on the root label feature to obtain a query vector; then, based on the attention layer, perform an attention calculation on the key vector, the value vector, and the query vector to obtain a fused text feature. Further, perform a feature dimension transformation on the fused text feature based on the feed-forward layer to obtain a dimension-transformed text feature. Then, perform a non-linear transformation on the dimension-transformed text feature based on the fully connected layer to obtain a text label feature. Finally, based on the activation function and the hierarchical label tree, perform label prediction on the text label feature to obtain the original hierarchical label of the target description text, where the original hierarchical label is the first-level label in the hierarchical label tree.

[0111] It should be noted that the original hierarchical label of the target description text can be understood as the first-level label of the target description text on the first level of the hierarchical label tree.

[0112] Please refer to Figure 3 , in some embodiments, the hierarchical label processing sub-network includes an embedding layer and a self-attention layer, and step S104 may include but is not limited to steps S301 to S302:

[0113] Step S301, perform vectorization processing on the original hierarchical label based on the embedding layer to obtain a hierarchical label embedding vector;

[0114] Step S302, perform self-attention calculation on the hierarchical label embedding vector based on the self-attention layer to obtain a hierarchical label feature.

[0115] The following will describe steps S301 to S302 in detail.

[0116] In step S301 of some embodiments, in the first recurrence round, the original hierarchical label is vectorized through an embedding layer to extract the label feature information of the original hierarchical label, and a hierarchical label embedding vector is obtained.

[0117] In step S302 of some embodiments, when performing self-attention calculation on the hierarchical label embedding vector based on the self-attention layer, first multiply the hierarchical label embedding vector by a preset query matrix parameter to obtain a query matrix; multiply the hierarchical label embedding vector by a preset key matrix parameter to obtain a key matrix; multiply the hierarchical label embedding vector by a preset value matrix parameter to obtain a value matrix. Further, multiply the query matrix by the transposed result of the key matrix to obtain a product result, and divide the product result by the square root of the feature dimension of the key matrix to obtain a division result. Finally, perform attention calculation on the division result using the softmax function to obtain an attention score, and multiply the attention score by the value matrix to obtain hierarchical label features.

[0118] Through the above steps S301 to S302, it is possible to use feature embedding to obtain the label feature information of the hierarchical label in this recurrence round, and for the label feature information, use the self-attention calculation method to achieve self-learning of the vector of the original hierarchical label, so as to obtain the deep feature information of the original hierarchical label and improve the feature depth and feature quality of the hierarchical label features.

[0119] Please refer to Figure 4 , in some embodiments, the classification sub-network includes an attention layer, a feed-forward layer, a fully-connected layer, and an activation function. Step S105 may include but is not limited to steps S401 to S403:

[0120] Step S401, perform linear projection on the target text semantic feature to obtain a key vector and a value vector, and perform linear projection on the hierarchical label feature to obtain a query vector;

[0121] Step S402, perform attention calculation on the key vector, value vector, and query vector based on the attention layer to obtain a fused text feature;

[0122] Step S403, perform feature dimension transformation on the fused text feature based on the feed-forward layer to obtain a dimension-transformed text feature;

[0123] Step S404, perform a non-linear transformation on the dimension-transformed text feature based on the fully-connected layer to obtain a text label feature;

[0124] Step S405, perform label prediction on the text label feature based on the activation function and the hierarchical label tree to obtain a predicted hierarchical label.

[0125] The following will describe steps S401 to S405 in detail.

[0126] In step S401 of some embodiments, first, a query projection matrix, a key projection matrix, and a value projection matrix are determined. The query projection matrix, the key projection matrix, and the value projection matrix are pre-determined. The query projection matrix, the key projection matrix, and the value projection matrix change with model training. When the model training ends, the query projection matrix, the key projection matrix, and the value projection matrix are fixed. Then, when linearly projecting the semantic features of the target text, the semantic features of the target text are multiplied by the key projection matrix to obtain a key vector; the semantic features of the target text are multiplied by the value projection matrix to obtain a value vector. When linearly projecting the hierarchical label features, the hierarchical label features are multiplied by the query projection matrix to obtain a query vector.

[0127] In step S402 of some embodiments, the process of performing attention calculation on the key vector, the value vector, and the query vector based on the attention layer is similar to the attention calculation process in step S302 above. To save space, it will not be elaborated here.

[0128] In step S403 of some embodiments, when performing feature dimensionality transformation on the fused text features based on the feed-forward layer, through linear transformation, the fused text features are first mapped to a high-dimensional vector space and then mapped to a low-dimensional vector space to extract deeper feature information, obtaining the dimensionality-transformed text features.

[0129] In step S404 of some embodiments, when performing non-linear transformation on the dimensionality-transformed text features based on the fully-connected layer, the dimensionality-transformed text features are first flattened, and the flattened dimensionality-transformed text features are non-linearly transformed using the activation function in the fully-connected layer to obtain text label features.

[0130] In step S405 of some embodiments, first, based on the activation function and the hierarchical label tree, label classification scores are calculated for the text label features to obtain classification score data for the target description text belonging to each hierarchical label. Then, based on the classification score data, a predicted hierarchical label is determined among multiple hierarchical labels. The predicted hierarchical label is often the hierarchical label with relatively larger classification score data among multiple hierarchical labels.

[0131] Through the above steps S401 to S405, various feature processing methods such as attention calculation, feature dimensionality transformation, and non-linear transformation can be used to process the semantic features of the target text and the hierarchical label features, improving the feature quality of the text label features for text classification. Further, using the activation function and the hierarchical label tree to perform label prediction on the text label features can improve the prediction efficiency and accuracy of the hierarchical labels.

[0132] Please refer to Figure 5 , in some embodiments, step S405 may include but is not limited to steps S501 to S502:

[0133] Step S501: Based on the activation function and the hierarchical label tree, perform label classification scoring on the text label features to obtain classification scoring data for the target description text belonging to each hierarchical label.

[0134] Step S502: Based on the classification scoring data, determine the predicted hierarchical label among multiple hierarchical labels.

[0135] The following provides a detailed description of Steps S501 to S502.

[0136] In Step S501 of some embodiments, the activation function is the sigmoid function. Specifically, input the text label features into the sigmoid function, and the sigmoid function outputs the probability of the target description text belonging to each hierarchical label. Use the probability output by the sigmoid function as the classification scoring data for the target description text belonging to each hierarchical label.

[0137] In Step S502 of some embodiments, when determining the predicted hierarchical label among multiple hierarchical labels, first, according to the hierarchical relationship between the labels of the hierarchical label tree, screen out the hierarchical labels that belong to the next level of the original hierarchical label from multiple hierarchical labels; then, among the screened hierarchical labels that belong to the next level of the original hierarchical label, select the hierarchical label with the largest classification scoring data as the predicted hierarchical label. At this time, the predicted hierarchical label is the second-level label of the hierarchical label tree.

[0138] Through the above Steps S501 to S502, label classification scoring can be performed based on the sigmoid function, improving the acquisition efficiency and accuracy of the classification scoring data. Further, based on the classification scoring data, among the hierarchical labels that belong to the next level of the original hierarchical label, select the one with the largest classification scoring data as the predicted hierarchical label, which can follow the subordinate relationship of each hierarchical label, quickly determine the hierarchical label of the current recursion round, and infer the text classification process from the original hierarchical label to the next-level label of the original hierarchical label, realizing the hierarchical inference of text classification and improving the depth of text classification.

[0139] In Step S106 of some embodiments, if it is determined that the label depth of the predicted hierarchical label reaches the preset label depth, determine the first path from the predicted hierarchical label to the root label feature in the hierarchical label tree.

[0140] The label depth of the predicted hierarchical label refers to the level to which the predicted hierarchical label belongs in the hierarchical label tree. The preset label depth refers to the total number of recursion rounds preset according to different business scenarios.

[0141] The first path is the path in the hierarchical label tree connecting the predicted hierarchical label and the root label.

[0142] In the specific implementation of this embodiment, first, query the predicted hierarchical label in the hierarchical label tree, and determine the label depth of the predicted hierarchical label according to the level where the predicted hierarchical label is located. Then, compare the label depth of the predicted hierarchical label with the preset label depth. If the label depth of the predicted hierarchical label is less than the preset label depth, it means that the label depth of the predicted hierarchical label has not reached the preset label depth, and hierarchical recursion needs to continue. The above steps S105 need to be repeated. When repeating the above steps S105, replace the original hierarchical label with the predicted hierarchical label of the current recursion round, repeat steps S104 to S105, and execute a new recursion round. If the label depth of the predicted hierarchical label obtained in the new recursion round is still less than the preset label depth, then continue to replace the predicted hierarchical label of the previous recursion round with the new predicted hierarchical label, and repeat steps S104 to S105 until the label depth of the predicted hierarchical label in a recursion round is equal to the preset label depth. At this time, stop the aforementioned recursive loop. Further, according to the predicted hierarchical label obtained in the last recursion round, depict the path from the predicted hierarchical label obtained in the last recursion round to the root label root in the hierarchical label tree, and use the depicted path as the first path.

[0143] It should be noted that for another case, when the label depth of the predicted hierarchical label is less than the preset label depth, but according to the hierarchical label tree, there are no hierarchical labels at the next level of the predicted hierarchical label, then take the current recursion round as the last round, no longer continue the next round of recursion, and use the current predicted hierarchical label as the predicted hierarchical label for determining the first path.

[0144] In step S107 of some embodiments, based on the hierarchical labels on the first path, determine the target hierarchical label of the target description text.

[0145] In the specific implementation of this embodiment, first, determine all the hierarchical labels passed by the first path; then, among all the passed hierarchical labels, use the hierarchical labels other than the root label as the target hierarchical label of the target description text.

[0146] Such as Figure 6AAs shown, it is a simple schematic diagram of the hierarchical label tree of the embodiment of the present application. Specifically, the highest level of the hierarchical label tree is the root label root; at the next level of the root label, there are hierarchical labels 1 and 2 belonging to the first level. Further, at the next level of hierarchical label 1, there are hierarchical labels 1.1, 1.2, and 1.3 belonging to the second level. At the next level of hierarchical label 2, there are hierarchical labels 2.1, 2.2, 2.3, and 2.4 belonging to the second level. At the next level of hierarchical label 1.2, there are hierarchical labels 1.2.1 and 1.2.2 belonging to the third level. At the next level of hierarchical label 2.1, there is hierarchical label 2.1.1 belonging to the third level. At the next level of hierarchical label 2.2, there is hierarchical label 2.2.1 belonging to the third level. At the next level of hierarchical label 2.3, there is hierarchical label 2.3.1 belonging to the third level. At the next level of hierarchical label 2.4, there is hierarchical label 2.4.1 belonging to the third level. At the next level of hierarchical label 1.2.1, there is hierarchical label 1.2.1.1 belonging to the fourth level.

[0147] As Figure 6BAs shown, it is the specific implementation process of text classification using a multi-label text classification model. It is set that the preset label depth is 4. Specifically, first, the target description text is input into the text encoding network of the multi-label text classification model, and the input target description text is processed successively through the position encoding layer, fully convolutional sub-network, and attention sub-network of the text encoding network, and the target text semantic features of the target description text are output by the last normalization layer of the attention sub-network. The specific implementation process is similar to the above steps S201 to S204. Further, the root label of the hierarchical label tree is input into the hierarchical label processing sub-network of the multi-label text classification model, and the root label is processed by the embedding layer and self-attention layer of the hierarchical label processing sub-network to obtain the root label features. Then, in the first recursion round, through the attention layer, feed-forward layer, fully connected layer, and activation function of the classification sub-network, the root label features and the target text semantic features are classified to obtain the original hierarchical label of the target description text. At this time, the original hierarchical label is the first-level label, the original hierarchical label is the hierarchical label 1 in the hierarchical label tree, and the label depth of the original hierarchical label is 1. The specific implementation process is similar to the above step S103. Further, in the second recursion round, based on the hierarchical label processing sub-network, feature extraction is performed on the original hierarchical label (hierarchical label 1) to obtain hierarchical label features, and based on the classification sub-network, classification processing is performed on the target text semantic features and the hierarchical label features to obtain the predicted hierarchical label of the target description text. At this time, the predicted hierarchical label is the second-level label, at this time the predicted hierarchical label is the hierarchical label 1.2 in the hierarchical label tree, and at this time the label depth of the predicted hierarchical label is 2. The specific implementation process is similar to the above steps S104 - S105. Since the label depth of the predicted hierarchical label at this time is 2, and 2 is less than 4, so the label depth of the predicted hierarchical label has not reached the preset label depth. At this time, the original hierarchical label is replaced with the hierarchical label 1.2, so that the original hierarchical label changes from the hierarchical label 1 to the hierarchical label 1.2. In the third recursion round, based on the hierarchical label processing sub-network, feature extraction is performed on the new original hierarchical label (hierarchical label 1.2) to obtain hierarchical label features, and based on the classification sub-network, classification processing is performed on the target text semantic features and the hierarchical label features to obtain the predicted hierarchical label of the target description text. At this time, the predicted hierarchical label is the third-level label, at this time the predicted hierarchical label is the hierarchical label 1.2.1 in the hierarchical label tree, and at this time the label depth of the predicted hierarchical label is 3. The specific implementation process is similar to the above steps S104 - S105. Since the label depth of the predicted hierarchical label at this time is 3, and 3 is less than 4, so the label depth of the predicted hierarchical label has not reached the preset label depth. At this time, the original hierarchical label of the previous recursion round is replaced with the hierarchical label 1.2.1, so that the original hierarchical label changes from the hierarchical label 1.2 to the hierarchical label 1.2.1.In the fourth recursive round, based on the hierarchical label processing sub-network, feature extraction is performed on the new original hierarchical label (hierarchical label 1.2.1) to obtain hierarchical label features. And based on the classification sub-network, classification processing is performed on the target text semantic features and the hierarchical label features to obtain the predicted hierarchical label of the target description text. At this time, the predicted hierarchical label is the fourth-level hierarchical label, which is hierarchical label 1.2.1.1 in the hierarchical label tree. At this time, the label depth of the predicted hierarchical label is 4. The specific implementation process is similar to the above steps S104 - S105. Since the label depth of the predicted hierarchical label at this time is 4, which is equal to the preset label depth, the recursive loop is ended, and hierarchical label 1.2.1.1 is used as the final predicted hierarchical label, and the first path from hierarchical label 1.2.1.1 to the root label is determined in the hierarchical label tree. Among them, the first path is [hierarchical label 1.2.1.1 → hierarchical label 1.2.1 → hierarchical label 1.2 → hierarchical label 1 → root label]. Therefore, the target hierarchical label of the target description text is determined as [hierarchical label 1, hierarchical label 1.2, hierarchical label 1.2.1, hierarchical label 1.2.1.1].

[0148] Next, a detailed description will be given of the training method for the multi-label text classification model in the embodiments of the present application.

[0149] Figure 7 is an optional flowchart of the training method for the multi-label text classification model provided by the embodiments of the present application, Figure 7 The method in may include but is not limited to steps S701 to S708:

[0150] Step S701, obtain a sample description text and the sample hierarchical label of the sample description text;

[0151] Step S702, based on the text encoding network of the preset multi-label text classification model, perform encoding processing on the sample description text to obtain the sample text semantic features of the sample description text. Among them, the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network;

[0152] Step S703, based on the classification sub-network, perform classification processing on the sample text semantic features and the root label features of the preset hierarchical label tree to obtain the first sample hierarchical label of the sample description text. The hierarchical label tree contains multiple hierarchical labels;

[0153] Step S704, extraction step: based on the hierarchical label processing sub-network, perform feature extraction on the first sample hierarchical label to obtain the sample hierarchical label features;

[0154] Step S705, Classification Step: Classify the semantic features of the sample text and the sample hierarchical label features based on the classification sub-network to obtain the second sample hierarchical label of the sample description text;

[0155] Step S706, Replace the first sample hierarchical label with the second sample hierarchical label, and repeat the extraction step to the classification step until the label depth of the second sample hierarchical label reaches the preset label depth, and determine the first sample path from the second sample hierarchical label to the root label feature in the hierarchical label tree;

[0156] Step S707, Determine the sample prediction hierarchical label of the sample description text based on the hierarchical labels on the first sample path;

[0157] Step S708, Based on the comparison between the sample prediction hierarchical label and the sample hierarchical label, adjust the model parameters of the multi-label text classification model to train the multi-label text classification model.

[0158] The following provides a detailed description of steps S701 to S708.

[0159] In step S701 of some embodiments, the sample description text is stored in the cloud or database by the relevant object. When it is necessary to train the model using the sample description text, with authorized permission, the sample description text and its sample hierarchical label are obtained by downloading from the cloud or by data calling from the database. Among them, the sample hierarchical label can be determined by manual annotation or machine annotation, etc. For example, when the sample hierarchical label of the sample description text is new employee training in training empowerment, the hierarchical label probabilities corresponding to training empowerment and new employee training are both 1, and the hierarchical label probabilities corresponding to other hierarchical labels are all 0.

[0160] In some embodiments, the specific implementation processes of steps S702 to S705 are similar to the specific implementation processes of the above steps S102 - S105. For the sake of brevity, they will not be elaborated here.

[0161] In step S706 of some embodiments, the specific implementation process of step S706 is similar to the specific implementation process of the above step S106.

[0162] It should be noted that, since in the first recursive round, the hierarchical label relied on is the root label of the hierarchical label tree, in the first recursive round, the obtained first sample hierarchical labels are the respective first-level labels of the hierarchical label tree and the probabilities of the respective first-level labels. In the second recursive round, the hierarchical labels relied on are the respective first-level labels of the hierarchical label tree. Therefore, in the second recursive round, the obtained second sample hierarchical labels are the second-level labels subordinate to the first-level labels obtained based on different first-level labels, and the classification scoring data of the second sample hierarchical labels are the probabilities of the respective second-level labels. Further, in the third recursive round, the new second sample hierarchical labels obtained are the third-level labels subordinate to the second-level labels obtained based on different second-level labels, and the classification scoring data of the second sample hierarchical labels are the probabilities of the respective third-level labels. By analogy, in the Nth recursive round, the obtained second sample hierarchical labels are the Nth-level labels subordinate to the (N - 1)th-level labels obtained based on different (N - 1)th-level labels, and the classification scoring data of the second sample hierarchical labels are the probabilities of the respective Nth-level labels.

[0163] Based on this, in the embodiment of the present application, the probability that the ith sample description text belongs to the jth hierarchical label in the hierarchical label tree can be expressed as shown in formula (1):

[0164] p ij = sigmoid(c) Formula (1)

[0165] where p ij is the probability that the ith sample description text belongs to the jth hierarchical label in the hierarchical label tree, and the value range of p ij is [0, 1]; c is the feature output by the fully connected layer of the classification sub-network for the ith sample description text in an iterative training round.

[0166] In some embodiments, the specific implementation process of step S707 is similar to the specific implementation process of the above-mentioned step S107. For the sake of brevity, it will not be elaborated here.

[0167] In step S708 of some embodiments, first, loss calculation is performed based on the sample predicted hierarchical label and the sample hierarchical label to obtain the target loss function; then, the model parameters of the multi-label text classification model are adjusted based on the target loss function, and its specific implementation process will be described in detail below.

[0168] The training method of the multi-label text classification model according to the embodiments of the present application designs the multi-label text classification model into an encoder-decoder network structure based on a transformer, and uses supervised learning to combine the idea of hierarchical recursion to implement model training, which can effectively improve the semantic understanding ability of the model. By fusing hierarchical label information and introducing auxiliary information, the model is more accurate in label classification, solves the problem of unclear semantic boundaries of fine-grained labels, and improves the text classification accuracy of the multi-label text classification model.

[0169] Please refer to Figure 8 , in some embodiments, step S708 may include but is not limited to steps S801 to S802:

[0170] Step S801, calculate the loss based on the sample predicted hierarchical label and the sample hierarchical label to obtain the target loss function;

[0171] Step S802, adjust the model parameters of the multi-label text classification model based on the target loss function.

[0172] The following will describe steps S801 to S802 in detail.

[0173] In step S701 of some embodiments, in each iteration training round, the specific process of calculating the loss based on the sample predicted hierarchical label and the sample hierarchical label can be expressed as shown in formula (2):

[0174]

[0175] Where L is the target loss function; m is the total number of sample description texts, and i refers to the i-th sample description text. N refers to the total number of hierarchical labels in the hierarchical label tree, and j refers to the j-th hierarchical label. y ij is the true probability that the i-th sample description text belongs to the j-th hierarchical label, and the value of y ij is 0 or 1. When the sample hierarchical label indicates that the i-th sample description text truly belongs to the j-th hierarchical label, y ij takes 1, otherwise, y ij takes 0. p ij is the predicted probability that the i-th sample description text belongs to the j-th hierarchical label, which is obtained by formula (1).

[0176] In step S702 of some embodiments, when adjusting the model parameters of the multi-label text classification model based on the target loss function, in each iteration training round, compare the output value of the target loss function with a preset loss threshold. If the output value of the target loss function is greater than the loss threshold, adjust the model parameters of the multi-label text classification model, and perform a new round of iterative training according to the above steps S702 - S707. Repeat this process until in a certain iteration training round, the output value of the target loss function is less than or equal to the preset loss threshold, stop the iterative training, and use the model parameters of this iteration training round as the final model parameters of the multi-label text classification model to complete the training of the multi-label text classification model.

[0177] Through the above steps S801 to S802, the training situation of the model in each iteration training round can be characterized by using the supervised training method and the cross-entropy loss function, and the iterative training of the model can be realized by quantifying the training situation and comparing it with the threshold, so as to standardize the training process of the model, thereby improving the model training efficiency and the model training effect.

[0178] Please refer to Figure 9 , the embodiment of the present application further provides a multi-label based text classification device, which can implement the above multi-label based text classification method. The device includes:

[0179] A text acquisition module 910, configured to acquire a target description text;

[0180] An encoding module 920, configured to perform encoding processing on the target description text based on the text encoding network of the preset multi-label text classification model to obtain the target text semantic features of the target description text, where the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network;

[0181] A first classification module 930, configured to perform classification processing on the target text semantic features and the root label features of the preset hierarchical label tree based on the classification sub-network to obtain the original hierarchical labels of the target description text, and the hierarchical label tree includes multiple hierarchical labels;

[0182] A feature extraction module 940, configured to perform feature extraction on the original hierarchical labels based on the hierarchical label processing sub-network to obtain hierarchical label features;

[0183] A second classification module 950, configured to perform classification processing on the target text semantic features and the hierarchical label features based on the classification sub-network to obtain the predicted hierarchical labels of the target description text;

[0184] A first determination module 960, configured to determine a first path from the predicted hierarchical label to the root label features in the hierarchical label tree if it is determined that the label depth of the predicted hierarchical label reaches the preset label depth;

[0185] A second determination module 970, configured to determine a target level label of a target description text based on level labels on a first path.

[0186] The specific implementation manner of the multi-label based text classification apparatus is substantially the same as the specific embodiments of the above multi-label based text classification method, and will not be elaborated herein.

[0187] Please refer to Figure 10 , an embodiment of the present application further provides a training apparatus for a multi-label text classification model, which can implement the above multi-label text classification model training method. The apparatus includes:

[0188] A sample acquisition module 1010, configured to acquire a sample description text and a sample level label of the sample description text;

[0189] A text encoding module 1020, configured to perform encoding processing on the sample description text based on a text encoding network of a preset multi-label text classification model to obtain a sample text semantic feature of the sample description text, where the multi-label text classification model further includes a level label processing sub-network and a classification sub-network;

[0190] A first feature classification module 1030, configured to perform classification processing on the sample text semantic feature and a root label feature of a preset level label tree based on the classification sub-network to obtain a first sample level label of the sample description text, where the level label tree includes multiple level labels;

[0191] An extraction module 1040, configured to perform an extraction step: perform feature extraction on the first sample level label based on the level label processing sub-network to obtain a sample level label feature;

[0192] A second feature classification module 1050, configured to perform a classification step: perform classification processing on the sample text semantic feature and the sample level label feature based on the classification sub-network to obtain a second sample level label of the sample description text;

[0193] A loop module 1060, configured to replace the first sample level label with the second sample level label, and repeatedly execute the extraction step to the classification step until the label depth of the second sample level label reaches a preset label depth, and determine a first sample path from the second sample level label to the root label feature in the level label tree;

[0194] A label determination module 1070, configured to determine a sample prediction level label of the sample description text based on level labels on the first sample path;

[0195] An adjustment module 1080 is configured to adjust the model parameters of the multi-label text classification model based on the comparison between the predicted hierarchical labels of the samples and the hierarchical labels of the samples, so as to train the multi-label text classification model.

[0196] The specific implementation manner of the training device of the multi-label text classification model is basically the same as the specific embodiments of the above-mentioned training method of the multi-label text classification model, and will not be elaborated herein.

[0197] An embodiment of the present application further provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the program is executed by the processor, it realizes the above-mentioned text classification method based on multi-labels and the training method of the multi-label text classification model. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0198] Please refer to Figure 11 , Figure 11 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0199] A processor 1101, which can be implemented by using a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0200] A memory 1102, which can be implemented in the form of a read-only memory (ReadOnly Memory, ROM), a static storage device, a dynamic storage device, or a random access memory (Random Access Memory, RAM), etc. The memory 1102 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1102 and are called by the processor 1101 to execute the text classification method based on multi-labels and the training method of the multi-label text classification model of the embodiments of the present application;

[0201] An input / output interface 1103, which is used to implement information input and output;

[0202] A communication interface 1104, which is used to implement the communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);

[0203] The bus 1105 transmits information between various components of the device (such as the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104);

[0204] Among them, the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 achieve communication connections with each other inside the device through the bus 1105.

[0205] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned multi-label text classification method and the training method of the multi-label text classification model.

[0206] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0207] The text classification method based on multi-labels, the training method of the multi-label text classification model, the text classification device based on multi-labels, the electronic device and the computer-readable storage medium provided by the embodiments of the present application obtain the target description text; encode the target description text through the text encoding network of the preset multi-label text classification model to obtain the target text semantic features of the target description text, which can improve the feature quality of the target text semantic features. Further, classify the target text semantic features and the root label features of the preset hierarchical label tree through the classification sub-network to obtain the original hierarchical labels of the target description text, and the label information of the root label of the hierarchical label tree can be used to assist classification and improve the determination accuracy of the original hierarchical labels. Further, extract the feature of the original hierarchical label through the hierarchical label processing sub-network to obtain the hierarchical label feature; classify the target text semantic features and the hierarchical label feature through the classification sub-network to obtain the predicted hierarchical label of the target description text, and the label information of the original hierarchical label of the hierarchical label tree can be used to assist classification and improve the determination accuracy of the predicted hierarchical label. Further, if it is determined that the label depth of the predicted hierarchical label reaches the preset label depth, determine the first path from the predicted hierarchical label to the root label feature in the hierarchical label tree, and the recursive relationship of the hierarchical label can be realized according to the relationship between the label depth and the preset label depth, and the same process can be executed repeatedly, and the timing of determining the first path can be accurately obtained, improving the accuracy and rationality of path determination. Finally, determine the target hierarchical label of the target description text based on the hierarchical labels on the first path, which can solve the problem of unclear semantic boundaries of label granularity and improve the multi-label classification accuracy of the text.

[0208] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0209] Those skilled in the art can understand that Figure 1 - 8 the technical solutions shown in do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine some steps, or different steps.

[0210] The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0212] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0213] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0214] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0215] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0216] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0217] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store programs.

[0218] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A multi-label based text classification method, characterized in that The method includes: Obtaining a target description text; Encoding the target description text based on a text encoding network of a preset multi-label text classification model to obtain target text semantic features of the target description text, where the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network; Classifying the target text semantic features and root label features of a preset hierarchical label tree based on the classification sub-network to obtain original hierarchical labels of the target description text, where the hierarchical label tree includes multiple hierarchical labels; Extracting features from the original hierarchical labels based on the hierarchical label processing sub-network to obtain hierarchical label features; Classifying the target text semantic features and the hierarchical label features based on the classification sub-network to obtain predicted hierarchical labels of the target description text; If it is determined that the label depth of the predicted hierarchical labels reaches a preset label depth, determining a first path from the predicted hierarchical labels to the root label features in the hierarchical label tree; Determining target hierarchical labels of the target description text based on the hierarchical labels on the first path.

2. The method according to claim 1, wherein The text encoding network includes a position encoding layer, a fully convolutional sub-network, and an attention sub-network. Encoding the target description text based on the text encoding network of the preset multi-label text classification model to obtain the target text semantic features of the target description text includes: Performing sentence splitting on the target description text to obtain multiple target description sentences; For each target description sentence, performing position encoding on the target description sentence based on the position encoding layer to obtain a position encoding vector of the target description sentence, where the position encoding vector is used to indicate the positions of each word in the target description sentence; Extracting features from the multiple target description sentences through the fully convolutional sub-network based on the position encoding vector of each target description sentence to obtain target description embedding features; Performing context feature extraction on the target description embedding features based on the attention sub-network to obtain the target text semantic features.

3. The method according to claim 1, characterized in that, The hierarchical label processing sub-network includes an embedding layer and a self-attention layer. Extracting features from the original hierarchical labels based on the hierarchical label processing sub-network to obtain hierarchical label features includes: Performing vectorization processing on the original hierarchical labels based on the embedding layer to obtain hierarchical label embedding vectors; Performing self-attention calculation on the hierarchical label embedding vectors based on the self-attention layer to obtain the hierarchical label features.

4. The method according to any one of claims 1 to 3, characterized in that The classification sub-network includes an attention layer, a feed-forward layer, a fully connected layer, and an activation function. Classifying the target text semantic features and the hierarchical label features based on the classification sub-network to obtain the predicted hierarchical labels of the target description text includes: Performing linear projection on the target text semantic features to obtain a key vector and a value vector, and performing linear projection on the hierarchical label features to obtain a query vector; Performing attention calculation on the key vector, the value vector, and the query vector based on the attention layer to obtain fused text features; Perform feature dimension transformation on the fused text features based on the feedforward layer to obtain transformed text features; Perform a non-linear transformation on the transformed text features based on the fully connected layer to obtain text label features; Perform label prediction on the text label features based on the activation function and the hierarchical label tree to obtain the predicted hierarchical label; 5. The method according to claim 4, wherein The performing label prediction on the text label features based on the activation function and the hierarchical label tree to obtain the predicted hierarchical label includes: Perform label classification scoring on the text label features based on the activation function and the hierarchical label tree to obtain classification scoring data for which hierarchical label the target description text belongs to each; Based on the classification scoring data, determine the predicted hierarchical label among multiple hierarchical labels; 6. A training method for a multi-label text classification model, characterized in that, The training method includes: Obtain a sample description text and the sample hierarchical label of the sample description text; Perform encoding processing on the sample description text based on the text encoding network of a preset multi-label text classification model to obtain the sample text semantic features of the sample description text, wherein the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network; Perform classification processing on the sample text semantic features and the root label features of a preset hierarchical label tree based on the classification sub-network to obtain the first sample hierarchical label of the sample description text, and the hierarchical label tree includes multiple hierarchical labels; Extraction step: perform feature extraction on the first sample hierarchical label based on the hierarchical label processing sub-network to obtain sample hierarchical label features; Classification step: perform classification processing on the sample text semantic features and the sample hierarchical label features based on the classification sub-network to obtain the second sample hierarchical label of the sample description text; Use the second sample hierarchical label to replace the first sample hierarchical label, and repeat the extraction step to the classification step until the label depth of the second sample hierarchical label reaches a preset label depth, and determine a first sample path from the second sample hierarchical label to the root label features in the hierarchical label tree; Based on the hierarchical labels on the first sample path, determine the sample predicted hierarchical label of the sample description text; Based on the comparison between the sample predicted hierarchical label and the sample hierarchical label, perform model parameter adjustment on the multi-label text classification model to train the multi-label text classification model.

7. The training method according to claim 6, wherein The performing model parameter adjustment on the multi-label text classification model based on the comparison between the sample predicted hierarchical label and the sample hierarchical label includes: Perform loss calculation based on the sample predicted hierarchical label and the sample hierarchical label to obtain a target loss function; Perform model parameter adjustment on the multi-label text classification model based on the target loss function.

8. A text classification device based on multi-labels, characterized in that, The device includes: A text acquisition module for acquiring a target description text; An encoding module, configured to encode the target description text based on the text encoding network of a preset multi-label text classification model to obtain the target text semantic features of the target description text, wherein the multi-label text classification model further includes a hierarchical label processing sub-network and a classification sub-network; A first classification module, configured to perform classification processing on the target text semantic features and the root label features of a preset hierarchical label tree based on the classification sub-network to obtain the original hierarchical labels of the target description text, where the hierarchical label tree includes multiple hierarchical labels; A feature extraction module, configured to extract features from the original hierarchical labels based on the hierarchical label processing sub-network to obtain hierarchical label features; A second classification module, configured to perform classification processing on the target text semantic features and the hierarchical label features based on the classification sub-network to obtain the predicted hierarchical labels of the target description text; A first determination module, configured to determine a first path from the predicted hierarchical labels to the root label features in the hierarchical label tree if it is determined that the label depth of the predicted hierarchical labels reaches a preset label depth; A second determination module, configured to determine the target hierarchical labels of the target description text based on the hierarchical labels on the first path.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the multi-label based text classification method according to any one of claims 1 to 5, or the training method of the multi-label text classification model according to any one of claims 6 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-label based text classification method according to any one of claims 1 to 5, or the training method of the multi-label text classification model according to any one of claims 6 to 7.