Multi-label Text Classification Method and Device
By extracting and encoding contextual words around keywords and using a Transformer-based model, the method enhances the accuracy of multiple label text classification by addressing the lack of contextual understanding in traditional methods.
Patent Information
- Application Number
- CN202210403778.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-18
AI Technical Summary
The multi-label text classification method of traditional machine learning is unable to correlate the context semantics because it is based on only individual keywords, resulting in inaccurate classification results, affecting the effect of multi-label text classification.
By obtaining the labeled dataset, the preset number of contextual words in the statement where the keywords are located are extracted, the keywords and their tags are encoded, and the text classification model is used for multi-label classification. The model includes the input layer, the calculation layer and the output layer. The TransformerEncoder structure and the multi-label classifier are used. The optimizer is Adam Optimizer, and the pre-trained parameters are Roberta models.
The accuracy of multi-label text classification has been improved and the effect of multi-label text classification has been improved. In particular, the classification effect on labels such as ORGANIZATION, GPE, and SUBSTANCE has been significantly improved.
Smart Images

Figure CN114722204B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of text classification, and specifically relates to a multi-label text classification method and device. Background Art
[0002] Text classification is a basic task in natural language processing. Traditional text classification techniques mainly focus on single-label classification. In single-label classification problems, each sample belongs to only one corresponding category, and there are obvious boundaries between each category. However, in some scenarios, for example, in the classification of academic papers, if a paper belongs to both the biological field and the artificial intelligence field at the same time, classifying it into only one category is incomplete, the classification granularity is relatively coarse, and it will also lead to the inability to correctly utilize and classify resource information. Therefore, multiple labels need to be set for classification. What multi-label classification needs to handle is the task where texts in real life have multiple categories. Compared with single-label classification, multi-label text classification is more common and more difficult in real life. The multi-label text classification method of traditional machine learning only extracts features based on individual keywords. Since the context semantics is not associated when extracting features, the classification result is inaccurate, affecting the multi-label text classification effect. Summary of the Invention
[0003] To at least overcome to some extent the problem that the multi-label text classification method of traditional machine learning only extracts features based on individual keywords, and since the context semantics is not associated when extracting features, the classification result is inaccurate, affecting the multi-label text classification effect, this application provides a multi-label text classification method and device.
[0004] In a first aspect, this application provides a multi-label text classification method, including:
[0005] Obtain an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords;
[0006] Extract a preset number of context words corresponding to the sentences where the keywords are located;
[0007] Encode the labels corresponding to the keywords;
[0008] Input the keywords, the preset number of context words corresponding to the sentences where the keywords are located, and the encoded labels corresponding to the keywords into a text classification model, and output a classification result.
[0009] Further, the obtaining of the annotated data set includes:
[0010] Split the original sentence into a list of single words;
[0011] Mark the order of each word in a list of single words;
[0012] Extract keywords from a list of single words and the position indexes of the keywords in the original sentence;
[0013] Annotate at least one classification label for the keywords.
[0014] Furthermore, extracting a preset number of context words corresponding to the sentence where the keyword is located includes:
[0015] Extract a preset number of context words corresponding to the sentence where the keyword is located according to the list of single words, the order of each word, and the position indexes of the keyword and the keyword in the original sentence.
[0016] Furthermore, it also includes:
[0017] Take each keyword as an independent keyword input sequence;
[0018] Starting from the first character in the keyword input sequence, and based on the order of each word and the position indexes of the keyword in the original sentence, extract the keyword left sequence sequentially to the left;
[0019] Starting from the last character in the keyword input sequence, and based on the order of each word and the position indexes of the keyword in the original sentence, extract the keyword right sequence sequentially to the right.
[0020] Furthermore, the keyword left sequence and / or the keyword right sequence includes punctuation marks.
[0021] Furthermore, the annotating at least one classification label for the keyword includes:
[0022] Use an NER program to annotate classification labels for the keywords, and the annotation categories include at least one of PERSON, ORGANIZATION, GPE, EVENT, SUBSTANCE, WORK_OF_ART, and LOCATION.
[0023] Furthermore, the text classification model includes:
[0024] An input layer, a computing layer, and an output layer;
[0025] The input layer is used to convert the keyword, the preset number of context words corresponding to the sentence where the keyword is located, and the label encoding corresponding to the keyword into the input format of the text classification model;
[0026] The computing layer is used to extract the features of the input data of the input layer and calculate the information of the input layer using multiple stacked TransformerEncoder structures;
[0027] The output layer is used to classify the results of the calculation layer through a multi-label classifier to obtain the final result.
[0028] Further, the parameter selection of the text classification model includes:
[0029] The multi-label classifier is multiple sigmoid functions;
[0030] The optimizer is Adam Optimizer, and the optimization parameters are β1 = 0.9 and β2 = 0.98;
[0031] The pre-training parameters of the model are initially trained using the parameters of the Roberta model.
[0032] Further, it further includes:
[0033] Use an evaluation effect model to evaluate the output results of the text classification model;
[0034] The text classification model corresponding to the output result whose evaluation score meets the preset requirements is used as the final text classification model.
[0035] In a second aspect, the present application provides a multi-label text classification device, including:
[0036] An acquisition module, configured to acquire an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords;
[0037] An extraction module, configured to extract a preset number of context words corresponding to the sentences where the keywords are located;
[0038] An encoding module, configured to encode the labels corresponding to the keywords;
[0039] An output module, configured to input the keywords, a preset number of context words corresponding to the sentences where the keywords are located, and the label encoding corresponding to the keywords into a text classification model, and output a classification result.
[0040] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0041] The multi-label text classification method and device provided by the embodiments of the present invention can improve the accuracy and effect of multi-label text classification by obtaining an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords, extracting a corresponding preset number of context words in the sentences where the keywords are located, encoding the labels corresponding to the keywords, and inputting the keywords, the corresponding preset number of context words in the sentences where the keywords are located, and the encoded labels corresponding to the keywords into a text classification model to output a classification result.
[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with this application and used together with the specification to explain the principles of this application.
[0044] Figure 1 It is a flowchart of a multi-label text classification method provided by an embodiment of this application.
[0045] Figure 2 It is a flowchart of another multi-label text classification method provided by an embodiment of this application.
[0046] Figure 3 It is a functional structure diagram of a multi-label text classification device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other implementation manners obtained by those of ordinary skill in the art without creative efforts based on the embodiments in this application belong to the scope protected by this application.
[0048] Figure 1 It is a flowchart of a multi-label text classification method provided by an embodiment of this application. As Figure 1 shown, the multi-label text classification method includes:
[0049] S11: Obtain an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords;
[0050] S12: Extract a corresponding preset number of context words in the sentences where the keywords are located;
[0051] S13: Encode the labels corresponding to the keywords;
[0052] S14: Input the keyword, the corresponding preset number of context words in the sentence where the keyword is located, and the label encoding corresponding to the keyword into the text classification model, and output the classification result.
[0053] Traditional multi-label text classification methods are based on machine learning and extract features for individual keywords. Since the context semantics is not associated when extracting features, the classification result is inaccurate, affecting the multi-label text classification effect.
[0054] In this embodiment, by obtaining an annotated data set, which includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords, extracting the corresponding preset number of context words in the sentences where the keywords are located, encoding the labels corresponding to the keywords, and inputting the keyword, the corresponding preset number of context words in the sentence where the keyword is located, and the label encoding corresponding to the keyword into the text classification model, and outputting the classification result, the accuracy of multi-label text classification can be improved, and the multi-label text classification effect can be enhanced.
[0055] Another multi-label text classification method is provided in an embodiment of the present invention, as Figure 2 shown in the flowchart. This multi-label text classification method includes:
[0056] S201: Split the original sentence into a list of single words;
[0057] S202: Mark the order of each word in the list of single words;
[0058] S203: Extract the keyword and the position index of the keyword in the original sentence from the list of single words;
[0059] S204: Annotate at least one classification label for the keyword;
[0060] In some embodiments, annotating at least one classification label for the keyword includes:
[0061] Use the NER program to annotate classification labels for the keyword, and the annotation categories include at least one of PERSON, ORGANIZATION, GPE, EVENT, SUBSTANCE, WORK_OF_ART, and LOCATION.
[0062] Run the NER program on the BBN dataset to annotate the BBN dataset. The NER program uses the Stanford Named Entity Recognizer framework based on the Java language to annotate the data. There are 7 annotation categories, namely: PERSON, ORGANIZATION, GPE, EVENT, SUBSTANCE, WORK_OF_ART, LOCATION. For a piece of mention data, the range of classification labels is at most 7 classification labels.
[0063] S205: Extract a preset number of context words corresponding to the keyword in the keyword-containing sentence according to the list of single words, the order of each word, the keyword, and the position index of the keyword in the original sentence.
[0064] In some embodiments, it further includes:
[0065] Take each keyword as an independent keyword input sequence;
[0066] Starting from the first character in the keyword input sequence, according to the order of each word and the position index of the keyword in the original sentence, extract the keyword left sequence sequentially to the left;
[0067] Starting from the last character in the keyword input sequence, according to the order of each word and the position index of the keyword in the original sentence, extract the keyword right sequence sequentially to the right.
[0068] Based on the annotated data, perform preprocessing on the annotated dataset, divide the data into training set and test set, and process the data into three parts in the format of left text, mention, and right text, including: randomly divide the annotated dataset, and the division ratio of the training set and the test set is 8:2; for the above data input format, it is divided into three parts: the mention part of the annotated data, 20 words on the left of the mention as the left text, and 20 words on the right of the mention as the right text.
[0069] In some embodiments, the keyword left sequence and / or the keyword right sequence include punctuation marks.
[0070] Not ignoring punctuation marks helps to understand the semantics, thereby further improving the accuracy of identifying classification labels.
[0071] S206: Encode the label corresponding to the keyword;
[0072] Encode the discrete feature label using one - hot encoding, including: for the input data of the same mention, there are multiple label classifications, and hot encoding is used for multiple labels of discrete features as the labels of the data.
[0073] S207: Input the keyword, the corresponding preset number of context words in the sentence where the keyword is located, and the label encoding corresponding to the keyword into the text classification model, and output the classification result.
[0074] In this embodiment, the text classification model includes:
[0075] An input layer, a computing layer, and an output layer;
[0076] The input layer is used to convert the keyword, the corresponding preset number of context words in the sentence where the keyword is located, and the label encoding corresponding to the keyword into the input format of the text classification model;
[0077] The computing layer is used to extract the features of the input data of the input layer and calculate the information of the input layer using multiple stacked Transformer Encoder structures;
[0078] The output layer is used to classify the result of the computing layer through a multi - label classifier to obtain the final result.
[0079] Build an algorithm model based on pytorch, input the training set data for training, adjust the model parameters, and save the training parameters. The input layer consists of three parts: token embedding, segment embedding, and position embedding; the computing layer is a model composed of multiple Transformer Encoders to calculate the information of the input layer; the output layer is used to classify the result of the computing layer through a multi - label classifier to obtain the final result. The value of the first node of the result calculated by the last layer is connected to a fully - connected layer, and then passed through a classifier. The classifier is changed to multiple sigmoid functions, which is equivalent to multiple binary classification tasks; the model optimizer is selected as Adam, and the optimization parameters are β1 = 0.9, β2 = 0.98; the model pre - training parameters are initially trained using the parameters of the Roberta model; save the final model parameters and provide them for testing on the test set.
[0080] S208: Use the evaluation effect model to evaluate the output result of the text classification model;
[0081] S209: Use the text classification model corresponding to the output result whose evaluation score meets the preset requirements as the final text classification model.
[0082] In this embodiment, a multi - label text classification method is provided, specifically including:
[0083] Step 1: Run the NER program on the BBN (Bilateral - Branch Network) dataset to annotate the BBN dataset.
[0084] Specifically, based on the original BBN public dataset, use the Stanford Named Entity Recognizer framework to annotate the original dataset. The original data format is JSON. An example of the annotation result: {"tokens": ["The", "harvest", "arrives", "in", "plenty", "after", "last", "year", "'s", "drought - ravaged", "effort", ":", "The", "government", "estimates", "corn", "output", "at", "7.45", "billion", "bushels", ",", "up", "51", "%", "from", "last", "fall", "."], "senid": 2, "mentions": [{"start": 12, "labels": [" / WORK_OF_ART", " / ORGANIZATION"], "end": 13}], "fileid": "WSJ1825"}, which is stored in the form of a dictionary. The value of "tokens" is a list that splits the original sentence into individual words; the value of "senid" represents the order of the sentence corresponding to the entity after each original sentence is annotated; the value of "mention" represents the mentioned keywords in the original sentence. The classification of the model is centered around "mention". The values of "start" and "end" represent the position indices of "mention" in the original sentence; the value of "labels" represents the label of "mention", which is the classification label of the model. One "mention" corresponds to multiple classification labels "labels"; the value of "fileid" represents the entity corresponding to "mention".
[0085] Step 2: Based on the annotated data, perform pre - processing on the annotated dataset, divide the data into training set and test set, and process the data into three parts: left text, mention, and right text.
[0086] Specifically, preprocess the labeled data in step 1. First, divide the labeled data into a training set and a test set according to the non-repetitive sampling method in a random order, with a ratio of 8:2. Second, process the data set into a preset format. For each mention in each line of the file, cut it into three parts: left text, mention, and right text. Concatenate the three parts in sequence and separate them with [SEP] as the preprocessed data. Explanation of the meaning of the preprocessed data format: Each mention part in the labeled and processed data is used as an independent mention input sequence; starting from the star-1 position of the mention word, take 20 words in sequence as the left text input sequence; starting from the end+1 position of the mention word, take 20 words in sequence as the right text input sequence, without ignoring the positions of punctuation marks. Example of preprocessed data: The authors, from Boston's [SEP] Beth Israel Hospital [SEP], say that 84% of the 50 births they followed occurred after only two in vitro cycles.
[0087] Step 3, encode the discrete feature label using the one-hot encoding method.
[0088] Specifically, perform data processing on the discrete feature label label to make it the output label of the model. The processing process is to convert PERSON into [1,0,0,0,0,0,0], ORGANIZATION into [0,1,0,0,0,0,0], GPE into [0,0,1,0,0,0,0], EVENT into [0,0,0,1,0,0,0], SUBSTANCE into [0,0,0,0,1,0,0], WORK_OF_ART into [0,0,0,0,0,1,0], LOCATION into [0,0,0,0,0,0,1]. When a mention has multiple labels, perform matrix addition on the involved labels to obtain the final multi-label encoding. For example: label [ORGANIZATION,GPE], perform matrix addition on ORGANIZATION [0,1,0,0,0,0,0] and GPE [0,0,1,0,0,0,0] to obtain the final result [0,1,1,0,0,0,0].
[0089] Step 4, build an algorithm model based on pytorch, input the training set data for training, adjust the model parameters, and save the training parameters.
[0090] Build an algorithm model based on Pytorch. The model includes: an input layer (input layer), a computing layer (computing layer), and an output layer (output layer). The input layer is used to convert the required preprocessed training text into the input format of the model. The computing layer is used to extract the features of the input data of the input layer and perform calculations using multiple stacked Transformer Encoder structures. The output layer is used to classify the results of the computing layer through a multi-label classifier to obtain the final result.
[0091] Specifically, the input layer consists of three parts: token embedding, segment embedding, and position embedding. First is token embedding. Use WordPiece tokenization to perform token changes on English words, and send the words after token changes into the token embedding layer to convert each word into a 768-dimensional digital vector. For example, n tokens are converted into a matrix of (n, 768); then is segment embedding. Assume that each input layer is n partial sentences, and each token of the nth sentence is marked as n - 1 as the digital vector of this layer; position embedding, learn the vector representation of each position to include the sequential features of the input sequence. So for a tokenized input sequence of length n, there will be three different representations, namely: token embedding, with shape (1, n, 768), the vector representation of the word; segment embedding, with shape (1, n, 768), which is the vector representation to help BERT distinguish paired input sequences; position embedding, with shape (1, n, 768), to let BERT know that its input has a temporal attribute and well simulates the order of word appearance. Sum these representations tensorially to generate a single representation with shape (1, n, 768). So the input of the input layer is the dataset composed of leftcontext + [seq] + rightcontext designed in this patent, and the output is the result of tensor summation of token embedding, segment embedding, and position embedding.
[0092] Specifically, the computing layer uses multiple stacked Transformer Encoder structures for computing. The computing model selected for this model is a model with 12 layers of this structure stacked. As the key of this structure, Multi-Head Attention defines the input data as X, and calculates Q, K, and V according to formulas (1), (2), and (3). Q, K, and V are the output results of each layer of the Transformer Encoder structure, and W Q W k W v are the weight parameters of each layer of the Transformer Encoder structure, and are substituted into the core calculation formula (4) of Multi-Head Attention as the output.
[0093] Q = X * W Q (1)
[0094] K = X * W k (2)
[0095] V = X * W v (3)
[0096]
[0097] Specifically, the output layer is used to classify the CLS result of the computing layer result through a multi-label classifier to obtain the final result. Specifically implement the multi-label text classifier, connect the output result of the last layer of CLS to the fully connected layer, and change the classifier to multiple sigmoid functions as in formula (5), which is equivalent to multiple binary classification tasks.
[0098]
[0099] Model parameter settings: The model optimizer is selected as Adam Optimizer, and the optimization parameters are β1 = 0.9 and β2 = 0.98; the model pre-training parameters are initially trained based on the Roberta model parameters; the training epoches are set to 100 times; the final model parameters are saved after training and provided for testing on the test set.
[0100] Step 5, use the test set and the training set to evaluate the model effect.
[0101] Specifically, the evaluation effect models are Precision, Recall, Accuracy, and F1-ScoreAccuracy. The Precision accuracy represents the proportion of examples predicted as positive that are actually positive. tp represents the number of positive samples correctly judged, and fp represents the number of negative classes predicted as positive classes, as shown in formula (6). The Recall represents a measure of coverage, measuring how many positive examples are classified as positive. fp represents the number of negative classes predicted as positive classes, and fn represents the number of positive classes predicted as negative classes, as shown in formula 7. Accuracy represents the proportion of correctly classified samples among all samples, as shown in formula 8. F1-Score is the harmonic mean of precision and recall.
[0102]
[0103]
[0104]
[0105]
[0106] The model effect is specifically shown in Table 1 as follows:
[0107] Table 1 Model Effect Evaluation Results
[0108] label precision recall f1-score PERSON 0.87507 0.87989 0.87747 ORGANIZATION 0.91922 0.88076 0.89958 GPE 0.80835 0.78520 0.79661 EVENT 0.53731 0.85714 0.66055 SUBSTANCE 0.89344 0.96035 0.92569 WORK_OF_ART 0.43750 0.50602 0.46927 LOCATION 0.44444 0.85106 0.58394 Accuracy - - 0.86471
[0109] It can be found from the model evaluation effect that the overall accuracy is 86.7%, indicating that the model is overall excellent. The classification effect is the best on the ORGANIZATION label, and the precision, recall, and f1-score are 0.91922, 0.88076, and 0.89958 respectively, all of which are the best scores in all label classifications. The model scoring criteria for LOCATION, ORGANIZATION, GPE, and SUBSTANCE are all higher than 80%.
[0110] The multi-label text classification method provided in this embodiment constructs a dataset in the format of keyword left sequence, keyword sequence, and keyword right sequence using the original data, and pre-trains deep bidirectional representations with a bidirectional encoder to learn the context meaning of the text and identify labels for text keywords. The purpose is to perform more accurate multi-label classification of text data in the existing public text dataset, solving the problem of poor classification effect of the existing multi-label text classification method in the public text dataset.
[0111] Figure 3 It is the functional structure diagram of the multi-label text classification device provided in an embodiment of this application, as Figure 3As shown in the figure, the multi-label text classification device includes:
[0112] An acquisition module 31, configured to acquire an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords;
[0113] An extraction module 32, configured to extract a preset number of context words corresponding to the sentences where the keywords are located;
[0114] An encoding module 33, configured to encode the labels corresponding to the keywords;
[0115] An output module 34, configured to input the keywords, a preset number of context words corresponding to the sentences where the keywords are located, and the label encodings corresponding to the keywords into a text classification model, and output a classification result.
[0116] In this embodiment, the acquisition module acquires an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords; the extraction module extracts a preset number of context words corresponding to the sentences where the keywords are located, and the encoding module encodes the labels corresponding to the keywords; the output module is configured to input the keywords, a preset number of context words corresponding to the sentences where the keywords are located, and the label encodings corresponding to the keywords into a text classification model, and output a classification result, which can improve the accuracy of multi-label text classification and the multi-label text classification effect.
[0117] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be referred to the same or similar content in other embodiments.
[0118] It should be noted that in the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "multiple" refers to at least two.
[0119] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of executable instructions including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a manner that is not shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the involved functions, which should be understood by those skilled in the technical field of the embodiments of the present application.
[0120] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0121] Those of ordinary skill in the art can understand that all or part of the steps carried by the method in the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0122] In addition, in each embodiment of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0123] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0124] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0125] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0126] It should be noted that the present invention is not limited to the above-mentioned optimal implementation manners. Those skilled in the art can obtain various other forms of products under the inspiration of the present invention. However, no matter what changes are made in its shape or structure, as long as it has a technical solution identical or similar to the present application, it falls within the protection scope of the present invention.
Claims
1. A multi-label text classification method, characterized in that, Including: Obtain an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords; Extract a preset number of context words corresponding to the sentences where the keywords are located; Encode the labels corresponding to the keywords; Input the keywords, a preset number of context words corresponding to the sentences where the keywords are located, and the encoded labels corresponding to the keywords into a text classification model, and output a classification result; The obtaining of the annotated data set includes: Split the original sentence into a list of individual words; Annotate the order of each word in the list of individual words; Extract the keywords and the position indexes of the keywords in the original sentence from the list of individual words; Annotate at least one classification label for the keywords; The extracting of a preset number of context words corresponding to the sentences where the keywords are located includes: Extract a preset number of context words corresponding to the sentences where the keywords are located according to the list of individual words, the order of each word, and the position indexes of the keywords and the keywords in the original sentence; Take each keyword as an independent keyword input sequence; Starting from the first character in the keyword input sequence, sequentially extract the keyword left sequence to the left according to the order of each word and the position indexes of the keywords in the original sentence; Starting from the last character in the keyword input sequence, sequentially extract the keyword right sequence to the right according to the order of each word and the position indexes of the keywords in the original sentence; Wherein, the classification result includes the labels corresponding to the keywords; The keyword left sequence and / or the keyword right sequence includes punctuation marks; The annotating of at least one classification label for the keywords includes: Use an NER program to annotate classification labels for the keywords, and the annotation categories include at least one of PERSON, ORGANIZATION, GPE, EVENT, SUBSTANCE, WORK_OF_ART, and LOCATION.
2. A multi-label text classification device, characterized in that Including: An obtaining module, configured to obtain an annotated data set, where the annotated data set includes keywords, the sentences where the keywords are located, and the labels corresponding to the keywords; An extracting module, configured to extract a preset number of context words corresponding to the sentences where the keywords are located; An encoding module, configured to encode the labels corresponding to the keywords; An output module, configured to input the keywords, a preset number of context words corresponding to the sentences where the keywords are located, and the encoded labels corresponding to the keywords into a text classification model, and output a classification result; The obtaining of the annotated data set includes: Split the original sentence into a list of individual words; Annotate the order of each word in the list of individual words; Extract the keywords and the position indexes of the keywords in the original sentence from the list of individual words; Annotate at least one classification label for the keywords; The extracting of a preset number of context words corresponding to the sentences where the keywords are located includes: Extract a preset number of context words corresponding to the sentences where the keywords are located according to the list of individual words, the order of each word, and the position indexes of the keywords and the keywords in the original sentence; Input each keyword as an independent keyword input sequence; Starting from the first character in the keyword input sequence, extract the left keyword sequences sequentially to the left according to the order of each word and the position index of the keyword in the original sentence: Starting from the last character in the keyword input sequence, extract the right keyword sequences sequentially to the right according to the order of each word and the position index of the keyword in the original sentence; Wherein, the classification result includes the label corresponding to the keyword; The left keyword sequence and / or the right keyword sequence includes punctuation marks; The step of annotating at least one classification label for the keyword includes: Use an NER program to annotate classification labels for the keyword, and the annotation categories include at least one of PERSON, ORGANIZATION, GPE, EVENT, SUBSTANCE, WORK_OF_ART, and LOCATION.
Citation Information
Patent Citations
Text data multi-label classification method and device
CN113297379A