Text Classification Method, Device, Equipment and Storage Medium for Multi-Level Tags

Through the text classification model for multi-level labels, the text is encoded and cosine similarity calculation is calculated automatically, and the text label is solved, which solves the dependence problem of a large number of manual labels in the prior art, and improves the accuracy of label labels and the accuracy of text classification.

CN114691866BActive Publication Date: 2025-05-30AVIATION IND INFORMATION CENT +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210225366.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-05-30
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

The prior art requires a large amount of labeled label data when classifying text multi-labels, and the label system has problems such as high subjectivity, high cost and slow update speed, resulting in low accuracy.

Method used

By obtaining the labels corresponding to the text and its keywords, using a preset text classification model for multi-level labels for encoding, the cosine similarity between the text feature vector and the label vector is calculated, and the label whose cosine similarity is greater than the preset threshold is the label of the text.

Benefits of technology

Reliance on a large number of manual labeling labels is reduced, the maintenance costs of manual labeling and labeling systems are reduced, the accuracy of labeling is improved, and the text classification results are more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691866B_ABST
    Figure CN114691866B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a text classification method, apparatus, device, and storage medium for multi-level tags. In the embodiments of the present disclosure, a text and tags corresponding to keywords in the text are obtained; the text is encoded by a text encoding model in a preset text classification model for multi-level tags to obtain a feature vector of the text, and the feature vector of the text sensitively represents the keywords of the text. Based on a tag encoding model in the preset text classification model for multi-level tags, the tags are encoded to obtain tag vectors; the cosine similarity between the feature vector of the text and the vector of each tag is calculated respectively; the tags with a cosine similarity greater than a preset threshold are determined as the tags of the text. By encoding the text and existing category tags and calculating the cosine similarity, tags that match the text content are selected, which can reduce the dependence on manual tag annotation, lower the maintenance costs of manual annotation and the tag system, improve the accuracy of tag annotation, and make the text classification result more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of natural language processing, and in particular, to a text classification method, apparatus, device, and storage medium for multi-level tags. Background Art

[0002] In recent years, in order to achieve efficient storage, management, and exchange of information resources, many industries have begun to carry out the construction of data platforms. With the aggregation of information resources, scientific management of resources has become increasingly important. Among them, information tag annotation is an important technology for efficient retrieval and management of information. Information tag annotation refers to the task of deeply analyzing the title and content of an article and finding one or more tags that reflect the text theme and topic in a defined tag system, which is also called a multi-label classification task.

[0003] Currently, the multi-label classification task is mainly based on the industry tag system. Each tag summarizes sub-tags or keywords, and by using the keyword matching method, the tags corresponding to the keywords are returned. In recent years, with the rise of deep learning algorithms, many deep learning methods have also been applied to the multi-label classification task.

[0004] However, when the related technology performs multi-label classification on text, a large amount of labeled tag data is required, and most of the tag systems rely on manual construction and maintenance, which have the problems of high subjectivity, high cost, and slow update speed; the tags obtained by keyword matching usually have a lot of noise and misjudgment, resulting in a weak correlation between the text and the tags, and the accuracy is low. On the other hand, the multi-label classification method based on deep learning requires a large amount of manually labeled tags. Due to the increase in the tag space, multi-label data is more difficult to label. Especially when the text length is relatively long, the annotator is likely to be unable to traverse all tags and only gives a subset of the true tags; at the same time, the multi-label classification related technology usually only considers tags and ignores keyword information, which is likely to cause the problem of low tag recall rate. Therefore, there is an urgent need for a simple text classification method for multi-level tags to improve the accuracy of tag annotation. Summary of the Invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, embodiments of the present disclosure provide a text classification method, apparatus, device, and storage medium for multi-level tags.

[0006] The first aspect of the embodiments of the present disclosure provides a text classification method for multi-level tags, and the method includes:

[0007] Obtain the text and the tags corresponding to the keywords in the text; perform encoding processing on the text based on the text encoding model in the preset text classification model for multi-level tags to obtain the feature vector of the text. The feature vector of the text sensitively represents the keywords in the text. Based on the tag encoding model in the preset text classification model for multi-level tags, perform encoding processing on the tags to obtain the vectors of the tags; calculate the cosine similarity between the feature vector of the text and the vector of each tag respectively; determine the tags with the cosine similarity greater than the preset threshold as the tags of the text.

[0008] The second aspect of the embodiments of the present disclosure provides a text classification device for multi-level tags, and the device includes:

[0009] An obtaining module, configured to obtain the text and the tags corresponding to the keywords in the text;

[0010] An encoding module, configured to perform encoding processing on the text based on the text encoding model in the preset text classification model for multi-level tags to obtain the feature vector of the text. The feature vector of the text sensitively represents the keywords in the text. Based on the tag encoding model in the preset text classification model for multi-level tags, perform encoding processing on the tags to obtain the vectors of the tags;

[0011] A calculation module, configured to calculate the cosine similarity between the feature vector of the text and the vector of each tag respectively;

[0012] A determination module, configured to determine the tags with the cosine similarity greater than the preset threshold as the tags of the text.

[0013] The third aspect of the embodiments of the present disclosure provides a computing device, and the device includes a memory and a processor. Among them, a computer program is stored in the memory. When the computer program is executed by the processor, the method in the first aspect above can be implemented.

[0014] The fourth aspect of the embodiments of the present disclosure provides a computer-readable storage medium, and a computer program is stored in the storage medium. When the computer program is executed by the processor, the method in the first aspect above can be implemented.

[0015] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:

[0016] In an embodiment of the present disclosure, by obtaining a text and tags corresponding to keywords in the text; encoding the text through a text encoding model in a preset text classification model for multi-level tags to obtain a feature vector of the text, the feature vector of the text sensitively represents the keywords of the text, and encoding the tags through a tag encoding model in the preset text classification model for multi-level tags to obtain a vector of the tags; calculating the cosine similarity between the feature vector of the text and the vector of each tag respectively; and determining the tags with the cosine similarity greater than a preset threshold as the tags of the text. In the embodiment of the present disclosure, by performing encoding processing and cosine similarity calculation processing on the text and existing category tags, the tags matching the text content are selected, reducing the dependence on a large number of manually labeled tags, reducing the maintenance costs of manual labeling and the tag system, improving the accuracy of tag labeling, and making the text classification result more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0018] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a flowchart of a method for training a text classification model for multi-level tags provided by an embodiment of the present disclosure;

[0020] Figure 2 is a schematic diagram of the embedding layer structure of a BERT model of a text classification model for multi-level tags provided by an embodiment of the present disclosure;

[0021] Figure 3 is a flowchart of a method for text classification for multi-level tags provided by an embodiment of the present disclosure;

[0022] Figure 4 is a schematic diagram of the structure of a text classification device for multi-level tags provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to more clearly understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0024] In the following description, many specific details are set forth to provide a thorough understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.

[0025] Figure 1 FIG. 4 is a flowchart of a method for training a text classification model for multi-level tags provided by an embodiment of the present disclosure. This method can be executed by a computing device, which can be understood as any device with computing functions and processing capabilities. As Figure 1 shown, the method for training a text classification model for multi-level tags provided by this embodiment includes the following steps:

[0026] Step 101: Obtain the text and the tags corresponding to the keywords in the text to form an original data set.

[0027] In the embodiment of the present disclosure, obtaining the text and the tags corresponding to the keywords in the text can be obtained through steps S11-S12:

[0028] S11: Based on the existing keyword table, match the keywords from the text.

[0029] The keyword table referred to in the embodiment of the present disclosure can be understood as a set of keywords predefined in a certain field. One or more keywords can be matched from a text. In one embodiment, no keywords may be matched from a text, and then the tag of the text can be set to "other".

[0030] S12: Based on the mapping relationship between the keywords and the tags, determine the tags corresponding to the keywords in the text.

[0031] Among them, the mapping relationship between the keywords and the tags can be obtained through the existing tag-keyword system. One tag can correspond to multiple keywords, and each text can correspond to one or more tags. The tag-keyword system can be constructed manually. The related construction techniques are existing technologies and will not be elaborated here.

[0032] For example, in the text "XXX Company manufactured the first 3D printed drone and successfully completed the flight test", the keywords "3D printing" and "flight test" appear. Therefore, in the existing mapping relationship between the keywords and the tags, the tags of this text are "structural strength technology" corresponding to the keyword "3D printing" and "flight test technology" corresponding to the keyword "flight test". This is only an exemplary illustration of obtaining the tags corresponding to the keywords in the text, rather than the only illustration.

[0033] Multiple texts and multiple labels corresponding to each text are obtained through the above method to form an original dataset. A text and its corresponding multiple labels form a sample, and the original dataset includes multiple samples. It should be noted that the texts in the above method do not require manual labeling of labels, but approximate labels are obtained through the text data itself, and there may be noise in the obtained labels, that is, labels that do not match the text semantics.

[0034] Step 102: Divide the original dataset into a training set and a validation set. Input the training set into a text classification model for multi-level labels for training, and construct positive samples and negative samples within each batch (Batch).

[0035] The text classification model for multi-level labels referred to in the embodiments of the present disclosure is a model trained based on contrastive learning. It is necessary to divide the original dataset into a training set and a validation set, input the training set into the text classification model for multi-level labels for training, and construct positive samples and negative samples within each batch (Batch).

[0036] In the embodiments of the present disclosure, positive samples and negative samples are constructed within each Batch. Specifically, N samples can be randomly selected from the dataset to form a Batch. Any sample within the Batch includes a text and k labels corresponding to the text. The text and the k labels corresponding to the text form k positive samples, and the text and the other K - k labels in the full label set form K - k negative samples. Here, the full label set can be understood as all the labels in the constructed label system, and K is the total number of labels in the full label set. Finally, each Batch obtains N·K samples, where N, K, and k are positive integers, and k ≤ K.

[0037] Step 103: Perform encoding processing on the text in the sample to obtain the feature vector of the text. The feature vector of the text sensitively represents the keywords of the text, and perform encoding processing on the labels to obtain the vector of the labels.

[0038] The text classification model for multi-level labels in the embodiments of the present disclosure requires positive samples and negative samples to learn in the feature space. Therefore, it is necessary to convert the text and labels in the sample into vector representations.

[0039] The feature vector of the text referred to in the embodiments of the present disclosure sensitively represents the keywords of the text, which can be understood as emphasizing the keyword information in the feature vector of the text and reflecting which keywords the text matches. In one implementation, non-keyword content and keyword content can be represented in different ways to sensitively represent the keywords.

[0040] In an embodiment of the present disclosure, when encoding text, an existing encoding model can be used to encode the text. To emphasize the keyword information of text matching, a keyword embedding layer can be established in the encoding model, and the matching keyword information is injected into the keyword embedding layer. Then, the text information is converted into a feature vector of the text. The feature vector of the text contains the keyword information of the text, and sensitive representation of the keywords of the text is performed.

[0041] In an implementation manner of the embodiment of the present disclosure, the encoding model for encoding text can use the BERT (Bidirectional Encoder Representations from Transformers) model to represent the feature vector of the text. Among them, BERT is a deep bidirectional pre-trained language model. By running the self-supervised learning method on a large-scale corpus, a large amount of language, syntactic, and semantic information can be learned, and the feature vector of the text that fuses the context semantics can be output through bidirectional representation.

[0042] On the basis of the original three embedding features in the BERT model embedding layer, that is, on the basis of the token embedding layer, the segment embedding layer, and the position embedding layer, a keyword embedding layer is added, and the keyword information extracted from the text is injected into the keyword embedding layer to form a sensitive representation of the keyword.

[0043] Exemplarily, as Figure 2 shown in the schematic diagram of the embedding layer structure of the BERT model, when the text "XXXX completed the flight test of the new radar" is input into the BERT model, CLS (classification) represents classification, [CLS] is located at the beginning of the input text sentence, indicating that subsequent classification tasks can be performed, SEP (separator) represents separation, [SEP] is located in the middle or at the end of the input text, and is used to separate two input sentences. The text hits the keyword "flight test". Therefore, in the keyword embedding layer, the position corresponding to the keyword "flight test" can be represented as E K , and other positions are represented as E N , forming a sensitive representation of the keyword "flight test". In the subsequent model training process, the parameters of E K and E N can be randomly initialized, and the model parameters are updated by minimizing the loss function.

[0044] After the embedding layer of the BERT model processes the text information, the processed data is input into the Transformer network architecture in the BERT model. This Transformer network architecture can enable the output text feature vector to express and contain the text semantic information. After being processed by the Transformer network architecture, the BERT model finally outputs the feature vector of the text, and this feature vector of the text contains the keyword information of the text.

[0045] In some embodiments of the present disclosure, the feature vector of the text can be the vector corresponding to the placeholder [CLS] in the last layer of the Transformer, or it can be the vector obtained by average pooling after accumulating the outputs of the first layer and the last layer of the Transformer. It is not limited here.

[0046] In some embodiments of the present disclosure, when encoding the text, the number of words that the encoding model can receive may be limited. For example, the BERT model can accept at most 512 words as input. In the case of a long text, the title, the content of the first paragraph and the last paragraph can be extracted from the text, and the title, the content of the first paragraph and the last paragraph can be spliced, and encoding processing is performed based on the spliced text content to obtain the feature vector of the text. In some other embodiments, the title and the abstract can be extracted from the text, and the title and the abstract can be spliced, and encoding processing is performed based on the spliced text content to obtain the feature vector of the text.

[0047] In the embodiments of the present disclosure, before encoding the text, in order to reduce the variation of words and reduce text noise, the text can also be preprocessed. The preprocessing methods include at least one of deleting Hyper Text Markup Language (HTML), converting traditional Chinese characters to simplified Chinese characters, unifying the English case, and deleting the content that conforms to the preset regular expression. For example, for information such as the author, source, and publication time in the text, it can be removed through regular expressions. Among them, the regular expression is an existing mature technology and will not be elaborated here.

[0048] In an embodiment of the present disclosure, when encoding a tag, in many scenarios, the tag is usually a general term with a wide semantic range and the expressed semantic information is not specific enough. The specific corresponding content needs to be understood in combination with domain knowledge and keywords, and it is difficult for a computing device to understand its content based on the literal meaning of the tag. For example, a tag is "overall integrated design technology of aircraft", and the corresponding keywords include "aircraft shape design technology", "aerodynamic layout technology", "stealth technology", etc. If the tag "overall integrated design technology of aircraft" is input into the computing device, it is difficult for the computing device to determine which specific semantic information this tag corresponds to. Therefore, in an embodiment of the present disclosure, in order to avoid the problem of increased semantic complexity in semantic representation of tags, an existing encoding model can be used to encode the tags, converting the tag information into a vector of the tag. The vector of this tag does not contain the semantic information of the tag, and the model only models the mapping relationship between the feature vector of the text and the vector of the tag.

[0049] In an implementation manner of an embodiment of the present disclosure, a discrete tag can be mapped into a vector representation through one-hot encoding, and then the dimension of the vector can be changed through matrix transformation so that the dimension of the obtained vector of the tag is the same as the dimension of the feature vector of the text, and the vector of the tag is output. During the model training process, a matrix is constructed and randomly initialized, and through backpropagation and minimizing the loss function, the parameters in the matrix are determined to achieve the purpose of modeling the mapping relationship between the feature vector of the text and the vector of the tag. Here, one-hot encoding is an existing mature technology and will not be elaborated here.

[0050] Exemplarily, assume the one-hot vector is where \(R\) represents the set of real numbers, that is, the set containing all rational and irrational numbers, and \(K\) is the total number of tags. During the model training process, a matrix \(A\) is constructed and randomly initialized, \(A\in R\) M*K , where \(M\) is the feature dimension of the BERT model, then the calculation formula for the vector of the tag is During the model training process, the parameters in matrix \(A\) are determined through backpropagation and minimizing the loss function.

[0051] Step 104: Input the feature vector of the text and the vector of each tag into the loss function to determine the loss value, and iteratively update the model parameters by minimizing the loss function.

[0052] The loss function referred to in the embodiments of the present disclosure can be understood as a function that maps the value of a random event or its related random variable to a non-negative real number to represent the "risk" or "loss" of the random event. In applications, the loss function is usually associated with the learning criterion and the optimization problem, that is, the accuracy of the model is solved and evaluated by minimizing the loss function.

[0053] The loss function in the embodiments of the present disclosure may adopt the InfoNCE (Information Noise Contrastive Estimation) loss function, which is a contrastive loss function for self-supervised learning, where NCE (Noise Contrastive Estimation) represents noise contrastive estimation. Specifically, the feature vectors of the text of the positive samples obtained in the above steps, the vectors of each label, and the feature vectors of the text of the negative samples corresponding to the positive samples and the vectors of each label are input into the loss function, and the loss function can be expressed in the following form:

[0054]

[0055] where, (q i , p i + ) represents the i-th sample and belongs to the positive sample set, where q i represents the feature vector of the article, and p i + represents the vector of the label;

[0056] (q j , p j + ) represents the j-th sample and belongs to the positive sample set, where q j represents the feature vector of the article, and p j + represents the vector of the label;

[0057] (q j , p j - ) represents the j-th sample and belongs to the negative sample set, where q j represents the feature vector of the article, and p j - represents the vector of the label;

[0058] Sim represents the cosine similarity;

[0059] τ represents the temperature parameter, which is a model hyperparameter;

[0060] i represents the input of the i-th sample;

[0061] j represents the input of the j-th sample;

[0062] K + represents the positive sample set, and K - represents the negative sample set.

[0063] The cosine similarity here can be understood as the cosine value of the angle between two vectors. It is based on the comparison of the angles of two vectors in space. The closer the angles of the two vectors are in space, the more similar the two vectors are. The cosine similarity can be used as a metric to compare the similarity between the feature vector of the text and the vector of the label. The range of the cosine similarity here is from 0 to 1. The larger the cosine similarity, the more similar the two vectors are.

[0064] The value of the loss function in the embodiments of the present disclosure can be defined as the loss value, which can be understood as a non - negative real number, representing the loss or error of the text classification model for multi - level labels. The smaller the loss value, the greater the cosine similarity between the feature vector of the text and the vector of the label, and the higher the accuracy of the model.

[0065] In one implementation manner of the embodiments of the present disclosure, during the backpropagation process, the Adam optimizer can be used to iteratively update the model parameters by minimizing the loss function until the stopping condition is reached. The Adam optimizer here is a mature existing technology and will not be elaborated here. Other related technologies can also be used to minimize the loss function, which is not limited here.

[0066] Step 105: Train the text classification model for multi - level labels based on the training set, verify the text classification model for multi - level labels based on the validation set, and calculate the loss value of the model on the validation set until the loss value of the model on the validation set is less than or equal to the first preset threshold, then stop training and determine the final parameters of the text classification model for multi - level labels.

[0067] In the embodiments of the present disclosure, the first preset threshold of the loss value can be set to the minimum value of the loss function, that is, when the value of the loss function reaches the minimum, it can be understood that the loss value converges and remains unchanged or no longer decreases, that is, the accuracy of the model reaches the highest, indicating that the probability that the feature vector of the text and the vector of the label are similar reaches the maximum. That is to say, the cosine similarity between the feature vector of the text and the vector of the label reaches the maximum. At this time, the label corresponding to the vector of the label is the label closest to the semantics of the text.

[0068] In other embodiments of the present disclosure, the first preset threshold of the loss value can be set by the user according to actual needs or can be set by default by the computing device, and this is not limited.

[0069] Train the text classification model for multi - level labels based on the training set, verify the text classification model for multi - level labels based on the validation set, and calculate the loss value of the model on the validation set. If the loss value is greater than the first preset threshold, continue to train the model; if the loss value is less than or equal to the first preset threshold, stop training and determine the parameters of the text classification model for multi - level labels with the loss value less than or equal to the first preset threshold at this time as the final parameters of the model.

[0070] In other embodiments of the present disclosure, the text classification model for multi-level labels is verified based on the verification set, and can also be verified within a preset number of cycles. The loss value of the model on the verification set is calculated in each preset cycle. When the loss value of the verification set of the model does not further decrease within the preset number of cycles, the training is stopped and the parameters in the iteration result of the previous cycle are determined as the final parameters of the model.

[0071] The text classification model for multi-level labels in the disclosed embodiment, through multiple training with a large amount of data, shortens the distance between positive samples (similar samples) in the feature space and lengthens the distance between negative samples in the feature space to characterize the feature representation of the samples, and learns a text encoding model for text and a label encoding model for labels. Under the condition that the labels contain noise, as long as the data volume is large enough, the model can still learn the correct correspondence between labels and texts.

[0072] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0073] In an embodiment of the present disclosure, an original data set is constructed by obtaining a text and tags corresponding to keywords in the text; the original data set is divided into a training set and a validation set, the training set is input into a text classification model for multi-level tags for training, and positive and negative samples are constructed within each batch; the text in the samples is encoded to obtain a feature vector of the text, the feature vector of the text sensitively represents the keywords in the text, the tags are encoded to obtain vectors of the tags; the feature vector of the text and the vectors of each tag are input into a loss function to determine a loss value; the text classification model for multi-level tags is trained based on the training set, the text classification model for multi-level tags is verified based on the validation set, the loss value of the model on the validation set is calculated, and training is stopped until the loss value of the model on the validation set is less than or equal to a first preset threshold, and the final parameters of the text classification model for multi-level tags are determined to obtain a text classification model for multi-level tags, which can be applied to text classification for multi-level tags. In the embodiment of the present disclosure, a text classification model for multi-level tags is obtained by training a text and existing category tags for a text classification model for multi-level tags. A keyword embedding layer is added to the input of the text classification model for multi-level tags to effectively inject the keywords in the text into the model, forming a sensitive representation of the keywords. A large amount of noise data is regarded as a training signal, and through the method of contrastive learning, a text encoding model for the text and a tag encoding model for the tags are learned, achieving the effect that the distance between the text and relevant tags is close and the distance between the text and irrelevant tags is far. The model is applied to text classification for multi-level tags, which can greatly reduce the dependence on a large number of manually labeled tags, improve the accuracy of tag labeling, make the text classification result more accurate, and in addition to learning the mapping relationship between existing keywords and tags, the model can also learn the mapping relationship between new keywords and tags, having good generalization ability and reducing the maintenance cost of manual labeling and the tag system.

[0074] Figure 3 FIG. is a flowchart of a method for text classification for multi-level tags provided by an embodiment of the present disclosure, and this method can be executed by a computing device. The computing device can be understood as any device with computing functions and processing capabilities. As Figure 3 shown, the method for text classification for multi-level tags provided in this embodiment includes the following steps:

[0075] Step 301, obtain a text and tags corresponding to keywords in the text.

[0076] The manner of obtaining the text and tags corresponding to keywords in the embodiment of the present disclosure is the same as steps S11-S12 in the above step 101, and will not be elaborated here.

[0077] Step 302: Encode the text using the text encoding model in the preset text classification model for multi-level labels to obtain the feature vector of the text. The feature vector of the text sensitively represents the keywords of the text. Then, encode the labels using the label encoding model in the preset text classification model for multi-level labels to obtain the label vector.

[0078] The so-called sensitive representation of the keywords of the text by the feature vector of the text in the embodiments of the present disclosure can be understood as emphasizing the keyword information in the feature vector of the text, reflecting which keywords the text matches. In one implementation, the non-keyword content and the keyword content can be represented in different ways to sensitively represent the keywords.

[0079] The preset text classification model for multi-level labels in the embodiments of the present disclosure is Figure 1 the text classification model for multi-level labels trained in [reference], input the obtained text into Figure 1 the text encoding model in the text classification model for multi-level labels trained in [reference], input the obtained labels into Figure 1 the label encoding model in the text classification model for multi-level labels trained in [reference]. Encode the text using this text encoding model to obtain the feature vector of the text. The feature vector of the text sensitively represents the keywords of the text. Encode the labels using this label encoding model to obtain the label vector.

[0080] In the embodiments of the present disclosure, before encoding the text, in order to reduce the variation of words and reduce text noise, the text can also be preprocessed. The preprocessing methods include at least one of deleting Hyper Text Markup Language (HTML), converting traditional Chinese characters to simplified Chinese characters, unifying the case of English, and deleting the content that conforms to the preset regular expression. For example, for information such as the author, source, and publication time in the text, it can be removed through regular expressions. The regular expression is a mature existing technology and will not be elaborated here.

[0081] In the embodiments of the present disclosure, when encoding the text, the number of characters that the text classification model for multi-level labels can receive may be limited. In the case of a long text, the title, the first paragraph, and the last paragraph can be extracted from the text, and the title, the first paragraph, and the last paragraph are concatenated. Then, encode the concatenated text content to obtain the feature vector of the text. In other implementations, the title and the abstract can be extracted from the text, and the title and the abstract are concatenated. Then, encode the concatenated text content to obtain the feature vector of the text.

[0082] Step 303: Calculate the cosine similarity between the feature vector of the text and the vector of each label respectively.

[0083] In the embodiments of the present disclosure, the so-called cosine similarity can be understood as evaluating the similarity between two vectors by calculating the cosine value of the included angle between them. The range of the cosine similarity in the embodiments of the present disclosure is from 0 to 1. The greater the cosine similarity, the more similar the feature vector of the text and the vector of the label can be understood. The calculation method of the cosine similarity has been disclosed in the related art, and it can be calculated with reference to the related art, which will not be elaborated here.

[0084] Step 304: Determine the label of the text as the label with a cosine similarity greater than the preset threshold.

[0085] In the embodiments of the present disclosure, the preset threshold of the cosine similarity can be set by the user according to actual needs, or can be set by the computing device by default, and this is not limited.

[0086] In the embodiments of the present disclosure, if the cosine similarity is less than or equal to the preset threshold, the label corresponding to the cosine similarity less than or equal to the preset threshold is discarded; if the cosine similarity is greater than the preset threshold, the label corresponding to the cosine similarity greater than the preset threshold is determined as the label of the text, and there may be one or more such labels, and all can be considered as the labels closest to the semantics of the text.

[0087] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art:

[0088] In the embodiments of the present disclosure, by obtaining the text and the labels corresponding to the keywords in the text; encoding the text through the text encoding model in the preset text classification model for multi-level labels to obtain the feature vector of the text, and the feature vector of the text sensitively represents the keywords of the text, and encoding the labels through the label encoding model in the preset text classification model for multi-level labels to obtain the vector of the labels; calculating the cosine similarity between the feature vector of the text and the vector of each label respectively; determining the label of the text as the label with a cosine similarity greater than the preset threshold. In the embodiments of the present disclosure, by encoding the text and the existing category labels and calculating the cosine similarity, the labels matching the text content are selected, greatly reducing the dependence on a large number of manually labeled labels, reducing the maintenance cost of manual labeling and the label system, improving the accuracy of label annotation, and making the text classification result more accurate.

[0089] Figure 4 It is a schematic structural diagram of a text classification device for multi-level labels provided by the embodiments of the present disclosure, and this device can be understood as the above-mentioned computing device or some functional modules in the above-mentioned computing device. As Figure 4 shown, the text classification device 400 for multi-level labels includes:

[0090] An acquisition module 410, configured to acquire a text and tags corresponding to keywords in the text;

[0091] An encoding module 420, configured to perform encoding processing on the text based on a text encoding model in a preset text classification model for multi-level tags to obtain a feature vector of the text. The feature vector of the text sensitively represents the keywords in the text. Based on a tag encoding model in the preset text classification model for multi-level tags, perform encoding processing on the tags to obtain a vector of the tags;

[0092] A calculation module 430, configured to calculate the cosine similarity between the feature vector of the text and the vector of each tag respectively;

[0093] A determination module 440, configured to determine the tags whose cosine similarity is greater than a preset threshold as the tags of the text.

[0094] Optionally, the above acquisition module 410 includes:

[0095] A first matching sub-module, configured to match keywords from the text based on an existing keyword table;

[0096] A first determination sub-module, configured to determine the tags corresponding to the keywords in the text based on the mapping relationship between the keywords and the tags.

[0097] Optionally, the text classification device 400 for multi-level tags further includes:

[0098] A preprocessing module, configured to preprocess the text. The preprocessing method includes at least one of deleting hypertext markup language, converting traditional Chinese characters to simplified Chinese characters, unifying English case, and deleting content that conforms to a preset regular expression.

[0099] Optionally, the above encoding module 420 includes:

[0100] A first splicing sub-module, configured to extract the title and the content of the first and last paragraphs from the text, and splice the title and the content of the first and last paragraphs, or extract the title and abstract from the text, and splice the title and the abstract;

[0101] A first encoding sub-module, configured to perform encoding processing on the spliced text content to obtain a feature vector.

[0102] The text classification device for multi-level tags provided in this embodiment can execute the methods in any of the above Figure 3 embodiments. Its execution manner and beneficial effects are similar and will not be elaborated here.

[0103] An embodiment of the present disclosure also provides a computing device, which includes a processor and a memory. Among them, a computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned Figure 3 methods of any one of the embodiments can be implemented. Their execution manners and beneficial effects are similar and will not be elaborated here.

[0104] An embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored in the storage medium. When the computer program is executed by a processor, the above-mentioned Figure 3 methods of any one of the embodiments can be implemented. Their execution manners and beneficial effects are similar and will not be elaborated here.

[0105] The above-mentioned computer-readable storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0106] The above-mentioned computer program can be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0107] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0108] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A text classification method for multi-level tags, characterized in that, the method includes: obtaining the text and the tags corresponding to the keywords in the text; encoding the text based on the text encoding model in the preset text classification model for multi-level tags to obtain the feature vector of the text, the feature vector of the text sensitively represents the keywords of the text, and encoding the tags based on the tag encoding model in the preset text classification model for multi-level tags to obtain the vector of the tags; calculating the cosine similarity between the feature vector of the text and the vector of each tag respectively; determining the tags with the cosine similarity greater than the preset threshold as the tags of the text; the encoding the text based on the text encoding model in the preset text classification model for multi-level tags to obtain the feature vector of the text, the feature vector of the text sensitively represents the keywords of the text, includes: injecting the keywords of the text into the keyword embedding layer in the text encoding model based on the text encoding model to form a sensitive representation of the keywords of the text; converting the text into the feature vector of the text based on the text encoding model, and the feature vector of the text sensitively represents the keywords of the text.

2. The method according to claim 1, characterized in that, the obtaining the text and the tags corresponding to the keywords in the text includes: matching the keywords from the text based on the existing keyword table; determining the tags corresponding to the keywords in the text based on the mapping relationship between the keywords and the tags.

3. The method according to claim 1, characterized in that, before the encoding the text based on the text encoding model in the preset text classification model for multi-level tags to obtain the feature vector of the text, the feature vector of the text sensitively represents the keywords of the text, and encoding the tags based on the tag encoding model in the preset text classification model for multi-level tags to obtain the vector of the tags, the method further includes: preprocessing the text, and the preprocessing method includes at least one of deleting hypertext markup language, converting traditional Chinese characters to simplified Chinese characters, unifying English case, and deleting the content that conforms to the preset regular expression.

4. The method according to claim 1, characterized in that, encoding the text includes: extracting the title and the content of the first and last paragraphs from the text, and splicing the title and the content of the first and last paragraphs, or extracting the title and the abstract from the text, and splicing the title and the abstract; encoding the text content obtained by splicing to obtain the feature vector of the text.

5. A text classification device for multi-level tags, characterized in that, the device includes: an obtaining module, configured to obtain the text and the tags corresponding to the keywords in the text; An encoding module, configured to perform encoding processing on the text based on a text encoding model in a preset text classification model for multi-level labels, to obtain a feature vector of the text, where the feature vector of the text sensitively represents keywords of the text, and based on a label encoding model in the preset text classification model for multi-level labels, perform encoding processing on the label to obtain a vector of the label; A calculation module, configured to calculate the cosine similarity between the feature vector of the text and the vector of each label respectively; A determination module, configured to determine a label whose cosine similarity is greater than a preset threshold as the label of the text; The encoding module is configured to inject the keywords of the text into a keyword embedding layer in the text encoding model based on the text encoding model, so as to form a sensitive representation of the keywords of the text; based on the text encoding model, convert the text into the feature vector of the text, and the feature vector of the text sensitively represents the keywords of the text.

6. The apparatus according to claim 5, wherein, the obtaining module includes: a first matching sub-module, configured to match keywords from the text based on an existing keyword table; a first determining sub-module, configured to determine a label corresponding to the keyword in the text based on a mapping relationship between the keyword and the label.

7. The apparatus according to claim 5, wherein, the apparatus further includes: a preprocessing module, configured to perform preprocessing on the text, and the preprocessing method includes at least one of deleting hypertext markup language, converting traditional Chinese characters to simplified Chinese characters, unifying English case, and deleting content that conforms to a preset regular expression.

8. The apparatus according to claim 5, wherein, the encoding module includes: a first splicing sub-module, configured to extract the title and the content of the first and last paragraphs from the text, and splice the title and the content of the first and last paragraphs, or extract the title and abstract from the text, and splice the title and abstract; a first encoding sub-module, configured to perform encoding processing on the spliced text content to obtain a feature vector of the text.

9. A computing device, wherein, it includes: a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the multi-level label-oriented text classification method according to any one of claims 1-4 is implemented.

10. A computer-readable storage medium, wherein, a computer program is stored in the storage medium, and when the computer program is executed by a processor, the multi-level label-oriented text classification method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Multi-label text classification method based on statistics and pre-trained language model

    CN112214599A

  • Text label mining method, device and equipment and storage medium

    CN112328655A