Internet text hierarchy multi-label classification method and system

Through comparative learning and two-way interaction technology, the problem of low-frequency label classification accuracy in the existing technology is solved, and more accurate multi-label classification at the Internet text level is achieved, which improves classification accuracy and content retrieval efficiency.

CN120030167APending Publication Date: 2025-05-23UNIV OF JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197525.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing Internet text-level multi-label classification method is difficult to accurately classify low-frequency labels when dealing with a large number of label dependencies and unbalanced label distributions, resulting in a decrease in classification accuracy.

Method used

Through comparative learning methods, the co-occurrence and hierarchical relationships between labels are mined separately, low-frequency and high-frequency labels are differentiated, semantically rich label embeddings are generated, and more semantic correlation information is extracted through the two-way interaction between the label and the text.

Benefits of technology

It improves the accuracy of text classification, can more accurately capture the complex relationships between labels and text diversity, and improves the efficiency of content retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030167A_ABST
    Figure CN120030167A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, and provides an internet text hierarchical multi-label classification method and system.The method comprises the steps that in the training process, an original text is enhanced to obtain an enhanced text, and the original text and the enhanced text serve as positive sample pairs to mine the co-occurrence relation between labels; taking the labels with the direct hierarchical relationship as positive label pairs to mine the hierarchical relationship between the labels; performing differentiation enhancement on original tag features, enhancing a low-frequency tag through high-frequency co-occurrence tag information, and enhancing a high-frequency tag through historical tag information; and finally, carrying out bidirectional interaction on the text features and the enhanced label features, and carrying out secondary enhancement by utilizing potential semantic association between the labels and the texts to obtain classification features. And classification is performed based on the classification features to obtain a classification result, so that the purpose of enriching semantic features of the tags and the texts is achieved, and meanwhile, the classification precision of hierarchical multi-tag classification is improved by utilizing a co-occurrence relationship and a hierarchical relationship between the tags.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a method and system for hierarchical multi-label classification of Internet texts. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the development of cloud computing and mobile computing technologies, users can browse and publish information through platforms such as Weibo and Zhihu, and the amount of Internet text has exploded. These texts cover multiple categories such as news information, blog posts, forum discussions, product reviews, etc. Each category contains multiple topics, and it is difficult to fully describe their content with a single label. For example, the content of an article involves multiple topics such as technology, education, and artificial intelligence, and the topic of artificial intelligence belongs to the topic of technology, so there is a hierarchical relationship between the topics. In order to quickly and accurately find the required information from massive data, the hierarchical multi-label classification method can capture the diversity of text topics and their hierarchical relationships, and improve the efficiency of content retrieval.

[0004] Existing Internet text hierarchical multi-label classification methods are mainly divided into the following categories:

[0005] (1) Planar method: The hierarchical multi-label classification problem is transformed into a planar multi-label classification problem. Any text classified as a sub-label will automatically be attributed to all its ancestor labels.

[0006] (2) Local method: According to the divide-and-conquer principle, the overall classification problem is decomposed into multiple local sub-problems, the label hierarchy is divided into several local regions, and the classifier is constructed by node-by-node, parent-by-node, or layer-by-layer. The node-by-node and parent-by-node construction methods use the label relationships at different levels, while the layer-by-layer construction method uses the label relationships at the same level.

[0007] (3) Global method: The hierarchical multi-label classification problem is considered as a whole, and only one classifier is built on this structure to classify all labels. Although this method considers the relationship between labels at a macro level, in practical applications, as the number of labels and the number of label co-occurrence combinations continue to increase, a large number of label dependencies need to be processed. At the same time, since an article is associated with multiple labels, the number of texts corresponding to common labels is much higher than that of uncommon labels. This unbalanced label distribution leads to excessive focus on high-frequency labels during training, making it difficult to accurately classify low-frequency labels. Summary of the invention

[0008] In order to solve the above problems, the present invention proposes a hierarchical multi-label classification method and system for Internet text. The present invention performs comparative learning on text and labels respectively, mines the co-occurrence relationship and hierarchical structure of labels; differentially processes low-frequency labels and high-frequency labels, generates semantically rich label embedding, strengthens the information interaction between labels and texts, captures more features related to classification, and improves the accuracy of text classification.

[0009] According to some embodiments, a first solution of the present invention provides an Internet text hierarchical multi-label classification method, which adopts the following technical solution:

[0010] A hierarchical multi-label classification method for Internet text, comprising:

[0011] Obtain Internet text and preprocess it to obtain the original text;

[0012] Based on the original text features, the trained label classification model is used to classify labels and determine the label category to which the Internet text belongs;

[0013] The label classification using the trained label classification model is specifically as follows:

[0014] Perform feature extraction based on the original text to obtain the original text features;

[0015] Using the semantic association between label features and original text features, a label correlation matrix and a text correlation matrix are determined;

[0016] The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, and the label-guided attention weight and the text-guided attention weight in the label correlation matrix and the text correlation matrix are calculated by using the optimized attention mechanism.

[0017] The original text features and label features are updated respectively using label-guided attention weights and text-guided attention weights;

[0018] The updated label features and the updated text features are merged to obtain the classification features;

[0019] Classification is performed based on classification features to determine the label category of the Internet text.

[0020] Furthermore, the scaling factor is introduced to optimize the attention mechanism by utilizing the label features and the original text features, specifically:

[0021] Based on the label features and the original text features, the linear rectification function is used to extract the positive feature information in the label correlation matrix and the text correlation matrix, and the label scaling factor is determined;

[0022] Based on the original text features and label features, the relative importance of each element in the text correlation matrix and the label correlation matrix is ​​smoothed and normalized using a nonlinear function to determine the text scaling factor;

[0023] The label scaling factor and the text scaling factor are added to the attention mechanism respectively to determine the optimized label attention mechanism and text label attention mechanism.

[0024] Furthermore, the training process of the label classification model is specifically as follows:

[0025] Obtain an Internet text dataset and a label set of corresponding categories, preprocess the Internet text to obtain the original text, and construct a label graph based on the label set;

[0026] Based on the label graph and the original text, the co-occurrence relationship and hierarchical relationship between labels are mined through a joint supervised contrastive learning method to determine the contrast loss;

[0027] Extract original tag features based on tag graph, determine low-frequency tags and high-frequency tags according to the frequency of tags in Internet text dataset, enhance low-frequency tag features by using high-frequency co-occurrence tag information, and enhance high-frequency tag features by using historical tag information, so as to achieve differentiated enhancement of tag features and obtain tag features;

[0028] Extract original text features based on the original text, perform bidirectional interaction between original text features and label features, use the potential semantic connection between labels and texts to perform feature enhancement and update, and concatenate the updated text features and updated label features to obtain classification features;

[0029] Classify based on classification features to obtain classification results;

[0030] The above process is iterated repeatedly until the contrast loss is minimized and a trained label classification model is obtained.

[0031] Furthermore, based on the label graph and the original text, the co-occurrence relationship and hierarchical relationship between the labels are mined through a joint supervised contrastive learning method to determine the contrast loss, specifically:

[0032] Perform data augmentation on the original text to generate enhanced text. Take the original text and its enhanced text as positive sample pairs, and the original text and other text as negative sample pairs. Use contrastive learning methods to mine the co-occurrence relationship between labels according to the similarity between sample pairs.

[0033] Based on the hierarchical structure between labels in the label graph, a hierarchical tree is constructed. Labels with parent-child or ancestor relationships are used as positive label pairs, and other labels are used as negative label pairs. The hierarchical relationship between labels is mined according to the similarity between label pairs through contrastive learning methods.

[0034] The contrast loss is determined based on the co-occurrence relationship between tags and the hierarchical relationship between tags.

[0035] Furthermore, the use of high-frequency co-occurrence tag information to enhance low-frequency tag features is specifically as follows:

[0036] When multiple tags are annotated on the same text, there is a co-occurrence relationship between these tags;

[0037] Filter out tags with a frequency higher than a set threshold from tags that have a co-occurrence relationship with the current low-frequency tag to form a tag group;

[0038] The co-occurrence weight is used to weight the high-frequency co-occurrence tag feature vectors in the tag group to obtain the co-occurrence tag features;

[0039] The co-occurrence tag features and the low-frequency tag features are linearly weighted to obtain the enhanced low-frequency tag features.

[0040] Furthermore, the high-frequency features are enhanced by using historical tag information, specifically:

[0041] The high-frequency label features of each round in the training process are used as historical label information, and the historical label information of all rounds before the current round are integrated to obtain historical label features;

[0042] The historical label features and the high-frequency label features of the current round are linearly weighted to obtain enhanced high-frequency label features.

[0043] According to some embodiments, a second solution of the present invention provides an Internet text hierarchical multi-label classification system, which adopts the following technical solution:

[0044] A hierarchical multi-label classification system for Internet texts, comprising:

[0045] The text processing module is configured to obtain Internet text and perform preprocessing to obtain original text;

[0046] The text label classification module is configured to classify labels based on original text features using a trained label classification model to determine the label category to which the Internet text belongs;

[0047] The label classification using the trained label classification model is specifically as follows:

[0048] Perform feature extraction based on the original text to obtain the original text features;

[0049] Using the semantic association between label features and original text features, a label correlation matrix and a text correlation matrix are determined;

[0050] The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, and the label-guided attention weight and the text-guided attention weight in the label correlation matrix and the text correlation matrix are calculated by using the optimized attention mechanism.

[0051] The original text features and label features are updated respectively using label-guided attention weights and text-guided attention weights;

[0052] The updated label features and the updated text features are merged to obtain the classification features;

[0053] Classification is performed based on classification features to determine the label category of the Internet text.

[0054] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium.

[0055] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the method for hierarchical multi-label classification of Internet text as described in the first aspect above.

[0056] According to some embodiments, a fourth aspect of the present invention provides a computer device.

[0057] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for hierarchical multi-label classification of Internet text as described in the first aspect above are implemented.

[0058] According to some embodiments, a fifth aspect of the present invention provides a computer program product or a computer program.

[0059] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the method for hierarchical multi-label classification of Internet texts as described in the first aspect above.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] The present invention provides a hierarchical multi-label classification method for Internet text, which compares the similarities and differences between the original text and the enhanced text through a text enhancement contrast learning method, and mines the co-occurrence relationship between labels; and mines the hierarchical relationship between labels using hierarchical information through a label contrast learning method, and improves the classification accuracy by utilizing the relationship between labels.

[0062] The present invention provides a hierarchical multi-label classification method for Internet text, which achieves the purpose of enriching label features by differentially enhancing low-frequency labels and high-frequency labels; extracts more semantic association information through two-way interaction between labels and texts, and utilizes the relationship between the two to make the classification effect more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0064] Figure 1 It is a training flow chart of a hierarchical multi-label classification method for Internet text in Embodiment 1 of the present invention;

[0065] Figure 2 Schematic diagram of the joint supervised contrast learning process in the first embodiment of the present invention;

[0066] Figure 3 Schematic diagram of the differential label feature enhancement process in the first embodiment of the present invention;

[0067] Figure 4 It is a schematic diagram of the label-text bidirectional interaction process in the first embodiment of the present invention. DETAILED DESCRIPTION

[0068] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0069] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0070] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0071] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0072] Terminology explanation:

[0073] Label embedding representation: a vector representation of the corresponding label of the text that is easy to operate.

[0074] Text embedding representation: Convert text into low-dimensional vectors to facilitate computer processing and analysis.

[0075] Label graph: For all labels in the training set, the number of co-occurrences between labels is counted, and a graph structure is constructed with the category corresponding to each label as the vertex and the number of label co-occurrences as the edge.

[0076] BERT model: A pre-trained language model based on the Transformer architecture for extracting features from text.

[0077] Graph Convolutional Network GCN: An application of convolutional networks in machine learning to graph data.

[0078] Scaling factor: Parameter used to perform targeted scaling of elements in the correlation matrix.

[0079] ReLU function: When the input is greater than 0, the output is equal to the input, and when it is less than 0, the output is 0, which can retain more positive information.

[0080] Sigmoid function: This function maps the input to the (0,1) interval with a smooth gradient, which can avoid jumping output values.

[0081] Concatenation: refers to adding one array to a specific dimension of another array along a specific direction.

[0082] Multilayer Perceptron: A feed-forward neural network consisting of an input layer, one or more hidden layers, and an output layer.

[0083] Embodiment 1

[0084] The present embodiment provides a hierarchical multi-label classification method for Internet text. The present embodiment uses the method applied to a server as an example for illustration. It is understandable that the method can also be applied to a terminal, and can also be applied to a system including a terminal, a server, and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In the present embodiment, the method includes the following steps:

[0085] Obtain Internet text and preprocess it to obtain the original text;

[0086] Based on the original text features, the trained label classification model is used to classify labels and determine the label category to which the Internet text belongs;

[0087] The label classification using the trained label classification model is specifically as follows:

[0088] Perform feature extraction based on the original text to obtain the original text features;

[0089] Using the semantic association between label features and original text features, a label correlation matrix and a text correlation matrix are determined;

[0090] The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, and the label-guided attention weight and the text-guided attention weight in the label correlation matrix and the text correlation matrix are calculated by using the optimized attention mechanism.

[0091] The original text features and label features are updated respectively using label-guided attention weights and text-guided attention weights;

[0092] The updated label features and the updated text features are merged to obtain the classification features;

[0093] Classification is performed based on classification features to determine the label category of the Internet text.

[0094] like Figure 1 As shown, this embodiment provides a training process for a hierarchical multi-label classification method for Internet text. First, the text data is enhanced to generate enhanced text. The original text and its enhanced text constitute positive and negative sample pairs, and the co-occurrence relationship between labels is mined through comparative learning; then a hierarchical label tree is constructed according to the label hierarchy to obtain hierarchical information, and hierarchical information is transmitted between neighbor nodes of the label graph. On this basis, positive and negative label pairs are constructed, and the hierarchical relationship between labels is mined through comparative learning; then, the label features are differentially enhanced, low-frequency labels are enhanced through high-frequency co-occurrence label information, and high-frequency labels are enhanced through historical label information; then, the text features and label features are bidirectionally interacted, and the potential semantic connection between labels and texts is applied to perform feature enhancement, and the updated text features and label features are spliced ​​to obtain classification features; finally, the classification results are obtained based on the classification features. Specifically, the following steps are included:

[0095] Step 1: Obtain Internet text data and the corresponding tag set, and perform data preprocessing on the Internet text data, including cleaning non-text data, removing stop words, removing low-frequency words, removing high-frequency words and word form restoration to obtain the original text.

[0096] Step 2: If Figure 2As shown in Figure 2, the joint supervised contrastive learning model includes two submodules: text-enhanced contrastive learning and label contrastive learning, which mine and utilize the co-occurrence relationship and hierarchical relationship between labels.

[0097] The co-occurrence relationship is: In the hierarchical multi-label classification task, multiple labels are labeled for the text according to the content of the text. When multiple labels are labeled on the same text, it means that there is a co-occurrence relationship between these labels;

[0098] The hierarchical relationship is a subordinate or inclusive relationship between tags, which is usually represented by a parent-child relationship. This relationship is the hierarchical relationship between tags.

[0099] Step (201): The text enhancement contrastive learning submodule first performs a text enhancement operation with grammatical structure constraints and context information constraints on the input original text to generate an enhanced text with similar semantics for the original text; the original text and its enhanced text are used as positive sample pairs, and the original text and other texts are used as negative sample pairs, and the co-occurrence relationship between labels is mined according to the similarity between the sample pairs through a contrastive learning method;

[0100] The text enhancement operation with grammatical structure constraints and context information constraints includes three text enhancement methods: context adaptation word replacement, context related word insertion and semantic equivalent structure transformation;

[0101] Among them, the context-adaptive word replacement selects words that are consistent with the context and have similar semantics for replacement based on word vector similarity and semantic analysis, without changing the core semantics of the text; specifically: constructing a context-aware model based on deep learning, capturing the contextual features of the current text, identifying key contextual information in the text, and combining the deep semantic knowledge base-worknet synonym dictionary to locate the key words in the original text, and selecting candidate words with a semantic similarity greater than a set threshold and adapted to the current context from the dictionary based on the similarity, and then selecting the candidate words with the highest semantic similarity as the replacement word; among them, the context-aware model includes but is not limited to using the keyLLM keyword extraction tool to extract keywords, and other keyword extraction tools can also be used.

[0102] Among them, the context-related word insertion determines the core theme of the text based on semantic topic analysis, inserts words closely related to the core theme, and enhances the semantic expression of the text; specifically: using the topic model LDA to perform latent semantic topic analysis on the original text, identify semantic information and topic relationships, and determine the core theme and key information nodes of the text. On this basis, the Cosine similarity based on the word vector is calculated, and words that are highly relevant to the text content and semantically consistent are selected for insertion to ensure that the inserted words are consistent with the semantics of the original text and improve the richness of the text.

[0103] Among them, the semantically equivalent structure transformation identifies the syntactic structure through syntactic dependency analysis, adjusts the syntactic structure, and generates a text that conforms to grammatical norms; first, the subject, predicate, object and other components of the sentence in the text and their interdependencies are identified through the dependency syntactic analysis tool spaCy. Then, the semantic roles in the sentence are further identified through semantic role labeling (SRL) to ensure the correctness of the grammatical structure and semantic roles. Finally, on this basis, the Transformer-based generative model is used to reorganize the sentence structure identified by the syntactic analysis tool, adjust the word order structure, generate semantically equivalent sentences, and then generate text that is semantically equivalent to the original text but different in form.

[0104] Use the BERT model to encode the original text and enhanced text into vector representations to obtain the original text features and enhanced text features;

[0105] For each original text, its corresponding enhanced text constitutes a positive sample pair, and other samples constitute a negative sample pair. Calculate the similarity between the original text and the positive sample and negative sample to obtain the similarity metric:

[0106]

[0107] Among them, v x is the original text vector representation, v xaug It is the enhanced text vector representation.

[0108] Calculate the loss based on the similarity metric and get the co-occurrence loss:

[0109]

[0110] Among them, τ is a temperature parameter used to adjust the scaling of similarity. The model learns the co-occurrence relationship between tags by minimizing this loss function.

[0111] Step (202): Label contrast learning submodule: first analyze the label graph G = (V, E), and construct a label tree according to the hierarchical structure between labels to generate original label features containing hierarchical information; labels with clear hierarchical relationships are used as positive label pairs, and other labels are used as negative label pairs, and the hierarchical relationship between labels is mined according to the similarity between the label pairs through a contrast learning method; the labels with clear hierarchical relationships are labels with a parent-child relationship or an ancestor relationship in the hierarchical tree;

[0112] For each node v i ∈V, which initializes the original label feature vector Then, the graph convolutional network GCN is used to propagate hierarchical information in the label map. The information propagation formula is:

[0113]

[0114] Among them, l represents the number of layers, N(i) is the set of neighbor nodes of node i, and c ij is the normalization constant, W(l) is the weight matrix of the lth layer, σ is the activation function, is the original label feature vector of node i at layer l+1, is the original label feature vector of node j at layer l.

[0115] Through multi-layer information propagation, the hierarchical information of neighboring nodes is integrated, and on this basis, positive and negative label pairs are constructed. The positive label pairs select labels with direct hierarchical relationships, and the negative label pairs are selected from labels at different levels that are not directly related. Define label Y i and label Y j The distance measure between:

[0116]

[0117] Among them, l i is the level of the i-th label in the hierarchy.

[0118] Calculate the layer loss based on the distance metric:

[0119]

[0120] Where N is the number of labels, f(*,*) is the exponential cosine similarity measure between two embedding vectors, and z ij is the feature of the jth label corresponding to the i-th sample, P ij is the set of positive label pairs, N ij is the set of negative label pairs, z p is the feature of label p in the set of positive label pairs, z a is the feature of label a in the positive label pair set, z k is the feature of label k in the set of negative label pairs, ρ pj is the distance measure between label p and the jth label, ρ aj is the distance measure between label a and the jth label, ρ kj is the distance metric between label k and the jth label.

[0121] By minimizing the loss function, the model brings labels with hierarchical relationships closer in the latent space and learns the hierarchical relationship between labels.

[0122] Step (203): Combine the co-occurrence loss and the layer loss to obtain the contrast loss L Com-loss :

[0123] L Com-loss =LCo-loss +L Hier-loss (6);

[0124] Through contrast loss, the co-occurrence information and hierarchical information between tags are comprehensively considered, so that the model can capture the regularity of simultaneous appearance of tags in the text based on the hierarchical structure of tags, thereby optimizing the model's understanding and representation of tag relationships.

[0125] Step 3: If Figure 3 As shown in the figure, the differential label feature enhancement model uses different processing methods to enhance low-frequency labels and high-frequency labels, and updates the original label features for the first time, so that the model can learn the features of labels with different frequencies more evenly;

[0126] Step (301): First, count the number of texts corresponding to each original label feature, calculate the frequency of each original label feature in the entire text data set, and obtain the label frequency information F:

[0127]

[0128] According to the label frequency information F, the original labels are divided into low-frequency labels, common labels and high-frequency labels;

[0129] Determine whether the original label is a high-frequency label or a low-frequency label, F>α 1 When it is a high-frequency label, F<α 2 is a low-frequency label, where α 1 is the preset high frequency threshold, α 2 It is the preset low frequency threshold;

[0130] Step (302): for low-frequency tag features, construct a high-frequency co-occurrence neighborhood, and filter out tags that have a high-frequency co-occurrence relationship with the low-frequency tag features, i.e., high-frequency co-occurrence tags, and group multiple high-frequency co-occurrence tags into a tag group. Set co-occurrence weights based on co-occurrence frequencies, generate tag group features, and perform linear weighting with low-frequency tag features to enhance low-frequency tags;

[0131] The high-frequency co-occurrence neighborhood is composed of tags that have a co-occurrence relationship with the current low-frequency tag feature;

[0132] The co-occurrence frequency is: given two tags, the proportion of the number of texts annotated with these two tags to the total number of texts;

[0133] The co-occurrence weight is the closeness of the co-occurrence relationship between two tags. The higher the co-occurrence frequency, the closer the relationship, and the higher the co-occurrence weight.

[0134] Specifically, first traverse the training data set, count the label combinations that appear in it, and build a high-frequency co-occurrence neighborhood based on the combination, which contains all labels that have a co-occurrence relationship with the current low-frequency label feature. Based on the neighborhood, the label co-occurrence matrix M is obtained, where the rows and columns of the matrix represent labels, and the values ​​of the matrix elements represent the co-occurrence frequency of the two labels in the text data. If the value of the matrix element exceeds the preset threshold, the two labels are considered to be high-frequency co-occurrence labels;

[0135] Select the high-frequency co-occurrence tags of the current low-frequency tag features to form a tag group G low Similarly, for each low-frequency tag feature, a tag group should be constructed, and the co-occurrence weight should be designed for the tag group G low Medium and high frequency co-occurrence label feature vector h 1 ,h 2 ,……,h n Weighted, we get the co-occurrence label features:

[0136]

[0137] Among them, w i is the co-occurrence weight, is the co-occurrence frequency between the current low-frequency tag and the high-frequency co-occurrence tag.

[0138] The co-occurrence weight calculation method is: the co-occurrence weight is the ratio of the co-occurrence frequency of a high-frequency co-occurrence tag in the tag group to the total co-occurrence frequency of all high-frequency co-occurrence tags in the entire tag group;

[0139] Co-occurrence label features With low-frequency label features The enhanced low-frequency label features are obtained by linear weighting:

[0140]

[0141] Among them, α is the weight coefficient.

[0142] Step (303): for high-frequency label features, extract the high-frequency label features of different rounds in the training process as historical label information, set the attenuation weight according to the training round, generate historical label features, and linearly weight them with the high-frequency label features to enhance the high-frequency label features;

[0143] The historical label information is: the high-frequency label information assigned to the text in different rounds during the training process, reflecting the classification of the text at different training stages;

[0144] The attenuation weight is the influence of high-frequency label information of different training rounds on the current high-frequency label feature. The closer to the current training round, the more important the information, the greater the influence, and the higher the attenuation weight.

[0145] Specifically, during the training process, the high-frequency label feature vector of each round is recorded. Design the decay weight according to the training round information, fuse the historical label information of the previous n rounds, and obtain the historical label features:

[0146]

[0147] ω t =(1-λ) T-t (12);

[0148] Among them, T is the current round, t represents the round, and λ is the decay factor.

[0149] Linearly weight the historical label features and the current round high-frequency label features to obtain the enhanced high-frequency label features:

[0150]

[0151] Among them, α is the weight coefficient.

[0152] The enhanced low-frequency label features and high-frequency label features are reconnected to form a complete enhanced label feature, which makes up for the defect of insufficient feature learning of low-frequency labels due to the small number of samples, and uses the high-frequency co-occurrence label information to better express its own semantics; it weakens the model's excessive dependence on high-frequency labels during training, so that the model can make accurate classification when facing different frequency labels.

[0153] Step 4: Figure 4 As shown in the figure, the label-text bidirectional interaction model first uses the local information of the label and the text to build a correlation matrix, and preliminarily measures the semantic association between the label and each part of the text. Then, the attention mechanism is used to introduce a scaling factor to optimize the attention calculation method, and the label features and text features are updated for the second time. After the bidirectional interaction, the text features contain the semantic focus related to the label, which can accurately reflect the content related to the label classification; the label features are adjusted according to the text content to better fit the content of the current text and avoid the generalization of the label semantics;

[0154] Step (401): inputting the original text features and the label features processed in step 3 into the two-way interaction module to implement text feature update guided by label information;

[0155] Obtain the label correlation matrix S based on the label feature H and the original text feature D HD :

[0156]

[0157] Among them, W Ht∈R k*d and W Dt ∈R k*d is the corresponding weight matrix, which will be continuously adjusted during the model training process to adapt to different text and label data.

[0158] According to the label correlation matrix S HD Calculate the label-guided attention weight A HD :

[0159] A HD =softmax(F⊙S HD ) (15);

[0160] F = ReLU((HW Hf )(DW Df ) T ) (16);

[0161] Where F is the label scaling factor.

[0162] The label scaling factor F uses the ReLU function as an activation function, focuses on extracting positive feature information in the correlation matrix, filters out negative values, emphasizes the positive association and interaction between the two, and mines the positively motivated feature interaction relationship between labels and texts, which helps to highlight the impact of some obvious strong correlation features on text attention;

[0163] According to this attention weight, we can get the text features guided by label information:

[0164]

[0165] Step (402): inputting the text features and the label features processed in step 3 into the two-way interaction module to implement label feature update guided by text information;

[0166] Obtain the text relevance matrix S based on the label features H and the original text features D DH :

[0167]

[0168] Among them, W Dt ∈R k*d and W Ht ∈R k*d is the corresponding weight matrix, which will be continuously adjusted during the model training process.

[0169] According to the text relevance matrix S DH Calculate the text-guided attention weight A DH :

[0170] A DH=softmax(G⊙S DH ) (19);

[0171] G=sigmoid((DW Df )(HW Hf ) T ) (20);

[0172] Where G is the text scaling factor.

[0173] The text scaling factor G uses the sigmoid function as an activation function, focuses on smoothing and normalization, and can finely adjust the relative importance of each element in the correlation matrix to avoid excessive adjustment of label features due to some extreme values, while making the scaling ratio of the text-label relationship more stable and controllable;

[0174] According to this attention weight, we get the label features guided by text information:

[0175]

[0176] Step (403): Update the features and Fusion is performed to obtain the final classification features:

[0177]

[0178] Step 5: Classification module, input the final classification feature Z into the multi-layer perceptron to obtain the final classification result.

[0179] The above-mentioned joint supervised contrastive learning model, differentiated label feature enhancement model, label-text bidirectional interaction model, classifier, etc. constitute a hierarchical multi-label classification model, which needs to be trained in advance using a manually labeled training set. The text to be tested is input into the classification model, and the hierarchical multi-label classification is performed on it through the trained classification model, and the final classification result is returned.

[0180] Embodiment 2

[0181] This embodiment provides an Internet text hierarchical multi-label classification system, including:

[0182] A text processing module is configured to obtain Internet text and perform preprocessing to obtain original text;

[0183] The text label classification module is configured to classify labels based on original text features using a trained label classification model to determine the label category to which the Internet text belongs;

[0184] The label classification using the trained label classification model is specifically as follows:

[0185] Perform feature extraction based on the original text to obtain the original text features;

[0186] Using the semantic association between label features and original text features, a label correlation matrix and a text correlation matrix are determined;

[0187] The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, and the label-guided attention weight and the text-guided attention weight in the label correlation matrix and the text correlation matrix are calculated by using the optimized attention mechanism.

[0188] The original text features and label features are updated respectively using label-guided attention weights and text-guided attention weights;

[0189] The updated label features and the updated text features are merged to obtain the classification features;

[0190] Classification is performed based on classification features to determine the label category of the Internet text.

[0191] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0192] The proposed system can be implemented in other ways. For example, the system embodiment described above is only illustrative, and the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0193] Embodiment 3

[0194] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the method for hierarchical multi-label classification of Internet text as described in the first embodiment above are implemented, including mining and utilizing the co-occurrence relationship and hierarchical relationship between tags, realizing high- and low-frequency tag feature enhancement, and performing secondary enhancement on features through bidirectional interaction between tags and text. By storing the above instructions in a computer-readable storage medium, it is possible to support efficient data processing and accurate operation of the model in the hierarchical multi-label classification process.

[0195] Embodiment 4

[0196] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the hierarchical multi-label classification method for Internet text as described in the first embodiment above are implemented. When the computer program is running, it can read the original text data and label set from the memory, call each module to complete data processing, feature generation and enhancement, obtain classification features, and finally call the classifier to output the hierarchical multi-label classification results based on the classification features. By combining the above storage medium and computer device, the specific steps of the hierarchical multi-label classification method for Internet text can be implemented to ensure the orderly processing of data and efficient operation during the classification process.

[0197] Embodiment 5

[0198] This embodiment provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the method for hierarchical multi-label classification of Internet texts described in the first embodiment.

[0199] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.

[0200] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0201] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0202] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0203] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0204] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A hierarchical multi-label classification method for Internet text, characterized in that: include: Obtain Internet text and preprocess it to obtain the original text; Based on the original text features, the trained label classification model is used to classify labels and determine the label category to which the Internet text belongs; The label classification using the trained label classification model is specifically as follows: Perform feature extraction based on the original text to obtain the original text features; Using the semantic association between label features and original text features, a label correlation matrix and a text correlation matrix are determined; The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, and the label-guided attention weight and the text-guided attention weight in the label correlation matrix and the text correlation matrix are calculated by using the optimized attention mechanism. The original text features and label features are updated respectively using label-guided attention weights and text-guided attention weights; The updated label features and the updated text features are merged to obtain the classification features; Classification is performed based on classification features to determine the label category of the Internet text.

2. A hierarchical multi-label classification method for Internet text as claimed in claim 1, characterized in that: The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, specifically: Based on the label features and the original text features, the linear rectification function is used to extract the positive feature information in the label correlation matrix and the text correlation matrix, and the label scaling factor is determined; Based on the original text features and label features, the relative importance of each element in the text correlation matrix and the label correlation matrix is ​​smoothed and normalized using a nonlinear function to determine the text scaling factor; The label scaling factor and the text scaling factor are added to the attention mechanism respectively to determine the optimized label attention mechanism and text label attention mechanism.

3. The method for hierarchical multi-label classification of Internet text according to claim 1, characterized in that: The training process of the label classification model is specifically as follows: Obtain an Internet text dataset and a label set of corresponding categories, preprocess the Internet text to obtain the original text, and construct a label graph based on the label set; Based on the label graph and the original text, the co-occurrence relationship and hierarchical relationship between labels are mined through a joint supervised contrastive learning method to determine the contrast loss; Extract original tag features based on tag graph, determine low-frequency tags and high-frequency tags according to the frequency of tags in Internet text dataset, enhance low-frequency tag features by using high-frequency co-occurrence tag information, and enhance high-frequency tag features by using historical tag information, so as to achieve differentiated enhancement of tag features and obtain tag features; Extract original text features based on the original text, perform bidirectional interaction between original text features and label features, use the potential semantic connection between labels and texts to perform feature enhancement and update, and concatenate the updated text features and updated label features to obtain classification features; Classify based on classification features to obtain classification results; The above process is iterated repeatedly until the contrast loss is minimized and a trained label classification model is obtained.

4. A hierarchical multi-label classification method for Internet text as claimed in claim 3, characterized in that: Based on the label map and the original text, the co-occurrence relationship and hierarchical relationship between the labels are mined through a joint supervised contrastive learning method to determine the contrast loss, which is specifically: Perform data augmentation on the original text to generate enhanced text. Take the original text and its enhanced text as positive sample pairs, and the original text and other text as negative sample pairs. Use contrastive learning methods to mine the co-occurrence relationship between labels according to the similarity between sample pairs. Based on the hierarchical structure between labels in the label graph, a hierarchical tree is constructed. Labels with parent-child or ancestor relationships are used as positive label pairs, and other labels are used as negative label pairs. The hierarchical relationship between labels is mined according to the similarity between label pairs through contrastive learning methods. The contrast loss is determined based on the co-occurrence relationship between tags and the hierarchical relationship between tags.

5. The method for hierarchical multi-label classification of Internet text according to claim 3, characterized in that: The method of enhancing low-frequency tag features by using high-frequency co-occurrence tag information is specifically as follows: When multiple tags are annotated on the same text, there is a co-occurrence relationship between these tags; Filter out tags with a frequency higher than a set threshold from tags that have a co-occurrence relationship with the current low-frequency tag to form a tag group; The co-occurrence weight is used to weight the high-frequency co-occurrence tag feature vectors in the tag group to obtain the co-occurrence tag features; The co-occurrence tag features and the low-frequency tag features are linearly weighted to obtain the enhanced low-frequency tag features.

6. A hierarchical multi-label classification method for Internet text as claimed in claim 3, characterized in that: The method of enhancing high-frequency features by using historical tag information is specifically as follows: The high-frequency label features of each round in the training process are used as historical label information, and the historical label information of all rounds before the current round are integrated to obtain historical label features; The historical label features and the high-frequency label features of the current round are linearly weighted to obtain enhanced high-frequency label features.

7. A hierarchical multi-label classification system for Internet text, characterized in that: include: A text processing module is configured to obtain Internet text and perform preprocessing to obtain original text; The text label classification module is configured to classify labels based on original text features using a trained label classification model to determine the label category to which the Internet text belongs; The label classification using the trained label classification model is specifically as follows: Perform feature extraction based on the original text to obtain the original text features; Using the semantic association between label features and original text features, a label correlation matrix and a text correlation matrix are determined; The scaling factor is introduced by using the label features and the original text features to optimize the attention mechanism, and the label-guided attention weight and the text-guided attention weight in the label correlation matrix and the text correlation matrix are calculated by using the optimized attention mechanism. The original text features and label features are updated respectively using label-guided attention weights and text-guided attention weights; The updated label features and the updated text features are merged to obtain the classification features; Classification is performed based on classification features to determine the label category of the Internet text.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the Internet text hierarchical multi-label classification method as described in any one of claims 1 to 6 are implemented.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the Internet text hierarchical multi-label classification method as described in any one of claims 1-6 are implemented.

10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the Internet text hierarchical multi-label classification method as described in any one of claims 1-6 are implemented.