Hybrid enhancement method based on local hierarchical tag sequence similarity guidance
Through the mixed enhancement method guided by local hierarchical label sequence similarity, the problem of difficult to model local hierarchical correlation in hierarchical text classification is solved, and a more efficient hierarchical text classification effect is achieved.
Patent Information
- Application Number
- CN202510403522.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
Existing hierarchical text classification methods are difficult to effectively model large-scale, unbalanced and structured label hierarchies, and lack clear mechanisms to capture the intrinsic correlations between local hierarchies.
A mixed enhancement method guided by local hierarchical tag sequence similarity is adopted. By introducing a depth identifier to represent the hierarchical depth, an intermediate sample is generated using a pre-trained language model, and combining the Mixup ratio guided by local hierarchical correlation, a more informative sample is generated.
Effectively capture the intrinsic correlation between local levels, improve the accuracy and efficiency of hierarchical text classification, and is better than traditional methods.
Smart Images

Figure CN120336533A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more particularly, to a hybrid enhancement method guided by local hierarchical label sequence similarity. Background Art
[0002] Hierarchical Text Classification (HTC) is a variant of multi-label classification tasks, characterized in that labels are organized in a predefined hierarchical structure, and each text will be assigned one or more labels in this hierarchy. This hierarchical structure captures the relationships and dependencies between labels, and deeper-level labels contain shallower-level labels.
[0003] The core challenge of HTC lies in how to effectively model large-scale, unbalanced, and structured label hierarchies. Some existing studies regard the hierarchy as a directed acyclic graph, and obtain label representations containing hierarchical information through a structure encoder. However, this strategy allows each text to share the entire static global hierarchical structure, thus introducing redundant information into the graph, and this redundancy becomes more obvious as the scale of the hierarchy increases. In contrast, another type of local hierarchical method extracts the local hierarchy related to the text from the global hierarchy, and the local hierarchy can be represented as a containment relationship in the latent space. As Figure 1 Figures (a) and (b) in [reference] show the differences between these two hierarchical structures. For example, "Software" and "Machine Learning" should exist in the same subspace of "CS (Computer Science)", while "Geometry" and "Statistics" should be located in a certain subspace of the "Math" category.
[0004] By treating the local hierarchy as a sequence, the language model can capture the inherent parent-child relationships in the hierarchy. In addition, prompted learning can effectively utilize sequence information. The present invention models the local hierarchy as a sequence through a hierarchical template, so as to encode and align its structure. At the same time, there may be intrinsic correlations between labels that transcend the parent-child relationship constraints in the hierarchy, such as Figure 1(c) shows the correlation between different local levels. Specifically, although "CS / Machine Learning" and "Math / Statistics" are in different subspaces in the hierarchy, they may exhibit a relatively close distance in the latent space, indicating a similarity relationship beyond the direct hierarchical constraints. In fact, due to the widespread correlations (including sibling and peer relationships) among tags in HTC, capturing the correlation between local levels becomes particularly important. However, existing methods lack a clear mechanism to address this specific challenge. The present invention utilizes the Mixup method to enhance the implicit correlation between input pairs by generating intermediate samples, and based on hierarchical prompt tuning, models the implicit relationship between local levels. Summary of the Invention
[0005] An object of the embodiments of the present disclosure is to provide a hybrid enhancement method guided by the similarity of local hierarchical label sequences, which solves the problems of effectively modeling local levels in a sequence and capturing the inherent correlation between local levels.
[0006] In one general aspect, a hybrid enhancement method guided by the similarity of local hierarchical label sequences is provided. For input text and labels, a method of local hierarchical association is adopted, depth identifiers are introduced to represent the hierarchical depth of each label, the local level is represented as a sequence of identifiers, and further converted into a soft identifier form through a pre-trained language model as a hierarchical prompt; at each depth of the hierarchy, both the input and output are mixed using the Mixup ratio guided by local hierarchical correlation to generate intermediate samples, and then hierarchical label representations are generated.
[0007] The specific method for introducing depth identifiers to represent the hierarchical depth of each label is as follows: for a given input X, its corresponding local level is represented as a sequence in a similar way by replacing the [MASK] identifier with the corresponding real label. When a local level contains multiple paths, the labels at the same depth level are concatenated together and placed after the corresponding depth identifier. By inputting the sentence into the pre-trained language model and extracting the hidden output of the [CLS] identifier, the embedding representation of the sentence is obtained, and this representation is used to calculate the sentence similarity;
[0008] Formally, assuming that the corresponding local level representations of two inputs X i and X j are h i CLS and h j CLS respectively, then the distance between these two levels is measured by normalized cosine similarity, and its calculation formula is:
[0009]
[0010] Among them, · represents the vector dot product operation, and |·| is the L 2 norm.
[0011] The method for calculating the sentence similarity is as follows: Design a heuristic function to characterize the relationship between the local hierarchical similarity s and the Mixup ratio λ:
[0012] λ = -(β - 0.5)s α + β
[0013] Among them, α > 0 controls the rate of change of λ with s, and β ∈ (0.5, 1] controls the upper limit value of λ.
[0014] The specific method of simultaneously mixing the input and output with the Mixup ratio guided by the local hierarchical correlation at each depth of the hierarchy is as follows:
[0015] For the input Mixup, interpolate the hidden output corresponding to the [MASK] identifier at each depth of the hierarchy:
[0016]
[0017] Subsequently, through the mixed input obtain the prediction score That is:
[0018]
[0019] For the mixing of the output, select the mixed loss term:
[0020]
[0021] When α = 1, there is a linear relationship between s and λ. When α < 1, as s increases, the rate of decrease of λ slows down; on the contrary, when α > 1, as s increases, the rate of decrease of λ speeds up; in addition, β determines the maximum value of λ, indicating the minimum impact that Mixup can exert.
[0022] Select α from [0.1, 0.3, 0.6, 1, 2, 5, 10] and select β from [0.7, 0.75, 0.8, 0.85, 0.9, 0.95, 1].
[0023] The technical effect to be achieved by the embodiments of the present invention is as follows:
[0024] First, prompt tuning is applied to hierarchical text classification (HTC), where the local hierarchy is regarded as a sequence. In this process, virtual tokens are introduced to represent each depth in the hierarchy, thereby disassembling the parent-child relationship and aligning the hierarchy according to the depth. Under this hierarchical prompt tuning framework, the Mixup technique is combined to capture the implicit correlations between sibling relationships and co-level relationships. Different from the Vanilla Mixup method that samples the ratio from the Beta distribution, the present invention adjusts the mixing degree according to instance similarity, assigns a higher mixing ratio to more similar instances to generate more informative samples. On the contrary, for less similar instances, a lower mixing degree is adopted to avoid generating out-of-distribution data.
[0025] Therefore, the present invention proposes a new method for controlling the Mixup ratio in the framework, namely Local Hierarchy Mixup (LH-Mix). LH-Mix evaluates the correlation between local hierarchy pairs by using the representations extracted from the hierarchical prompts, and can learn the correlation between hierarchical labels more effectively compared with the basic Mixup method. Brief Description of the Drawings
[0026] The above and other objects and features of the present disclosure will become more apparent from the following description in conjunction with the drawings.
[0027] Figure 1 is a diagram of hierarchical text classification in the prior art, where (a) is the global layer, (b) is the local hierarchy, and (c) is the label inclusion relationship;
[0028] Figure 2 is a schematic diagram of the architecture of the hybrid augmentation method guided by local hierarchy label sequence similarity according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram showing the relationship between s and λ according to an embodiment of the present invention. Detailed Description of the Embodiments
[0030] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of the present application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0031] The features described herein can be implemented in various forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.
[0032] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.
[0033] Although terms such as "first", "second", and "third" may be used herein to describe various components, components, regions, layers, or parts, these components, components, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, component, region, layer, or part from another. Thus, a first component, first component, first region, first layer, or first part referred to in the examples described herein may also be referred to as a second component, second component, second region, second layer, or second part without departing from the teachings of the examples.
[0034] In the specification, when an element (such as a layer, region, or substrate) is described as "on", "connected to", or "coupled to" another element, the element may be directly "on", directly "connected to", or "coupled to" the other element, or there may be one or more other elements in between. In contrast, when an element is described as "directly on", "directly connected to", or "directly coupled to" another element, there may be no other elements in between.
[0035] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0036] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and should not be interpreted in an idealized or overly formal manner.
[0037] In addition, in the description of the examples, when a detailed description of a related structure or function that is considered to be well-known will cause an ambiguous interpretation of the present disclosure, such a detailed description will be omitted.
[0038] Figure 2 It is a schematic diagram showing a hybrid enhancement method guided by local hierarchical label sequence similarity according to an embodiment of the present disclosure.
[0039] To achieve the above-mentioned invention object, the technical framework adopted by the present invention is as Figure 2 shown.
[0040] HTC setting:
[0041] Given a training data set where X i represents the input text, Y i represents the label corresponding to X i and N is the size of the training data set. Let be the label set. Then, in the HTC problem setting, is divided into D different subsets, denoted as where D is the total depth of the entire hierarchical structure, is the label set at the d-th depth. can be further organized into a tree structure, representing the global hierarchical structure of the entire data set. In addition, each input X i contains multiple labels from , and these labels can form an independent subtree, called the local hierarchical structure. The goal of HTC is to assign corresponding labels to each text in the test data set from .
[0042] Hierarchical hint of HTC:
[0043] Generally, a hint contains a template for a specific task to utilize the knowledge embedded in a pre-trained language model (PLM). In HTC, the basic hint box creates D soft 'tokens' in the template, and each 'token' specifically corresponds to one depth in the hierarchical structure. For example, the text input can be constructed in the following form:
[0044] [CLS][Dth 1 [MASK]...[Dth d [MASK]...[Dth D [MASK][SEP]X[SEP]
[0045] where each 'token' [Dth dused to prompt the prediction of the label in its corresponding d - layer hierarchy and use the hidden output of the [MASK] 'token' immediately following it for label prediction. Specifically, let denote the hidden output of [MASK] in the d - th layer, then the classification process is defined as:
[0046]
[0047] where is the prediction score vector, is the mapping function that maps the hidden output to the prediction scores at the d - th layer. It should be noted that contains a classifier trained through the Masked Language Model (MLM) task and a mapper for the label words that conform to Through this hierarchical hint of HTC, the classification process is assigned to each layer of the hierarchical structure, rather than predicting all labels through a single classifier. This method encodes the local hierarchy as a sequence, thus allowing the alignment of the hierarchical structure for each input at each depth level, which greatly facilitates the subsequent Mixup operation.
[0048] In the HTC task, many existing models regard the task as a multi - label classification problem and adopt the traditional Binary Cross Entropy (BCE) loss. However, some studies have pointed out that the BCE loss ignores the correlation between labels and proposed the Zero - bounded Multi - label Cross Entropy (ZMLCE) loss. The ZMLCE loss specifically considers that the score of the positive label should be greater than 0, while the score of the negative label should be less than 0. The formal definition of the ZMLCE loss is as follows:
[0049]
[0050] where and represent the set of positive labels and the set of negative labels in the d - th layer respectively, and ·v represents the v - th element of the vector (for example, refers to the v - th element of the prediction score vector). In the inference stage, if the prediction score is greater than 0, it is regarded as a positive label; otherwise, it is regarded as a negative label.
[0051] LH - Mix framework:
[0052] The Mixup method generates intermediate samples by linearly interpolating the input and its corresponding labels, thereby enhancing the internal structure in the latent space. The present invention integrates Mixup into the hierarchical prompting framework to better capture the correlation between hierarchical structure labels in HTC. Specifically, the present invention proposes a new adaptive Mixup ratio strategy guided by local hierarchical correlation, called LH-Mix, to further optimize the basic Mixup method (VanillaMixup). The LH-Mix framework diagram is as shown in Figure 2 shown.
[0053] Local hierarchical association:
[0054] An important feature of the hierarchical structure is its strict hierarchical relationship, where the prediction of lower-level labels depends on the prediction of higher-level labels. This feature allows the local hierarchical sequence to be represented as a sequence. Therefore, the present invention introduces a depth identifier [Dth d to represent the hierarchical depth of each label. In this way, the local hierarchy can be re-represented as a sequence of identifiers and further transformed into a soft identifier form through a pre-trained language model as a hierarchical prompt. This soft prompt can combine the hierarchical depth and label information, thereby allowing the hierarchical structure to be represented as a sequence.
[0055] More specifically, for a given input X, its corresponding local hierarchy can also be represented as a sequence in a similar way by replacing the [MASK] identifier with the corresponding true label. For example, Figure 1 the two local hierarchies "CS / Machine Learning" (Computer Science / Machine Learning) and "Math / Statistics" (Mathematics / Statistics) in can be represented as:
[0056] [CLS][Dth 1 CS[Dth 2 Machine Learning[SEP]
[0057] [CLS][Dth 1 Math[Dth 2 Statistics[SEP]
[0058] In this way, the local hierarchy can be transformed into a form similar to a "sentence". In addition, when a local hierarchy contains multiple paths, the labels at the same depth level are concatenated together and placed in the corresponding [Dth 2After the identifier. Based on this, by inputting the sentence into the pre-trained language model and extracting the hidden output of the [CLS] identifier, the embedding representation of the sentence can be obtained, which can be used to calculate the sentence similarity. Similarly, in the HTC task, the [CLS] output corresponding to the "local hierarchical sequence" can be used as the local hierarchical representation, effectively used to calculate the correlation between local hierarchies.
[0059] Formally, assume that the corresponding local hierarchical representations of two inputs X i and X j are h i CLS and h j CLS , respectively. Then the distance between these two hierarchies is measured by the normalized cosine similarity, and its calculation formula is:
[0060]
[0061] where, · represents the vector dot product operation, and |·| is the L 2 norm. It should be noted that since the hierarchical encoding format is shared, when calculating the local hierarchical representation, the same pre-trained model encoder is used, but the relevant encoder is not affected by the gradient update during the similarity calculation.
[0062] Mixup ratio guided by local hierarchical correlation:
[0063] Mixup captures the correlation between samples by generating intermediate samples. By adjusting the size of the Mixup ratio, "difficult samples" of different degrees can be generated, thereby controlling the influence range of Mixup on the data. Research shows that samples with different similarity degrees should be mixed with different intensities. Therefore, for the problem of the correlation difference between different local hierarchical pairs, different from the way of sampling the Mixup ratio from a fixed distribution (such as the Beta distribution) in the Vanilla Mixup method, LH-Mix applies different Mixup ratios to each local hierarchical sample pair.
[0064] Although many studies have explored the relationship between sample correlation and Mixup ratio, there is currently no perfect theoretical framework to precisely characterize the numerical relationship between the two. In fact, strictly verifying the effectiveness of Mixup remains an open question. Therefore, to reflect the relationship between Mixup ratio and similarity, the present invention intuitively designs a heuristic function: when the similarity between two local levels increases, it is expected that Mixup can better capture the potential correlation between them. Specifically, for highly correlated local levels, it is hoped that the impact of Mixup is greater (i.e., the Mixup ratio is close to 0.5); while for less correlated local levels, the impact of Mixup should be weakened (i.e., the Mixup ratio is close to 1).
[0065] Based on this idea, the present invention heuristically designs the following function to characterize the relationship between local level similarity s and Mixup ratio λ:
[0066] λ = -(β - 0.5)s α + β
[0067] where α > 0 controls the rate of change of λ with s, and β ∈ (0.5, 1] controls the upper limit value of λ.
[0068] The advantage of this function is that by simply adjusting the values of α and β, common linear or non-linear relationships between s and λ can be covered.
[0069] To better understand the role of the above formula, Figure 3 the influence of parameters β and α on the final mixing ratio is visualized. Specifically, when α = 1, a linear relationship exists between s and λ. When α < 1, as s increases, the rate of decrease of λ slows down. On the contrary, when α > 1, as s increases, the rate of decrease of λ speeds up. In addition, β determines the maximum value of λ, indicating the minimum impact that Mixup can exert.
[0070] Local level Mixup:
[0071] At each depth of the level, both the input and output are mixed using the Mixup ratio guided by local level correlation. For input Mixup, referring to some Mixup variants in previous text classification, at each depth of the level, interpolation is performed on the hidden output corresponding to the [MASK] identifier, and its formula is as follows:
[0072]
[0073] Subsequently, through the mixed input the prediction score is obtained That is:
[0074]
[0075] For the output mixture, Vanilla Mixup directly mixes the labels. There is also research showing that directly mixing the losses is also an effective method, that is, in the context of the cross-entropy loss, the gradients of label mixing and loss mixing are equivalent.
[0076] In the context of the ZMLCE loss, since it separates positive and negative labels centered around 0, both positive and negative labels are regarded as two independent combinations. Although this design enables ZMLCE to focus on the correlation between labels, it ignores the relative magnitude relationship between positive or negative labels, resulting in the interpretability of label mixing becoming less clear. Therefore, in the LH-Mix framework, the mixing loss term is selected, and its definition is as follows:
[0077]
[0078] In the context of cross-entropy, Vanilla Mixup also generates the gradient of the mixed prediction score through linear combination. Therefore, it is more in line with the design logic of the basic Mixup to use loss mixing for ZMLCE.
[0079] Experimental results:
[0080] Datasets and evaluation methods:
[0081] The present invention evaluates the model on three widely used datasets: WebOfScience (WOS), NYTimes (NYT), and RCV1-V2. The statistical information of these datasets is shown in Table 1. In terms of evaluation metrics, following previous research, Macro-F1 and Micro-F1 are used to measure the results. Micro-F1 calculates the overall precision and recall rate of all instances, while Macro-F1 represents the average of the F1 scores between labels.
[0082] Table 1: Dataset statistical information
[0083]
[0084] Experimental details:
[0085] Consistent with previous research, the present invention uses the pre-trained model bert-base-uncased provided by Hugging Face Transformer to train the model. In the hierarchical prompting mechanism, the newly added depth identifier [Dth d is randomly initialized, while the label verbalizer is initialized with the average value represented by the label name. All parameters are fine-tuned using Adam optimization with a learning rate of 3e-5.
[0086] Regarding the parameters related to LH-Mix, such as Figure 3 As shown, select α from [0.1, 0.3, 0.6, 1, 2, 5, 10] and select β from [0.7, 0.75, 0.8, 0.85, 0.9, 0.95, 1]. This configuration can cover most common relationships between s and λ. To accelerate the convergence of the model, a two-step training strategy is adopted. First, train the model for 5 epochs without using Mixup, and then introduce Mixup in the subsequent training process. In addition, an early stopping strategy is also adopted, that is, if Macro-F1 does not improve in 5 epochs, stop training.
[0087] Baseline models:
[0088] The present invention compares LH-Mix with three groups of models: hierarchical perception models, large language models, and pre-trained language models. Among the hierarchical perception models, the present invention selects 4 strong baseline models for comparison: TextRNN, HiAG, HTCInfoMax, and HiMatch. In terms of large language models, the present invention mainly reports the instruction-tuned results of ChatGPT and the supervised fine-tuned results of LLaMA-2-7B. In terms of pre-trained language models, in addition to replacing the encoders of the above 4 hierarchical perception models with BERT for comparison, the present invention also incorporates 5 newly proposed models: HGCL, HP, HBGL, HJC, and HiTIN. Among these baseline models, HPT and HBGL generally achieve the current state-of-the-art performance.
[0089] Main results:
[0090] The main experimental results of Micro-F1 and Macro-F1 on three datasets are shown in Table 2. From the data in the table, it can be seen that LH-Mix achieved the best performance in five out of a total of six metrics, which demonstrates the effectiveness of LH-Mix.
[0091] Table 2: Model comparison results
[0092]
[0093] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.
Claims
1. A hybrid enhancement method guided by local hierarchical tag sequence similarity, characterized in that, For the input text and tags, a local hierarchical association method is adopted. Depth identifiers are introduced to represent the hierarchical depth of each tag. The local hierarchy is represented as a sequence of identifiers and further transformed into a soft identifier form through a pre-trained language model as a hierarchical hint. At each depth of the hierarchy, both the input and output are mixed using a Mixup ratio guided by local hierarchical correlation to generate intermediate samples, and then hierarchical label representations are generated.
2. The hybrid enhancement method guided by local hierarchical tag sequence similarity as described in claim 1, wherein The specific method for introducing depth identifiers to represent the hierarchical depth of each tag is as follows: For a given input X, its corresponding local hierarchy is represented as a sequence in a similar way by replacing the [MASK] identifier with the corresponding real label. When a local hierarchy contains multiple paths, the labels at the same depth level are concatenated together and placed after the corresponding depth identifier. By inputting the sentence into the pre-trained language model and extracting the hidden output of the [CLS] identifier, the embedding representation of the sentence is obtained, and this representation is used to calculate the sentence similarity. Formally, assume two inputs X i and X j whose corresponding local hierarchical representations are h i CLS and h j CLS , respectively. Then the distance between these two hierarchies is measured by the normalized cosine similarity, and its calculation formula is: where, · represents the vector dot product operation, and |·| is the L 2 norm.
3. The hybrid enhancement method guided by local hierarchical tag sequence similarity as claimed in claim 2, wherein The method for calculating the sentence similarity is: Design a heuristic function to characterize the relationship between the local hierarchical similarity s and the Mixup ratio λ: λ = -(β - 0.5)s α + β where α > 0 controls the rate of change of λ with s, and β ∈ (0.5, 1] controls the upper limit value of λ.
4. The hybrid enhancement method guided by local hierarchical label sequence similarity as described in claim 3, wherein The specific method for mixing both the input and output using a Mixup ratio guided by local hierarchical correlation at each depth of the hierarchy is as follows: For input Mixup, interpolation is performed on the hidden output corresponding to the [MASK] identifier at each depth of the hierarchy. Subsequently, by obtaining a prediction score for the mixed input That is: For the mixing of the output, a mixed loss term is selected.
5. The hybrid enhancement method based on local hierarchical tag sequence similarity guidance according to claim 3, characterized in that, When α = 1, there is a linear relationship between s and λ. When α < 1, as s increases, the rate of decrease of λ slows down; on the contrary, when α > 1, as s increases, the rate of decrease of λ speeds up. In addition, β determines the maximum value of λ, representing the minimum influence that Mixup can exert.
6. The hybrid enhancement method guided by local hierarchical tag sequence similarity as claimed in claim 5, wherein, Select α from [0.1, 0.3, 0.6, 1, 2, 5, 10] and select β from [0.7, 0.75, 0.8, 0.85, 0.9, 0.95, 1].