A dynamic graph disambiguation-based partial multi-label data classification method and system

By constructing a triple reliable label set and a multi-scale fusion graph, and dynamically adjusting the label confidence, the problems of high-frequency noise interference and insufficient influence of rare labels in existing technologies are solved, thereby improving the classification accuracy of data with multiple labels.

CN121705850BActive Publication Date: 2026-05-12GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG UNIV OF TECH
Filing Date
2026-02-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively suppress high-frequency noise interference and enhance the influence of rare discriminative labels when processing multi-labeled data. This results in graph structures that are biased towards association patterns of common labels, limiting the upper limit of the accuracy of classification models.

Method used

By constructing a triple reliable label set, a weighted similarity is built based on the occurrence frequency of candidate biased labels, discriminative scores, and dependencies between instances. A multi-scale fusion graph is used, and a graph propagation mechanism with dynamic adjustment of label confidence is introduced to perform progressive multi-stage training to purify the label matrix.

Benefits of technology

It significantly improved the accuracy of the classification model, with an average accuracy increase of 9.2% and a ranking loss reduction of 8.7%, effectively solving the problems of feature-semantic ambiguity, structured noise interference, and uneven label contribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121705850B_ABST
    Figure CN121705850B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, more particularly, to a kind of dynamic graph disambiguation-based partial multi-label data classification method and system, three reliable label sets are constructed according to output, confidence, consistency, weighted similarity between instances is constructed based on the occurrence frequency of candidate partial multi-label in three reliable label sets, the discriminative score of candidate partial multi-label, the dependency between instances, structured noise is suppressed using weighted similarity, multi-scale similarity graph is constructed and adaptively fused, a graph propagation mechanism is introduced by introducing label confidence dynamic adjustment, and the collaborative evolution of classifier and label quality is realized by progressive multi-stage training. The method effectively solves the problems of feature-semantic ambiguity, structured noise interference, label contribution imbalance, single-scale graph limitation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for classifying partially labeled data based on dynamic graph disambiguation. Background Technology

[0002] In data processing tasks, multi-label learning aims to identify multiple semantic labels from a single sample. However, training data in real-world scenarios often contains a large amount of noise due to the subjectivity, omissions, or errors in manual annotation, resulting in "biased multi-label" data, where the candidate labels associated with each sample are mixed with some erroneous labels. Efficiently cleaning, correcting, and extracting information from this incomplete and inaccurate data is a key preprocessing step for improving the performance of downstream classification models, and belongs to a typical and challenging area of ​​data processing technology.

[0003] Existing multi-label learning methods often employ graph-based data processing frameworks. They construct similarity graphs between samples and propagate label confidence along the edges to filter out noisy labels and recover true labels. However, these methods suffer from a fundamental flaw: the construction of their graph structure and subsequent similarity calculations are often dominated by frequently occurring (but potentially low-discriminative, or even structured noise) labels. During data processing, the contribution of rare labels, which may have higher class discriminative power but appear less frequently, is severely weakened. This results in a graph structure that essentially reflects the association patterns of common labels rather than the true semantic structure of the data. Consequently, this leads to a systematic bias in the label cleansing and confidence estimation results, limiting the upper limit of the final classification model's accuracy. Therefore, effectively suppressing the interference of high-frequency noise and enhancing the influence of rare discriminative labels during multi-labeling has become a critical technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies in effectively suppressing high-frequency noise interference and fully enhancing the influence of rare discriminative labels when performing biased labeling on data. This invention provides a biased labeling data classification method and system based on dynamic graph disambiguation, which can effectively suppress high-frequency noise interference and enhance the influence of rare discriminative labels when performing biased labeling on data.

[0005] According to one aspect of the present invention, a method for classifying partially labeled data based on dynamic graph disambiguation is provided, the method comprising the following steps:

[0006] S1: Obtain a partially labeled dataset for training, and obtain an instance feature matrix and a candidate partially labeled matrix based on the partially labeled dataset; the partially labeled dataset is derived from natural scene images or medical images;

[0007] S2: Based on the instance feature matrix and the candidate partial multi-label matrix, pre-train the initial classifier;

[0008] S3: Obtain the initial prediction model based on the initial classifier and the first activation function;

[0009] S4: The initial prediction model is trained iteratively multiple times. During the training process, the candidate biased label matrix is ​​purified to obtain the optimal prediction model after training. This includes at least the following sub-steps:

[0010] The initial prediction model is trained iteratively multiple times. In the t-th iteration, a triple reliable label set is constructed based on the prediction model of the (t-1)-th iteration, according to the output, confidence, and consistency. The dynamic sparse weight of the candidate biased label is obtained based on the occurrence frequency of the candidate biased label in the triple reliable label set and the discriminative score of the candidate biased label. The weighted similarity between instances is constructed based on the triple reliable label set, the dynamic sparse weight, and the dependency between instances. A multi-scale fusion graph is constructed based on the weighted similarity. A cleaned soft label matrix is ​​obtained based on the multi-scale fusion graph. The cleaned soft label matrix is ​​used as the output of the t-th iteration to adjust the parameters of the prediction model in the t-th iteration, thus obtaining the prediction model in the t-th iteration. After training, the optimal prediction model after training is obtained.

[0011] S5: Obtain test instances, which are derived from natural scene images or medical images, and classify the test instances based on the optimal prediction model.

[0012] In an alternative approach, in step S1, the feature vector of each instance in the instance feature matrix is ​​mapped to the candidate biased label vector at the corresponding position in the candidate biased label matrix; the candidate biased label matrix includes true labels and erroneous inaccurate labels.

[0013] In an alternative approach, step S4, which involves constructing a triple reliable label set based on the t-1 round prediction model according to the output, confidence level, and consistency, includes at least the following sub-steps:

[0014] Based on the t-1 round prediction model, the multiple candidate biased labels with the highest output scores in the t-1 round prediction model are taken as a local relatively reliable set;

[0015] Based on the t-1 round prediction model, the multiple candidate biased labels with the highest confidence scores in the t-1 round prediction model are used as the global high-confidence anchor set;

[0016] Based on the t-1 round prediction model, the multiple candidate biased labels with the highest consistency scores in the t-1 round prediction model are taken as the label consistency set;

[0017] The union of the locally relatively reliable set, the globally high-confidence anchor set, and the label consistency set yields the triple reliable label set.

[0018] In an alternative approach, in step S4, the discriminative score of the candidate biased label is obtained based on the importance analysis of classifier instance features.

[0019] In an alternative approach, obtaining the dependencies between instances in step S4 includes at least the following sub-steps:

[0020] A co-occurrence matrix is ​​obtained based on the triple reliable label set and the occurrence frequency of candidate predominant labels within the triple reliable label set; the label co-occurrence matrix is ​​used to capture the pairwise correlation strength between candidate predominant labels.

[0021] The dependencies between instances are obtained based on the marker co-occurrence matrix and the triple reliable marker set; when two instances share candidate predominant markers... j If these two instances also share additional candidate-biased tags... j Other candidate markers l The greater the dependency between the two instances, the stronger the relationship between them.

[0022] In one alternative approach, step S4, which involves constructing a multi-scale fusion graph based on the weighted similarity, includes at least the following sub-steps:

[0023] A local fine-grained map is constructed based on the weighted similarity using a proximity algorithm.

[0024] A global manifold graph is constructed based on the weighted similarity using optimal transport theory;

[0025] A multi-scale fusion graph is constructed based on the local fine-grained graph and the global manifold graph.

[0026] In an alternative approach, step S4, which involves constructing the global manifold graph based on the weighted similarity using optimal transport theory, includes at least the following sub-steps:

[0027] The cost matrix is ​​obtained based on the weighted similarity.

[0028] The dynamic supply distribution is obtained based on the t-1 round prediction model, the candidate biased label matrix, the triple reliable label set, and the dynamic sparse weight.

[0029] The dynamic demand distribution is obtained based on the t-1 round prediction model, the candidate biased label matrix, the triple reliable label set, and the dynamic sparse weight.

[0030] Based on the cost matrix, the dynamic supply distribution, and the dynamic demand distribution, a global manifold is constructed using entropy regularization for optimal transmission.

[0031] In an alternative approach, step S4, obtaining the cleaned soft label matrix based on the multi-scale fusion graph, includes at least the following sub-steps:

[0032] The confidence adjustment factor is obtained based on the t-1 round prediction model, the confidence nonlinear adjustment parameter, the reliability label enhancement coefficient, and the triple reliable label set;

[0033] The purified soft label matrix is ​​obtained based on the multi-scale fusion graph, the t-1 round prediction model, the confidence adjustment factor, and the candidate biased label matrix.

[0034] In an alternative approach, in step S4, the parameters of the prediction model for the t-th round are adjusted using the purified soft-label matrix as the output of the t-th iteration to obtain the t-th round prediction model, which includes at least the following sub-steps:

[0035] Using the purified soft-label matrix as the output of the t-th iteration, the parameters of the prediction model in the t-th iteration are adjusted using the gradient descent method to obtain the t-th fine-tuned prediction model;

[0036] The cleanup hard label matrix is ​​obtained based on the cleanup soft label matrix;

[0037] The t-round fine-tuning prediction model is trained using the binary cross-entropy loss of the cleaned soft-label matrix and the cleaned hard-label matrix to obtain the t-round prediction model.

[0038] According to a second aspect of the present invention, a partial label data classification system based on dynamic graph disambiguation is provided, for implementing the above-described partial label data classification method based on dynamic graph disambiguation, comprising:

[0039] Data processing unit: used to obtain instance feature matrix based on the partially labeled dataset;

[0040] Initial classifier construction unit: used to pre-train an initial classifier based on the instance feature matrix and the candidate partial multi-label matrix;

[0041] Initial prediction model building unit: used to obtain an initial prediction model based on the initial classifier and the first activation function;

[0042] Model Iteration Training Unit: Used to perform multiple iterations of training on the initial prediction model, and to purify the candidate biased label matrix during the training process to obtain the optimal prediction model after training;

[0043] Label prediction unit: used to classify the test instance based on the optimal prediction model.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] This invention constructs a triple reliable label set based on output, confidence, and consistency. A weighted similarity between instances is built based on the frequency of occurrence of candidate overlabeled labels within this set, their discriminative scores, and dependencies. This weighted similarity is used to suppress structured noise, and a multi-scale similarity graph is constructed and adaptively fused. A graph propagation mechanism with dynamic adjustment of label confidence is introduced, and progressive multi-stage training achieves the co-evolution of classifier and label quality. This method effectively solves problems such as feature-semantic ambiguity, structured noise interference, imbalanced label contributions, and limitations of single-scale graphs. It significantly outperforms state-of-the-art methods on multiple benchmark datasets, improving average accuracy by 9.2% and reducing ranking loss by 8.7%. This invention is applicable to various overlabeled data classification scenarios, including natural scene images, medical images, and e-commerce products.

[0046] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating the partial labeling data classification method based on dynamic graph disambiguation of the present invention.

[0048] Figure 2 This is a flowchart illustrating the sub-step of step S4 in the partial label data classification method based on dynamic graph disambiguation of the present invention.

[0049] Figure 3 This is a schematic diagram of the sub-step process for constructing a triple reliable label set in the partial multi-label data classification method based on dynamic graph disambiguation of the present invention. Detailed Implementation

[0050] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0051] Example 1

[0052] This embodiment is the first embodiment of a partially labeled data classification method based on dynamic graph disambiguation, such as... Figure 1 As shown, it includes the following steps:

[0053] S1: Obtain a partially labeled dataset for training. The partially labeled dataset is derived from natural scene images or medical images. Obtain the instance feature matrix based on the partially labeled dataset. and candidate biased label matrix The feature vector of each instance in the instance feature matrix maps to the candidate biased label vector at the corresponding position in the candidate biased label matrix; the candidate biased label matrix includes true labels and erroneous inaccurate labels.

[0054] Among them, the instance feature matrix Specifically, it can be expressed as the following formula:

[0055]

[0056] In the formula, ; Indicates the first The feature vector of each instance n represents the number of instance vectors; d represents the dimension of the feature vectors.

[0057] Candidate biased label matrix Specifically, it can be expressed as the following formula:

[0058]

[0059] In the formula, ; Indicates the first The candidate predominantly labeled vectors of the nth instance, further... The first instance j The class candidate predominant marker value is , Indicates the first The instance was labeled as the first. j Candidate predominant tags for each tag category, Then it means the first One instance was not labeled; q represents the number of candidate biased label categories; This indicates transpose.

[0060] S2: Pre-train the initial classifier f based on the instance feature matrix and the candidate partial label matrix.

[0061] Specifically, the initial classifier uses a multilayer perceptron (MLP) structure, comprising an input layer (dimension equal to the dimension of the feature vectors corresponding to the instance feature matrix), an output layer (dimension equal to the number of candidate biased label classes corresponding to the candidate biased label matrix), and two hidden layers (dimensions of 512 and 256, respectively). The hidden layer activation function is ReLU. The initial classifier can be pre-trained for up to 200 epochs using a weighted binary cross-entropy loss function. During training, different weights are assigned to candidate biased labels of different classes to handle class imbalance. Specifically, the weight for each class is the total number of instances divided by the total number of 1-valued labels in that class. The initial classifier parameters are obtained after training. ;

[0062] S3: Obtain the initial prediction model based on the initial classifier and the first activation function. Specifically, the first activation function is the sigmoid activation function. Initial prediction model Specifically, it can be expressed as the following formula:

[0063]

[0064] In the formula, This represents the sigmoid activation function; .

[0065] S4: Iterate the training of the initial prediction model multiple times, and clean up the candidate biased label matrix during the training process to obtain the optimal prediction model after training. This includes at least the following sub-steps, such as... Figure 2 As shown:

[0066] The initial prediction model is trained iteratively multiple times. In the t-th iteration, a triple reliable label set is constructed based on the prediction model from the (t-1)-th iteration, according to the output, confidence, and consistency. Dynamic sparse weights of candidate predominant labels are obtained based on their frequency of occurrence within the triple reliable label set and their discriminative scores. Weighted similarity between instances is constructed based on the triple reliable label set, dynamic sparse weights, and dependencies between instances. A multi-scale fusion graph is constructed based on the weighted similarity. A cleaned soft label matrix is ​​obtained based on the multi-scale fusion graph. The cleaned soft label matrix is ​​used as the output of the t-th iteration to adjust the parameters of the prediction model in the t-th iteration, resulting in the t-th prediction model. After training, the optimal prediction model model(x) is obtained. ;

[0067] S5: Obtain test instances, which are derived from natural scene images or medical images, and classify the test instances based on the optimal prediction model.

[0068] Specifically, feature extraction is performed to obtain the feature vector of the test instance. By inputting the feature vector of the instance to be tested into the optimal prediction model, its soft label vector can be predicted. Specifically, it can be expressed as the following formula:

[0069]

[0070] Finally, a binary label is obtained based on its soft label vector. Specifically, it can be expressed as the following formula:

[0071]

[0072] This embodiment presents a method for classifying partially labeled data based on dynamic graph disambiguation:

[0073] This method constructs a triple reliable label set based on output, confidence, and consistency. It then builds a weighted similarity between instances based on the frequency of candidate overlabeled labels within this set, their discriminative scores, and dependencies. This weighted similarity is used to suppress structured noise. A multi-scale similarity graph is constructed and adaptively fused. A graph propagation mechanism with dynamically adjusted label confidence is introduced. Progressive multi-stage training enables the co-evolution of classifier and label quality. This method effectively addresses issues such as feature-semantic ambiguity, structured noise interference, imbalanced label contributions, and limitations of single-scale graphs. It significantly outperforms state-of-the-art methods on multiple benchmark datasets, achieving a 9.2% improvement in average accuracy and an 8.7% reduction in ranking loss. This invention is applicable to various overlabeled data classification scenarios, including natural scene images, medical images, and e-commerce products.

[0074] Example 2

[0075] This embodiment is a second embodiment of a partially labeled data classification method based on dynamic graph disambiguation. This embodiment is similar to the first embodiment, except that the detailed steps of step S4 are as follows:

[0076] S41: Train the initial prediction model iteratively multiple times. In the t-th iteration, based on the prediction model from the (t-1)-th iteration... Construct a triple reliable tag set for the t-th iteration based on output, confidence level, and consistency. Specifically, such as Figure 3 As shown:

[0077] S411: Based on the intra-instance relative reliability assumption, each instance contains at least one true label. Based on the t-1 round prediction model, the highest-scoring output from the t-1 round prediction model is selected. The candidate biased labels are used as a locally relatively reliable set. Specifically, it can be expressed as the following formula:

[0078]

[0079] In the formula, This indicates that the K candidate labels with the highest output scores are selected. , This indicates rounding down. This indicates the local selection scaling parameter. Locally Relatively Reliable Sets Ensure that each sample contributes at least one reliable label, and select multiple high-scoring labels for samples with a large number of candidate labels.

[0080] S412: Based on the global high confidence assumption. Based on the t-1 round prediction model, the N candidate biased labels with the highest confidence scores in the t-1 round prediction model are used as the global high confidence anchor set. Specifically, it can be expressed as the following formula:

[0081]

[0082] In the formula, This indicates that the N candidate labels with the highest confidence scores are selected as the most frequently used labels. , This indicates the global selection of the scaling parameter. .

[0083] S413: Introduce a label consistency evaluation mechanism. Based on the t-1 round prediction model, the multiple candidate biased labels with the highest consistency scores in the t-1 round prediction model are used as the label consistency set. Specifically, it can be expressed as the following formula:

[0084]

[0085] In the formula, KN For example of NN (K-Nearest Neighbor) nearest neighbor set, Represents the first feature-based The cosine similarity between the z-th instance and the z-th instance. Represents the normalization constant; Indicates the similarity threshold. It can generally be set to 0.6.

[0086] S414: Locally Reliable Sets Global high-confidence anchor set and the set of labels consistent The union of the three reliable labels in the t-th iteration is obtained. Specifically, it can be expressed as the following formula:

[0087]

[0088] In the formula, In the t-th iteration, the first... The first instance j Triple reliable label values ​​for class candidate biased labels.

[0089] S42: Obtain the dynamic sparse weights and weighted similarities of the candidate predominantly labeled tags. Specifically:

[0090] S421: Obtain discriminative scores for candidate predominant labels based on classifier instance feature importance analysis; specifically, for the first... j The class candidate has a bias towards multiple labels, and the classifier for round t-1 is obtained based on the classifier parameters obtained in round t-1 iteration. Calculate the importance gradient of the classifier in round t-1 to the instance feature matrix. The average L2 norm of the importance gradients of all instances is used as the feature importance. , feature importance After normalization, the first number is obtained. j Discriminative score of class with more candidate labels .

[0091] Among them, the importance gradient of the classifier in round t-1 to the instance feature matrix Specifically, it can be expressed as the following formula:

[0092]

[0093] In the formula, X represents the instance feature matrix.

[0094] Among them, feature importance Specifically, it can be expressed as the following formula:

[0095]

[0096] In the formula, Indicates the first The first instance j The importance gradient of class candidate labels with more labels; n represents the number of instance vectors.

[0097] S422: The dynamic sparse weight of candidate-dominant markers is obtained based on the frequency of occurrence of candidate-dominant markers within the triple reliable marker set and the discriminative score of candidate-dominant markers. Specifically, it is expressed as follows:

[0098]

[0099] In the formula, In the t-th iteration, the first... j Dynamic sparse weights for class-preferred-candidate labels; In the t-th iteration, the first... j Class candidate predominance markers Frequency of occurrence within, ; This represents the frequency penalty intensity parameter. The penalty is stronger for high-frequency marking; Indicates the first j Discriminative score of class with more candidate labels.

[0100] By designing dynamic rarity weights, rare but highly discriminative tags are given higher weights, while high-frequency but low-discriminative tags (which may be structured noise) are effectively suppressed.

[0101] S423: Obtain the marker co-occurrence matrix for the t-th iteration based on the triple reliable marker set and the occurrence frequency of candidate predominant markers within the triple reliable marker set. The co-occurrence matrix is ​​used to capture the pairwise association strength between candidate predominantly labeled tags; specifically, it is expressed as follows:

[0102]

[0103] In the formula, In the t-th iteration, the first... j The co-occurrence matrix of candidate markers of class m and candidate markers of class m. The candidate markers of class m are those excluding the first... j Other candidate tags besides the class with a high proportion of candidate tags.

[0104] S424: Obtain the dependencies between instances based on the tag co-occurrence matrix and the triple reliable tag set; when two instances share candidate predominant tags... j If these two instances also share additional candidate-biased tags... j Other candidate markers l The greater the dependency between the two instances, the stronger the dependence. Specifically, this can be expressed as follows:

[0105]

[0106] In the formula, Indicates that the t-th iteration shares the th... j The first class of candidate-dominated markers i The dependency between the first instance and the kth instance; Indicates the dependency enhancement coefficient; In the t-th iteration, the first... j Class candidate biased markers and the first l The co-occurrence matrix of candidate markers with a high prevalence; In the t-th iteration, the first... i The first instance lTriple reliable label values ​​for class with a high proportion of candidate labels; This represents the k-th instance in the t-th iteration. l Triple reliable label values ​​for class candidate biased labels.

[0107] S425: Construct a weighted similarity between instances in the t-th iteration based on a triple reliable tag set, dynamic sparse weights, and dependencies between instances. Specifically, it can be expressed as follows:

[0108]

[0109] In the formula, In the t-th iteration, the first... i The weighted similarity between the kth instance and the kth instance; This represents the k-th instance in the t-th iteration. j Triple reliable label values ​​for class candidate biased labels.

[0110] S43: Constructing a multi-scale fusion graph based on weighted similarity. Specifically:

[0111] S431: Constructing a fine-grained local map for the t-th iteration using a neighbor-to-neighbor algorithm based on weighted similarity. Specifically, it can be expressed as the following formula:

[0112]

[0113] In the formula, In the t-th iteration, the first... i A detailed local graph of the kth instance and the kth instance.

[0114] S432: Construct a global manifold graph based on weighted similarity using optimal transport theory.

[0115] Specifically, the cost matrix for the t-th iteration is obtained based on weighted similarity. Specifically, it can be expressed as follows:

[0116]

[0117] In the formula, In the t-th iteration, the first... i Cost matrix between the kth instance and the kth instance.

[0118] The dynamic supply distribution of the t-th iteration is obtained based on the t-1 round prediction model, the candidate biased label matrix, the triple reliable label set, and the dynamic sparse weight. Specifically, it can be expressed as follows:

[0119]

[0120] In the formula, In the t-th iteration, the first... i Dynamic supply distribution for each instance; Indicates proportional to; In the (t-1)th iteration, the... i The first instance j The output value of the prediction model with a predominance of candidate labels; Let represent the set of unreliable tags in the t-th iteration. , In the t-th iteration, the first... i The first instance j Unreliable label values ​​for classes with a predominance of candidate labels.

[0121] The dynamic demand distribution of the t-th iteration is obtained based on the t-1 round prediction model, the candidate biased label matrix, the triple reliable label set, and the dynamic sparse weight. Specifically, it can be expressed as follows:

[0122]

[0123] In the formula, In the t-th iteration, the first... i Dynamic demand distribution for each instance.

[0124] Normalize the dynamic supply distribution and the dynamic demand distribution respectively, so that , .

[0125] Based on the cost matrix, dynamic supply distribution, and dynamic demand distribution, a global manifold graph for the t-th iteration is constructed using entropy regularization for optimal transport. Specifically, it can be expressed as the following formula:

[0126]

[0127] In the formula, This represents the global manifold graph in the t-th iteration; arg min represents the parameter that takes the minimum value. Represents the set of transmission plan constraints. P represents the dynamic distribution; Represents the Frobenius inner product; Represents the regularization parameter. Solve efficiently using the Sinkhorn algorithm; Represents the entropy regularization term. , Indicates the first i The dynamic distribution of the kth instance and the kth instance.

[0128] S433: Constructing a multi-scale fused graph based on local fine-grained graphs and global manifold graphs Specifically, it can be expressed as the following formula:

[0129]

[0130] In the formula, This represents the multi-scale fusion graph in the t-th iteration; Indicates the fusion weight. ,generally Take 0.2;

[0131] S44: Obtaining the cleaned soft-label matrix based on the multi-scale fusion graph. Specifically:

[0132] S441: Based on the t-1 round prediction model, confidence nonlinear adjustment parameter, reliable label enhancement coefficient, and triple reliable label set, the confidence adjustment factor for the t-th round iteration is obtained. Specifically, it can be expressed as follows:

[0133]

[0134] In the formula, In the t-th iteration, the first... i The first instance j Confidence adjustment factor for class with a high proportion of candidate labels; This indicates a non-linear adjustment parameter for the confidence level; This represents the reliability enhancement factor.

[0135] S442: A purified soft label matrix is ​​obtained based on a multi-scale fusion graph, a t-1 round prediction model, a confidence adjustment factor, and a candidate biased label matrix. This ensures that high-confidence true labels have a stronger influence during propagation, labels in the reliable label set receive additional enhancement, and propagation only occurs within the candidate label range.

[0136] Specifically, the multi-scale fused image is normalized, as shown in the following formula:

[0137]

[0138] In the formula, In the t-th iteration, the first... i Normalized multiscale fusion graph between the kth instance and the kth instance; In the t-th iteration, the first... i A multi-scale fusion graph between the kth instance and the kth instance;

[0139] The confidence-weighted label propagation is specifically expressed as follows:

[0140]

[0141] In the formula, This represents the mark propagation in the t-th iteration; Represents the normalized multi-scale fusion graph of the t-th iteration; This represents the prediction model in the (t-1)th iteration; It represents the Hadamardi (or Hadama) stack.

[0142] The cleanup soft-label matrix is ​​generated by applying candidate label constraints, as shown in the following formula:

[0143]

[0144] In the formula, Y represents the cleaned soft label matrix in the t-th iteration; Y represents the candidate biased label matrix.

[0145] S45: Using the purified soft-label matrix as the output of the t-th iteration, the parameters of the prediction model in the t-th iteration are adjusted to obtain the prediction model for the t-th iteration. Specifically:

[0146] Using the purified soft-label matrix as the output of the t-th iteration, the parameters of the prediction model in the t-th iteration are adjusted using gradient descent to obtain the t-th fine-tuned prediction model. During gradient descent, a learning rate that decays according to a cosine function over iteration is used to scale the gradient to determine the parameter update step size for each iteration. The learning rate is specifically expressed as follows:

[0147]

[0148] In the formula, This represents the learning rate during the e-th gradient descent adjustment. This represents the minimum learning rate, typically set to 0.0001. This represents the maximum learning rate, typically set to 0.01; e represents the current gradient descent iteration. This represents the total number of adjustments made during gradient descent.

[0149] After obtaining the t-round fine-tuning prediction model, the t-round fine-tuning prediction model is jointly trained using soft and hard labels.

[0150] Specifically, the purified hard label matrix for the t-th iteration is obtained based on the purified soft label matrix. Specifically, it can be expressed as follows:

[0151]

[0152] In the formula, In the t-th iteration, the first... i The first instance j Purification of hard markers for classes with a high number of candidate markers; In the t-th iteration, the first... i The first instance j Clean up soft tags with a large number of candidate tags.

[0153] The t-round fine-tuning prediction model is trained using binary cross-entropy loss with purified soft-labeled and hard-labeled matrices. Specifically, the output of the t-round fine-tuning prediction model is adjusted using the total cross-entropy loss, and then the parameters of the t-round fine-tuning prediction model are adjusted based on the adjusted output until the output of the t-round fine-tuning prediction model converges.

[0154] Among them, the total cross-entropy loss Specifically, it can be expressed as the following formula:

[0155]

[0156] In the formula, denoted by the total cross-entropy loss; p represents the output of the prediction model during the joint training phase using soft and hard labels in the t-th iteration. Indicates the hardening parameters. The value gradually increases with each iteration of joint training, typically starting at 0.3 and increasing by 0.01 times with each iteration.

[0157] Finally, the t-round prediction model was obtained. .

[0158] S46: After t rounds of iterative training, the output of the prediction model gradually converges. Training is complete when the output of the prediction model converges. After training, the optimal prediction model model(x) is obtained. .

[0159] Example 3

[0160] This embodiment is a first embodiment of a partially labeled data classification system based on dynamic graph disambiguation, used to implement the partially labeled data classification method based on dynamic graph disambiguation provided in Embodiment 1 or Embodiment 2, including:

[0161] Data processing unit: used to obtain instance feature matrices based on a partially labeled dataset;

[0162] Initial classifier building unit: used to pre-train an initial classifier based on the instance feature matrix and the candidate partial multi-label matrix;

[0163] Initial prediction model building unit: used to obtain the initial prediction model based on the initial classifier and the first activation function;

[0164] Model Iteration Training Unit: Used to train the initial prediction model multiple times, and to clean up the candidate biased label matrix during the training process to obtain the optimal prediction model after training.

[0165] Label prediction unit: used to classify the test instance based on the optimal prediction model.

[0166] The above system constructs a triple reliable label set based on output, confidence, and consistency. It then builds a weighted similarity between instances based on the frequency of candidate overrepresented labels within this set, their discriminative scores, and dependencies. This weighted similarity is used to suppress structured noise. A multi-scale similarity graph is constructed and adaptively fused. A graph propagation mechanism with dynamically adjusted label confidence is introduced. Progressive multi-stage training enables the co-evolution of classifier and label quality. This method effectively addresses issues such as feature-semantic ambiguity, structured noise interference, imbalanced label contributions, and limitations of single-scale graphs.

[0167] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Furthermore, the embodiments of this invention are not directed to any particular programming language.

[0168] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. Similarly, for the sake of brevity and to aid in understanding one or more aspects of the invention, in the description of exemplary embodiments of the invention above, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.

[0169] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.

[0170] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A method for classifying biased multi-labeled data based on dynamic graph disambiguation, characterized by, Includes the following steps: S1: Obtain a partially labeled dataset for training, and obtain an instance feature matrix and a candidate partially labeled matrix based on the partially labeled dataset; the partially labeled dataset is derived from natural scene images or medical images; S2: Based on the instance feature matrix and the candidate partial multi-label matrix, pre-train the initial classifier; S3: Obtain the initial prediction model based on the initial classifier and the first activation function; S4: The initial prediction model is trained iteratively multiple times. During the training process, the candidate biased label matrix is ​​purified to obtain the optimal prediction model after training. This includes at least the following sub-steps: The initial prediction model is trained iteratively multiple times. In the t-th iteration, a triple reliable label set is constructed based on the t-1 round prediction model according to the output, confidence, and consistency. The dynamic sparse weight of the candidate abundant markers is obtained based on the occurrence frequency of the candidate abundant markers in the triple reliable marker set and the discriminative score of the candidate abundant markers. The weighted similarity between instances is constructed based on the triple reliable marker set, the dynamic sparse weight, and the dependency between instances. A multi-scale fusion graph is constructed based on the weighted similarity. A cleaned soft-label matrix is ​​obtained based on the multi-scale fusion graph; the cleaned soft-label matrix is ​​used as the output of the t-th iteration to adjust the parameters of the prediction model in the t-th iteration, thus obtaining the prediction model in the t-th iteration; after training, the optimal prediction model after training is obtained. S5: Obtain test instances, which are derived from natural scene images or medical images, and classify the test instances based on the optimal prediction model.

2. The method of claim 1, wherein, In step S1, the feature vector of each instance in the instance feature matrix is ​​mapped to the candidate biased label vector at the corresponding position in the candidate biased label matrix; the candidate biased label matrix includes true labels and erroneous inaccurate labels.

3. The method of claim 1, wherein, In step S4, the construction of a triple reliable label set based on the t-1 round prediction model according to the output, confidence, and consistency includes at least the following sub-steps: Based on the t-1 round prediction model, the multiple candidate biased labels with the highest output scores in the t-1 round prediction model are taken as a local relatively reliable set; Based on the t-1 round prediction model, the multiple candidate biased labels with the highest confidence scores in the t-1 round prediction model are used as the global high-confidence anchor set; Based on the t-1 round prediction model, the multiple candidate biased labels with the highest consistency scores in the t-1 round prediction model are taken as the label consistency set; The union of the locally relatively reliable set, the globally high-confidence anchor set, and the label consistency set yields the triple reliable label set.

4. The method for classifying partially labeled data based on dynamic graph disambiguation according to claim 1, characterized in that, In step S4, the discriminative score of the candidate biased label is obtained based on the importance analysis of classifier instance features.

5. The method for classifying partially labeled data based on dynamic graph disambiguation according to claim 1, characterized in that, In step S4, obtaining the dependencies between instances includes at least the following sub-steps: Based on the triple reliable label set and the occurrence frequency of candidate over-probable labels within the triple reliable label set, a label co-occurrence matrix is ​​obtained; The marker co-occurrence matrix is ​​used to capture the correlation strength between pairs of candidate predominant markers; The dependencies between instances are obtained based on the tag co-occurrence matrix and the triple reliable tag set; When two instances share candidate-biased tags j If these two instances also share additional candidate-biased tags... j Other candidate markers l The greater the dependency between the two instances, the stronger the relationship between them.

6. The method for classifying partially labeled data based on dynamic graph disambiguation according to claim 1, characterized in that, In step S4, constructing the multi-scale fusion graph based on the weighted similarity includes at least the following sub-steps: A local fine-grained map is constructed based on the weighted similarity using a proximity algorithm. A global manifold graph is constructed based on the weighted similarity using optimal transport theory; A multi-scale fusion graph is constructed based on the local fine-grained graph and the global manifold graph.

7. The method for classifying partially labeled data based on dynamic graph disambiguation according to claim 6, characterized in that, In step S4, the construction of the global manifold graph based on the weighted similarity using optimal transport theory includes at least the following sub-steps: The cost matrix is ​​obtained based on the weighted similarity. The dynamic supply distribution is obtained based on the t-1 round prediction model, the candidate biased label matrix, the triple reliable label set, and the dynamic sparse weight. The dynamic demand distribution is obtained based on the t-1 round prediction model, the candidate biased label matrix, the triple reliable label set, and the dynamic sparse weight. Based on the cost matrix, the dynamic supply distribution, and the dynamic demand distribution, a global manifold is constructed using entropy regularization for optimal transmission.

8. The method for classifying partially labeled data based on dynamic graph disambiguation according to claim 1, characterized in that, In step S4, obtaining the cleaned soft label matrix based on the multi-scale fusion graph includes at least the following sub-steps: The confidence adjustment factor is obtained based on the t-1 round prediction model, the confidence nonlinear adjustment parameter, the reliability label enhancement coefficient, and the triple reliable label set; The purified soft label matrix is ​​obtained based on the multi-scale fusion graph, the t-1 round prediction model, the confidence adjustment factor, and the candidate biased label matrix.

9. The method for classifying partially labeled data based on dynamic graph disambiguation according to claim 1, characterized in that, In step S4, the purified soft-label matrix is ​​used as the output of the t-th iteration to adjust the parameters of the prediction model in the t-th iteration, thus obtaining the t-th prediction model. This includes at least the following sub-steps: Using the purified soft-label matrix as the output of the t-th iteration, the parameters of the prediction model in the t-th iteration are adjusted using the gradient descent method to obtain the t-th fine-tuned prediction model; A cleaned hard-label matrix is ​​obtained based on the cleaned soft-label matrix; the t-round fine-tuning prediction model is trained using the binary cross-entropy loss of the cleaned soft-label matrix and the cleaned hard-label matrix to obtain the t-round prediction model.

10. A partial label classification system based on dynamic graph disambiguation, used to implement the partial label classification method based on dynamic graph disambiguation as described in any one of claims 1 to 9, characterized in that, include: Data processing unit: used to obtain instance feature matrix based on the partially labeled dataset; Initial classifier construction unit: used to pre-train an initial classifier based on the instance feature matrix and the candidate partial multi-label matrix; Initial prediction model building unit: used to obtain an initial prediction model based on the initial classifier and the first activation function; Model Iteration Training Unit: Used to perform multiple iterations of training on the initial prediction model, and to purify the candidate biased label matrix during the training process to obtain the optimal prediction model after training; Label prediction unit: used to classify the test instance based on the optimal prediction model.