Complex task-based high-quality pseudo-annotation data set construction method
By constructing a fusion of cross-modal causal graphs and dynamic knowledge graphs, inter-modal confusion variables in pseudo-labels are identified and corrected, non-causal paths are cut off, and high-quality pseudo-notation data sets are generated, which solves the noise accumulation and causal confusion problems caused by cross-modal interference and domain knowledge loss, and achieves the semantic consistency and stability of the label.
Patent Information
- Application Number
- CN202510461536.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing pseudo-notation datasets have not been effectively solved by the noise accumulation, causal confusion and semantic drift caused by cross-modal interference and lack of domain knowledge. Traditional methods rely on static threshold filtering low confidence samples, and cannot adapt to dynamic semantic offsets and lack an interpretable causal decoupling mechanism.
By constructing a cross-modal causal graph, loading a domain knowledge graph, identifying confusing variables between modals, generating initial pseudo-labels, performing distance calculations of semantic embedding and domain knowledge graphs, dynamically adjusting semantic correction weights, cutting off non-causal paths, building a dynamic optimization closed loop, and generating high-quality pseudo-notation datasets.
Effectively suppress cross-modal co-occurrence noise interference, enhance the causal correlation robustness between multimodal features, realize label semantic consistency and domain entity alignment, and ensure the stability and reliability of the pseudo-label generation process.
Smart Images

Figure CN120297445A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal learning, and particularly to a method for constructing a high-quality pseudo-labeled data set based on complex tasks. Background Art
[0002] Currently, semi-supervised learning of pseudo-labeled data has become an important way to solve the problem of too high annotation cost for complex tasks. The existing technologies mainly adopt a self-training framework, and generate pseudo-labels through cross-modal feature alignment and confidence threshold filtering. However, with the successive proposal of cross-modal association modeling methods based on graph neural networks (such as CM-GNN) and knowledge graph enhancement methods (such as KG-BERT), the label generation process is optimized by introducing inter-modal attention mechanisms and entity embedding alignment. But there are still significant defects in aspects such as elimination of multi-modal heterogeneous data interference, integration of domain knowledge structured constraints, and dynamic optimization of causal paths.
[0003] The current technical bottlenecks are concentrated on the problems of cumulative confounding bias and semantic drift in the cross-modal pseudo-label generation process. Traditional methods rely on static thresholds to filter low-confidence samples, but fail to effectively distinguish causally related features from cross-modal co-occurrence noise, resulting in pseudo-labels being interfered by false paths. Although knowledge graph fusion methods can introduce domain priors, entity alignment uses fixed similarity thresholds, which cannot adapt to dynamic semantic shifts, and lacks an interpretable causal decoupling mechanism at the level of counterfactual intervention. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for constructing a high-quality pseudo-labeled data set based on complex tasks to solve the problems of noise accumulation, causal confusion, and semantic drift caused by cross-modal interference and lack of domain knowledge in existing pseudo-labeled data sets.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for constructing a high-quality pseudo-labeled dataset based on complex tasks, which includes constructing a cross-modal causal graph based on multi-modal raw data, loading a domain knowledge graph, identifying confounding variables between modalities, and generating initial pseudo-labels; calculating the distance between the semantic embedding of the initial pseudo-labels and the nodes of the domain knowledge graph, generating semantic correction weights to correct the semantics of the initial pseudo-labels, and outputting semantically consistent pseudo-labels; generating counterfactual samples by forcibly cutting off non-causal paths in the cross-modal causal graph, comparing the pseudo-label differences between the original samples and the counterfactual samples, and generating cross-modal debiased pseudo-labels; constructing a dynamic optimization closed-loop to monitor the gradient direction of the cross-modal debiased pseudo-labels and the conflict frequency with the domain knowledge graph in real time, and dynamically adjusting the cross-modal causal graph structure, semantic correction weights, and the intervention intensity in the generation process of counterfactual samples; combining the cross-modal debiased pseudo-labels and the semantically consistent pseudo-labels to fuse and generate a standardized pseudo-labeled dataset with multi-modal alignment, clear entity relationships, and semantic consistency.
[0007] As a preferred embodiment of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: constructing the cross-modal causal graph includes using the PC algorithm combined with conditional independence testing to extract causal relationships, and generating a directed acyclic causal graph that labels the set of confounding variables after optimization by the causal structure equation model and adversarial training.
[0008] As a preferred embodiment of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: generating the initial pseudo-labels includes weighted fusion of spatio-temporally aligned multi-modal features through a cross-modal attention network, using the confidence dynamic filtering threshold of entity relationships in the domain knowledge graph as a prior constraint, and performing backdoor path correction on modal features with confounding variables in combination with conditional independence testing; The confidence dynamic filtering threshold is dynamically calculated from the confidence scores of the sample pseudo-labels in the current iteration batch.
[0009] As a preferred embodiment of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: the semantic correction refers to using the knowledge graph node embedding method to map the pseudo-labels and the entity nodes of the domain knowledge graph to a unified semantic space, calculating the cosine similarity between the two, and dynamically adjusting the semantic correction weights based on the piecewise function and hierarchical relationship constraints.
[0010] As a preferred embodiment of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: generating the counterfactual samples includes the following steps, Traverse the non-causal paths in the cross-modal causal graph that satisfy the d-separation criterion, and screen the paths to be intervened based on the causal effect strength formula; Perform intervention operations on non-causal paths through a cross-modal variational autoencoder to cut off the dependence relationship of confounding variables and reconstruct the counterfactual feature distribution; Use Jensen-Shannon divergence to quantify the pseudo-label differences between the original samples and the counterfactual samples, generate cross-modal bias weights based on an exponential decay function, and perform probability normalization and debiasing on the pseudo-labels.
[0011] As a preferred solution of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: the construction of the dynamic optimization closed-loop includes the following steps: Extract the gradient direction features of the cross-modal debiased pseudo-labels, and calculate the cosine similarity conflict frequency between the gradient direction features and the knowledge graph node embeddings; According to the dynamic adjustment strategy of the conflict threshold, delete the high-frequency conflict paths in the causal graph, and add causal edges verified by the knowledge graph relationship paths; Based on the sliding window statistical value of the cosine similarity conflict frequency, apply exponential decay to the semantic correction weight, and adaptively adjust the intervention intensity coefficient in the adversarial loss function of CM-VAE.
[0012] As a preferred solution of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: the standardized pseudo-labeled dataset includes probabilistically fusing the debiased pseudo-labels and the semantic consistency pseudo-labels through a weighted average strategy, performing entity alignment verification based on traversing the knowledge graph relationship paths, and screening the labels using a confidence dynamic filtering threshold.
[0013] As a preferred solution of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to the present invention, wherein: the confidence of the sample pseudo-labels in the current iteration batch refers to the probability score generated after the initial pseudo-labels are semantically aligned by the knowledge graph and dynamically weighted and corrected; Each sample in the batch is a set of multi-modal data units that have completed spatio-temporal alignment processing.
[0014] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and wherein: when the computer program is executed by the processor, it implements any step of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks as described in the first aspect of the present invention.
[0015] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, it implements any step of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks as described in the first aspect of the present invention.
[0016] The beneficial effects of the present invention are as follows: By constructing a cross-modal causal graph and integrating it with a dynamic knowledge graph, the quality of pseudo-labeled data is improved. The detection of confounding variables based on the backdoor criterion, combined with an adversarial training strategy, effectively suppresses the interference of cross-modal co-occurrence noise on the causal path and enhances the robustness of the causal association between multi-modal features. The dynamic weight correction mechanism of knowledge graph embedding relies on semantic space projection and hierarchical relationship constraints to achieve precise control of label semantic drift and improve the semantic consistency of domain entity alignment. The counterfactual intervention framework drives debiasing of non-causal paths through probability distribution differences, eliminating the influence of spurious associations in cross-modal interactions. The dynamic optimization closed-loop realizes online adaptive adjustment of the causal graph structure based on gradient conflict monitoring, ensuring the stability and reliability of the pseudo-label generation process in complex inference scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of a method for constructing a high-quality pseudo-labeled dataset based on complex tasks in Embodiment 1; Figure 2 It is a flowchart of cross-modal causal annotation generation in Embodiment 1; Figure 3 It is a flowchart of knowledge graph semantic correction in Embodiment 1; Figure 4 It is a flowchart of counterfactual intervention and debiasing processing in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.
[0020] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0021] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.
[0022] Example 1, referring to Figures 1 to 4 , this example provides a method for constructing a high-quality pseudo-labeled dataset based on complex tasks, including the following steps: S1. Construct a cross-modal causal graph based on multi-modal raw data, load the domain knowledge graph, identify the confounding variables between modalities, and generate initial pseudo-labels.
[0023] Specifically, it includes the following steps: Receive the input multi-modal raw data (text, image, audio, etc.) and perform spatio-temporal alignment. Compensate for time offsets through dynamic time warping, use a cross-modal attention network to align semantic regions, and output a multi-modal feature representation with a unified spatio-temporal reference. Based on the aligned multi-modal features, use the PC algorithm combined with conditional independence testing to learn the causal relationships between modalities. Optimize the edge weights through a causal structural equation model, construct a directed acyclic causal graph, and use adversarial training to enhance the robustness of the causal relationships. It should be noted that the directed acyclic causal graph (DAG) is a specific manifestation of the cross-modal causal graph, and the cross-modal causal graph is a causal relationship graph constructed in the context of multi-modal data.
[0024] Load the predefined domain knowledge graph (such as Wikidata), use the TransE algorithm to embed the graph entities and relationships into vectors, map the multi-modal features and the knowledge graph entities to the same semantic space through contrastive learning, and force the alignment of high-confidence entities. Detect cross-modal common cause paths in the directed acyclic causal graph based on the backdoor criterion, verify the confounding variables (such as "lighting affects images and text") in combination with the domain relationships of the knowledge graph, eliminate pseudo-related paths through conditional independence testing, and output a set of confounding variables. Propagate the confidence of multi-modal features along the causal paths in the directed acyclic causal graph, infer labels in combination with the entity relationship constraints of the knowledge graph, use the confidence dynamic filtering threshold of the entity relationships in the domain knowledge graph as a prior constraint, filter out low-confidence samples through the confidence dynamic filtering threshold, and clean the noise labels based on entity consistency conflict detection, and output the denoised initial pseudo-labels.
[0025] It should be noted that the confidence dynamic filtering threshold is dynamically calculated based on the confidence scores of the pseudo-labels of the current batch of samples. Specifically, the confidence scores of all samples are collected and their mean and standard deviation are calculated; the confidence dynamic filtering threshold is calculated by subtracting several times the standard deviation from the mean, where the adjustable parameter usually takes 2 or 3 to control the strictness of filtering; the confidence score of each sample is compared with the confidence dynamic filtering threshold, and the samples with a confidence score lower than the confidence dynamic filtering threshold are marked as low-confidence and filtered out, so as to ensure the quality of the pseudo-labels. Among them, the batch of samples is a set of multi-modal data units divided according to a preset scale and having completed spatio-temporal alignment processing.
[0026] S2. Calculate the distance between the semantic embedding of the initial pseudo-label and the nodes of the domain knowledge graph, generate semantic correction weights to correct the semantics of the initial pseudo-label, and output semantically consistent pseudo-labels.
[0027] Specifically, it includes the following steps: Encode the initial pseudo-label into a low-dimensional semantic vector through a pre-trained semantic model, and at the same time extract the embedding representation of the corresponding nodes in the domain knowledge graph to ensure that the label and the graph entity are represented in a unified semantic space; It should be noted that the pre-training of the semantic model adopts a two-stage optimization strategy. Based on a large-scale general corpus for basic semantic representation learning, train a deep Transformer architecture through a masked language modeling task, so that the semantic model masters the lexical-level and syntactic-level semantic representation capabilities, and the contrastive learning objective enhances the fine-grained semantic discrimination. Use the professional corpus of the target domain for adaptive fine-tuning, construct an entity-relationship classification task, and use the knowledge distillation technology to integrate the structural information of the domain knowledge graph into the semantic model representation. The training process uses a 12-24 layer Transformer network, and the hierarchical normalization strategy ensures the training stability, and the mixed-precision training accelerates the convergence. The semantic model introduces an orthogonal constraint to avoid dimensional collapse, and the adversarial training enhances the cross-domain generalization ability. The prototype contrast loss and the dynamic negative sampling strategy strengthen the intra-class aggregation and the clarity of the semantic boundary.
[0028] The extraction of the knowledge graph node embedding representation adopts a translation-based model (such as TransE / RotatE), and the structured representation is obtained through triple facts and negative sampling training. For entities with text descriptions, the text embedding generated by the language model is fused with the graph structure embedding, and a gating mechanism is used for dynamic weighting. The adaptive margin loss and the curriculum learning strategy are introduced in the training process to hierarchically capture direct relationships and multi-hop semantics. The finally generated embedding is verified through tasks such as link prediction, retaining the graph topology structure and semantic relationships, and providing a stable knowledge representation basis for downstream tasks.
[0029] Calculate the cosine similarity between the semantic embedding of the pseudo-label and the node embedding of the knowledge graph, filter out the matching nodes with a similarity higher than the preset similarity threshold, and record the label-node pairs with low similarity as the semantic conflict candidate set; Among them, the preset similarity threshold refers to a dynamically optimized boundary value used to determine the semantic matching quality between the pseudo-label and the nodes of the knowledge graph. Specifically, this similarity threshold is determined in the following way: In the verification stage, the similarity distribution of the annotated entity pairs in the knowledge graph is statistically analyzed, and the critical value that can achieve the best balance between precision and recall is selected (usually the 85th percentile of the similarity distribution). In practical applications, a hierarchical strategy is adopted: a similarity threshold of 0.7 - 0.8 is set for core entities, and a similarity threshold of 0.6 - 0.7 is used for marginal entities. The finally determined similarity threshold needs to ensure that the semantic matching accuracy on the validation set is not less than 90%. The correction cases triggered by the similarity threshold will be recorded and regularly used to optimize the similarity threshold parameters.
[0030] Dynamically assign weights according to the similarity value between the label and the node. High-similarity labels are assigned high weights, and low-similarity labels introduce a penalty factor to reduce the weight. At the same time, combine the hierarchical relationship constraints of the knowledge graph to correct the label semantics; Specifically, in the process of weight assignment, "high-similarity labels" are defined as label-node pairs with a cosine similarity value greater than the preset similarity threshold (such as 0.8), indicating that their semantics are highly consistent with the nodes of the knowledge graph; "low-similarity labels" refer to label-node pairs with a cosine similarity value lower than the threshold (such as 0.6), indicating significant semantic deviation.
[0031] It should be noted that the weight assignment is implemented using a piecewise function: for high-similarity labels, the weight increases linearly, with a maximum value of 1; for low-similarity labels, the weight is reduced by a penalty factor (such as 0.5), with a minimum value of 0.2. At the same time, correct the weight in combination with the hierarchical relationship of the knowledge graph (such as is-a or part-of): if there is a direct hierarchical relationship between the label and the node, the weight is additionally increased by 0.1; if there is an indirect hierarchical relationship, the weight is increased by 0.05 to ensure that the semantic correction process makes full use of the structured information of the graph.
[0032] Re-weight the probability distribution of the initial pseudo-label according to the weight, enhance the probability of high-weight labels and suppress low-weight labels. For the semantic conflict candidate set, correct the semantic definition of the conflicting labels through the entity relationship path reasoning of the knowledge graph; Specifically, when re-weighting the probability distribution of the initial pseudo-label, the probability value of each label should be multiplied by its corresponding weight. For high-weight labels (weight > 0.8), use an exponential amplification strategy to increase the probability value to 1.2 times the original value; for low-weight labels (weight < 0.4), use a linear suppression strategy to reduce the probability value to 0.8 times the original value.
[0033] Furthermore, for the semantic conflict candidate set, it is corrected through entity relationship path reasoning of the knowledge graph: retrieve the k-hop neighbor nodes (k = 2) of the conflict label in the knowledge graph, calculate its semantic similarity with the label, and select the node with the highest similarity as the correction target; then propagate the correction signal backward along the graph path to update the semantic definition of the label to make it consistent with the semantics of the correction target node, and at the same time adjust the weight distribution of relevant entities.
[0034] Fuse the corrected label probability distribution, filter and retain high-quality labels through dynamic statistical thresholds, and output a semantic consistency pseudo-label set that is strictly aligned with the entity relationships of the knowledge graph.
[0035] Among them, when fusing the corrected label probability distribution, a weighted average strategy is adopted: for each label, its original probability value and the corrected probability value are weighted and summed according to the ratio of 0.6:0.4 to ensure a smooth transition during the correction process.
[0036] It should be noted that the setting of the dynamic statistical threshold is based on the standard deviation of the label probability distribution: specifically, calculate the mean and standard deviation of all label probability values, and initially set the dynamic statistical threshold to the mean plus 1.5 times the standard deviation; if the recall rate is lower than 90%, then reduce the dynamic statistical threshold by 0.1 times the standard deviation; if the recall rate is higher than 95%, then increase the dynamic statistical threshold by 0.1 times the standard deviation. Finally, filter out low-probability labels through this dynamic statistical threshold, retain labels with probability values higher than the threshold, and output a semantic consistency pseudo-label set that is strictly aligned with the entity relationships of the knowledge graph to ensure that each label has a clear semantic correspondence with at least one knowledge graph node.
[0037] S3. Generate counterfactual samples by forcibly cutting off non-causal paths in the cross-modal causal graph, and compare the pseudo-label differences between the original samples and the counterfactual samples to generate cross-modal debiased pseudo-labels.
[0038] Specifically, it includes the following steps: Based on the directed acyclic causal graph, identify all non-causal paths that satisfy the backdoor criterion through the d-separation criterion (such as the path "image background color → text sentiment label"), and filter the paths that contain confounding variables (such as lighting conditions, cross-modal co-occurrence noise) in the non-causal paths; among them, the paths may be real causal or false non-causal paths, and need to be further verified through the causal effect strength.
[0039] Specifically, select the path to be intervened based on the causal effect strength formula, which is expressed as: , In the formula, Y represents the original sample pseudo-label (such as classification probability distribution or regression value), represents the pathP The causal effect strength value for Y is used to quantify the net impact of this path (whether causal or non-causal) on Y . X represents a confounding variable that is intervened on and lies on a non-causal path. and respectively represent two different intervention values applied to X , and they need to satisfy comparability. represents the expected value of X when an intervention is applied to (forcing it to be the value Y ). represents the expected value. is a causal intervention (rather than a conditional probability), that is, it forcibly cuts off all incoming edges of X (eliminating confounding effects). represents the expected value of X when another intervention value is applied to Y ; It should be noted that the causal path P refers to an indirect causal path connecting two modalities in a cross-modal causal graph (such as "image background → text label").
[0040] Retain the path of to generate an intervention strategy table, define uniform sampling interventions for discrete variables (such as replacing the image background with a random color), and mean interventions for continuous variables (such as fixing the audio spectrum energy to the training set mean); Furthermore, generate counterfactual samples through cross-modal decoupling and intervention to eliminate the influence of non-causal paths. Specifically, use a cross-modal variational autoencoder (CM-VAE) to decouple multimodal features into causal factors and irrelevant factors , and reconstruct counterfactual features through a structural equation model, expressed as: . In the formula, Decoder represents a cross-modal decoder responsible for reconstructing latent variables (causal factors and non-causal factors) into multimodal features. represents a causal intervention operator used to cut off all input dependencies of (eliminating confounding). represents an intervention on the non-causal factor , forcing it to take a specific value ; Among them, the assignment method for the specific value includes, for continuous variables: set to the training set mean or random sampling (such as ; Discrete variable: Forced to be set to a certain category (such as setting the background color to a solid color).
[0041] Preferably, the data features are decomposed into causal and non-causal parts through a cross-modal variational autoencoder (CM-VAE). Adversarial training ensures that only the non-causal part is changed during intervention, without affecting the true causal relationship. Thus, it can accurately eliminate spurious associations, such as modifying the background of a picture while keeping the main content; maintain semantic consistency between different modalities; and can also intuitively see which features are intervened, being more accurate and reliable, and especially suitable for dealing with complex multi-modal data.
[0042] It should be noted that the causal intervention operator , needs to constrain the causal factor to be consistent with the original distribution through an adversarial loss function . The adversarial loss function is expressed as: ; ; In the formula, D represents the discriminator, which is a binary classification neural network used to distinguish the original causal factor and the counterfactual causal factor, represents the causal factor of the counterfactual sample, represents the causal factor encoder; Furthermore, for the original sample pseudo-label Y and the counterfactual sample pseudo-label , the Jensen-Shannon divergence is used to quantify the difference value , which is expressed as: , In the formula, JSD represents the Jensen-Shannon divergence, represents the Kullback-Leibler divergence, represents the intermediate distribution (mean distribution) of the two distributions, which is used for symmetric JSD calculation, represents the normalization coefficient to ensure that the JSD value is within range; According to the difference value , the debiasing weight is dynamically allocated, and the weight function is defined as: ; In the formula, w is the debiasing weight, which represents controlling the label correction intensity, and the value range is , is the attenuation coefficient (default value 5), and exp represents the exponential function, which is used to ensure that the weight increases smoothly with the increase of JSD. Indicates the reverse output, such that when the JSD is larger w it is closer to 1 (fully corrected); Fuse the original pseudo-labels and counterfactual differences, generate debiased pseudo-labels, and perform probability normalization, expressed as: , ; In the formula, is the debiased pseudo-label, represents the difference vector between the original and counterfactual labels, reflecting the interference direction of the non-causal path, represents the weighted correction term of the difference vector, with the weight w dynamically determined by the JSD, represents the sum of the scores for all classes to achieve normalization, represents the Softmax function.
[0043] S4. Construct a dynamic optimization closed-loop to monitor in real time the conflict frequency between the gradient direction of the cross-modal debiased pseudo-labels and the domain knowledge graph, and dynamically adjust the cross-modal causal graph structure, semantic correction weights, and intervention intensity during the counterfactual sample generation process.
[0044] Specifically, it includes the following steps: Calculate the difference between the gradient direction of the cross-modal debiased pseudo-labels and the semantic direction of the knowledge graph nodes.
[0045] Specifically, let the semantic direction of the entity nodes in the domain knowledge graph be the embedding vector , and the pseudo-label gradient direction be (generated by the backpropagation process of the cross-modal joint training framework, reflecting the optimization direction of the current pseudo-labels); The conflict frequency C is the proportion of the mismatch between the pseudo-label gradient direction and the knowledge graph semantic direction, expressed as: ; In the formula, N represents the total number of entity nodes participating in the calculation in the domain knowledge graph, i represents the entity node index in the domain knowledge graph, j represents the sample index of the pseudo-labels in the current batch, M represents the number of pseudo-labels to be evaluated in the current batch, represents calculating the pseudo-label gradient and the knowledge graph node cosine similarity, is the conflict threshold, the critical value for determining whether there is a conflict; It should be noted that the pseudo-label gradient Semantic direction projected onto knowledge graph nodes , if the projection value is lower than , it is determined as a conflict; initially set to 0.5, if the conflict frequency C rises for three consecutive batches, then decrease (step size 0.05); otherwise increase .
[0046] Among them, the conflict threshold is obtained through multi-stage dynamic optimization: based on pre-calibration of the validation set, the gradient cosine similarity distribution between knowledge graph nodes and true labels is statistically calculated, and the 85th percentile value is taken as the initial threshold (typical range 0.5 - 0.7); during the running stage, the moving window mean of the conflict frequency C (window size 10 batches) is monitored in real time. When it rises for three consecutive batches C , decrease it by a step size of 0.05 to relax the determination criteria, otherwise increase the threshold to strengthen the alignment requirements. The whole process is restricted to prevent extreme adjustments; after each adjustment, it needs to be rechecked by the validation set. If the accuracy drops by more than 2%, it will be rolled back to the previous valid threshold. The finally output needs to satisfy that the standard deviation of the fluctuation of the conflict frequency C is less than 0.1 in the recent 20 batches to ensure the stability of the conflict threshold .
[0047] If the conflict frequency corresponding to a certain causal path (such as "image → text") is continuously higher than the mean value, delete this causal path and recalculate the causal effect (that is, recalculate the CausalEffect formula); For the pseudo-labels with frequent conflicts, retrieve the unconnected entity relationships in the domain knowledge graph (such as "drug - disease"). If the causal effect intensity of its path , then add a new edge in the causal graph; For the labels with high conflict frequency, reduce their semantic correction weights , forcing the model to rely more on data features rather than the knowledge graph, expressed as: ; In the formula, represents the new weight after conflict frequency adjustment, represents the semantic correction weight at the t th iteration, t represents the number of iteration steps, represents the hyperparameter controlling the weight descent speed, , C represents the conflict index, , is the decay factor, and the value range is ; If the conflict frequency continuously decreases, then Restore linearly (e.g., +0.05 per batch) Set the initial value to 1.0 (full intervention), and dynamically adjust the intervention intensity coefficient according to the conflict frequency. The expression is: , In the formula, is the intervention intensity coefficient, which is used to control the intervention intensity when generating counterfactual samples; It should be noted that in the CM-VAE training, if C is too high, then increase the adversarial loss weight of the causal factor and the irrelevant factor to force more thorough decoupling; Preset the pseudo-label error rate (e.g., classification error rate < 5%). Calculate the error rate for each batch. If it meets the standard for 10 consecutive batches, terminate; if the fluctuation range of the error rate (standard deviation) is < 1% for 5 consecutive batches, force termination and output the current optimal parameter set.
[0048] It should be noted that the optimal parameters output when the dynamic optimization closed-loop terminates include four core parts: First, the optimized cross-modal causal graph structure, including the finally retained causal edges and their strength values, as well as the deleted confounding paths and their corresponding conflict frequencies; second, the dynamically adjusted and converged weight parameters, covering the semantic correction weight, the adversarial loss coefficient, and the intervention intensity; the third part is the key parameters of the feature decoupling model, including the encoder-decoder weights, the discriminator parameters, and the distribution characteristics of the causal factor and the non-causal factor; finally, the verification indicators are output, including the pseudo-label error rate, the stable conflict frequency, and the alignment score of each modality. These parameters together constitute the basic configuration for generating the standardized dataset.
[0049] Preferably, by comparing the real-time directions of the pseudo-label gradients and the semantic vectors of the knowledge graph, an adaptive conflict threshold mechanism is established to achieve intelligent correction of multi-modal data, thereby automatically identifying high-frequency conflict paths and performing causal graph reconstruction, and dynamically balancing the semantic correction weight and the counterfactual intervention intensity. Furthermore, during the optimization process, continuously correct the initial causal hypothesis, precisely coordinate the data features and domain knowledge, and output the core parameter set including the optimized causal structure, stable parameter configuration, and verification indicators to form an adaptive and interpretable multi-modal annotation solution.
[0050] S5. Combine the cross-modal debiased pseudo-labels and the semantic consistency pseudo-labels to fuse and generate a standardized pseudo-annotated dataset with multi-modal alignment, clear entity relationships, and semantic consistency.
[0051] Specifically, it includes the following steps: The cross-modal debiased pseudo-labels (causality) and semantically consistent pseudo-labels (knowledge constraints) are fused with a weight of 0.6:0.4, and conflicting labels are forced to be unified through the synonym relationship of the knowledge graph; Verify the existence of entity relationship paths and remove labels that are not supported by the knowledge graph; A pre-trained semantic model is used to align the multimodal embedding space (cosine similarity < 0.2), and misaligned samples trigger counterfactual intervention regeneration; Invalid labels are removed based on knowledge graph node mapping, and logical conflict labels are removed based on the causal graph pruning status to complete double filtering; Output standardized datasets in the formats of COCO (images), BIO (text), and TIMIT (audio), with additional knowledge graph entity IDs and causal path metadata.
[0052] This embodiment also provides a computer device, which is suitable for the method of constructing a high-quality pseudo-annotated dataset based on complex tasks, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the method of constructing a high-quality pseudo-annotated dataset based on complex tasks proposed in the above embodiment.
[0053] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0054] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for constructing a high-quality pseudo-labeled dataset based on complex tasks as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0055] In summary, the present invention improves the quality of pseudo-labeled data through the cross-modal causal graph construction and dynamic knowledge graph fusion mechanism. The detection of confounding variables based on the backdoor criterion combined with the adversarial training strategy effectively suppresses the interference of cross-modal co-occurrence noise on the causal path and enhances the robustness of the causal association between multi-modal features. The dynamic weight correction mechanism of knowledge graph embedding relies on semantic space projection and hierarchical relationship constraints to achieve precise control of label semantic drift and improve the semantic consistency of domain entity alignment. The counterfactual intervention framework drives the debiasing of non-causal paths through the probability distribution difference to eliminate the influence of spurious associations in cross-modal interactions. The dynamic optimization closed-loop realizes the online adaptive adjustment of the causal graph structure based on gradient conflict monitoring to ensure the stability and reliability of the pseudo-label generation process in complex reasoning scenarios.
[0056] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for constructing a high-quality pseudo-labeled data set based on complex tasks, characterized in that: including Construct a cross-modal causal graph based on multi-modal raw data, load the domain knowledge graph, identify confounding variables between modalities, and generate initial pseudo-labels; Calculate the distance between the semantic embedding of the initial pseudo-labels and the nodes of the domain knowledge graph, generate semantic correction weights to correct the semantics of the initial pseudo-labels, and output semantically consistent pseudo-labels; Generate counterfactual samples by forcibly cutting off non-causal paths in the cross-modal causal graph, and compare the pseudo-label differences between the original samples and the counterfactual samples to generate cross-modal debiased pseudo-labels; Construct a dynamic optimization closed-loop, monitor the conflict frequency between the gradient direction of the cross-modal debiased pseudo-labels and the domain knowledge graph in real time, and dynamically adjust the cross-modal causal graph structure, semantic correction weights, and intervention intensity in the counterfactual sample generation process; Combine the cross-modal debiased pseudo-labels and the semantically consistent pseudo-labels to fuse and generate a standardized pseudo-annotation dataset with multi-modal alignment, clear entity relationships, and consistent semantics.
2. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 1, wherein: The construction of the cross-modal causal graph includes using the PC algorithm combined with conditional independence testing to extract causal relationships, and generating a directed acyclic causal graph that annotates the set of confounding variables after optimization by the causal structure equation model and adversarial training.
3. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 2, characterized in that: The generation of the initial pseudo-labels includes weighted fusion of multi-modal features after spatio-temporal alignment through a cross-modal attention network, using the confidence dynamic filtering threshold of entity relationships in the domain knowledge graph as a prior constraint, and performing backdoor path correction on modal features with confounding variables in combination with conditional independence testing; The confidence dynamic filtering threshold is dynamically calculated from the confidence scores of the sample pseudo-labels in the current iteration batch.
4. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 3, characterized in that: The semantic correction refers to using the knowledge graph node embedding method to map the pseudo-labels and the entity nodes of the domain knowledge graph to a unified semantic space, calculating the cosine similarity between the two, and dynamically adjusting the semantic correction weights based on a piecewise function and hierarchical relationship constraints.
5. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 1, characterized in that: The generation of counterfactual samples includes the following steps Traverse the non-causal paths in the cross-modal causal graph that satisfy the d-separation criterion, and screen the paths to be intervened based on the causal effect strength formula; Perform intervention operations on the non-causal paths through a cross-modal variational autoencoder, cut off the dependence relationship of confounding variables, and reconstruct the counterfactual feature distribution; Use Jensen-Shannon divergence to quantify the pseudo-label differences between the original samples and the counterfactual samples, generate cross-modal bias weights based on an exponential decay function, and perform probability normalization debiasing on the pseudo-labels.
6. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 5, characterized in that: The construction of the dynamic optimization closed-loop includes the following steps: Extract the gradient direction features of the cross-modal debiased pseudo-labels, and calculate the conflict frequency of the cosine similarity between the gradient direction features and the knowledge graph node embeddings; According to the dynamic adjustment strategy of the conflict threshold, delete the high-frequency conflict paths in the causal graph, and add causal edges verified by the knowledge graph relationship paths; Based on the sliding window statistical value of the cosine similarity conflict frequency, apply exponential decay to the semantic correction weights, and adaptively adjust the intervention intensity coefficient in the adversarial loss function of CM-VAE.
7. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 1, wherein: The standardized pseudo-labeled dataset includes probabilistically fusing debiased pseudo-labels and semantic consistency pseudo-labels through a weighted average strategy, performing entity alignment verification based on traversing the relationship paths of the knowledge graph, and screening labels using a confidence dynamic filtering threshold.
8. The method for constructing a high-quality pseudo-labeled data set based on complex tasks according to claim 3, wherein: The confidence of the sample pseudo-labels in the current iteration batch refers to the probability score generated after the initial pseudo-labels are semantically aligned by the knowledge graph and dynamically weighted and corrected; Each sample in the batch is a set of multimodal data units that have completed spatio-temporal alignment processing.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the method for constructing a high-quality pseudo-labeled dataset based on complex tasks according to any one of claims 1 to 8.
Citation Information
Patent Citations
Infrared weak target detection method based on anti-fact causal learning
CN114972869A
Pseudo tag data construction method and device, terminal and medium
CN116956935A
Node classification graph neural network model based on neighborhood label distribution and global label relation
CN119810556A
Cited By
Fragmented data cross-modal label generation system and method based on deep transfer learning
CN120744707A
Fragment data cross-modal label generation system and method based on deep transfer learning
CN120744707B
Statistical analysis method and system for traditional Chinese medicine clinical research based on syndrome logic network
CN121011313A
Semantic enhancement processing system based on large model agent RAG database
CN121072782A
Data label generation method based on directed acyclic graph
CN121117811A