Large model data fusion method and system based on multi-source heterogeneous data

Through modular designs such as multimodal causal decoupling and dynamic subspace routing, the cross-modal consistency and bias problems of multi-source heterogeneous data are solved, efficient fusion and bias isolation of multi-source data are achieved, and the accuracy and robustness of data representation are improved.

CN120470520AActive Publication Date: 2025-08-12TUBA (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510518863.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-12
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Traditional methods are difficult to effectively mine the deep cross-modal consistency of multi-source heterogeneous data, and are susceptible to noise interference and implicit bias superposition, resulting in difficulty in eliminating dimensional disasters and conflicts, and lack of end-to-end collaborative optimization.

Method used

The multimodal causal decoupling module, dynamic subspace routing module, heterogeneous meta-learning adaptation module, sparse activation gated module, counterfactual obfuscation pooling module, small sample contradiction dissolution module and modular parameter isolation module are adopted to realize the implicit association of multi-source data, decoupling of biased paths and physical isolation of parameter activation through causal graph construction, dynamic subspace routing, meta-learning and differential topological constraints.

Benefits of technology

The full-link fusion control of multi-source data is realized, and the cascaded diffusion of bias propagation and long-tail conflicts is accurately suppressed, improving the accuracy and robustness of data representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470520A_ABST
    Figure CN120470520A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data fusion, in particular to a large model data fusion method and system based on multi-source heterogeneous data. Comprising a multi-modal causal decoupling module, a dynamic subspace routing module, a heterogeneous element learning adaptation module, a sparse activation gating module, an anti-fact confusion pooling module, a small sample contradiction resolution module, a multi-modal memory bank module and a modular parameter isolation module. According to the method, dependence of a traditional method on explicit alignment signals is broken through, a meta-learning mechanism and differential topology constraints are introduced, multi-source noise separation of a gradient space and physical isolation of parameter activation are achieved, the defect of dimension coupling of static fusion is overcome, contradiction arbitration, anti-fact correction and hardware-level isolation are seamlessly integrated through systematic modular design, and the method has the advantages of being simple in structure and convenient to operate. While complementarity of multi-source data is dynamically maintained, cascade diffusion of prejudice propagation and long-tail conflicts is accurately inhibited, and a full-link fusion control capability from data representation to computing hardware is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and in particular to a large-model data fusion method and system based on multi-source heterogeneous data. Background Art

[0002] Multi-source heterogeneous data refers to diverse data from different collection sources, storage formats, and semantic structures. It encompasses multiple modalities, including text, images, video, and time series signals. Its inherent representational spaces, distribution characteristics, and semantic granularity vary significantly. This type of data exhibits dynamic evolution across time and space, and the implicit connections between data from different sources are often obscured by surface heterogeneity, making it difficult for traditional methods to effectively mine deep cross-modal consistency and suppress noise interference.

[0003] Traditional methods typically rely on static alignment rules or shallow statistical matching, failing to model the nonlinear causal relationships and dynamic coupling effects between multi-source data. They are susceptible to majority class dominance when dealing with long-tail distributions, lack robustness to the accumulation of implicit biases, and their parameter update mechanisms struggle to distinguish between semantic representations and bias propagation paths. Furthermore, rigid fusion strategies can easily lead to the curse of dimensionality, while discrete conflict resolution modules hinder end-to-end collaborative optimization.

[0004] Based on this, the present invention provides a large-scale model data fusion method and system based on multi-source heterogeneous data to solve the technical problems raised above. Summary of the Invention

[0005] The purpose of the present invention is to provide a large-scale model data fusion method and system based on multi-source heterogeneous data to solve the problems raised by the above-mentioned background technology.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] The first aspect of the present invention:

[0008] It provides a large-model data fusion system based on multi-source heterogeneous data, including a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous meta-learning adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, a multimodal memory library module, and a modular parameter isolation module.

[0009] The multimodal causal decoupling module is used for causal graph construction and bias path separation;

[0010] The dynamic subspace routing module is used for multi-armed bandit routing and subspace conflict detection;

[0011] The heterogeneous meta-learning adaptation module is used for implicit meta-feature extraction and meta-gradient reweighting;

[0012] The sparse activation gating module is used to provide Top-K dynamic gating mechanism and conflict-aware sparsification;

[0013] The counterfactual confusion pooling module is used for adversarial confusion generation and pooling attention decoupling;

[0014] The small sample contradiction resolution module is used for contradiction graph reasoning and implicit vote alignment;

[0015] The multimodal memory library module is used for temporal memory slicing and conflict sample playback;

[0016] The modular parameter isolation module is used for parameter cluster discovery and hard isolation gating.

[0017] Preferably, the multimodal causal decoupling module further comprises a causal graph construction unit and a bias path separation unit;

[0018] The causal graph construction unit is used to conduct intervention experiments on latent variables of multimodal data, construct cross-modal causal graphs, and locate bias collaborative paths;

[0019] The bias path separation unit is used to apply an intervention mask to the high-risk paths identified in the causal graph, and separate the coupling parameters of bias propagation and semantic representation.

[0020] Preferably, the dynamic subspace routing module further includes a multi-armed bandit routing unit and a subspace conflict detection unit;

[0021] The multi-armed bandit routing unit is used to dynamically select a sparse subspace related to the current sample to avoid bias parameters from contaminating small sample features;

[0022] The subspace conflict detection unit is used to monitor the update direction conflicts of different subspace parameters in real time, and block the influence of long-tail error labels on the majority class.

[0023] Preferably, the heterogeneous meta-learning adaptation module further includes an implicit meta-feature extraction unit and a meta-gradient reweighting unit;

[0024] The implicit meta-feature extraction unit is used to extract meta-features related to bias and small samples from multi-source heterogeneous data and generate dynamic weights;

[0025] The meta-gradient reweighting unit is used to perform latent variable decomposition on the gradient of small sample conflict instances and eliminate abnormal gradient components generated by bias collaboration.

[0026] Preferably, the sparse activation gating module further includes a Top-K dynamic gating unit and a conflict-aware sparsification unit;

[0027] The Top-K dynamic gating unit only activates a subset of K parameters related to the current task, which is used to enforce isolation of high-dimensional channels for bias propagation;

[0028] The conflict-aware sparsification unit is used to impose L0 constraints on neurons with frequent long-tail conflicts, so that they are naturally eliminated or weakened during training.

[0029] Preferably, the counterfactual confusion pooling module further includes an adversarial confusion generation unit and a pooling attention decoupling unit;

[0030] The adversarial obfuscation generation unit is used to generate counterfactual samples containing known bias combinations to decouple the superposition effect of multi-source biases;

[0031] The pooled attention decoupling unit is used to separate the feature responses of counterfactual samples and real samples in the multi-head attention layer, blocking the cross-head propagation of biased features.

[0032] Preferably, the small sample contradiction resolution module further includes a contradiction graph reasoning unit and an implicit voting alignment unit;

[0033] The contradiction graph inference unit is used to construct a long-tail small sample labeled contradiction graph and infer the potential true category;

[0034] The implicit voting alignment unit is used to map multi-source annotations to the optimal transmission space, calculate implicit voting weights, and avoid majority suppression problems in explicit voting.

[0035] Preferably, the multimodal memory library module further includes a temporal memory slicing unit and a conflict sample playback unit;

[0036] The temporal memory slicing unit is used to store the historical states of bias patterns of different modalities according to the time dimension, so as to detect the evolution path of collaborative bias;

[0037] The conflict sample replay unit is used to dynamically retain high-information instances of long-tail conflict samples to prevent them from being submerged during training.

[0038] Preferably, the modular parameter isolation module further comprises a parameter cluster discovery unit and a hard isolation gating unit;

[0039] The parameter cluster discovery unit is used to perform spectral clustering on the model parameters according to the gradient propagation path to identify parameter clusters associated with specific biases or conflicts;

[0040] The hard isolation gating unit is used to impose physical isolation on high-risk parameter clusters to prevent them from being activated in specific tasks.

[0041] Based on the above system, the present invention also proposes a large model data fusion method, which is applied to a large model data fusion system based on multi-source heterogeneous data, and includes the following steps:.

[0042] S1. Apply structured intervention to multi-source heterogeneous data, construct a cross-modal causal graph, and define the intervention process, as shown in formula (1):

[0043]

[0044] Where, v m is the original feature of mode m, is the feature after intervention on mode n, W int is the learnable intervention weight matrix, represents vector concatenation, where:

[0045] The causal alignment loss function quantifies the strength of the causal relationship between modalities through the second-order derivative, identifying the superposition path of hidden biases between multiple modalities;

[0046] S2. Causal strength matrix based on step S1 The improved adversarial multi-armed bandit is used to select the subspace, see formula (2):

[0047]

[0048] Where, t k is the number of historical selections of subspace k, ξ is the distribution shift penalty coefficient, JS is the Jensen-Shannon divergence, p current is the current batch data distribution, For the historical data distribution of subspace k, dynamically select the subspace that best matches the current data distribution and has stable gradient updates, blocking the distribution drift caused by incorrect labeling of long-tail samples;

[0049] S3. Perform gradient latent variable decomposition on the selected subspace data, see formula (3):

[0050]

[0051] in, is the k-nearest neighbor set of sample i, MMD is the maximum mean difference metric, is the cleaned latent variable space, κ is the suppression coefficient, and by comparing the difference between the sample gradient and the neighborhood average gradient and combining the degree of deviation of the latent variable distribution, the abnormal gradient components caused by the superposition of multi-source bias are eliminated;

[0052] S4. Based on the purified gradient of step S3, a sparse activation with differential topology constraints is applied, as shown in Equation (4):

[0053]

[0054] Where ST is the Straight-Through estimator, h l-1 is the activation value of the previous layer, Ric l is the Ricci curvature of the l-th layer parameter, ζ is the curvature threshold, and the high curvature parameters are identified by using the differential geometry properties. Only the parameters in the low curvature area are allowed to participate in the update, thus achieving bias isolation at the physical level;

[0055] S5. Generate combined counterfactual samples and reconstruct the attention mechanism, see formula (5):

[0056]

[0057] Where, is the key-value vector of the counterfactual sample, and d is the vector dimension. This splits the traditional Softmax into a two-way competition between real and counterfactual, forcing the model to explicitly distinguish biased combination features at the attention layer. For example, a counterfactual sample containing both the text "female" and the image "kitchen" will suppress the spurious association between gender and scene in the original data.

[0058] S6. Implicit arbitration of long-tail conflict annotations, see formula (6):

[0059]

[0060] Where, is the transmission matrix constraint, μ yj For category y j The prototype vector, D IPM is the integrated probability metric distance, T i,j It is an entropy regularization term that solves the majority suppression problem in small sample annotation conflicts by jointly optimizing feature space alignment and distribution matching.

[0061] S7. Dynamically update the memory library and strengthen the key samples, see formula (7):

[0062]

[0063] Where Conv1D is a one-dimensional convolution in the time dimension, and Attn is the attention weight based on the memory bank. The convolution operation captures the temporal evolution of bias patterns, such as the propagation dynamics of emerging bias terms in social media data, ensuring that the memory bank can retain key historical states.

[0064] S8. Finally, perform spectral clustering isolation on the parameters, see formula (8):

[0065]

[0066] Where Tr is the matrix trace operation, Σ j is the covariance matrix of cluster j parameters, η is the gradient covariance penalty coefficient, and clusters with strong correlation with bias propagation in parameter space are identified. The physical isolation of video memory is achieved by modifying the CUDA kernel. k Map it to an independent video memory page, completely blocking its participation in the forward propagation.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] The present invention explicitly decouples the implicit correlations and bias paths of multimodal data through the collaboration of causal reasoning and dynamic subspace routing, breaking through the traditional method's reliance on explicit alignment signals. The introduced meta-learning mechanism and differential topology constraints achieve multi-source noise separation in gradient space and physical isolation of parameter activation, overcoming the dimensional coupling defects of static fusion. At the same time, the systematic modular design of the present invention seamlessly integrates contradiction arbitration, counterfactual correction and hardware-level isolation, while dynamically maintaining the complementarity of multi-source data, accurately suppressing the cascade diffusion of bias propagation and long-tail conflicts, and forming a full-link fusion control capability from data representation to computing hardware. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a topology diagram of the large-scale data fusion system based on multi-source heterogeneous data of the present invention;

[0070] Figure 2 This is a flow chart of the large model data fusion method based on multi-source heterogeneous data of the present invention. DETAILED DESCRIPTION

[0071] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0072] Example 1, please refer to Figure 1 , the present invention proposes a large-model data fusion system based on multi-source heterogeneous data, including a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous meta-learning adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, a multimodal memory library module and a modular parameter isolation module;

[0073] Among them, it should be noted that the multimodal causal decoupling module is used for causal graph construction and bias path separation, the dynamic subspace routing module is used for multi-armed bandit routing and subspace conflict detection, the heterogeneous meta-learning adaptation module is used for implicit meta-feature extraction and meta-gradient reweighting, the sparse activation gating module is used to provide Top-K dynamic gating mechanism and conflict-aware sparsification, the counterfactual confusion pooling module is used for adversarial confusion generation and pooled attention decoupling, the small sample contradiction resolution module is used for contradiction graph reasoning and implicit voting alignment, the multimodal memory library module is used for temporal memory slicing and conflict sample replay, and the modular parameter isolation module is used for parameter cluster discovery and hard isolation gating.

[0074] In this embodiment, it should also be noted that the multimodal causal decoupling module also includes a causal graph construction unit and a bias path separation unit;

[0075] The causal graph construction unit is used to conduct intervention experiments on latent variables of multimodal data, construct cross-modal causal graphs, and locate bias collaborative paths;

[0076] The bias path separation unit is used to apply an intervention mask to the high-risk paths identified in the causal graph, and separate the coupling parameters of bias propagation and semantic representation.

[0077] In this embodiment, it should also be noted that the dynamic subspace routing module further includes a multi-armed bandit routing unit and a subspace conflict detection unit;

[0078] The multi-armed bandit routing unit is used to dynamically select a sparse subspace related to the current sample to avoid bias parameters from contaminating small sample features;

[0079] The subspace conflict detection unit is used to monitor the update direction conflicts of different subspace parameters in real time, and block the influence of long-tail error labels on the majority class.

[0080] In this embodiment, it should also be noted that the heterogeneous meta-learning adaptation module further includes an implicit meta-feature extraction unit and a meta-gradient reweighting unit;

[0081] The implicit meta-feature extraction unit is used to extract meta-features related to bias and small samples from multi-source heterogeneous data and generate dynamic weights;

[0082] The meta-gradient reweighting unit is used to perform latent variable decomposition on the gradient of small sample conflict instances and eliminate abnormal gradient components generated by bias collaboration.

[0083] In this embodiment, it should also be noted that the sparse activation gating module also includes a Top-K dynamic gating unit and a conflict-aware sparsification unit;

[0084] The Top-K dynamic gating unit only activates a subset of K parameters related to the current task, which is used to enforce isolation of high-dimensional channels for bias propagation;

[0085] The conflict-aware sparsification unit is used to impose L0 constraints on neurons with frequent long-tail conflicts, so that they are naturally eliminated or weakened during training.

[0086] In this embodiment, it should also be noted that the counterfactual confusion pooling module also includes an adversarial confusion generation unit and a pooling attention decoupling unit;

[0087] The adversarial obfuscation generation unit is used to generate counterfactual samples containing known bias combinations to decouple the superposition effect of multi-source biases;

[0088] The pooled attention decoupling unit is used to separate the feature responses of counterfactual samples and real samples in the multi-head attention layer, blocking the cross-head propagation of biased features.

[0089] In this embodiment, it should also be noted that the small sample contradiction resolution module also includes a contradiction graph reasoning unit and an implicit voting alignment unit;

[0090] The contradiction graph inference unit is used to construct a long-tail small sample labeled contradiction graph and infer the potential true category;

[0091] The implicit voting alignment unit is used to map multi-source annotations to the optimal transmission space, calculate implicit voting weights, and avoid majority suppression problems in explicit voting.

[0092] In this embodiment, it should also be noted that the multimodal memory library module further includes a temporal memory slicing unit and a conflict sample playback unit;

[0093] The temporal memory slicing unit is used to store the historical states of bias patterns of different modalities according to the time dimension, so as to detect the evolution path of collaborative bias;

[0094] The conflict sample replay unit is used to dynamically retain high-information instances of long-tail conflict samples to prevent them from being submerged during training.

[0095] In this embodiment, it should also be noted that the modular parameter isolation module further includes a parameter cluster discovery unit and a hard isolation gating unit;

[0096] The parameter cluster discovery unit is used to perform spectral clustering on the model parameters according to the gradient propagation path to identify parameter clusters associated with specific biases or conflicts;

[0097] The hard isolation gating unit is used to impose physical isolation on high-risk parameter clusters to prevent them from being activated in specific tasks.

[0098] Example 2, please refer to Figure 2 In this embodiment, based on the above system, the present invention also proposes a large model data fusion method, which is applied to a large model data fusion system based on multi-source heterogeneous data, including the following steps:.

[0099] S1. Apply structured intervention to multi-source heterogeneous data, construct a cross-modal causal graph, and define the intervention process, as shown in formula (1):

[0100]

[0101] Where, v m is the original feature of mode m, is the feature after intervention on mode n, W int is the learnable intervention weight matrix, represents vector concatenation, where:

[0102] The causal alignment loss function quantifies the strength of the causal relationship between modalities through the second-order derivative, identifying the superposition path of hidden biases between multiple modalities;

[0103] In practical applications, when v m For the word vector of "nurse", By applying intervention to the image modality, replacing the gender of the person in the image with male, and calculating the causal strength between the text features and the intervention image, we found a strong correlation path between the "nurse" text and female images (causal weight > 0.8), locating the bias propagation link.

[0104] S2. Causal strength matrix based on step S1 The improved adversarial multi-armed bandit is used to select the subspace, see formula (2):

[0105]

[0106] Where, t k is the number of historical selections of subspace k, ξ is the distribution shift penalty coefficient, JS is the Jensen-Shannon divergence, p current is the current batch data distribution, For the historical data distribution of subspace k, dynamically select the subspace that best matches the current data distribution and has stable gradient updates, blocking the distribution drift caused by incorrect labeling of long-tail samples;

[0107] In practical applications, current The current batch data distribution is 80% common occupations and 20% unpopular occupations. The historical distribution of subspace k is that 90% of the medical subspace has been "doctor" data in the past. The medical subspace (k=5) is dynamically selected because its gradient stability is the highest, isolating the impact of incorrect labeling of long-tail unpopular occupations such as "underwater welder" on the mainstream data.

[0108] S3. Perform gradient latent variable decomposition on the selected subspace data, see formula (3):

[0109]

[0110] in, is the k-nearest neighbor set of sample i, MMD is the maximum mean difference metric, is the cleaned latent variable space, κ is the suppression coefficient, and by comparing the difference between the sample gradient and the neighborhood average gradient and combining the degree of deviation of the latent variable distribution, the abnormal gradient components caused by the superposition of multi-source bias are eliminated;

[0111] In practical applications, The k nearest neighbors of sample i are 10 samples related to "nurse", and the unbiased latent space after cleaning is To ensure a balanced distribution of male and female nurse features, for the "nurse + female image" sample, the difference between its gradient and the neighborhood (angle > 60°) is calculated. Combined with the latent space deviation (MMD = 1.4), 70% of abnormal gradients are suppressed to prevent bias from propagating in multimodal fusion.

[0112] S4. Based on the purified gradient of step S3, a sparse activation with differential topology constraints is applied, as shown in Equation (4):

[0113]

[0114] Where ST is the Straight-Through estimator, h l-1 is the activation value of the previous layer, Ric l is the Ricci curvature of the l-th layer parameter, ζ is the curvature threshold, and the high curvature parameters are identified by using the differential geometry properties. Only the parameters in the low curvature area are allowed to participate in the update, thus achieving bias isolation at the physical level;

[0115] In actual application, Ric l The curvature of the parameters in the first layer is 3.2 for a neuron in the fully connected layer. The ζ curvature threshold (2.5) filters out high curvature parameters, shuts down neurons with excessive curvature, and retains only low curvature parameters, thus blocking the coupling between “gender” and “occupation” in the parameter space, such as the feature extraction node for medical equipment.

[0116] S5. Generate combined counterfactual samples and reconstruct the attention mechanism, see formula (5):

[0117]

[0118] Where, is the key-value vector of the counterfactual sample, d is the vector dimension, and the traditional Softmax is split into a real-counterfactual two-way competition, forcing the model to explicitly distinguish biased combination features at the attention layer, for example:

[0119] a. The counterfactual sample containing both “female” text and “kitchen” images will suppress the spurious gender-scene association in the original data;

[0120] b. When the key value of the counterfactual sample is the visual feature of "nurse + male image", d is 768 dimensions. In the attention layer, the model is forced to compare the real sample ("nurse + female image") with the counterfactual sample ("nurse + male image"), and the weight of the gender-related attention head is increased from 0.3 to 0.7 to strengthen the occupation-related features.

[0121] S6. Implicit arbitration of long-tail conflict annotations, see formula (6):

[0122]

[0123] Where, is the transmission matrix constraint, μ yj For category y j The prototype vector, D IPM is the integrated probability metric distance, T i,j It is an entropy regularization term that solves the majority suppression problem in small sample annotation conflicts by jointly optimizing feature space alignment and distribution matching.

[0124] In practical applications, μ yj The prototype of the category "Underwater Welder" is welding tools + diving equipment features, D IPM The distribution distance metric compares the distribution differences between annotation sources A and B. The optimal transfer method aligns the distributions of annotation sources A (correct annotation) and B (incorrect annotation). The output arbitration weight is 0.8, and “underwater welder” is determined to be the true category.

[0125] S7. Dynamically update the memory library and strengthen the key samples, see formula (7):

[0126]

[0127] Where Conv1D is a one-dimensional convolution in the time dimension, Attn is the attention weight based on the memory bank, and the temporal evolution of the bias pattern is captured through the convolution operation, for example:

[0128] The dynamics of the spread of emerging biased terms in social media data, ensuring that memory banks can preserve historically critical states;

[0129] In practical applications, for example, the significant gradient of key samples, such as the gradient of the emerging bias "Programmer → Male with Glasses," Conv1D captures a temporal pattern of bias growth of 12% per week, stores highly significant samples (gradient norm > 1.5 and entropy > 1.2), identifies bias evolution patterns through temporal convolution, and dynamically updates the model to cover emerging bias scenarios.

[0130] S8. Finally, perform spectral clustering isolation on the parameters, see formula (8):

[0131]

[0132] Where Tr is the matrix trace operation, Σ j is the covariance matrix of cluster j parameters, η is the gradient covariance penalty coefficient, and clusters with strong correlation with bias propagation in parameter space are identified. The physical isolation of video memory is achieved by modifying the CUDA kernel. k Mapped to an independent video memory page, completely blocking its participation in the forward propagation;

[0133] In practical application, Σ j The covariance matrix of the parameter cluster, such as the gender-related parameter covariance of cluster C2, is greater than 0.7. The Tr matrix trace operation quantifies the gradient correlation, isolates the identified bias parameter cluster C2 to an independent video memory page, and skips the calculation of this area in the forward propagation, which reduces the standard deviation of gender prediction for the "nurse" image by 43%.

[0134] Through the above steps, after step S1 of the present invention locates the multimodal bias path through causal discovery, long-tail noise is isolated through subspace routing in step S2, and then multi-source data gradients are purified by gradient decomposition in step S3. Physical blocking bias parameters are sparsely activated in step S4, and then counterfactual correction is performed to reconstruct cross-modal attention in step S5. Conflicting annotations are fused through the arbitration mechanism in step S6. After the temporal patterns are continuously fused through the memory bank in S7, finally, hardware isolation in step S8 is performed to ensure that the fusion result is not contaminated, thereby improving the accuracy of gender-occupation bias detection, and improving the conflict resolution rate of unpopular occupation annotations, with higher arbitration accuracy.

[0135] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0136] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A large-scale data fusion system based on multi-source heterogeneous data, characterized by: It includes a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous meta-learning adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, a multimodal memory bank module, and a modular parameter isolation module; The multimodal causal decoupling module is used for causal graph construction and bias path separation; The dynamic subspace routing module is used for multi-armed bandit routing and subspace conflict detection; The heterogeneous meta-learning adaptation module is used for implicit meta-feature extraction and meta-gradient reweighting; The sparse activation gating module is used to provide Top-K dynamic gating mechanism and conflict-aware sparsification; The counterfactual confusion pooling module is used for adversarial confusion generation and pooling attention decoupling; The small sample contradiction resolution module is used for contradiction graph reasoning and implicit vote alignment; The multimodal memory library module is used for temporal memory slicing and conflict sample playback; The modular parameter isolation module is used for parameter cluster discovery and hard isolation gating.

2. The large-scale model data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The multimodal causal decoupling module further includes a causal graph construction unit and a bias path separation unit; The causal graph construction unit is used to conduct intervention experiments on latent variables of multimodal data, construct cross-modal causal graphs, and locate bias collaborative paths; The bias path separation unit is used to apply an intervention mask to the high-risk paths identified in the causal graph, and separate the coupling parameters of bias propagation and semantic representation.

3. The large-scale data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The dynamic subspace routing module also includes a multi-armed bandit routing unit and a subspace conflict detection unit; The multi-armed bandit routing unit is used to dynamically select a sparse subspace related to the current sample to avoid bias parameters from contaminating small sample features; The subspace conflict detection unit is used to monitor the update direction conflicts of different subspace parameters in real time, and block the influence of long-tail error labels on the majority class.

4. The large-scale data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The heterogeneous meta-learning adaptation module further includes an implicit meta-feature extraction unit and a meta-gradient reweighting unit; The implicit meta-feature extraction unit is used to extract meta-features related to bias and small samples from multi-source heterogeneous data and generate dynamic weights; The meta-gradient reweighting unit is used to perform latent variable decomposition on the gradient of small sample conflict instances and eliminate abnormal gradient components generated by bias collaboration.

5. The large-scale data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The sparse activation gating module also includes a Top-K dynamic gating unit and a conflict-aware sparsification unit; The Top-K dynamic gating unit only activates a subset of K parameters related to the current task, which is used to enforce isolation of high-dimensional channels for bias propagation; The conflict-aware sparsification unit is used to impose L0 constraints on neurons with frequent long-tail conflicts, so that they are naturally eliminated or weakened during training.

6. The large-scale model data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The counterfactual confusion pooling module also includes an adversarial confusion generation unit and a pooling attention decoupling unit; The adversarial obfuscation generation unit is used to generate counterfactual samples containing known bias combinations to decouple the superposition effect of multi-source biases; The pooled attention decoupling unit is used to separate the feature responses of counterfactual samples and real samples in the multi-head attention layer, blocking the cross-head propagation of biased features.

7. The large-scale data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The small sample contradiction resolution module further includes a contradiction graph reasoning unit and an implicit voting alignment unit; The contradiction graph inference unit is used to construct a long-tail small sample labeled contradiction graph and infer the potential true category; The implicit voting alignment unit is used to map multi-source annotations to the optimal transmission space, calculate implicit voting weights, and avoid majority suppression problems in explicit voting.

8. The large-scale data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The multimodal memory library module further includes a temporal memory slicing unit and a conflict sample playback unit; The temporal memory slicing unit is used to store the historical states of bias patterns of different modalities according to the time dimension, so as to detect the evolution path of collaborative bias; The conflict sample replay unit is used to dynamically retain high-information instances of long-tail conflict samples to prevent them from being submerged during training.

9. The large-scale model data fusion system based on multi-source heterogeneous data according to claim 1 is characterized in that: The modular parameter isolation module also includes a parameter cluster discovery unit and a hard isolation gating unit; The parameter cluster discovery unit is used to perform spectral clustering on the model parameters according to the gradient propagation path to identify parameter clusters associated with specific biases or conflicts; The hard isolation gating unit is used to impose physical isolation on high-risk parameter clusters to prevent them from being activated in specific tasks.

10. A large-scale model data fusion method, applied to a large-scale model data fusion system based on multi-source heterogeneous data according to any one of claims 1 to 9, characterized in that: The following steps are involved: 。 S1. Apply structured intervention to multi-source heterogeneous data, construct a cross-modal causal diagram, define the intervention process, quantify the strength of causal relationships between modalities through second-order derivatives, and identify the superposition path of hidden biases across multiple modalities; S2. Based on the causal strength matrix from step S1, an improved adversarial multi-armed bandit is used to select subspaces. This dynamically selects the subspace that best matches the current data distribution and has stable gradient updates, thereby blocking distribution drift caused by incorrect labeling of long-tail samples. S3. Perform gradient latent variable decomposition on the selected subspace data. By comparing the difference between the sample gradient and the neighborhood average gradient and combining the degree of latent variable distribution deviation, the abnormal gradient components caused by the superposition of multiple sources of bias are eliminated. S4. Based on the purified gradient in step S3, sparse activation with differential topology constraints is applied. High curvature parameters are identified using differential geometric properties, and only parameters in low curvature regions are allowed to participate in the update, thus achieving bias isolation at the physical level. S5. Generate combined counterfactual samples and restructure the attention mechanism, splitting the traditional Softmax into a two-way competition between truth and counterfactual. This forces the model to explicitly distinguish biased combined features at the attention layer, suppressing spurious gender-scene correlations in the original data. S6. Implicit arbitration of long-tail conflicting annotations, solving the majority suppression problem in small sample annotation conflicts by jointly optimizing feature space alignment and distribution matching; S7. Dynamically update the memory bank and strengthen key samples, capturing the temporal evolution of bias patterns through convolution operations; S8. Finally, spectral clustering is performed on the parameters to identify clusters in the parameter space that are strongly correlated with bias propagation, and physical isolation of the video memory is achieved by modifying the CUDA kernel.

Citation Information

Patent Citations

  • Multi-source heterogeneous data fusion architecture system based on additive model

    CN107798137A

  • Feature-adaptive mutual-guiding multi-source information fusion classification method and system

    CN114187526A

  • Training method and apparatus for human-factor intelligence state monitoring model, and human-factor intelligence state monitoring method and apparatus

    WO2025030989A1