A method and system for large-scale model data fusion based on multi-source heterogeneous data
By employing modular designs such as multimodal causal decoupling and dynamic subspace routing, the cross-modal fusion challenge of multi-source heterogeneous data is solved, achieving decoupling of implicit associations and biased paths, and improving the accuracy and robustness of data fusion.
Patent Information
- Application Number
- CN202510518863.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Traditional methods struggle to effectively uncover deep cross-modal consistency in multi-source heterogeneous data, are susceptible to noise interference, lack robustness against implicit biases, and lead to difficulties in resolving the dimensionality curse and conflicts.
By employing a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous meta-learning adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, and a multimodal memory module, combined with causal graph construction, dynamic subspace routing, meta-learning adaptation, and differential topological constraints, we can achieve the decoupling of implicit associations and biased paths in multi-source data and the physical isolation of parameters.
It achieves end-to-end collaborative optimization of multi-source heterogeneous data, accurately suppresses the propagation of bias and long-tail conflicts, and improves the accuracy and robustness of data fusion.
Smart Images

Figure CN120470520B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data fusion technology, specifically to a method and system for large-scale model data fusion based on multi-source heterogeneous data. Background Technology
[0002] Multi-source heterogeneous data refers to diverse data from different acquisition sources, storage formats, and semantic structures, encompassing multiple modalities such as text, images, videos, and time-series signals. These data exhibit significant differences in their inherent representational space, distribution characteristics, and semantic granularity. Such data possesses dynamic evolutionary characteristics in the spatiotemporal dimensions, and the implicit correlations between different source data are often obscured by surface heterogeneity, making it difficult for traditional methods to effectively uncover deep cross-modal consistency and suppress noise interference.
[0003] Traditional methods typically rely on static alignment rules or shallow statistical matching, failing to model nonlinear causal relationships and dynamic coupling effects among multi-source data. They are susceptible to majority class dominance when dealing with long-tailed distributions, lack robustness against the accumulation of implicit biases, and their parameter update mechanisms struggle to distinguish between semantic representations and bias propagation paths. Furthermore, rigid fusion strategies are prone to the curse of dimensionality, while discrete conflict resolution modules struggle to achieve end-to-end collaborative optimization.
[0004] Based on this, the present invention provides a method and system for large model data fusion based on multi-source heterogeneous data to solve the above-mentioned technical problems. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for large-scale model data fusion based on multi-source heterogeneous data, thereby solving the problems mentioned in the background.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] First aspect of the invention:
[0008] It provides a large model data fusion system based on multi-source heterogeneous data, including a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous meta-learning adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, a multimodal memory module, and a modular parameter isolation module;
[0009] The multimodal causal decoupling module is used for causal graph construction and biased path separation;
[0010] The dynamic subspace routing module is used for multi-armed gambling machine routing and subspace conflict detection;
[0011] The heterogeneous meta-learning adaptation module is used for implicit meta-feature extraction and meta-gradient reweighting;
[0012] The sparse activation gating module is used to provide a Top-K dynamic gating mechanism and conflict-aware sparsification.
[0013] The counterfactual obfuscation pooling module is used for adversarial obfuscation generation and pooling attention decoupling;
[0014] The small-sample conflict resolution module is used for conflict graph reasoning and implicit voting alignment;
[0015] The multimodal memory module is used for temporal memory slicing and conflict sample playback;
[0016] The modular parameter isolation module is used for parameter cluster discovery and hard isolation gating.
[0017] Preferably, the multimodal causal decoupling module further includes a causal graph construction unit and a biased path separation unit;
[0018] The causal graph construction unit is used to conduct intervention experiments on latent variables of multimodal data, construct cross-modal causal graphs, and locate biased collaborative paths;
[0019] The bias path separation unit is used to apply an intervention mask to the high-risk paths identified in the causal graph, thereby separating the coupling parameters between bias propagation and semantic representation.
[0020] Preferably, the dynamic subspace routing module further includes a multi-armed gambling machine routing unit and a subspace conflict detection unit;
[0021] The multi-armed gambling machine routing unit is used to dynamically select the sparse subspace related to the current sample to avoid bias parameters contaminating small sample features.
[0022] The subspace conflict detection unit is used to monitor the conflict of update directions of different subspace parameters in real time, and to block the impact of long-tailed error labeling on the majority class.
[0023] Preferably, the heterogeneous meta-learning adaptation module further includes an implicit meta-feature extraction unit and a meta-gradient reweighting unit;
[0024] The implicit meta-feature extraction unit is used to extract meta-features related to bias and small sample size from multi-source heterogeneous data and generate dynamic weights.
[0025] The meta-gradient reweighting unit is used to perform latent variable decomposition on the gradients of small-sample conflict instances, eliminating abnormal gradient components generated by bias collaboration.
[0026] Preferably, the sparse activation gating module further includes a Top-K dynamic gating unit and a conflict-aware sparsification unit;
[0027] The Top-K dynamic gating unit activates only a subset of K parameters related to the current task, which is used to forcibly isolate the high-dimensional channels of bias propagation;
[0028] The conflict-aware sparsification unit is used to impose L0 constraints on neurons with frequent long-tail conflicts, so that they are naturally eliminated or weakened during training.
[0029] Preferably, the antifactual obfuscation pooling module further includes an adversarial obfuscation generation unit and a pooling attention decoupling unit;
[0030] The adversarial confusion generation unit is used to generate counterfactual samples containing combinations of known biases, in order to decouple the cumulative effect of multi-source biases;
[0031] The pooling attention decoupling unit is used to separate the feature responses of counterfactual samples and true samples in the multi-head attention layer, thereby blocking the cross-head propagation of biased features.
[0032] Preferably, the small-sample contradiction resolution module further includes a contradiction graph reasoning unit and an implicit voting alignment unit;
[0033] The contradiction graph reasoning unit is used to construct a labeled contradiction graph of long-tailed small samples and reason about the potential true category;
[0034] The implicit voting alignment unit is used to map multi-source labels to the optimal transmission space, calculate implicit voting weights, and avoid the majority suppression problem in explicit voting.
[0035] Preferably, the multimodal memory module further includes a temporal memory slicing unit and a conflict sample playback unit;
[0036] The temporal memory slicing unit is used to store the historical state of bias patterns of different modalities in a time dimension, and is used to detect the evolution path of collaborative bias.
[0037] The conflict sample replay unit is used to dynamically retain high-information instances of long-tail conflict samples to prevent them from being overwhelmed during training.
[0038] Preferably, the modular parameter isolation module further includes a parameter cluster discovery unit and a hard isolation gating unit;
[0039] The parameter cluster discovery unit is used to perform spectral clustering of model parameters based on gradient propagation paths to identify parameter clusters related to specific biases or conflicts.
[0040] The hard isolation gating unit is used to apply physical isolation to high-risk parameter clusters, preventing them from being activated in specific tasks.
[0041] Based on the above system, this invention also proposes a large model data fusion method, applied to a large model data fusion system based on multi-source heterogeneous data, comprising the following steps:
[0042] S1. Apply structured intervention to multi-source heterogeneous data, construct a cross-modal causal graph, and define the intervention process, as shown in equation (1):
[0043]
[0044] In the formula, v m For the original features of mode m, To describe the characteristics of mode n after intervention, W int The intervention weight matrix is a learnable matrix. This represents vector concatenation, where:
[0045] As a causal alignment loss function, the strength of causal relationships between modes is quantified by the second derivative, and the superposition path of hidden biases among multiple modes is identified.
[0046] S2. Causality strength matrix based on step S1 An improved selection subspace for adversarial multi-armed gambling machines is adopted, as shown in Equation (2):
[0047]
[0048] In the formula, t k Let ξ be the number of historical choices for subspace k, ξ be the distribution offset penalty coefficient, JS be the Jensen-Shannon divergence, and p be the number of choices for subspace k. current This represents the data distribution for the current batch. Given the historical data distribution of subspace k, dynamically select the subspace that best matches the current data distribution and has stable gradient updates to block distribution drift caused by mislabeling of long-tail samples;
[0049] S3. Perform gradient latent variable decomposition on the selected subspace data, as shown in equation (3):
[0050]
[0051] in, Let be the set of k nearest neighbors of sample i, and MMD be the maximum mean difference measure. The cleaned latent variable space is κ, which is the suppression coefficient. By comparing the difference between the sample gradient and the neighborhood average gradient, and combining the degree of deviation of the latent variable distribution, abnormal gradient components caused by the superposition of multi-source bias are eliminated.
[0052] S4. Based on the purification gradient of step S3, apply sparse activation with differential topological constraints, as shown in equation (4):
[0053]
[0054] In the formula, ST is the Straight-Through estimator, h l-1 Ric is the activation value of the previous layer. l Let ζ be the Ricci curvature of the parameters in the l-th layer, and ζ be the curvature threshold. High curvature parameters are identified using differential geometry properties, and only parameters in low curvature regions are allowed to participate in the update, thus achieving physical bias isolation.
[0055] S5. Generate combined counterfactual samples and reconstruct the attention mechanism, as shown in Equation (5):
[0056]
[0057] In the formula, d is the key-value vector of the counterfactual sample, and d is the vector dimension. The traditional Softmax is split into a real-counterfactual dual-path competition, which forces the model to explicitly distinguish biased combination features at the attention layer. For example, counterfactual samples that contain both "woman" text and "kitchen" image will suppress the false association between gender and scene in the original data.
[0058] S6. Implicit arbitration is performed on long-tail conflict annotations, as shown in equation (6):
[0059]
[0060] In the formula, For the transfer matrix constraint, μ yj For category y j The prototype vector, D IPM The distance is measured by integral probability, H(T) = -∑ i,j T i,j logT i,j As an entropy regularization term, it solves the majority suppression problem in small sample labeling conflicts by jointly optimizing feature space alignment and distribution matching;
[0061] S7. Dynamically update the memory and strengthen key samples, as shown in equation (7):
[0062]
[0063] In the formula, Conv1D is a one-dimensional convolution in the time dimension, and Attn is an attention weight based on the memory bank. The convolution operation captures the temporal evolution of bias patterns, such as the propagation dynamics of emerging bias terms in social media data, and ensures that the memory bank can retain key historical states.
[0064] S8. Finally, the parameters are isolated by spectral clustering, as shown in equation (8):
[0065]
[0066] In the formula, Tr represents the trace operation of the matrix, and Σ j Let η be the covariance matrix of cluster j parameters, and η be the gradient covariance penalty coefficient. This identifies clusters in the parameter space that are strongly correlated with bias propagation, and achieves physical memory isolation by modifying the CUDA kernel. k Mapped to a separate memory page, completely blocking its participation in forward propagation.
[0067] Compared with the prior art, the beneficial effects of the present invention are:
[0068] This invention explicitly decouples implicit associations and biased paths in multimodal data through the synergy of causal reasoning and dynamic subspace routing. It breaks through the dependence of traditional methods on explicit alignment signals. The introduced meta-learning mechanism and differential topological constraints realize the separation of multi-source noise in gradient space and the physical isolation of parameter activation, overcoming the dimensional coupling defects of static fusion. At the same time, the systematic modular design of this invention seamlessly integrates contradiction arbitration, counterfactual correction and hardware-level isolation. While dynamically maintaining the complementarity of multi-source data, it accurately suppresses the cascading diffusion of bias propagation and long-tail conflicts, forming a full-link fusion control capability from data representation to computing hardware. Attached Figure Description
[0069] Figure 1 This is a topology diagram of the large model data fusion system based on multi-source heterogeneous data according to the present invention;
[0070] Figure 2 This is a flowchart of the large model data fusion method based on multi-source heterogeneous data of the present invention. Detailed Implementation
[0071] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0072] Example 1, please refer to Figure 1 This invention proposes a large model data fusion system based on multi-source heterogeneous data, including a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous element learning and adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, a multimodal memory module, and a modular parameter isolation module.
[0073] It should be noted that the multimodal causal decoupling module is used for causal graph construction and biased path separation; the dynamic subspace routing module is used for multi-armed gambling machine routing and subspace conflict detection; the heterogeneous meta-learning adaptation module is used for implicit meta-feature extraction and meta-gradient reweighting; the sparse activation gating module is used to provide a Top-K dynamic gating mechanism and conflict-aware sparsity; the counterfactual confusion pooling module is used for adversarial confusion generation and pooling attention decoupling; the few-shot conflict resolution module is used for conflict graph reasoning and implicit voting alignment; the multimodal memory module is used for temporal memory slicing and conflict sample replay; and the modular parameter isolation module is used for parameter cluster discovery and hard isolation gating.
[0074] In this embodiment, it should also be noted that the multimodal causal decoupling module further includes a causal graph construction unit and a biased path separation unit;
[0075] The causal graph construction unit is used to conduct intervention experiments on latent variables of multimodal data, construct cross-modal causal graphs, and locate biased collaborative paths;
[0076] The bias path separation unit is used to apply an intervention mask to the high-risk paths identified in the causal graph, thereby separating the coupling parameters between bias propagation and semantic representation.
[0077] In this embodiment, it should also be noted that the dynamic subspace routing module further includes a multi-armed gambling machine routing unit and a subspace conflict detection unit;
[0078] The multi-armed gambling machine routing unit is used to dynamically select the sparse subspace related to the current sample to avoid bias parameters contaminating small sample features.
[0079] The subspace conflict detection unit is used to monitor the conflict of update directions of different subspace parameters in real time, and to block the impact of long-tailed error labeling on the majority class.
[0080] In this embodiment, it should also be noted that the heterogeneous meta-learning adaptation module further includes an implicit meta-feature extraction unit and a meta-gradient reweighting unit.
[0081] The implicit meta-feature extraction unit is used to extract meta-features related to bias and small sample size from multi-source heterogeneous data and generate dynamic weights.
[0082] The meta-gradient reweighting unit is used to perform latent variable decomposition on the gradients of small-sample conflict instances, eliminating abnormal gradient components generated by bias collaboration.
[0083] In this embodiment, it should also be noted that the sparse activation gating module further includes a Top-K dynamic gating unit and a conflict-aware sparsification unit;
[0084] The Top-K dynamic gating unit activates only a subset of K parameters related to the current task, which is used to forcibly isolate the high-dimensional channels of bias propagation;
[0085] The conflict-aware sparsification unit is used to impose L0 constraints on neurons with frequent long-tail conflicts, so that they are naturally eliminated or weakened during training.
[0086] In this embodiment, it should also be noted that the antifactual obfuscation pooling module further includes an adversarial obfuscation generation unit and a pooling attention decoupling unit;
[0087] The adversarial confusion generation unit is used to generate counterfactual samples containing combinations of known biases, in order to decouple the cumulative effect of multi-source biases;
[0088] The pooling attention decoupling unit is used to separate the feature responses of counterfactual samples and true samples in the multi-head attention layer, thereby blocking the cross-head propagation of biased features.
[0089] In this embodiment, it should also be noted that the small sample contradiction resolution module further includes a contradiction graph reasoning unit and an implicit voting alignment unit;
[0090] The contradiction graph reasoning unit is used to construct a labeled contradiction graph of long-tailed small samples and reason about the potential true category;
[0091] The implicit voting alignment unit is used to map multi-source labels to the optimal transmission space, calculate implicit voting weights, and avoid the majority suppression problem in explicit voting.
[0092] In this embodiment, it should also be noted that the multimodal memory module further includes a temporal memory slicing unit and a conflict sample playback unit;
[0093] The temporal memory slicing unit is used to store the historical state of bias patterns of different modalities in a time dimension, and is used to detect the evolution path of collaborative bias.
[0094] The conflict sample replay unit is used to dynamically retain high-information instances of long-tail conflict samples to prevent them from being overwhelmed during training.
[0095] In this embodiment, it should also be noted that the modular parameter isolation module further includes a parameter cluster discovery unit and a hard isolation gating unit;
[0096] The parameter cluster discovery unit is used to perform spectral clustering of model parameters based on gradient propagation paths to identify parameter clusters related to specific biases or conflicts.
[0097] The hard isolation gating unit is used to apply physical isolation to high-risk parameter clusters, preventing them from being activated in specific tasks.
[0098] Example 2, please refer to Figure 2 In this embodiment, based on the above system, the present invention also proposes a large model data fusion method, applied to a large model data fusion system based on multi-source heterogeneous data, including the following steps:
[0099] S1. Apply structured intervention to multi-source heterogeneous data, construct a cross-modal causal graph, and define the intervention process, as shown in equation (1):
[0100]
[0101] In the formula, v m For the original features of mode m, To describe the characteristics of mode n after intervention, W int The intervention weight matrix is a learnable matrix. This represents vector concatenation, where:
[0102] As a causal alignment loss function, the strength of causal relationships between modes is quantified by the second derivative, and the superposition path of hidden biases among multiple modes is identified.
[0103] In practical applications, when v m When generating the word vector for "nurse", After intervening in the image modality, the gender of the people in the image was replaced with male. By intervening in the gender features of the image, the causal strength between the text features and the image after the intervention was calculated. A strong correlation path (causal weight > 0.8) was found between the text "nurse" and the image of a woman, and the bias propagation link was located.
[0104] S2. Causality strength matrix based on step S1 An improved selection subspace for adversarial multi-armed gambling machines is adopted, as shown in Equation (2):
[0105]
[0106] In the formula, t k Let ξ be the number of historical choices for subspace k, ξ be the distribution offset penalty coefficient, JS be the Jensen-Shannon divergence, and p be the number of choices for subspace k. current This represents the data distribution for the current batch. Given the historical data distribution of subspace k, dynamically select the subspace that best matches the current data distribution and has stable gradient updates to block distribution drift caused by mislabeling of long-tail samples;
[0107] In practical applications, p current The current batch of data shows that 80% are common occupations and 20% are less common occupations. The historical distribution of subspace k shows that the medical subspace has historically contained 90% "doctor" data. The medical subspace (k=5) is dynamically selected because it has the highest gradient stability, thus isolating the impact of erroneous labeling of long-tail, niche occupations such as "underwater welder" on the mainstream data.
[0108] S3. Perform gradient latent variable decomposition on the selected subspace data, as shown in equation (3):
[0109]
[0110] in, Let be the set of k nearest neighbors of sample i, and MMD be the maximum mean difference measure. The cleaned latent variable space is κ, which is the suppression coefficient. By comparing the difference between the sample gradient and the neighborhood average gradient, and combining the degree of deviation of the latent variable distribution, abnormal gradient components caused by the superposition of multi-source bias are eliminated.
[0111] In practical applications, The k-nearest neighbors of sample i are 10 "nurse" related samples, and the cleaned unbiased latent space... To ensure a balanced distribution of male / female nurse features, for "nurse + female image" samples, the difference between its gradient and the neighborhood (angle > 60°) is calculated. Combined with the latent space deviation (MMD = 1.4), 70% of abnormal gradients are suppressed to avoid the propagation of bias in multimodal fusion.
[0112] S4. Based on the purification gradient of step S3, apply sparse activation with differential topological constraints, as shown in equation (4):
[0113]
[0114] In the formula, ST is the Straight-Through estimator, h l-1 Ric is the activation value of the previous layer. l Let ζ be the Ricci curvature of the parameters in the l-th layer, and ζ be the curvature threshold. High curvature parameters are identified using differential geometry properties, and only parameters in low curvature regions are allowed to participate in the update, thus achieving physical bias isolation.
[0115] In practical applications, Ric l If the curvature of the parameters in the l-th layer is 3.2, then the ζ curvature threshold (2.5) filters out high curvature parameters, shuts down neurons with excessive curvature, and retains only low curvature parameters, thus blocking the coupling between "gender" and "occupation" in the parameter space, such as the feature extraction node of medical equipment.
[0116] S5. Generate combined counterfactual samples and reconstruct the attention mechanism, as shown in Equation (5):
[0117]
[0118] In the formula, Let d be the key-value vector of the counterfactual samples, and d be the vector dimension. This splits the traditional Softmax model into a true-counterfactual dual-path competition, forcing the model to explicitly distinguish biased combined features at the attention layer. For example:
[0119] a. Counterfactual samples containing both "woman" text and "kitchen" image suppress spurious gender-scene associations in the original data;
[0120] b. When the key value of the counterfactual sample is the visual feature of "nurse + male image", d is 768-dimensional. In the attention layer, the model is forced to compare the real sample ("nurse + female image") with the counterfactual sample ("nurse + male image"), which increases the weight of the gender-related attention head from 0.3 to 0.7 and strengthens the occupation-related features.
[0121] S6. Implicit arbitration is performed on long-tail conflict annotations, as shown in equation (6):
[0122]
[0123] In the formula, For the transfer matrix constraint, μ yj For category y j The prototype vector, D IPM The distance is measured by integral probability, H(T) = -∑ i,j T i,j logT i,j As an entropy regularization term, it solves the majority suppression problem in small sample labeling conflicts by jointly optimizing feature space alignment and distribution matching;
[0124] In practical applications, μ yj The prototype of the category "underwater welder" is characterized by welding tools + diving equipment features, D IPM The distribution distance metric compares the distribution differences of annotation sources A and B. By aligning the distributions of annotation sources A (correct annotation) and B (incorrect annotation) through optimal transmission, the arbitration weight is set to 0.8, and "underwater welder" is determined to be the true category.
[0125] S7. Dynamically update the memory and strengthen key samples, as shown in equation (7):
[0126]
[0127]
[0128] In the formula, Conv1D is a one-dimensional convolution in the time dimension, and Attn is the attention weight based on the memory bank. The convolution operation captures the temporal evolution of bias patterns, for example:
[0129] The propagation dynamics of emerging biased terms in social media data ensure that memory banks can preserve key historical states;
[0130] In practical applications, the significant gradients of key samples, such as the gradient of emerging bias "programmer → bespectacled men", are captured by Conv1D. The temporal pattern is a weekly bias growth rate of 12%. Highly significant samples (gradient norm > 1.5 and entropy > 1.2) are stored. The evolution of bias is identified through temporal convolution, and the model is dynamically updated to cover emerging bias scenarios.
[0131] S8. Finally, the parameters are isolated by spectral clustering, as shown in equation (8):
[0132]
[0133] In the formula, Tr represents the trace operation of the matrix, and Σ j Let η be the covariance matrix of cluster j parameters, and η be the gradient covariance penalty coefficient. This identifies clusters in the parameter space that are strongly correlated with bias propagation, and achieves physical memory isolation by modifying the CUDA kernel. k Mapped to a separate memory page, completely blocking its participation in forward propagation;
[0134] In practical applications, Σ j The covariance matrix of the parameter cluster, such as the gender-related parameter covariance of cluster C2, is >0.7. The trace operation of the Tr matrix is performed to quantify the gradient correlation. The identified bias parameter cluster C2 is isolated to an independent memory page. The calculation of this region is skipped in the forward propagation, which reduces the gender prediction standard deviation of the "nurse" image by 43%.
[0135] Through the above steps, after locating the multimodal bias path through causal discovery in step S1, long-tail noise is isolated through subspace routing in step S2. Then, gradient decomposition in step S3 purifies the gradient of multi-source data. S4 physically blocks bias parameters through sparse activation. Then, counterfactual correction is performed in step S5 to reconstruct cross-modal attention. Conflict annotations are fused through arbitration mechanism in step S6. After continuous fusion of temporal patterns through the memory bank in step S7, the fusion result is finally ensured to be uncontaminated through hardware isolation in step S8. This improves the accuracy of gender-occupation bias detection, increases the conflict resolution rate of unpopular occupation annotations, and achieves higher arbitration accuracy.
[0136] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0137] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A large-scale model data fusion method, characterized in that, Includes the following steps: S1. Apply structured interventions to multi-source heterogeneous data, construct cross-modal causal graphs, define the intervention process, quantify the strength of intermodal causal relationships through second derivatives, and identify the superposition path of hidden biases among multimodal data. S2. Based on the causal strength matrix of step S1, an improved adversarial multi-armed gambling machine selection subspace is adopted to dynamically select the subspace that best matches the current data distribution and has stable gradient updates, thereby blocking the distribution drift caused by mislabeling of long-tailed samples. S3. Perform gradient latent variable decomposition on the selected subspace data. By comparing the difference between the sample gradient and the neighborhood average gradient, and combining the degree of deviation of the latent variable distribution, eliminate abnormal gradient components caused by the superposition of multi-source bias. S4. Based on the purification gradient of step S3, apply sparse activation with differential topological constraints, use differential geometric properties to identify high curvature parameters, and only allow low curvature region parameters to participate in the update, thereby achieving physical bias isolation. S5. Generate combined counterfactual samples and reconstruct the attention mechanism. The traditional Softmax is split into a true-counterfactual dual-path competition, which forces the model to explicitly distinguish biased combined features at the attention layer and suppress false associations between gender and scene in the original data. S6. Implicit arbitration is performed on long-tail conflict annotations. The majority suppression problem in small sample annotation conflicts is solved by jointly optimizing feature space alignment and distribution matching. S7. Dynamically update the memory and strengthen key samples, and capture the temporal evolution of bias patterns through convolution operations; S8. Finally, spectral clustering is performed on the parameters to isolate them, identify clusters in the parameter space that are strongly correlated with bias propagation, and implement physical isolation of the video memory by modifying the CUDA kernel.
2. A large model data fusion system based on multi-source heterogeneous data, used to implement the large model data fusion method as described in claim 1, characterized in that, It includes a multimodal causal decoupling module, a dynamic subspace routing module, a heterogeneous learning and adaptation module, a sparse activation gating module, a counterfactual confusion pooling module, a small sample contradiction resolution module, a multimodal memory module, and a modular parameter isolation module; The multimodal causal decoupling module is used for causal graph construction and biased path separation; The dynamic subspace routing module is used for multi-armed gambling machine routing and subspace conflict detection; The heterogeneous meta-learning adaptation module is used for implicit meta-feature extraction and meta-gradient reweighting; The sparse activation gating module is used to provide a Top-K dynamic gating mechanism and conflict-aware sparsification. The counterfactual obfuscation pooling module is used for adversarial obfuscation generation and pooling attention decoupling; The small-sample conflict resolution module is used for conflict graph reasoning and implicit voting alignment; The multimodal memory module is used for temporal memory slicing and conflict sample playback; The modular parameter isolation module is used for parameter cluster discovery and hard isolation gating.
3. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The multimodal causal decoupling module also includes a causal graph construction unit and a biased path separation unit; The causal graph construction unit is used to conduct intervention experiments on latent variables of multimodal data, construct cross-modal causal graphs, and locate biased collaborative paths; The bias path separation unit is used to apply an intervention mask to the high-risk paths identified in the causal graph, thereby separating the coupling parameters between bias propagation and semantic representation.
4. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The dynamic subspace routing module also includes a multi-armed gambling machine routing unit and a subspace conflict detection unit; The multi-armed gambling machine routing unit is used to dynamically select the sparse subspace related to the current sample to avoid bias parameters contaminating small sample features. The subspace conflict detection unit is used to monitor conflicts in the update directions of different subspace parameters in real time, and to block the impact of long-tailed erroneous annotations on the majority class.
5. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The heterogeneous meta-learning adaptation module also includes an implicit meta-feature extraction unit and a meta-gradient reweighting unit; The implicit meta-feature extraction unit is used to extract meta-features related to bias and small sample size from multi-source heterogeneous data and generate dynamic weights. The meta-gradient reweighting unit is used to perform latent variable decomposition on the gradients of small-sample conflict instances, eliminating abnormal gradient components generated by bias collaboration.
6. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The sparse activation gating module also includes a Top-K dynamic gating unit and a conflict-aware sparsification unit; The Top-K dynamic gating unit activates only a subset of K parameters related to the current task, which is used to forcibly isolate the high-dimensional channels of bias propagation; The conflict-aware sparsification unit is used to impose L0 constraints on neurons with frequent long-tail conflicts, so that they are naturally eliminated or weakened during training.
7. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The counterfactual obfuscation pooling module also includes an adversarial obfuscation generation unit and a pooling attention decoupling unit; The adversarial confusion generation unit is used to generate counterfactual samples containing combinations of known biases, in order to decouple the cumulative effect of multi-source biases; The pooling attention decoupling unit is used to separate the feature responses of counterfactual samples and true samples in the multi-head attention layer, thereby blocking the cross-head propagation of biased features.
8. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The small-sample contradiction resolution module also includes a contradiction graph reasoning unit and an implicit voting alignment unit; The contradiction graph reasoning unit is used to construct a labeled contradiction graph of long-tailed small samples and reason about the potential true category; The implicit voting alignment unit is used to map multi-source labels to the optimal transmission space, calculate implicit voting weights, and avoid the majority suppression problem in explicit voting.
9. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The multimodal memory module also includes a temporal memory slicing unit and a conflict sample playback unit; The temporal memory slicing unit is used to store the historical state of bias patterns of different modalities in a time dimension, and is used to detect the evolution path of collaborative bias. The conflict sample replay unit is used to dynamically retain high-information instances of long-tail conflict samples to prevent them from being overwhelmed during training.
10. The large model data fusion system based on multi-source heterogeneous data according to claim 2, characterized in that, The modular parameter isolation module also includes a parameter cluster discovery unit and a hard isolation gating unit; The parameter cluster discovery unit is used to perform spectral clustering of model parameters based on gradient propagation paths to identify parameter clusters related to specific biases or conflicts. The hard isolation gating unit is used to apply physical isolation to high-risk parameter clusters, preventing them from being activated in specific tasks.
Citation Information
Patent Citations
Multi-source heterogeneous data fusion architecture system based on additive model
CN107798137A
Feature-adaptive mutual-guiding multi-source information fusion classification method and system
CN114187526A