Multi-modal multi-source heterogeneous data fusion method and system

By building multimodal prototype networks, memory modules, diversity sampling, modal balance processing and knowledge distillation, the challenges of multimodal data fusion in small sample learning, catastrophic forgetting, modal imbalance and rare mode protection are solved, and efficient data fusion and model performance improvement are achieved.

CN120197141AActive Publication Date: 2025-06-24FUJIAN YANGTENG INNOVATION INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510687617.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing multimodal data fusion methods have challenges in small sample learning, catastrophic forgetting, modal imbalance and rare pattern protection, making it difficult to learn quickly and effectively under limited data conditions and maintain model performance.

Method used

By building a multimodal prototype network, different modal features are extracted and semantic mapping relationships are established between modals; building a memory module to store multimodal representations of historical key samples through sample value evaluation; performing diversity sampling to ensure that a wide range of data patterns are included and rare anomaly patterns are retained; implementing modal balance processing, dynamically adjusting the contribution weights of different modal gradients; combining knowledge distillation, extracting general fusion knowledge from data-rich scenarios to assist small sample decision-making.

Benefits of technology

It significantly enhances the learning ability of small samples, alleviates catastrophic forgetting problems, improves modal balance performance, and effectively protects rare modes, improves computing efficiency and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197141A_ABST
    Figure CN120197141A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data fusion, and discloses a multi-modal multi-source heterogeneous data fusion method and system.The multi-modal multi-source heterogeneous data fusion method comprises the steps that a multi-modal prototype network is constructed, different modal features are extracted, and a semantic mapping relation between modals is established; a memory module is constructed, and multi-modal characterization of historical key samples is stored through sample value evaluation; executing diversity sampling to reserve a rare abnormal mode; modal balance processing is implemented, and contribution weights of different modal gradients are dynamically adjusted; knowledge distillation is applied, and general fusion knowledge is extracted from a data-rich scene to assist small sample decision making; the method solves the problems of low fusion efficiency of small sample and cold start scenes, disastrous forgetting in incremental learning, unbalanced modal learning rate and rare mode retention under long-tail distribution, and is suitable for predictive maintenance, personalized recommendation and medical health monitoring in the manufacturing industry and scenes requiring small sample learning and incremental updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and more specifically, to a multi-modal multi-source heterogeneous data fusion method and system. Background Art

[0002] With the rapid development of Internet of Things, big data and artificial intelligence technologies, the fusion analysis of multi-source heterogeneous data has become the core technology of intelligent decision-making systems. In the fields of manufacturing predictive maintenance, personalized recommendation services, medical health monitoring, etc., it is necessary to simultaneously process multi-modal data from different sensors or channels, such as machine vibration signals, temperature data, sound recordings, image data, text records, etc., in order to obtain more comprehensive and accurate information.

[0003] Existing multi-modal data fusion methods mainly include techniques such as feature-level fusion based on deep neural networks, selective fusion based on attention mechanisms, cross-modal mapping based on transfer learning, etc. These methods perform well in data-rich scenarios, but still face multiple technical challenges in practical applications: in cold start scenarios such as new device deployment or new user access, the available labeled data is extremely limited, and traditional fusion methods are difficult to adapt quickly; in continuously running systems, the continuously generated new data will cause the model to over-adapt to new patterns and forget the important knowledge learned from history; the complexity and learning difficulty of different modal data (such as images, texts, sounds) vary significantly, which easily leads to certain modalities dominating the decision-making process during fusion and ignoring the key information of other modalities; real-world data usually follows a long-tail distribution, and samples of key abnormal patterns (such as device fault characteristics) are scarce and difficult to be effectively learned and recognized.

[0004] The above problems seriously restrict the application effect of multi-modal data fusion technology in small sample learning and incremental update scenarios, and a new fusion method that can quickly and effectively learn under limited data conditions, prevent catastrophic forgetting, balance the contributions of different modalities, and effectively protect rare patterns is needed. Summary of the Invention

[0005] The present invention provides a multi-modal multi-source heterogeneous data fusion method and system, which solves the technical problems of small sample learning, catastrophic forgetting, modal imbalance, and rare pattern protection in related technologies.

[0006] The present invention provides a multi-modal multi-source heterogeneous data fusion method, including the following steps: Construct a multi-modal prototype network, extract different modal features and establish semantic mapping relationships between modalities; Based on the multi-modal prototype network, construct a memory module, and store the multi-modal representations of historical key examples through sample value evaluation; Utilize the memory module to perform diversity sampling to ensure the inclusion of a wide range of data patterns and retain rare abnormal patterns; Based on the diversity sampling results, perform modal balance processing and dynamically adjust the contribution weights of different modal gradients; Combined with the results of modal balance processing, apply knowledge distillation to extract general fusion knowledge from data-rich scenarios to assist small-sample decision-making.

[0007] In a preferred embodiment, the construction of the multi-modal prototype network includes: Construct a corresponding feature extraction network for each type of modal data; Calculate the prototype representation of each category for each modality; Learn the mapping function between different modal feature spaces; Optimize the entire prototype network using a model-agnostic meta-learning algorithm.

[0008] In a preferred embodiment, the construction of the memory module includes: Initialize the memory storage structure for saving historical data samples and their feature representations; Calculate the retention value of each candidate sample; Implement a memory update mechanism based on sample value; Introduce a time-based decay mechanism to adjust the sample value.

[0009] In a preferred embodiment, the execution of diversity sampling includes: Define the diversity metric function for the sample set; Implement a diversity sampling strategy based on entropy regularization; Construct an abnormal pattern protection mechanism; Assign importance weights to the selected samples.

[0010] In a preferred embodiment, the implementation of modal balance processing includes: Evaluate the learning difficulty of each modality; Calculate the contribution weights of each modal gradient; Implement balanced gradient updates; Introduce a modality-specific adaptive learning rate.

[0011] In a preferred embodiment, the application of knowledge distillation includes: Construct a teacher-student knowledge distillation architecture; Implement feature-level, relational knowledge, attention, and output distillation; Use a generative adversarial network to synthesize rare pattern samples; Comprehensively apply the results of knowledge distillation and adversarial generation to update the target model.

[0012] In a preferred embodiment, the sample value evaluation is determined by calculating a weighted combination of the difference between the sample feature representation and the current prototype set and the information gain of the sample for the task, where the weight between the difference and the information gain can be dynamically adjusted.

[0013] In a preferred embodiment, the modal gradient contribution weight is calculated by a normalized exponential function, which converts the learning difficulty index of each modality into the corresponding contribution weight to ensure that modalities with higher difficulty obtain more learning resources.

[0014] In a preferred embodiment, diversity sampling is achieved by optimizing an objective function, which comprehensively considers the diversity measure of the sample batch, the entropy value of the sample batch, and the abnormality degree of the sample to ensure that the sampling result contains both a wide range of data distributions and preserves rare abnormal patterns.

[0015] In a preferred embodiment, a multi-modal multi-source heterogeneous data fusion system for performing a multi-modal multi-source heterogeneous data fusion method includes: A multi-modal prototype network module for extracting different modal features and establishing an inter-modal semantic mapping relationship; A memory storage module for storing the multi-modal representations of historical key examples through sample value evaluation; A diversity sampling module for ensuring the inclusion of a wide range of data patterns and preserving rare abnormal patterns; A modal balance processing module for dynamically adjusting the modal gradient contribution weights; A knowledge distillation module for extracting general fusion knowledge from data-rich scenarios to assist small-sample decision-making.

[0016] The beneficial effects of the present invention are as follows: The small-sample learning ability is significantly enhanced: The prototype network and meta-learning optimization strategy adopted by the present invention, combined with the hierarchical knowledge distillation technology, enable the model to quickly learn the inter-modal mapping relationship from a small number of samples.

[0017] The catastrophic forgetting problem is effectively alleviated: Through the construction of the memory module and the time decay mechanism, the present invention can retain historical key knowledge during the continuous learning process.

[0018] The modal balance performance is significantly improved: Based on the modal balance processing mechanism of modal difficulty evaluation and dynamic adjustment of gradient contribution weights, different modalities can be equally treated during the fusion process.

[0019] The protection effect of rare patterns is prominent: The diversity sampling strategy and the abnormal pattern protection mechanism, combined with the method of synthesizing rare samples by the adversarial generative network, improve the sensitivity of the model to rare patterns.

[0020] Significantly improved computational efficiency: Compared with traditional deep fusion models, the student model of the present invention has a reduced number of parameters while maintaining or enhancing the fusion performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of a multi-modal multi-source heterogeneous data fusion method of the present invention; Figure 2 is a line graph comparing the performance of few-shot learning of the present invention; Figure 3 is a multi-bar graph and line graph of the historical knowledge retention rate during the incremental learning process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and the functions and arrangements of the elements discussed can be changed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0023] In at least one embodiment of the present invention, a multi-modal multi-source heterogeneous data fusion method is disclosed, as Figure 1 shown, including the following steps: Step 1, construct a multi-modal prototype network, extract different modal features, and establish semantic mapping relationships between modalities; Specifically, it includes the following sub-steps: Step 1.1, construct a feature extraction network; For each type of modal data, construct a corresponding feature extraction network to map heterogeneous data of different modalities to a unified feature space. Suppose there are types of different modalities, then construct corresponding feature extraction networks , where , , respectively represent the feature extraction networks designed for modalities , , , and represents the total number of modalities.

[0024] For the input data of modality , through the feature extraction network , the feature representation is obtained: ; where is the modality The feature representation vector represents the features obtained after converting the original data; It indicates modality The corresponding feature extraction network includes convolutional layers, recurrent layers, or Transformer layers. The specific structure is selected according to the data modality characteristics and is used to map the original samples to the feature space. is modal The original input data can be different types of data such as images, text, audio, etc. is the index of the modality, used to distinguish different types of data sources. For example, for the image modality, It can be a pre-trained convolutional neural network; for text modality, It can be a BERT-based encoder.

[0025] Step 1.2, modal prototype calculation; For each category or task of each modality, calculate its prototype representation. Assume that for the modality ,category have Support samples ,in, , , Respectively represent the mode Medium Category No. , , Support samples, Indicates support for centralized categories The number of samples is , then the prototype representation of the category is calculated as follows: ; in, is modal Medium Category The prototype representation of is obtained by averaging the feature representations of all supporting samples of the category; Indicates support for centralized categories The number of samples; Representing modality Medium Category No. Support samples; Representing modality The corresponding feature extraction network is used to map the original samples into the feature space; Indicates that the The feature representation obtained after the support samples pass through the feature extraction network. This method can effectively capture the core features of the category when the number of samples is limited, and provide a stable category representation for small sample learning.

[0026] Step 1.3, Cross-modal mapping function learning; Construct a cross-modal mapping function to establish a correspondence between feature spaces of different modalities.

[0027] For the mapping between modality and modality , learn the mapping function : ; where is the feature representation of modality , predicted from the feature representation of modality , representing the target modality feature obtained after transformation by the mapping function; is the feature representation of the source modality , that is, the feature vector input to the mapping function; is the mapping function from modality to modality , used to transform the source modality feature space to the target modality feature space.

[0028] The mapping function can be implemented as a multi-layer perceptron or an attention mechanism network, and its parameters are optimized by minimizing the distance between the predicted representation and the true representation: ; where is the loss function of cross-modal mapping, used to measure the quality of the mapping effect; represents a pair of corresponding samples from different modalities, is the sample of modality , is the sample of modality is the dataset containing the corresponding sample pairs of modality and modality ; and are the feature extraction networks of modality and modality respectively; is the mapping function from modality to modality ; is the distance metric function in the feature space, such as Euclidean distance or cosine distance, used to calculate the difference between the predicted feature and the true feature. This loss function optimizes the parameters of the cross-modal mapping function by minimizing the feature differences after mapping for all sample pairs.

[0029] Step 1.4, Meta - learning Optimization; The model - agnostic meta - learning algorithm is used to optimize the entire prototype network, improving its generalization ability under few - shot conditions. For each meta - learning task, the dataset is divided into a support set and a query set . First, the prototype representation is calculated based on the support set , and then the performance is evaluated on the query set and the model parameters are updated.

[0030] The objective function of meta - learning is: ; where, represents the overall loss function of meta - learning, which is used to optimize the overall performance of the model on multiple tasks; represents the mathematical expectation symbol, indicating the average over all sampled tasks; is the task distribution, representing the set of all possible learning tasks; represents a specific task sampled from the task distribution , which contains a support set and a query set; represents the support set, which contains a small number of labeled samples for the model to quickly adapt; represents the query set, which is used to evaluate the performance of the model on the current task; represents the feature extraction network with parameters , which is responsible for mapping the input data into the feature space; represents the set of learnable parameters of the model, which are optimized through the meta - learning process; represents the loss function calculated on the query set, which measures the performance of the model on the current task.

[0031] By optimizing this objective function, the model can learn how to quickly adapt to new tasks using a small number of samples, which is especially suitable for cold - start scenarios. This method enables the model to have the ability to "learn how to learn" and can quickly adapt to new data distributions or task requirements with only a small number of samples.

[0032] The output of Step 1 is a trained multi - modal prototype network, which includes feature extractors for each modality, prototype calculation methods, and cross - modal mapping functions, laying a foundation for subsequent memory enhancement and diversity protection.

[0033] Step 2, Based on the multi - modal prototype network, construct a memory module to store the multi - modal representations of historical key examples through sample value evaluation; Specifically, it includes the following sub - steps: Step 2.1, Memory Storage Structure Initialization; Establish an explicit memory storage structure , used to store historical data samples and their feature representations.

[0034] The memory storage structure can be represented as a set of tuples: ; where represents the entire memory storage structure, represents the original multimodal data sample, containing the original data content from different modalities (such as images, text, audio, etc.); represents the feature representation of the sample, generated by the feature extraction network, which is a high-dimensional feature vector converted from the original data; represents the label or task-related information of the sample, which is the target value for supervised learning and evaluating the model performance; represents the timestamp when the sample is added to the memory, used to implement the time decay mechanism and track the "age" of the sample; represents the usage frequency counter of the sample, recording the number of times the sample is accessed or used for training by the model, which affects the retention value of the sample; represents the upper limit of the memory capacity, set according to the application scenario and computational resource limitations, controlling the maximum number of samples that the memory module can store.

[0035] To ensure the efficient use of the memory module, initialize the upper limit of the memory capacity and set an index-based fast retrieval mechanism, using approximate nearest neighbor search algorithms such as Locality-Sensitive Hashing (LSH) to accelerate the retrieval process.

[0036] Step 2.2, sample value evaluation; Calculate the retention value for each candidate sample , which measures the potential contribution of the sample to the model performance: ; where represents the retention value score of the sample , represents the retention value weight, used to evaluate the potential contribution of the sample to the model performance; represents the sample 's feature representation, a vector generated by the feature extraction network; represents the set of prototypes stored in the current memory, representing the learned class centers; represents the sample 's information gain for the task, measuring the improvement in the model's prediction ability after adding this sample; represents the label space, containing all possible class labels; represents the sample feature representation The difference degree from the prototype set in the current memory is calculated as the distance to the nearest prototype: ; where, represents the sample feature representation The difference degree from the prototype set in the current memory is the Euclidean distance between the sample feature and the prototype; represents the sample whose feature representation is the vector generated by the feature extraction network; represents the prototype set stored in the current memory, representing the learned class center; represents the operation operator for finding the minimum value in the prototype set ; represents the prototype set a single prototype vector in; The information gain of the sample for the task, measuring the improvement of the model's prediction ability after adding this sample: ; where, represents the information gain of the sample for the task, is the entropy of the label distribution, representing the uncertainty of the label distribution; is the conditional entropy under the condition of the given sample feature representation, representing the uncertainty of the label distribution after knowing the sample feature; represents the label space, containing all possible class labels; represents the sample whose feature representation is the vector generated by the feature extraction network.

[0037] Step 2.3, implement the memory update mechanism; Implement the memory update mechanism, including strategies for adding new samples and eliminating old samples: Add new samples: For the new sample , if the memory is not full or its retention value is higher than the sample with the lowest retention value in the memory, then add it to the memory.

[0038] Elimination strategy: When the memory reaches the capacity limit , adopt the elimination strategy based on the sample value, and eliminate the sample with the lowest retention value: ; where, represents the sample to be removed; represents the independent variable when the objective function reaches the minimum value; represents the samples in the memory module represents the entire memory storage structure; represents a sample The retention value is jointly determined by the degree of difference and information gain.

[0039] Memory integration: Regularly perform clustering analysis on the samples in the memory module, merge highly similar samples, and release space to save more diverse samples: ; Among them, represents the integrated memory module, which contains the set of samples after clustering processing; represents the clustering algorithm function, which is used to merge or group similar samples; represents the original memory module, which contains all samples to be integrated; represents the sample similarity threshold, which is used to control the clustering granularity. A smaller threshold will generate more clusters, and a larger threshold will group more samples into the same class.

[0040] Step 2.4, time decay mechanism; Introduce a time-based decay mechanism to enable the memory module to adapt to changes in the data distribution while retaining long-term valuable knowledge.

[0041] For each sample in the memory, its value is adjusted over time: ; Among them, represents the adjusted retention value of the sample after considering the time factor, which is updated over time; represents the initial retention value of the sample which is jointly determined by the degree of difference and information gain; represents the current timestamp, indicating the time point when the system is currently running; represents the timestamp when the sample was added to the memory, which is used to calculate the duration of the sample in the memory; represents the time decay factor, which controls the rate of sample value decay over time. The larger the value, the faster the decay; represents the time decay term, which decreases as the sample exists in the memory for an increasing time; represents the usage frequency counter of the sample which records the number of times the sample is accessed or used for training by the model; represents the usage frequency influence factor, which controls the gain effect of the usage frequency on the sample value. The larger the value, the more significant the frequency influence; represents the usage frequency gain term, where frequently used samples gain value, and the logarithmic function ensures that the gain does not grow infinitely; represents the natural logarithm function.

[0042] This time decay mechanism ensures that the memory module can gradually adapt to changes in the data distribution while retaining long-term valuable knowledge.

[0043] Step 3: Use the memory module to perform diversity sampling to ensure that a wide range of data patterns are included and rare abnormal patterns are retained; Specifically, it includes the following sub-steps: Step 3.1: Definition of diversity metric; Define the diversity metric function of the sample set , which is used to evaluate the extent to which the sample set covers the feature space: ; where represents the sample set to be evaluated, which contains multiple data samples; represents the sample set in the th sample; represents the sample set in the th sample, and represents a sample that is not the same as ; represents the feature representation of the sample , that is, the vector that maps the original sample to the feature space; represents the feature representation of the sample ; represents the similarity function between feature representations, which is used to calculate the similarity degree of two feature vectors; represents the natural logarithm function, which is used to convert the similarity into a diversity metric; represents the diversity metric value of the sample set , and the larger the value, the higher the diversity of the sample set.

[0044] The similarity function can be the cosine similarity: ; where represents the first feature vector, corresponding to , representing the representation of the i-th sample in the feature space; represents the second feature vector, corresponding to , representing the representation of the j-th sample in the feature space; represents the dot product of two vectors, representing the correlation between vectors, and the larger the value, the more similar the directions of the two feature vectors; , respectively represent the vectors and the vector of The norm (Euclidean norm), which is the vector length, is used to normalize the eigenvector; represents the cosine similarity function, with a value range of [-1, 1]. A value of 1 indicates that the directions of two vectors are exactly the same, a value of -1 indicates that the directions are exactly opposite, and a value of 0 indicates orthogonality (no correlation).

[0045] This metric function encourages the selection of samples that are significantly different from each other, thereby increasing the diversity of the sample set.

[0046] Step 3.2, implementation of the entropy regularization sampling strategy; Implement a diversity sampling strategy based on entropy regularization to select a subset of samples from the memory module such that this subset maximizes diversity while maintaining the representation of the original distribution: ; where, represents the sampled sample batch, which is the subset of samples selected from the memory module; represents the optimal sample batch, that is, the subset of samples that maximizes the objective function; represents the entire memory storage structure, which contains all stored samples; represents the batch size, specifying the number of samples to be selected; represents the diversity metric of the sample batch, used to evaluate the dispersion degree of the sample set in the feature space; represents the entropy of the sample batch, measuring the balance of the sample class distribution. The higher the value, the more balanced the class distribution; represents the first balance factor, used to adjust the relative importance of the diversity metric and entropy in the objective function; represents the independent variable when the objective function reaches the maximum value, that is, to find the subset of samples that maximizes the expression within the square brackets.

[0047] ; where, represents the entropy value of the sample batch used to measure the balance of the sample class distribution; represents the set of all possible classes, indicating all classes existing in the data; represents the sample 's class label; represents the batch in the class the number of samples; represents the batch the total number of samples; represents the class The proportion in the batch, i.e., the category probability.

[0048] ; Among them, represents the sample of the optimal choice, i.e., the sample that maximizes the objective function; represents the independent variable when the objective function reaches the maximum value; represents the sample from the memory module but not in the current batch ; represents adding the sample to the current batch to form a new batch; represents the diversity measure of the batch after adding the sample ; represents the entropy value of the batch after adding the sample ; represents the diversity measure of the current batch ; represents the entropy value of the current batch ; represents the second balance factor; Add the selected sample to the batch ; Repeat step 2 to step 3 until the batch size reaches .

[0049] Step 3.3, Abnormal Pattern Protection Mechanism; To ensure that rare but important abnormal patterns are protected, the following mechanism is introduced: Error-Sensitive Resampling: For samples misclassified by the model, increase their weights in the sampling process: ; Among them, represents the importance weight of the sample used to adjust the probability of the sample being selected in the sampling process; represents the th sample in the memory module; represents the error sensitivity parameter, which controls the degree of increasing the weight of misclassified samples. The larger the value, the stronger the protection for misclassified samples; represents the indication operation, taking the value of 1 when the model prediction result is inconsistent with the true label and 0 when they are consistent; represents the predicted label of the model for the sample ; represents the sample true label.

[0050] Cluster density awareness: Based on the sample density in the feature space, preferentially select samples in low-density regions (which may contain abnormal patterns): ; Among them, represents the cluster density index of the sample , and the larger the value, the more likely the sample is in a low-density region (more likely to be an abnormal sample); represents the representation vector of the sample in the feature space; represents the representation vector of the sample in the feature space; represents the Euclidean distance between the sample and the sample in the feature space; represents the index set of the nearest neighbor samples of the sample in the feature space; represents the number of nearest neighbor samples considered, which is a preset hyperparameter; represents the sum over all nearest neighbor samples of the sample ; represents normalizing the sum result to calculate the average distance.

[0051] Anomaly score calculation: Calculate the anomaly score for each sample: ; Among them, represents the anomaly score of the sample , which is used to quantify the anomaly degree and importance of the sample; represents the cluster density value of the sample , reflecting the sparsity of the sample in the feature space. The larger the value, the more likely the sample is in a low-density region (abnormal region); represents the error-sensitive resampling weight of the sample , which assigns higher weights to samples misclassified by the model.

[0052] Modify the objective function of diverse sampling to incorporate an anomaly protection mechanism: ; Among them, represents the optimal sample batch, i.e., the subset of samples that maximizes the objective function; represents the sampled sample batch, which is a subset of samples selected from the memory module; represents the entire memory storage structure, which contains all stored samples; Denotes the batch size, specifying the number of samples to be selected; Represents the diversity measure of the sample batch, used to evaluate the dispersion degree of the sample set in the feature space; Represents the entropy of the sample batch, measuring the balance of the sample class distribution; Represents the third balance factor, used to adjust the relative importance of the diversity measure and entropy in the objective function; Represents the outlier protection intensity parameter, controlling the importance weight of outlier samples in the sampling process; Represents the sample 's outlier score, reflecting the rarity and importance of the sample; Represents the sum of the outlier scores of all samples in the batch, used to ensure that rare but important outlier patterns are selected into the batch; Represents the independent variable when the objective function reaches the maximum value, that is, to find the sample subset that maximizes the expression in the square brackets.

[0053] Step 3.4, Sample importance weighting; According to the results of diversity sampling, assign importance weights to each selected sample, which are used to adjust its contribution in the subsequent learning process: ; Among them, Represents the sample 's importance weight, indicating the contribution degree of this sample in the subsequent learning process; Represents the sample 's outlier score, calculated from the previous steps, reflecting the rarity and importance of the sample; Represents the temperature parameter, controlling the smoothness of the weight distribution. A smaller value will make the weight distribution steeper, highlighting the importance of outlier samples; a larger value will make the weight distribution smoother, ensuring that all samples can receive a certain degree of attention; Represents the exponential function, used to convert the outlier score to a non - negative value; Represents the currently selected sample batch, containing all samples participating in the calculation; Represents the batch Sum of all samples in, used for normalization to ensure that the sum of all weights is 1.

[0054] The output of Step 3 is a diversity - optimized sample batch and the corresponding sample importance weights , which maximizes the feature space coverage and class distribution balance, while protecting rare outlier patterns, providing high - quality training data for subsequent modal balance processing.

[0055] Step 4: Based on the diversity sampling results, perform modal balance processing and dynamically adjust the contribution weights of different modal gradients. Specifically, it includes the following sub-steps: Step 4.1: Modal difficulty assessment Evaluate the learning difficulty of each modality based on the performance and convergence speed of the model on that modality: ; Where, represents the learning difficulty index of modality , which is the ratio of the current learning difficulty of the modality to the historical average difficulty; represents the loss of the current model on modality , represents the part of the model on modality , represents the sample set of the current batch; represents the model in the previous training steps, the historical loss value at the step on modality ; represents the size of the historical time window considered for calculating the average historical loss of modality ; represents the average loss value of modality in the historical training steps; represents a specific modality index, indicating a specific modality in the multimodal data (such as image, text, audio, etc.).

[0056] The learning difficulty index reflects the ratio of the current loss to the historical average loss. A value greater than 1 indicates an increase in the current learning difficulty, and a value less than 1 indicates a decrease in the learning difficulty.

[0057] Step 4.2: Gradient contribution weight calculation Calculate the contribution weights of the gradients of each modality based on the modal learning difficulty: ; Where, represents the gradient contribution weight of modality , indicating the contribution ratio of modality to the total gradient during the multimodal fusion process; , respectively represent the learning difficulty indices of modality and modality , reflecting the current learning difficulty of the modality. The larger the value, the more difficult the learning; Represents the temperature parameter, which controls the smoothness of the weight distribution. A smaller value makes the weight distribution steeper (modes that are difficult to learn get higher weights), and a larger value makes the weight distribution more uniform; Represents the total number of modes, which refers to the number of different data modes in the system (such as images, text, audio, etc.); Represents the exponential function, which is used to non-linearly map the difficulty index to the positive value range; Represents the summation variable, which traverses all modes for normalization to ensure that the sum of the weights of all modes is 1.

[0058] This way of calculating weights assigns a larger gradient contribution weight to modes with higher learning difficulty, prompting the model to pay more attention to these difficult-to-learn modes, thus balancing the learning progress of different modes.

[0059] Step 4.3, Modal balance gradient update; Based on the calculated modal contribution weights, achieve balanced gradient update: ; Among them, Represents the total gradient of the model parameters, which is the final gradient value used to update the model parameters after integrating all modes; Represents the gradient contribution weight of mode , which controls the influence degree of this mode in the total gradient; Represents the gradient generated by mode , which is the partial derivative of the model parameters with respect to the loss function of mode ; Represents the total number of modes, which is the number of different data modes in the system; Represents the operation of summing the weighted gradients of all modes; Represents the mode index, ranging from 1 to , which represents a specific data mode (such as images, text, audio, etc.).

[0060] In specific implementation, the loss functions of each mode are weighted and summed to obtain the total loss function: ; Among them, Represents the total loss function, which is used to integrate the losses of all modes and guide the update of model parameters; Represents the operation of summing over all modes; Represents the mode index, ranging from 1 to , which represents a specific data mode (such as images, text, audio, etc.); Represents the total number of modalities, i.e., the number of different data modalities included in the system; Represents a modality 's gradient contribution weight, which determines the importance of this modality in the total loss; Represents a modality 's loss function value, which reflects the performance of the model on this specific modality.

[0061] Then, calculate the gradient based on the total loss function and update the model parameters: ; Among them, Represents the updated model parameters; Represents the model parameters before update, indicating the parameter values in the current training iteration; Represents the base learning rate, which controls the step size of parameter update; Represents the total loss function with respect to the model parameters 's gradient, which represents the change direction and magnitude of the loss function in the parameter space; Represents the weighted total loss function, which is composed of the loss functions of each modality combined according to their contribution weights.

[0062] Step 4.4, Adaptive learning rate adjustment; To further balance the learning processes of different modalities, introduce modality-specific adaptive learning rates: ; Among them, Represents the adaptive learning rate of modality , a specific learning rate dynamically adjusted according to the learning difficulty of this modality; Represents the base learning rate, the initial learning rate value shared by all modalities, serving as the benchmark for adaptive adjustment; Represents the learning rate adjustment factor, which controls the magnitude of learning rate adjustment according to the learning difficulty, and the larger the value, the more obvious the adjustment; Represents modality 's learning difficulty index, indicating the ratio of the current modality's learning difficulty to the historical average difficulty. A value greater than 1 indicates an increase in difficulty, and a value less than 1 indicates a decrease in difficulty.

[0063] When , the learning rate of modality will increase to accelerate the learning of difficult modalities; When , the learning rate will decrease to avoid overfitting of simple modalities.

[0064] To ensure balanced learning among modalities, introduce a regularization term to limit the performance differences among modalities: ; Among them, represents the modal balance regularization loss, which is used to reduce the difference in the loss function values between different modalities; represents the regularization strength parameter, which controls the influence degree of the regularization term on the total loss. The larger the value, the more emphasis is placed on the balance between modalities; represents modality 's loss function value, which reflects the current learning state of modality ; represents modality 's loss function value, and calculates the difference paired with ; represents the total number of modalities, that is, the number of different data modalities included in the system; represents modality and modality 's squared difference of loss function values, which penalizes the performance difference between different modalities; represents double summation, which calculates the total sum of loss differences between all modality pairs.

[0065] Add this regularization term to the total loss function: ; Among them, represents the final total loss function, which is the optimization objective for model parameter update; represents the sum of weighted losses of all modalities; represents the modal balance regularization term, which is used to limit the performance difference between different modalities.

[0066] The output of Step 4 is a gradient adjustment mechanism that realizes modal balance, including a modal difficulty evaluation method, a gradient contribution weight calculation formula, and an adaptive learning rate adjustment strategy, enabling the model to balance the processing of modal data with different difficulties and learning rates, avoiding a single modality from dominating the fusion process, and improving the comprehensiveness and accuracy of the fusion result.

[0067] Step 5: Combine the modal balance processing results and apply knowledge distillation to extract general fusion knowledge from data-rich scenarios to assist small-sample decision-making; Specifically, it includes the following sub-steps: Step 5.1: Construct a teacher-student model; Construct a teacher-student knowledge distillation architecture, where the teacher model is pre-trained on a data-rich source domain, and the student model is fine-tuned on a data-scarce target domain: Teacher model training: Use the data-rich source domain data to train the teacher model so that it masters rich modal fusion knowledge: ; Among them, Denote the optimal parameters of the teacher model; Denote finding the parameters that minimize the following expression ; Denote the loss function for a specific task; Denote the model function with parameters ; Denote the data-rich source domain dataset.

[0068] Student model initialization: Initialize the student model using the same network structure as the teacher model but with fewer parameters , or use the same structure as the teacher model but apply regularization techniques such as dropout.

[0069] Step 5.2, Implement hierarchical knowledge distillation; Implement hierarchical knowledge distillation to comprehensively transfer the knowledge of the teacher model from low-level features to high-level semantics: Feature-level knowledge distillation: Align the feature representations of the teacher model and the student model at each level: ; Where, Denote the feature-level knowledge distillation loss function; Denote the layer index of the network; Denote the total number of layers of the network; Denote the weight coefficient for feature distillation at the Denote the feature output of the teacher model at the layer for the input ; Denote the feature output of the student model at the layer for the input ; Denote the feature adaptation function for adjusting the feature dimension of the teacher model to match that of the student model; Denote the distance metric function for calculating the difference between the feature representations of the teacher model and the student model.

[0070] Relational knowledge distillation: Preserve the distance relationships between samples in the teacher model: ; Where, Denote the relational knowledge distillation loss for preserving the similarity relationships between samples; Denote the Gram matrix generated by the teacher model; Denote the Gram matrix generated by the student model; Denote the distance metric function.

[0071] is the Gram matrix calculated based on sample features: ; Among them, represents the similarity between samples and ; represents the similarity calculation function, which is used to measure the similarity degree of two feature vectors; , respectively represent the feature representations extracted by the model for the input samples and .

[0072] Attention knowledge distillation: Transfer the attention distribution of the teacher model: ; Among them, represents the attention knowledge distillation loss, represents the attention weight distribution of the teacher model for the input , represents the attention weight distribution of the student model for the same input , is the Kullback-Leibler divergence, which is used to measure the difference between two probability distributions. By minimizing this divergence, the student model can learn the attention mechanism of the teacher model.

[0073] Output distillation: Learn the soft label output of the teacher model: ; Among them, represents the output distillation loss, which is used to measure the difference between the output distributions of the teacher model and the student model; represents the Kullback-Leibler divergence, which is used to measure the difference between two probability distributions; represents the logits output of the teacher model (the raw prediction value without softmax); represents the logits output of the student model (the raw prediction value without softmax); represents the temperature parameter, which controls the smoothness of the soft label. A higher temperature value will generate a smoother probability distribution, which helps to transfer the similarity relationship between classes in the teacher model; represents the function that converts logits into a probability distribution.

[0074] Total distillation loss: ; Among them, represents the total distillation loss; Represents the feature-level distillation loss function; Represents the relational knowledge distillation loss; Represents the attention knowledge distillation loss; Represents the output distillation loss; 、 、 、 Represent the weight coefficients of the feature-level distillation loss, relational knowledge distillation loss, attention knowledge distillation loss, and output distillation loss respectively. These weight coefficients jointly determine the relative importance of knowledge distillation at different levels.

[0075] The total training objective of the student model is: ; Among them, is the total loss function of the student model, is the loss function for a specific task, represents the student model, is the limited data in the target domain, is the weight coefficient of the distillation loss (used to balance the importance of the task loss and the distillation loss), is the knowledge distillation loss function defined above (including four parts: feature-level distillation, relational knowledge distillation, attention knowledge distillation, and output distillation).

[0076] Step 5.3, the generative adversarial network synthesizes rare samples; Introduce a generative adversarial network (GAN), and generate synthetic samples based on existing rare pattern samples to augment the training data: Generator construction: Construct a conditional generator , which accepts random noise and conditional information (such as class labels or key features) and generates synthetic samples.

[0077] Discriminator construction: Construct a discriminator , which judges whether the sample is a real sample or a generated synthetic sample.

[0078] Adversarial training: Optimize the generator and discriminator through a min-max game: ; Among them, represents the generator whose goal is to minimize the objective function, while the discriminator whose goal is to maximize the objective function, reflecting the game process of adversarial training; represents sampling samples from the true data distribution in Find the expectation; Denote the log - probability output of the discriminator for real samples The discriminator hopes to maximize this term; Denote the expectation of joint sampling from the noise distribution and the conditional distribution ; Denote the random noise vector input to the generator, which serves to generate diverse samples; Denote the conditional information, such as class labels or key features, used to control specific attributes of the generated samples; Denote the synthetic samples generated by the generator based on the noise and conditional information; Denote the probability score of the discriminator for the generated samples, representing the probability that the discriminator thinks the sample is a real sample; Denote that the generator hopes to minimize this term, that is, hopes the discriminator misclassifies the generated samples as real samples.

[0079] Feature matching: Ensure that the feature distribution of the generated samples is similar to that of the real samples: ; Among them, Denote the feature - matching loss, which is used to ensure the similarity between the generated samples and the real samples in the feature space; Denote the feature extraction function of the teacher model, which is used to extract feature representations from the input samples; Denote the real samples, coming from the original dataset; Denote the generator network, which is used to generate synthetic samples; Denote the random noise vector, serving as the input source of the generator; Denote the conditional information, such as class labels or key features, used to guide the generator to generate specific types of samples; Denote the square of the L2 norm, which is used to calculate the square of the Euclidean distance between two feature vectors, measuring the degree of difference in the feature space.

[0080] Diversity enhancement: Encourage the generator to produce diverse samples: ; Among them, Denote the diversity loss function, which is used to encourage the generator to produce diverse samples; Denote the number of samples in the memory bank; Denote the number of all possible sample pairs; Denote the th random noise vector; Denote the th random noise vector; Represents conditional information (such as class labels or key features); Represents a sample generated by the generator based on the th noise vector and conditional information; Represents a sample generated by the generator based on the th noise vector and conditional information; Represents a distance metric function between samples, used to calculate the difference between two generated samples; the negative sign indicates that minimizing this loss is equivalent to maximizing the distance between generated samples, thereby increasing sample diversity.

[0081] Semantic consistency ensures that for multimodal data, the generated samples are semantically consistent across different modalities: ; Among them, Represents the semantic consistency loss, used to ensure that the samples generated in different modalities are semantically consistent; Represents the total number of modalities; and are the generators of modality and modality respectively; Represents the random noise input to the generator; Represents conditional information, such as class labels or key features; and are the feature extractors of modality and modality respectively; Represents a mapping function from modality to modality used to transform the features of one modality into the feature space of another modality; Represents the square of the L2 norm, used to calculate the Euclidean distance between two feature vectors.

[0082] Step 5.4, Knowledge Integration and Model Update; Integrate the results of knowledge distillation and adversarial generation and apply them to update the target model: Fusion dataset construction: Combine the original data, the data obtained by diversity sampling, and the generated synthetic data into an enhanced dataset: ; Among them, Represents the enhanced fusion dataset, which contains the original data and various augmented data; Represents the original dataset in the target domain, which contains a limited number of real samples; Represents a subset of data selected by the diversity sampling strategy, ensuring that it contains various patterns and boundary cases; Represents a rare sample dataset synthesized by a generative adversarial network, used to enhance the model's ability to recognize rare patterns; Represents the union operation.

[0083] Comprehensive training objective: ; Among them, Represents the final comprehensive training loss function, used to optimize the student model; Represents the main loss function for a specific task, measuring the model's performance on the enhanced dataset; Represents the student model, that is, the target model to be optimized; Represents the enhanced dataset, including original data, diversity sampled data, and generated synthetic data; Represents the weight coefficient of the knowledge distillation loss, controlling the intensity of knowledge transfer from the teacher model; Represents the knowledge distillation loss, including the integration of feature level, relational knowledge, attention, and output distillation; Represents the weight coefficient of the modality balance regularization term, regulating the influence degree of the modality balance mechanism; Represents the modality balance regularization term, used to balance the contributions of different modalities.

[0084] Progressive knowledge distillation: First, distill general knowledge on a broader task, and then gradually focus on a specific task: ; Among them, Represents the knowledge distillation loss weight that changes dynamically during training; Represents the initial weight value of the knowledge distillation loss, controlling the intensity of the distillation process at the beginning of training; Represents the base of the natural logarithm; Represents the decay factor, controlling the rate at which the weight decreases over time, the larger the value, the faster the decay; Represents the number of training steps, indicating the progress of model training; This formula implements progressive knowledge distillation, enabling the model to rely more on the knowledge of the teacher model at the beginning of training. As training progresses, it gradually reduces its dependence on teacher knowledge and relies more on target domain data for learning, thus achieving a smooth transition from teacher guidance to autonomous learning.

[0085] Model fine-tuning and update: Use the final training objective to optimize the student model, and regularly update and save the model parameters: ; Among them, Represents the parameters of the updated student model; Represents the parameters of the student model before update; represents the base learning rate, which controls the step size of parameter updates; represents the final loss function for the model parameters gradient, indicating the direction and magnitude of parameter updates.

[0086] The output of Step 5 is a fused model enhanced by knowledge distillation and adversarial generation techniques. This model can effectively fuse multi-modal heterogeneous data in a small-sample scenario, while maintaining the ability to recognize rare patterns, providing reliable support for the final decision-making analysis.

[0087] Application example of this embodiment: To verify the effectiveness of the present invention, the following provides a detailed description of this method through an application example in a predictive maintenance scenario of the manufacturing industry.

[0088] A smart manufacturing factory needs to monitor the status and predict faults of core production equipment to avoid production losses caused by unplanned downtime. The core equipment of this factory is equipped with a variety of sensors to collect the following multi-modal data: Vibration signal: data from a triaxial acceleration sensor with a sampling rate of 20 kHz; Sound data: ambient sound recordings sampled at 16 kHz; Temperature data: temperatures of various parts of the equipment sampled once per second; Infrared thermal imaging: thermal images of the equipment at one frame per minute; Operation logs: equipment operation records and status information This scenario faces the following challenges: Lack of sufficient fault samples during the deployment of new equipment (cold start problem); The fault types show a long-tailed distribution, and key fault examples are scarce; There are significant differences in the characteristics and acquisition rates of different sensor data (modal imbalance problem); As the equipment operation parameters are adjusted, the data distribution continuously changes (catastrophic forgetting problem).

[0089] Method implementation process: Prototype network construction: For different modal data, corresponding feature extraction networks are constructed: Vibration signal: a network combining 1D-CNN and BiLSTM is used to map the original vibration signal into a 128-dimensional feature vector; Sound data: Mel spectrogram combined with ResNet18 structure is used to extract 96-dimensional sound features; Temperature data: 1D-TCN (temporal convolutional network) is applied to extract 64-dimensional temperature features; Infrared thermal imaging: Using the lightweight MobileNetV3 to extract 128-dimensional image features; Operation log: Using a variant model of BERT to extract text features and compress them to 64 dimensions.

[0090] For the newly deployed Type A devices, there are only 5 failure samples as the support set. Taking vibration and temperature two modalities as examples, first calculate the prototype representation for each failure type (normal, bearing failure, gear failure, etc.): For the bearing failure class, calculate the vibration modality prototype: ; Then learn the mapping function from the vibration modality to the temperature modality , so that the temperature features predicted from the vibration features are as close as possible to the actual temperature features.

[0091] Through meta-learning optimization, the prototype network can identify different failure types with only 5 samples, achieving a detection accuracy of 85%, while the traditional method only has an accuracy of 52% under the same conditions.

[0092] Memory module construction and diversity sampling: The initial capacity of the memory module is 500 samples, and it preferentially saves the examples with high retention value. Taking the bearing failure class as an example, a special early failure sample, due to its high difference degree ( ) and high information gain ( ) with the existing prototypes, calculates the retention value , and is preferentially saved in the memory.

[0093] When adopting the entropy regularization diversity sampling strategy, a sample set with a batch size of 32 is selected from the historical data for training. By maximizing the sample diversity and class distribution balance, it is ensured that the sample batches cover the entire failure evolution process from slight wear to severe damage.

[0094] For rare failure modes (such as pitting failure of the bearing inner ring, which only accounts for 2% of the total failure samples), through the abnormal mode protection mechanism, its sampling probability is increased from the original 0.02 to 0.15, significantly enhancing the model's recognition ability for this type of failure.

[0095] Modal balance processing: In the initial stage of training, the learning difficulty index of the model for the vibration modality is significantly higher than that of the temperature modality , indicating that the vibration signal is more difficult to learn.

[0096] Through modal balance gradient adjustment, calculate the gradient contribution weight of the vibration modality, which is much higher than that of the temperature modality , making the model pay more attention to the vibration modes that are difficult to learn.

[0097] At the same time, an adaptive learning rate is applied. The learning rate of the vibration mode is adjusted from the base value of 0.001 to 0.00145, while the learning rate of the temperature mode is adjusted to 0.00083, further balancing the learning progress of different modes.

[0098] The results show that after the modal balance processing, the difference in the utilization degree of the vibration and temperature modal features by the model is reduced from the original 76% to 21%, and the fusion accuracy is improved by 24%.

[0099] Knowledge distillation application: Train a powerful teacher model from data-rich Class B devices (with more than 2000 fault samples), and transfer the knowledge to the student model of the target Class A devices through hierarchical knowledge distillation.

[0100] In feature-level distillation, the student model successfully replicated 98.7% of the intermediate layer feature distributions of the teacher model. In relation knowledge distillation, 94.3% of the similarity relationships between samples were retained.

[0101] For the extremely rare bearing inner ring fault mode (only 2 samples in Class A devices), 15 high-quality synthetic samples were synthesized through a generative adversarial network, and the detection rate of this type of fault was increased from the original 43% to 92%.

[0102] The generated synthetic samples maintained a high semantic consistency among multiple modalities, and the mutual prediction error between the vibration modality and the temperature modality was reduced by 78%, ensuring the authenticity and effectiveness of the synthetic data.

[0103] Verification of technical effects: Verification of few-shot learning ability: As shown in Table 1, the fault detection accuracies of this method and three mainstream fusion methods under different sample numbers are compared: Table 1: Comparison of the fault detection accuracies of this method and three mainstream fusion methods under different sample numbers;

[0104] This method achieved an accuracy of 87.3% with only 5 samples, which is 65.9% higher than the best comparison method, verifying its excellent few-shot learning ability. More importantly, this method can achieve the performance level that other methods need 15 to 20 samples to reach with only 5 samples, greatly reducing the data collection cost.

[0105] Verification of the protection effect of rare modes: For four different fault types (sorted from the most to the least number of samples), as shown in Table 2, the detection performance of each method was tested: Table 2: Detection performance test of each method;

[0106] It can be seen that this method still maintains a high detection rate of 91.4% for rare fault modes (pitting of the inner ring of the bearing) that only account for 2%. Compared with the average value of the comparative methods, it has increased by 40.0%, verifying its excellent rare mode protection ability. More importantly, while protecting rare modes, this method does not significantly sacrifice the detection performance for common modes.

[0107] As Figure 2 , Figure 3 shown, the comparison results of few-shot learning performance and the retention rate of historical knowledge during incremental learning are respectively presented.

[0108] The actual application effect shows that this method can establish an effective fault prediction model only 5 days after the initial deployment of the equipment, while traditional methods require more than 30 days of data accumulation. In addition, this method successfully predicted a key bearing fault, issued a warning 15 days in advance, and avoided a sudden shutdown accident that was expected to cause production losses of 2 million yuan, fully verifying the practical value of the present invention.

[0109] In summary, the multi-modal multi-source heterogeneous data fusion method proposed in this embodiment comprehensively solves key technical problems such as few-shot learning, catastrophic forgetting, modal imbalance, and rare mode protection by innovatively combining meta-learning, memory enhancement, and experience replay techniques, greatly improving the performance and application scope of multi-modal data fusion, and providing strong technical support for intelligent decision-making in fields such as industry, healthcare, and personalized recommendation.

[0110] The above describes the embodiments of the present invention, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A multi-modal multi-source heterogeneous data fusion method, characterized in that It includes the following steps: Construct a multi-modal prototype network, extract different modal features and establish semantic mapping relationships between modalities; Based on the multi-modal prototype network, construct a memory module to store the multi-modal representations of historical key examples through sample value evaluation; Utilize the memory module to perform diversity sampling to ensure the inclusion of a wide range of data patterns and the retention of rare abnormal patterns; Based on the results of diversity sampling, implement modal balance processing to dynamically adjust the weight of the gradient contribution of different modalities; Combined with the results of modal balance processing, apply knowledge distillation to extract general fusion knowledge from data-rich scenarios to assist small-sample decision-making.

2. The multimodal multi-source heterogeneous data fusion method according to claim 1, wherein The construction of the multi-modal prototype network includes: Construct a corresponding feature extraction network for each type of modal data; Calculate the prototype representation of each category for each modality; Learn the mapping function between different modal feature spaces; Adopt a model-agnostic meta-learning algorithm to optimize the entire prototype network.

3. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that The construction of the memory module includes: Initialize the memory storage structure for saving historical data samples and the feature representations of historical data samples; Calculate the retention value of each candidate sample; Implement a memory update mechanism based on sample value; Introduce a time-based decay mechanism to adjust the sample value.

4. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that The execution of diversity sampling includes: Define a diversity metric function for the sample set; Implement a diversity sampling strategy based on entropy regularization; Construct an abnormal pattern protection mechanism; Assign importance weights to the selected samples.

5. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that The implementation of modal balance processing includes: Evaluate the learning difficulty of each modality; Calculate the contribution weight of the gradient of each modality; Implement balanced gradient updates; Introduce a modality-specific adaptive learning rate.

6. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that The application of knowledge distillation includes: Construct a teacher-student knowledge distillation architecture; Implement feature-level, relational knowledge, attention, and output distillation; Utilize a generative adversarial network to synthesize rare pattern samples; Comprehensively apply the results of knowledge distillation and adversarial generation to update the target model.

7. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that The sample value evaluation is determined by calculating the weighted combination of the difference between the sample feature representation and the current prototype set and the information gain of the sample for the task, where the weight between the difference and the information gain can be dynamically adjusted.

8. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that The modal gradient contribution weight is calculated through a normalized exponential function, which converts the learning difficulty index of each modality into the corresponding contribution weight to ensure that modalities with higher difficulty obtain more learning resources.

9. A multimodal multi-source heterogeneous data fusion method according to claim 1, characterized in that Diversity sampling is achieved by optimizing the objective function, which comprehensively considers the diversity metric of the sample batch, the entropy value of the sample batch, and the abnormality degree of the sample to ensure that the sampling results include both a wide range of data distributions and retain rare abnormal patterns.

10. A multimodal multi-source heterogeneous data fusion system for performing a multimodal multi-source heterogeneous data fusion method according to any one of claims 1-9, characterized in that, It includes: A multi-modal prototype network module for extracting different modal features and establishing semantic mapping relationships between modalities; A memory storage module for storing the multi-modal representations of historical key examples through sample value evaluation; A diversity sampling module for ensuring the inclusion of a wide range of data patterns and the retention of rare abnormal patterns; A modal balance processing module for dynamically adjusting the weight of the gradient contribution of different modalities; A knowledge distillation module for extracting general fusion knowledge from data-rich scenarios to assist small-sample decision-making.

Citation Information

Patent Citations

  • Safety knowledge generation method and system based on large language model

    CN118797077A

  • Decision-making method and model for offline reinforcement learning and continuous online fine tuning

    CN119249360A

  • Multi-source heterogeneous data fusion and processing method based on big data

    CN119783037A

  • Knowledge data classification multi-modal reasoning evolution system based on depth model

    CN119862961A

  • Automatic compression method and platform for multilevel knowledge distillation-based pre-trained language model

    WO2022126797A1

Cited By

  • Enhanced retrieval-based agent rapid construction method and system

    CN120407751A

  • Nuclear power safety assessment method and system based on machine learning

    CN120952526A

  • A method and system for nuclear power safety assessment based on machine learning

    CN120952526B

  • Remote sensing data fusion method based on locality sensitive hashing and cross-modal attention

    CN122200258A

  • Remote sensing data fusion method based on locality-sensitive hashing and cross-modal attention

    CN122200258B