A causal-driven unsupervised multi-modal small sample double-cycle data fusion method

By preprocessing medical data and constructing causal graphs, the problem of multimodal data fusion in small sample scenarios was solved, the cross-scenario adaptability and interpretability of the model were improved, and the efficient utilization of multimodal data was achieved.

CN121051708BActive Publication Date: 2026-02-03湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511607702.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-03
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

In small-sample scenarios in the medical field, traditional methods struggle to capture stable modal correlation patterns, leading to model overfitting and weak cross-scenario transferability. This makes it difficult to meet interpretability and adaptability requirements, especially under the complex heterogeneity of multimodal data.

Method used

By acquiring medical images, electronic medical records, and physiological signal data, preprocessing and feature dimension alignment are performed to construct a causal graph and conduct causal consistency tests. Kernel independent component analysis and dynamic fusion weight calculation are used to generate causal enhancement features, and an unsupervised few-sample fusion model is constructed through global fine-tuning.

Benefits of technology

It enhances the generalization ability across scenarios, provides an efficient path for utilizing multimodal data, and improves the interpretability and adaptability of the model, making it suitable for the fusion of small sample medical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121051708B_ABST
    Figure CN121051708B_ABST
Patent Text Reader

Abstract

The application relates to a causal driving unsupervised multi-modal small sample double cycle data fusion method. Each modality data of each lesion sample is respectively preprocessed and feature dimension alignment is performed to obtain an aligned feature matrix. An undirected graph is constructed based on the aligned feature matrix. A causal consistency sample pair set is constructed by using kernel independent component analysis, medical priori and causal consistency test technology, comparative learning of a feature extraction network is performed on the set, causal enhancement features are output, and an updated causal graph is constructed. The updated causal graph is input into a dynamic fusion weight calculation model to calculate dynamic fusion weights, the causal enhancement features after weighted summation and standardization are calculated, preliminary fusion features are generated, the preliminary fusion features are screened, the screened effective fusion features are input into a downstream diagnosis model to correct and update the causal graph, and global fine tuning is performed on the feature extraction network, the dynamic fusion weight calculation model and the downstream diagnosis model to construct an unsupervised small sample fusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data fusion technology, and in particular to a causal-driven unsupervised multimodal small-sample dual-loop data fusion method. Background Technology

[0002] In small-sample scenarios in the medical field, the amount of data is usually only a few dozen cases or even less. Traditional methods face the problem that data scarcity makes it difficult for models to capture stable modal correlation patterns, failing to meet the core requirement of interpretability in medical diagnosis and easily overfitting to irrelevant features of individual patients. At the same time, the ability to transfer data across scenarios is weak. Differences in equipment between different hospitals and different medical record writing standards may cause the performance of trained models to drop sharply in new scenarios. Existing small-sample transfer methods (such as meta-learning and domain adaptation) mostly rely on single-modal data or externally labeled knowledge, making it difficult to adapt to the complex heterogeneity of multimodal data.

[0003] In summary, traditional fusion methods mainly suffer from insufficient robustness due to neglecting causal relationships, lack of targeted design for multimodal data in few-shot learning methods, and difficulty in balancing interpretability and few-shot adaptability. Summary of the Invention

[0004] Therefore, it is necessary to provide a causal-driven unsupervised multimodal small-sample dual-loop data fusion method, including:

[0005] S1: Obtain medical imaging data, electronic medical record text data, and physiological signal data for each lesion sample, and preprocess and align the feature dimensions of each modality data to obtain the aligned feature matrix;

[0006] S2: Construct an undirected graph based on the aligned feature matrix, use kernel independent component analysis to filter undirected edges in the undirected graph, and combine medical priors to determine the direction of the undirected edges to form a causal graph; based on the causal graph and each modality data, construct a set of causal consistency sample pairs after causal consistency test; perform comparative learning of the feature extraction network based on the set of causal consistency sample pairs, output causal enhancement features, and construct and update the causal graph based on the causal enhancement features.

[0007] S3: Input the updated causal graph into the dynamic fusion weight calculation model to calculate the dynamic fusion weight. Based on the causal enhancement features after weighted summation and standardization of the dynamic fusion weight, generate preliminary fusion features and filter the preliminary fusion features to obtain effective fusion features. Input the effective fusion features into the downstream diagnostic model to calculate the prediction error and correct and update the causal graph. Feed the corrected causal graph back to step S2 to guide the next round of sample generation and comparative learning.

[0008] S4: Global fine-tuning of the feature extraction network, dynamic fusion weight calculation model, and downstream diagnostic model based on contrastive loss and prediction error;

[0009] S5: Repeat steps S2-S4 until the mean difference between the contrast loss, the prediction loss, and the edge weights in the causal graph satisfies the set convergence condition, thus constructing an unsupervised small sample fusion model.

[0010] S6: Input the multimodal data to be fused into the unsupervised few-sample fusion model to obtain the final fused features.

[0011] Beneficial Effects: This method preprocesses and aligns the feature dimensions of each modal data of each lesion sample to obtain an aligned feature matrix. Based on the aligned feature matrix, an undirected graph is constructed. Kernel Independent Component Analysis, medical priors, and causal consistency testing techniques are used to construct a causal consistency sample pair set for comparative learning of the feature extraction network, outputting causal enhancement features and constructing an updated causal graph. The updated causal graph is input into a dynamic fusion weight calculation model to calculate dynamic fusion weights, thereby weighted summing and standardizing the causal enhancement features to generate preliminary fusion features. The preliminary fusion features are then screened, and the selected effective fusion features are input into the downstream diagnostic model to correct and update the causal graph. Global fine-tuning is performed on the feature extraction network, the dynamic fusion weight calculation model, and the downstream diagnostic model to construct an unsupervised small-sample fusion model. The multimodal data to be fused is input into the unsupervised small-sample fusion model to obtain the final fusion features. This method improves cross-scenario generalization ability by dynamically fusing multimodal features, providing a new technical path for the fusion and efficient utilization of small-sample medical multimodal data. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of the causal-driven unsupervised multimodal small-sample dual-loop data fusion method in the embodiments of this application. Detailed Implementation

[0014] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0015] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0016] like Figure 1 As shown, this embodiment provides a causal-driven unsupervised multimodal small-sample dual-loop data fusion method, including:

[0017] S1: Obtain medical imaging data, electronic medical record text data, and physiological signal data for each lesion sample, and preprocess and align the feature dimensions of each modality data to obtain the aligned feature matrix.

[0018] Specifically, the preprocessing of medical image data includes:

[0019] The mean and standard deviation of pixel values ​​in all medical image data are calculated. The pixel values ​​of each pixel are normalized based on the mean and standard deviation to obtain normalized medical image data. The size of each normalized medical image data is uniformly adjusted to 224×224 using bilinear interpolation, and noise in each medical image data after size adjustment is removed by Gaussian filtering.

[0020] Each denoised medical image data is input into a pre-trained RseNet-50 network with the fully connected layers removed, and the initial image features corresponding to each medical image data are extracted.

[0021] Preprocessing of electronic medical record text data includes:

[0022] Each electronic medical record text data was cleaned and segmented, and special characters were removed using regular expressions. The characters were then uniformly encoded to UTF-8 to obtain string sequences. Each string sequence was split using a word segmenter, and "[CLS]" was added before each string sequence and "[SEP]" was added after each string sequence to obtain the corresponding word list. Each word list was converted into a corresponding number using a word list conversion function. All numbers were input into the pre-trained BioBERT model, and the features output from the hidden layer corresponding to the "[CLS]" mark were taken as the initial text features for each electronic medical record text data.

[0023] Preprocessing of physiological signal data includes:

[0024] The physiological signal data is subjected to discrete Fourier transform to convert it into a frequency domain signal. A 50Hz band-stop filter is used to remove power frequency interference from the frequency domain signal to obtain the physiological signal data after frequency domain denoising.

[0025] Inverse Fourier transforms were performed on each physiological signal data after frequency domain denoising to obtain the denoised physiological signal data.

[0026] The Pan-Tompkins algorithm is used to detect the R-wave in each denoised physiological signal data. The corresponding interval is calculated based on the R-wave position of adjacent R-waves. All intervals in any physiological signal data are combined to obtain the time-domain feature vector corresponding to the physiological signal data.

[0027] The denoised physiological signal data are mapped to the Mel spectrum to obtain the corresponding filter energy; each filter energy is logarithmically processed, and the discrete cosine transform is used to transform each logarithmically processed filter energy into the frequency domain feature vector corresponding to the physiological signal data.

[0028] By splicing the time-domain feature vectors and frequency-domain feature vectors corresponding to the physiological signal data, the initial physiological features corresponding to each physiological signal data are obtained.

[0029] Furthermore, for subsequent multimodal causal analysis, it is necessary to perform class-normalized mapping on the initial features of different modalities. Features of different dimensions need to be mapped to a unified space. Feature dimension alignment includes:

[0030] The initial image features, initial text features, and initial physiological features extracted from the medical imaging data, electronic medical record text data, and physiological signal data of each lesion sample are respectively passed through 3 independent fully connected layers, and each feature is mapped to 256 dimensions.

[0031] All initial image features, initial text features, and initial physiological features after dimension alignment are concatenated, and features belonging to the same modality are placed in the same column to obtain the aligned feature matrix.

[0032] S2: Construct an undirected graph based on the aligned feature matrix, use kernel independent component analysis to filter the undirected edges in the undirected graph, and combine medical priors to determine the direction of the undirected edges to form a causal graph; based on the causal graph and each modal data, construct a set of causal consistency sample pairs after causal consistency test; perform comparative learning of the feature extraction network based on the causal consistency sample pairs to output causal enhancement features, and construct and update the causal graph based on the causal enhancement features.

[0033] Specifically, the process of constructing a cause-effect graph includes:

[0034] By aligning each feature in the feature matrix as a node and creating an undirected edge between any two nodes, an undirected graph is constructed.

[0035] Kernel independent component analysis is used to perform conditional independence tests on any two nodes, and undirected edges that have no connection between nodes are removed based on the independence test results;

[0036] Based on prior knowledge in the medical field, undirected edges are extracted and retained. Then, any two nodes connected by a retained undirected edge are matched with the causal logic of the prior knowledge to determine the causal relationship between the two nodes. Based on this causal relationship, the undirected edges are modified into directed edges, resulting in a directed graph. Features of "imaging lesions", If the feature is "textual symptom description", then the causal direction is determined based on the prior knowledge that "lesion morphology determines symptom presentation". And change the undirected edge to a directed edge according to the direction;

[0037] The quantization method based on kernel mutual information quantifies the causal relationship strength between any two nodes in a directed graph and performs normalization processing to obtain the edge weight of the corresponding directed edge.

[0038] A causal graph is formed based on nodes, edges, and their corresponding edge weights.

[0039] In this embodiment, the process of performing the conditional independence test includes:

[0040] For any two nodes and We select a condition set S and map the feature matrix to a high-dimensional space using a Gaussian kernel function, which is expressed as:

[0041] ;

[0042] in, This represents the input feature vector. The kernel width parameter represents the Gaussian kernel inner product matrix of the mapped samples after feature mapping. :

[0043] ;

[0044] in, Representation of the characteristic matrix The The feature vector of each sample is then quantized. and In a given The independence of time was determined by calculating the conditional independence index using kernel independent component analysis (KICA). ,in The trace of a matrix is ​​the sum of its diagonal elements. Represents the identity matrix. Indicates the regularization parameter to be taken The calculation formula is:

[0045] ;

[0046] The larger the index, the better the performance under a given set of conditions. At that time, characteristics With features The stronger the independence between them, the more likely a threshold can be set. ,when At that time, the characteristics were considered With features Given a set of conditions The undirected edge between the two nodes is removed by making them independent.

[0047] In this embodiment, a quantization method based on kernel mutual information is used to quantify the causal relationship strength between any two nodes in a directed graph and then normalize it to obtain the edge weight of the corresponding directed edge, including:

[0048] The quantization method based on Kernel Mutual Information (KMI) further maps high-dimensional features to the Reproducing Kernel Hilbert Space (RKHS) through a kernel function, compared to Mutual Information (MI), and calculates mutual information in the high-dimensional space. This solves the problems of traditional MI, such as strong assumptions about data distribution and large estimation bias under small sample conditions. The calculation formula is:

[0049] ;

[0050] in, Let N be an all-one vector, and N be the number of lesion samples. , Representing features respectively , The Gaussian kernel matrix. After quantizing the directed graph, the KMI values ​​are normalized to avoid excessively large weight values:

[0051] ;

[0052] in Indicates causal edge edge weights, This represents all retained causal edges; the normalization result maps all edge weights to... Within the range.

[0053] Furthermore, the construction process of the causal consistency sample set includes:

[0054] Edges with weights greater than or equal to a threshold in a causal graph are considered strong causal edges.

[0055] If the modal data corresponding to the nodes connected by strong causal edges is medical image data, the Faster-R-CNN object detection algorithm is used to locate the core pixel region of the lesion sample in the medical image data, and the pixel value of the core pixel region is fixed. The background of the medical image data with fixed pixel value is randomly perturbed by Gaussian noise or by ±10% brightness perturbation, and the perturbed medical image data is used as a positive sample image.

[0056] If the modal data corresponding to the nodes connected by strong causal edges is electronic medical record text data, the BioBERT segmenter is used to extract the semantic description of the lesion samples in the electronic medical record text data. The semantic description is kept unchanged, and other redundant descriptions are perturbed by synonym replacement or sentence structure transformation. The perturbed electronic medical record text data is used as positive sample text.

[0057] If the modal data corresponding to the nodes connected by strong causal edges is physiological signal data, the Pan-Tompkins algorithm is used to detect the R-wave time series segments in the physiological signal data. The intervals in the R-wave time series segments are kept constant, and random perturbations of ±5% of the sampling values ​​are added to the non-R-wave time series segments. The perturbed physiological signal data is then used as the positive sample physiological signal.

[0058] Edges with weights less than a threshold (set to 0.6 in this embodiment) in the causal graph are considered weak causal edges.

[0059] If the modal data corresponding to the nodes connected by weak causal edges is medical image data, a generative adversarial network is used to modify the size of the lesion sample region in the medical image data while keeping the size of the remaining regions unchanged, thus constructing a counterfactual intervention negative sample image. The first strong causal edge pointing from the node corresponding to the medical image data in the causal graph to the node corresponding to the electronic medical record text data is queried. The medical image data corresponding to the nodes of any lesion sample, the electronic medical record text data, and the two nodes connected by the first strong causal edge whose descriptions are different are used as conflict-filtered negative sample images.

[0060] If the modal data corresponding to the nodes connected by the weak causal edge is electronic medical record text data, the semantic description of the lesion sample in the electronic medical record text data is rewritten as the opposite, while keeping the other semantic descriptions unchanged, to construct the counterfactual intervention negative sample text; query the second strong causal edge in the causal graph that points from the node corresponding to the electronic medical record text data to the node corresponding to the physiological signal data, and use the electronic medical record text data corresponding to the electronic medical record text data and the node corresponding to the physiological signal data in any lesion sample whose descriptions are different from the descriptions of the two nodes connected by the second strong causal edge as the conflict-filtered negative sample text.

[0061] If the modal data corresponding to the nodes connected by the weak causal edge is physiological signal data, all intervals in the fixed R-wave time sequence of the physiological signal data are compressed to construct counterfactual intervention negative sample physiological signals. The third strong causal edge pointing from the node corresponding to the medical image data in the causal graph to the node corresponding to the physiological signal data is queried. The corresponding physiological signal data whose descriptions differ between the nodes corresponding to the medical image data and physiological signal data in any lesion sample and the two nodes connected by the third strong causal edge are used as conflict screening negative sample physiological signals.

[0062] It is worth noting that, to explain the "difference in description," this embodiment provides an example as follows:

[0063] If the descriptions of the two nodes connected by the first strong causal edge are respectively: cause node (image modality: lung nodule diameter > 3cm) → effect node (text modality: risk of lung cancer);

[0064] However, the actual descriptions recorded in a certain lesion sample were as follows: Imaging modality: The diameter of the lung nodule is 1cm; Text modality: There is a risk of lung cancer.

[0065] At this point, if the feature description corresponding to the image modality in the lesion sample is different from the description of the causal node connected by the first strong causal edge, it is considered that the original data in the lesion sample violates the causal logic of the first strong causal edge, so it is filtered out as a negative sample image for conflict screening.

[0066] The same principle applies to the screening methods for negative sample text and negative sample physiological signals in conflict screening.

[0067] Each modality of data corresponds to one positive sample, one counterfactual intervention negative sample, and one conflict screening negative sample. Based on the original data of each modality, a set of sample pairs is constructed.

[0068] A second causal graph is constructed based on the modal data in the sample pair set. The weight of the second edge of the strong causal edge in the second causal graph is calculated, and the weight difference between the edge weight corresponding to the strong causal edge and the weight of the second edge is calculated. In each positive sample, positive samples whose weight difference of each strong causal edge is less than the second threshold (set to 0.1 in this embodiment) are selected as valid positive samples. In each negative sample, negative samples whose weight difference of each strong causal edge is greater than the third threshold (set to 0.5 in this embodiment) are selected as valid negative samples. The remaining samples are removed to obtain the causal consistency sample pair set.

[0069] The causal consistency sample pair set includes several causal sample pairs. Each causal sample pair includes positive samples, counterfactual intervention negative samples, conflict screening negative samples, and original data corresponding to any data of the same modality.

[0070] Furthermore, the contrastive learning of feature extraction networks includes:

[0071] The positive sample, the counterfactual intervention negative sample, the conflict screening negative sample, and the original data in any sample pair are respectively input into the feature extraction network corresponding to the modality to obtain the corresponding initial features.

[0072] A third causal graph is constructed based on four initial features in each sample pair;

[0073] The basic contrastive loss is calculated based on the four initial features of each sample pair, using the following formula:

[0074] ;

[0075] in, Indicates the baseline contrast loss; N represents the lesion sample size; Indicates the first In the lesion sample, the first The initial features corresponding to the raw data of each modality; Indicates the first In the lesion sample, the first The initial features corresponding to positive samples of each modality; Indicates the first In the lesion sample, the first Initial features corresponding to counterfactual intervention negative samples or conflict screening negative samples in each modality; This represents the temperature coefficient, taken as 0.1; This represents the cosine similarity.

[0076] The causal constraint loss is calculated based on the weights of the third edge in the third causal graph to avoid the model learning spurious associations with the causal graph. The formula for calculating the causal constraint loss is as follows:

[0077] ;

[0078] in, Indicates loss due to causal constraint; This represents the constraint coefficient weight, which is set to 0.5. This represents the core mutual information; the larger the value, the higher the consistency. This represents the third cause-and-effect graph; Represents a cause-and-effect graph; Represents a possible third causal graph;

[0079] Contrast loss is constructed based on basic contrast loss and causality constraint loss. The Adam optimizer is used, with an initial learning rate set to 1. The weight decay coefficient is Minimize the contrast loss to optimize the network parameters of the feature extraction network corresponding to the modality;

[0080] In every 10 iterations, the positive sample, the counterfactual intervention negative sample, the conflict screening negative sample, and the original data from any sample pair are input into the feature extraction network corresponding to the optimized modality, and the corresponding causal enhancement features are obtained respectively.

[0081] An updated causal graph is constructed based on the four causal enhancement features in each sample pair.

[0082] Until 15 consecutive iterations, the rate of change of total loss And predict causal graph With the true cause-and-effect diagram Training can be stopped when the KMI metric reaches 0.9. In the current optimized feature extraction network, features extracted from the original samples by each modality serve as causal enhancement features. Output. Compared to the initial feature extraction network, the weights of causal enhancement features and features related to strong causal edges in the causal graph are effectively improved, while the weights of non-causal dimensions are further suppressed.

[0083] S3: Input the updated causal graph into the dynamic fusion weight calculation model to calculate the dynamic fusion weight. Based on the causal enhancement features after weighted summation and standardization of the dynamic fusion weight, generate preliminary fusion features and filter the preliminary fusion features to obtain effective fusion features. Input the effective fusion features into the downstream diagnostic model to calculate the prediction error and correct and update the causal graph. Feed the corrected causal graph back to step S2 to guide the next round of sample generation and comparative learning.

[0084] Specifically, the process of obtaining effective fusion features includes:

[0085] Step 1: In the dynamic fusion weight calculation model, extract and update the weights of the causal paths between any two nodes in the causal graph. Causal paths include direct causal paths and indirect causal paths. A direct causal path is a directed edge that directly connects two nodes, and an indirect causal path is a path that indirectly connects two nodes, containing at least two directed edges. The weight of a direct causal path is the weight of the fourth edge of the corresponding directed edge, and the weight of an indirect causal path is the product of the weights of the fourth edges of all directed edges on the indirect causal path. When there are multiple indirect causal paths between two nodes, the path with the largest product of the weights of the fourth edges of all directed edges on the path is taken as the indirect causal path.

[0086] Step 2: Calculate the sum of the weights of the causal paths from any node to the other nodes, which serves as the total causal influence score for that node; normalize the total causal influence scores of each node to obtain the dynamic fusion weights corresponding to the nodes. The normalization formula is:

[0087] ;

[0088] in, Indicates the first The first lesion sample Dynamic fusion weights of nodes corresponding to each modality; Indicates the first The first lesion sample The total causal influence score of the nodes corresponding to each modality; Indicates the first The first lesion sample The total causal influence score of the nodes corresponding to each modality;

[0089] The dynamic fusion weights of nodes corresponding to the three modalities belonging to the same lesion sample are used as the weight set corresponding to the current circulating lesion sample.

[0090] Step 3: Standardize each causal enhancement feature corresponding to the same modality based on the mean and standard deviation of the causal enhancement features corresponding to the same modality;

[0091] Step 4: The standardized causal enhancement features corresponding to the three modalities of the same lesion sample are weighted and summed with the weight set corresponding to the current circulating lesion sample to obtain the corresponding preliminary fusion features. The calculation formula is:

[0092] ;

[0093] in, Indicates the first Preliminary fusion characteristics of individual lesion samples; Indicates the current loop number. The weight set corresponding to each lesion sample; Indicates transpose; , , They respectively represent belonging to the first Standardized causal enhancement features corresponding to the image modality, text modality, and physiological signal modality of each lesion sample; , , They represent , , The corresponding dynamic fusion weights;

[0094] Step 5: Perform L2 normalization on the preliminary fusion features corresponding to each lesion sample to obtain the normalized fusion features corresponding to each lesion sample. The calculation formula is as follows:

[0095] ;

[0096] in, Indicates the first Normalized fusion features corresponding to each lesion sample; Represents the L2 norm;

[0097] Step 6: Calculate the kernel mutual information of the normalized fusion features with the causal enhancement features corresponding to the three modalities belonging to the same lesion sample, respectively, to obtain the kernel mutual information sequence corresponding to the lesion sample; Nuclear mutual information sequence corresponding to each lesion sample Represented as:

[0098] ;

[0099] The cosine similarity between the nuclear mutual information sequence corresponding to the lesion sample and the weight set corresponding to the current circulating lesion sample is calculated using the following formula:

[0100] ; ;

[0101] in, Indicates the first The cosine similarity between the nuclear mutual information sequence corresponding to each lesion sample and the weight set corresponding to the current circulating lesion sample; Indicates the current loop number. The weight set corresponding to each lesion sample;

[0102] The higher the cosine similarity value, the more consistent the directions of the two sequences are. When the cosine similarity is greater than the fourth threshold (set to 0.65 in this embodiment), the fusion is determined to be successful, and the preliminary fusion feature is taken as the effective fusion feature. Otherwise, steps 1-6 are repeated until the fusion is successful.

[0103] Furthermore, the process of updating the cause-effect graph includes:

[0104] Step 1: Input the effective fusion features corresponding to the lesion sample into the downstream diagnostic model. The downstream diagnostic model includes a binary classification model and outputs the corresponding prediction label, which includes healthy or unhealthy.

[0105] Step 2: Calculate the cross-entropy loss for each lesion sample based on the predicted label and the corresponding true label.

[0106] Step 3: Classify lesion samples with cross-entropy loss higher than the error threshold as high-error samples, and classify lesion samples with cross-entropy loss less than or equal to the error threshold as low-error samples.

[0107] Step 4: For any high-error sample, call the weight set corresponding to the current cyclic lesion sample, and select the two nodes with the largest dynamic fusion weights from the weight set as core nodes;

[0108] In updating the causal graph, the direct causal path between two core nodes is extracted as the core edge corresponding to the high error sample. The first frequency of the core edge corresponding to the high error sample in all causal paths between the two core nodes in all high error samples is counted. If the first frequency is higher than the frequency threshold, the corresponding core edge is marked as an unreliable path; otherwise, it is marked as a reliable path.

[0109] For example, in a set of three modal nodes A, B, and C, the core nodes are A and C. In the high-error sample set, if the first frequency of the direct causal path (core edge) from A to B among all paths from A to C in all high-error samples is less than a frequency threshold, then it is marked as a reliable path (core edge from A to B); otherwise, it is marked as unreliable. A higher frequency in high-error samples indicates a stronger correlation between the path and the error, and therefore, it is less reliable. Therefore, this method is used to intuitively and clearly identify the optimization target in the path optimization-causal graph correction stage.

[0110] Step 5: For any low-error sample, execute step 4 to obtain the second frequency of the core edge appearing in all low-error samples. If the second probability is higher than the frequency threshold, mark the corresponding core edge as an unreliable path; otherwise, mark it as a reliable path.

[0111] Step 6: Adjust and update the weight of the fourth edge of the corresponding directed edge in the causal graph based on the first and second frequencies. The adjustment formula is as follows:

[0112] ;

[0113] in, Indicates the adjusted core edge The weight of the fourth side; This indicates updating the core edges in the causal graph. The weight of the fourth side; This represents the adjustment factor, which is 0.2. Represents the core edge The first frequency; Represents the core edge The second frequency; and the adjusted fourth side weights are normalized;

[0114] Step 7: In updating the causal graph, delete the core edges marked as unreliable paths and those that conflict with medical priors; reverse the core edges in high-error samples where the first frequency is lower than the second frequency of their reverse edges in low-error samples, and retain the original weight of the fourth edge.

[0115] Step 8: Feed the corrected causal graph back to step S2 to guide the next round of sample generation and comparative learning.

[0116] S4: Based on the contrast loss and prediction error, perform global fine-tuning of the feature extraction network, dynamic fusion weight calculation model, and downstream diagnostic model; repeat steps S2-S4 until the mean difference between the contrast loss, prediction loss, and edge weights in the causal graph meets the set convergence condition, and construct an unsupervised small sample fusion model.

[0117] Specifically, in the overall global model optimization, the first loop (corresponding to S2) and the second loop (corresponding to S3) constitute a collaborative iteration until the global convergence condition is met, thus completing the global optimization. After the second loop in the first round, the corrected causal graph is... Feedback is sent to the first cycle of the next round to guide the generation of causal consistency samples and subsequent steps.

[0118] The global model is considered convergent when both of the following conditions are met:

[0119] (1) Extract the contrast loss corresponding to the first loop in each round of double loop. Cross-entropy loss with the second loop part Let the total global loss be... , A value of 0.3 represents the loss weight of the downstream task in the second loop, and this is used until 15 consecutive iterations to determine the rate of change of the total loss. .

[0120] (2) Cause-effect graph after each round of optimization Cause-and-effect diagram from the previous round The mean difference of the weights on the fourth side satisfies:

[0121] ;

[0122] Where E represents the set of directed edges; Indicates the adjusted core edge in the current iteration round. The weight of the fourth side; This represents the core edge after adjustment in the previous iteration. The weight of the fourth side;

[0123] This ensures that the causal graph is stable during multiple iterations, and that no new or deleted directed edges are added.

[0124] In each iteration, the network parameters are updated by performing a global fine-tuning of the network parameters for each feature extraction network, the dynamic fusion weight calculation model, and the downstream task model after loss calculation. The AdamW optimizer is used, with an initial learning rate of [missing information]. Every 5 rounds, the value decays to 0.8 times the previous value, with the total global loss... The goal is to minimize the impact of parameters. L2 regularization is added to the feature extraction network to constrain parameter tuning, and causal consistency constraints are added to the dynamic weight calculation to prevent parameter fine-tuning from disrupting the initial causal structure.

[0125] To avoid overfitting during training with small sample data, the loop can be stopped earlier if the convergence condition is difficult to meet. After the loop converges, the final causal graph, which can be used for clinical causal logic interpretation, is output. It can be used as a causal augmentation feature set for other downstream tasks. A deployable, adaptively adjustable dynamic fusion weight calculation model for modal contributions, and a lightweight downstream diagnostic model that can be used in primary hospitals to assist in diagnosis and reduce the use of medical resources and data dependence.

[0126] The overall system achieves a closed loop for unsupervised multimodal small sample models and proposes a causal-driven dual-loop solution for medical scenarios with small-scale labels, small sample size, and fluctuating modal quality.

[0127] S5: Input the multimodal data to be fused into the unsupervised few-sample fusion model to obtain the final fused features.

[0128] The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method provided in this embodiment has the following beneficial effects:

[0129] This method can uncover the essential causal relationships between modalities in small sample scenarios to filter out false associations, and improves the cross-scenario generalization ability by dynamically fusing multimodal features, providing a new technical path for the fusion and efficient utilization of small sample medical multimodal data.

[0130] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0131] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A causal-driven unsupervised multimodal small-sample dual-cycle data fusion method, characterized in that, include: S1: Obtain medical imaging data, electronic medical record text data, and physiological signal data for each lesion sample, and preprocess and align the feature dimensions of each modality data to obtain the aligned feature matrix; S2: Construct an undirected graph based on the aligned feature matrix, use kernel independent component analysis to filter undirected edges in the undirected graph, and combine medical priors to determine the direction of the undirected edges to form a causal graph; based on the causal graph and each modality data, construct a set of causal consistency sample pairs after causal consistency test; perform comparative learning of the feature extraction network based on the set of causal consistency sample pairs, output causal enhancement features, and construct and update the causal graph based on the causal enhancement features. S3: Input the updated causal graph into the dynamic fusion weight calculation model to calculate the dynamic fusion weight. Based on the causal enhancement features after weighted summation and standardization of the dynamic fusion weight, generate preliminary fusion features and filter the preliminary fusion features to obtain effective fusion features. Input the effective fusion features into the downstream diagnostic model to calculate the prediction error and correct and update the causal graph. Feed the corrected causal graph back to step S2 to guide the next round of sample generation and comparative learning. The process of calculating dynamic fusion weights includes: In the dynamic fusion weight calculation model, the weights of the causal paths between any two nodes in the causal graph are extracted and updated. Causal paths include direct causal paths and indirect causal paths. A direct causal path is a directed edge that directly connects two nodes, and an indirect causal path is a path that indirectly connects two nodes, containing at least two directed edges. The weight of a direct causal path is the weight of the fourth edge of the corresponding directed edge, and the weight of an indirect causal path is the product of the weights of the fourth edges of all directed edges on the indirect causal path. When there are multiple indirect causal paths between two nodes, the path with the largest product of the weights of the fourth edges of all directed edges on the path is taken as the indirect causal path. Calculate the sum of the weights of the causal paths from any node to the other nodes, which serves as the total causal influence score for that node. Normalize the total causal influence scores of each node to obtain the dynamic fusion weights corresponding to those nodes. The normalization formula is as follows: ; in, Indicates the first The first lesion sample Dynamic fusion weights of nodes corresponding to each modality; Indicates the first The first lesion sample The total causal influence score of the nodes corresponding to each modality; Indicates the first The first lesion sample The total causal influence score of the nodes corresponding to each modality; S4: Based on the contrast loss and prediction error, perform global fine-tuning of the feature extraction network, dynamic fusion weight calculation model, and downstream diagnostic model; repeat steps S2-S4 until the mean difference between the contrast loss, prediction loss, and edge weights in the causal graph meets the set convergence condition, and construct an unsupervised small sample fusion model. S5: Input the multimodal data to be fused into the unsupervised few-sample fusion model to obtain the final fused features.

2. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 1, characterized in that, Preprocessing of medical imaging data includes: The mean and standard deviation of pixel values ​​in all medical image data are calculated. The pixel values ​​of each pixel are normalized based on the mean and standard deviation to obtain normalized medical image data. The size of each normalized medical image data is uniformly adjusted to 224×224 using bilinear interpolation, and noise in each medical image data after size adjustment is removed by Gaussian filtering. Each denoised medical image data is input into a pre-trained RseNet-50 network with the fully connected layers removed, and the initial image features corresponding to each medical image data are extracted.

3. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 1, characterized in that, Preprocessing of electronic medical record text data includes: Each electronic medical record text data was cleaned and segmented, and special characters were removed using regular expressions to obtain string sequences. Each string sequence was split using a word segmenter, and "[CLS]" was added before each string sequence and "[SEP]" was added after each string sequence to obtain the corresponding word list. Each word list was converted into a corresponding number using a word list conversion function, and all numbers were input into the pre-trained BioBERT model. The features output by the hidden layer corresponding to the "[CLS]" mark were taken as the initial text features for each electronic medical record text data.

4. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 1, characterized in that, Preprocessing of physiological signal data includes: The physiological signal data is subjected to discrete Fourier transform to convert it into a frequency domain signal. A 50Hz band-stop filter is used to remove power frequency interference from the frequency domain signal to obtain the physiological signal data after frequency domain denoising. Inverse Fourier transforms were performed on each physiological signal data after frequency domain denoising to obtain the denoised physiological signal data. The Pan-Tompkins algorithm is used to detect the R-wave in each denoised physiological signal data. The corresponding interval is calculated based on the R-wave position of adjacent R-waves. All intervals in any physiological signal data are combined to obtain the time-domain feature vector corresponding to the physiological signal data. The denoised physiological signal data are mapped to the Mel spectrum to obtain the corresponding filter energy; each filter energy is logarithmically processed and then transformed into the frequency domain feature vector corresponding to the physiological signal data using discrete cosine transform. By splicing the time-domain feature vectors and frequency-domain feature vectors corresponding to the physiological signal data, the initial physiological features corresponding to each physiological signal data are obtained.

5. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 1, characterized in that, Feature dimension alignment includes: The initial image features, initial text features, and initial physiological features extracted from the medical imaging data, electronic medical record text data, and physiological signal data of each lesion sample are passed through three independent fully connected layers, and each feature is mapped to 256 dimensions. All initial image features, initial text features, and initial physiological features after dimension alignment are concatenated, and features belonging to the same modality are placed in the same column to obtain the aligned feature matrix.

6. The causal-driven unsupervised multimodal small-sample dual-loop data fusion method according to claim 5, characterized in that, In S2, the process of constructing the cause-effect graph includes: By aligning each feature in the feature matrix as a node and creating an undirected edge between any two nodes, an undirected graph is constructed. Kernel independent component analysis is used to perform conditional independence tests on any two nodes, and undirected edges that have no connection between nodes are removed based on the independence test results; Based on prior knowledge in the medical field, undirected edges are extracted and retained. Then, the two nodes connected by any retained undirected edge are matched with the causal logic of prior knowledge to determine the causal relationship between the two nodes. Based on the causal relationship, the undirected edge is modified into a directed edge to obtain a directed graph. The quantization method based on kernel mutual information quantifies the causal relationship strength between any two nodes in a directed graph and performs normalization processing to obtain the edge weight of the corresponding directed edge. A causal graph is formed based on nodes, edges, and their corresponding edge weights.

7. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 6, characterized in that, In S2, the process of constructing the set of causal consistent sample pairs includes: Edges with weights greater than or equal to a threshold in a causal graph are considered strong causal edges. If the modal data corresponding to the nodes connected by strong causal edges is medical image data, the Faster-R-CNN object detection algorithm is used to locate the core pixel region of the lesion sample in the medical image data, and the pixel value of the core pixel region is fixed. The background of the medical image data with fixed pixel value is randomly perturbed by Gaussian noise or by ±10% brightness perturbation, and the perturbed medical image data is used as a positive sample image. If the modal data corresponding to the nodes connected by strong causal edges is electronic medical record text data, the BioBERT segmenter is used to extract the semantic description of the lesion sample in the electronic medical record text data. The semantic description is fixed, and other redundant descriptions are perturbed by synonym replacement or sentence structure transformation. The perturbed electronic medical record text data is used as positive sample text. If the modal data corresponding to the nodes connected by strong causal edges is physiological signal data, the Pan-Tompkins algorithm is used to detect the R-wave time series segments in the physiological signal data. The interval in the R-wave time series segments is fixed, and a random perturbation of ±5% of the sampling value is added to the non-R-wave time series segments. The perturbed physiological signal data is used as the positive sample physiological signal. Edges with weights less than a threshold in the causal graph are considered weak causal edges. If the modal data corresponding to the nodes connected by weak causal edges is medical image data, a generative adversarial network is used to modify the size of the lesion sample region in the medical image data while keeping the size of the remaining regions unchanged, thus constructing a counterfactual intervention negative sample image. The first strong causal edge pointing from the node corresponding to the medical image data in the causal graph to the node corresponding to the electronic medical record text data is queried. The medical image data corresponding to the nodes of any lesion sample, the electronic medical record text data, and the two nodes connected by the first strong causal edge whose descriptions are different are used as conflict-filtered negative sample images. If the modal data corresponding to the nodes connected by the weak causal edge is electronic medical record text data, the semantic description of the lesion sample in the electronic medical record text data is rewritten as the opposite, while keeping the other semantic descriptions unchanged, to construct the counterfactual intervention negative sample text; query the second strong causal edge in the causal graph that points from the node corresponding to the electronic medical record text data to the node corresponding to the physiological signal data, and use the electronic medical record text data corresponding to the electronic medical record text data and the node corresponding to the physiological signal data in any lesion sample whose descriptions are different from the two nodes connected by the second strong causal edge as the conflict-filtered negative sample text; If the modal data corresponding to the nodes connected by the weak causal edge is physiological signal data, all intervals in the fixed R-wave time sequence of the physiological signal data are compressed to construct counterfactual intervention negative sample physiological signals; query the third strong causal edge of the node corresponding to the medical image data in the causal graph that points to the node corresponding to the physiological signal data, and use the corresponding physiological signal data whose descriptions are different between the nodes corresponding to the medical image data and physiological signal data in any lesion sample and the two nodes connected by the third strong causal edge as conflict screening negative sample physiological signals; Each modality of data corresponds to one positive sample, one counterfactual intervention negative sample, and one conflict screening negative sample. Based on the original data of each modality, a set of sample pairs is constructed. A second causal graph is constructed based on the modal data in the sample pair set. The weight of the second edge of the strong causal edge in the second causal graph is calculated, and the weight difference between the edge weight of the strong causal edge and the weight of the second edge is calculated. In each positive sample, the positive samples in which the weight difference of each strong causal edge is less than the second threshold are selected as valid positive samples. In each negative sample, the negative samples in which the weight difference of each strong causal edge is greater than the third threshold are selected as valid negative samples. The remaining samples are removed to obtain the causal consistency sample pair set. The causal consistency sample pair set includes several causal sample pairs. Each causal sample pair includes positive samples, counterfactual intervention negative samples, conflict screening negative samples, and original data corresponding to any data of the same modality.

8. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 7, characterized in that, In S2, the contrastive learning of the feature extraction network includes: The positive sample, the counterfactual intervention negative sample, the conflict screening negative sample, and the original data in any sample pair are respectively input into the feature extraction network corresponding to the modality to obtain the corresponding initial features. A third causal graph is constructed based on four initial features in each sample pair; The basic contrast loss is calculated based on the four initial features in each sample pair, and the causal constraint loss is calculated based on the third edge weight in the third causal graph. A contrastive loss is constructed based on the basic contrastive loss and the causal constraint loss. The network parameters of the feature extraction network corresponding to the mode are optimized by minimizing the contrastive loss. The positive sample, the counterfactual intervention negative sample, the conflict screening negative sample, and the original data from any sample pair are respectively input into the feature extraction network corresponding to the optimized modality to obtain the corresponding causal enhancement features. An updated causal graph is constructed based on the four causal enhancement features in each sample pair.

9. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 6, characterized in that, In S3, the process of obtaining effective fusion features includes: Step 1: Calculate the dynamic fusion weights corresponding to each node, and use the dynamic fusion weights of the nodes corresponding to the three modalities of the same lesion sample as the weight set corresponding to the current circulating lesion sample. Step 2: Standardize each causal enhancement feature corresponding to the same modality based on the mean and standard deviation of the causal enhancement features corresponding to the same modality; Step 3: The standardized causal enhancement features corresponding to the three modalities of the same lesion sample are weighted and summed with the weight set corresponding to the current circulating lesion sample to obtain the corresponding preliminary fusion features; Step 4: Perform L2 normalization on the preliminary fusion features corresponding to each lesion sample to obtain the normalized fusion features corresponding to each lesion sample; Step 5: Calculate the kernel mutual information of the normalized fusion feature with the causal enhancement features corresponding to the three modalities belonging to the same lesion sample, respectively, to obtain the kernel mutual information sequence corresponding to the lesion sample; calculate the cosine similarity between the kernel mutual information sequence corresponding to the lesion sample and the weight set corresponding to the current circulating lesion sample. When the cosine similarity is greater than the fourth threshold, the fusion is determined to be successful, and the preliminary fusion feature is taken as the effective fusion feature. Otherwise, repeat steps 1-5 until the fusion is successful.

10. The causal-driven unsupervised multimodal small-sample dual-cycle data fusion method according to claim 9, characterized in that, The process of updating the cause-effect graph includes: Step 1: Input the effective fusion features corresponding to the lesion sample into the downstream diagnostic model. The downstream diagnostic model includes a binary classification model and outputs the corresponding prediction label, which includes healthy or unhealthy. Step 2: Calculate the cross-entropy loss for each lesion sample based on the predicted label and the corresponding true label. Step 3: Classify lesion samples with cross-entropy loss higher than the error threshold as high-error samples, and classify lesion samples with cross-entropy loss less than or equal to the error threshold as low-error samples. Step 4: For any high-error sample, call the weight set corresponding to the current cyclic lesion sample, and select the two nodes with the largest dynamic fusion weights from the weight set as core nodes; In updating the causal graph, the direct causal path between two core nodes is extracted as the core edge corresponding to the high error sample. The first frequency of the core edge corresponding to the high error sample in all causal paths between the two core nodes in all high error samples is counted. If the first frequency is higher than the frequency threshold, the corresponding core edge is marked as an unreliable path; otherwise, it is marked as a reliable path. Step 5: For any low-error sample, execute step 4 to obtain the second frequency of the core edge appearing in all low-error samples. If the second probability is higher than the frequency threshold, mark the corresponding core edge as an unreliable path; otherwise, mark it as a reliable path. Step 6: Adjust and update the weight of the fourth edge of the corresponding directed edge in the causal graph based on the first and second frequencies; and normalize the adjusted weight of the fourth edge. Step 7: In updating the causal graph, delete the core edges marked as unreliable paths and those that conflict with medical priors; reverse the core edges in high-error samples where the first frequency is lower than the second frequency of their reverse edges in low-error samples, and retain the original weight of the fourth edge. Step 8: Feed the corrected causal graph back to step S2 to guide the next round of sample generation and comparative learning.

Citation Information

Patent Citations

  • Commercial credit evaluation and supervision method based on multi-modal coevolution algorithm

    CN119250963A

  • Multi-modal neural networks with decoder-only language models

    WO2024206231A1