Multi-modal data fusion method and device, equipment and storage medium
By using attention gating mechanism for adaptive fusion in multimodal data fusion, the problem of inability to adaptively adjust modal weights in the prior art is solved, and the accuracy of disease diagnosis and the applicability of the model are improved.
Patent Information
- Application Number
- CN202510149376.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-23
AI Technical Summary
The existing multimodal data fusion method cannot adjust the weight of different modal information according to pathological adaptation in pathological diagnosis, resulting in inaccurate multimodal features and affecting the accuracy of disease diagnosis results.
The attention gating mechanism is used to control data of different modes for adaptive fusion, dynamically adjust the weight of clinical feature information and picture feature information to achieve adaptive fusion.
The reliability of multimodal features and the accuracy of disease diagnosis results are improved, and the problem of inaccurate diagnosis results caused by fixed weights is avoided, while the applicability of the model is improved.
Smart Images

Figure CN120030496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data fusion, and in particular to a multimodal data fusion method, device, equipment and storage medium. Background Art
[0002] With the development of deep learning technology, significant progress has been made in the field of medical image analysis, especially in computer-aided diagnosis. Multimodal fusion has opened up new ways to improve diagnostic accuracy by integrating different types of medical data. In pathological diagnosis tasks, the joint analysis of WSI (Whole Slide Imaging) and clinical data is crucial for accurate diagnosis of diseases.
[0003] Current joint analysis and diagnosis methods usually use fixed weights to fuse whole-slice images and clinical data. This fusion method cannot adaptively adjust the weights of different modal information according to the pathology for information fusion, resulting in inaccurate multimodal features, thereby affecting the accuracy of disease diagnosis results. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a multimodal data fusion method, device, equipment and storage medium, which can use the attention gating mechanism to control the adaptive fusion of data of different modalities, so that the data of different modalities can be adaptively adjusted for weight fusion, thereby improving the accuracy of disease diagnosis results. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a multimodal data fusion method, comprising:
[0006] Acquire a target full-slice image and target clinical data corresponding to the target full-slice image from a target data source, and acquire corresponding clinical feature information based on the target clinical data;
[0007] Using the target multimodal fusion diagnosis model and the preset transformation functions to obtain each initial feature information of different scales corresponding to the target full-slice image, and performing a weighted fusion operation on each of the initial feature information to obtain the image feature information corresponding to the target full-slice image;
[0008] The weights of the clinical feature information and the image feature information are dynamically adjusted based on a preset attention gating mechanism, and the clinical feature information and the image feature information are adaptively fused according to the corresponding adjustment results to obtain corresponding target multimodal features.
[0009] Optionally, before acquiring the target full-slice image and the target clinical data corresponding to the target full-slice image from the target data source, the method further includes:
[0010] A target multimodal fusion framework is constructed according to the target analysis task corresponding to the target full-slice image and the target clinical data, and the target multimodal fusion diagnosis model is acquired based on the target multimodal fusion framework.
[0011] Optionally, the using the target multimodal fusion diagnosis model and the preset transformation functions to obtain each initial feature information of different scales corresponding to the target full slice image includes:
[0012] The initial feature information of different scales corresponding to the target full slice image is calculated in parallel using the preset transformation functions corresponding to the feature extraction layers in the target multimodal fusion diagnosis model.
[0013] Optionally, the performing a weighted fusion operation on each of the initial feature information includes:
[0014] Calculating the attention weight corresponding to each of the initial feature information to obtain each weighted initial feature information according to the corresponding calculation result;
[0015] The preset attention mechanism is used to analyze each of the weighted initial feature information to obtain the importance distribution result of each of the weighted initial feature information, and a weighted fusion operation is performed on each of the initial feature information according to the importance distribution result.
[0016] Optionally, the dynamically adjusting the weights of the clinical feature information and the image feature information based on a preset attention gating mechanism includes:
[0017] Acquire target gating values corresponding to the clinical feature information and the picture feature information according to the clinical feature information, the picture feature information and the preset attention gating mechanism;
[0018] The weights of the clinical feature information and the image feature information are dynamically adjusted based on the target gating value.
[0019] Optionally, the acquiring, according to the clinical feature information, the image feature information and the preset attention gating mechanism, a target gating value corresponding to the clinical feature information and the image feature information includes:
[0020] The clinical feature information and the image feature information are mapped to a target feature space using a preset feature transformation network, and the clinical feature information and the image feature information are spliced in the target feature space using the preset attention gating mechanism to obtain corresponding spliced feature information, and the target gating value is obtained based on the spliced feature information.
[0021] Optionally, the model structure of the target multimodal fusion diagnosis model includes a residual learning network, and the target multimodal fusion diagnosis model is pre-configured with a residual learning mechanism.
[0022] In a second aspect, the present application provides a multimodal data fusion device, comprising:
[0023] A data acquisition module, used to acquire a target full-slice image and target clinical data corresponding to the target full-slice image from a target data source, and acquire corresponding clinical feature information based on the target clinical data;
[0024] A feature fusion module is used to obtain each initial feature information of different scales corresponding to the target full-slice image by using the target multimodal fusion diagnosis model and each preset transformation function, and perform a weighted fusion operation on each initial feature information to obtain the image feature information corresponding to the target full-slice image;
[0025] A multimodal feature acquisition module is used to dynamically adjust the weights of the clinical feature information and the image feature information based on a preset attention gating mechanism, and adaptively fuse the clinical feature information and the image feature information according to the corresponding adjustment results to obtain the corresponding target multimodal features.
[0026] In a third aspect, the present application provides an electronic device, including:
[0027] Memory, used to store computer programs;
[0028] A processor is used to execute the computer program to implement the aforementioned multimodal data fusion method.
[0029] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which implements the aforementioned multimodal data fusion method when executed by a processor.
[0030] In the present application, firstly, a target full-slice image and target clinical data corresponding to the target full-slice image are obtained from a target data source, and corresponding clinical feature information is obtained based on the target clinical data. Then, the target multimodal fusion diagnosis model and each preset transformation function are used to obtain each initial feature information of different scales corresponding to the target full-slice image, and each initial feature information is weighted fused to obtain the image feature information corresponding to the target full-slice image. Finally, the weights of the clinical feature information and the image feature information are dynamically adjusted based on a preset attention gating mechanism, and the clinical feature information and the image feature information are adaptively fused according to the corresponding adjustment results to obtain the corresponding target multimodal features. It can be seen that the present application avoids the problem of using only one modality of data for diagnosis and analysis of the disease by integrating data of two different modalities, namely, images and clinical data, and improves the accuracy of the diagnosis and analysis results. By using the attention gating mechanism to control the adaptive fusion of data of different modalities, the data of different modalities can adaptively adjust the fusion weights according to the case, avoiding the problem of inaccurate diagnosis results caused by using fixed weights, and at the same time improving the applicability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0032] Figure 1 A flow chart of a multimodal data fusion method provided in an embodiment of the present application;
[0033] Figure 2 A schematic diagram of a multimodal data fusion method flow chart provided in an embodiment of the present application;
[0034] Figure 3 A multi-scale feature information fusion flow chart provided in an embodiment of the present application;
[0035] Figure 4 A multi-scale feature information extraction effect diagram provided in an embodiment of the present application; wherein (a) is a small-scale feature information extraction effect diagram, (b) is a medium-scale feature information extraction effect diagram, and (c) is a large-scale feature information extraction effect diagram;
[0036] Figure 5A schematic diagram of multi-scale feature activity provided in an embodiment of the present application; wherein (a) is a schematic diagram of small-scale feature activity, (b) is a schematic diagram of medium-scale feature activity, and (c) is a schematic diagram of large-scale feature activity;
[0037] Figure 6 A schematic diagram of attention distribution provided in an embodiment of the present application;
[0038] Figure 7 A feature information fusion flow chart provided in an embodiment of the present application;
[0039] Figure 8 A schematic diagram of a gating value distribution provided in an embodiment of the present application; wherein (a) and (b) are schematic diagrams of gating value distribution corresponding to different diseases, respectively;
[0040] Fig. 9 A schematic diagram of the structure of a multimodal data fusion device provided in an embodiment of the present application;
[0041] Fig.10 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] At present, in pathological diagnosis and analysis, the fusion method of multimodal data cannot adaptively adjust the weights of different modal information according to the pathology to perform information fusion, resulting in the problem of inaccurate multimodal features. To this end, an embodiment of the present application provides a multimodal data fusion method, which improves the reliability of multimodal features and thus improves the accuracy of disease diagnosis results by controlling the adaptive fusion of data of different modalities using an attention gating mechanism.
[0044] See also Figure 1 As shown, an embodiment of the present invention discloses a multimodal data fusion method, comprising:
[0045] Step S11: acquiring a target full-slice image and target clinical data corresponding to the target full-slice image from a target data source, and acquiring corresponding clinical feature information based on the target clinical data.
[0046] In this embodiment, before obtaining the target full-slice image and the target clinical data corresponding to the target full-slice image, it also includes: constructing a target multimodal fusion framework according to the target analysis tasks corresponding to the target full-slice image and the target clinical data, and obtaining a target multimodal fusion diagnosis model based on the target multimodal fusion framework; specifically, for the joint analysis task of WSI images and clinical data, this embodiment proposes a multimodal fusion framework based on an optimized gating mechanism, namely, the MSGF-DiagNet (Multi-Scale Gated Fusion Diagnostic Network) multimodal fusion framework, which aims to improve diagnostic performance through dynamic feature fusion while maintaining good interpretability; and the framework adopts a two-stream network structure to process images and clinical data respectively; and according to the framework, a multimodal fusion diagnosis model is obtained: a multi-level attention mechanism is designed, and at the feature extraction level, multi-scale feature adaptive fusion is adopted to enhance the representation ability of the model, especially the recognition ability of subtle lesions; at the modality fusion level, a gated fusion mechanism is used to dynamically adjust the weights of image and clinical features, maintain the balance of feature contribution, realize adaptive fusion of multimodal features, and improve feature utilization. It should be noted that the model structure of the above-mentioned target multimodal fusion diagnosis model includes a residual learning network, and the target multimodal fusion diagnosis model is pre-configured with a residual learning mechanism, that is, the framework of the above-mentioned target multimodal fusion diagnosis model includes three core components: a modality-specific feature transformation network, an attention gating mechanism, and a residual learning structure; it should be noted that in the residual learning structure, this embodiment can use different residual connection methods, such as using residual connection methods such as dense connections (Dense Connections) or shortcut connections.
[0047] This embodiment can use API (Application Programming Interface) or other data extraction tools to automatically extract data from the target database. In this embodiment, after obtaining the target full-slice image and clinical data from the target data source, in order to improve the reliability of the data, the data can also be screened and judged. For example, after obtaining the full-slice image and clinical data from the target data source, the correspondence of the two data is judged, and the availability of the above-mentioned full-slice image and clinical data is judged separately. If there are errors in the above-mentioned data, the erroneous data is discarded, and the full-slice data and clinical data are obtained again from the target data source. By screening and judging the data, the reliability of the data is improved; by configuring a multimodal fusion diagnosis model including an attention gating mechanism, the data of different modalities can be adaptively adjusted in weight for fusion, thereby improving the reliability of the multimodal features. This modular design not only improves the flexibility of the model, but also facilitates subsequent expansion and optimization.
[0048] Step S12: using the target multimodal fusion diagnosis model and each preset transformation function to obtain each initial feature information of different scales corresponding to the target full slice image, and performing a weighted fusion operation on each of the initial feature information to obtain the image feature information corresponding to the target full slice image.
[0049] Depend on Figure 2 As shown in the multimodal data fusion process shown, it is first necessary to obtain the initial feature information corresponding to the target full-slice image, and fuse the initial feature information to obtain the image feature information corresponding to the full-slice image, and finally fuse the image feature information with the clinical feature information corresponding to the clinical data; in the multimodal fusion diagnosis model, this embodiment proposes a novel multi-scale feature extraction module. This module effectively captures the multi-granular pathological features in the WSI image by combining multi-scale feature learning and hierarchical attention mechanism; accordingly, the above-mentioned process of obtaining the initial feature information of different scales corresponding to the full-slice image can specifically include: using the preset transformation functions corresponding to each feature extraction layer in the target multimodal fusion diagnosis model to parallelly calculate the initial feature information of different scales corresponding to the target full-slice image; specifically, since the pathological image contains both microscopic cytological features and macroscopic histological features, three parallel feature extraction branches are designed in this example, namely the above-mentioned feature extraction layers. As Figure 3 As shown in the multi-scale feature information fusion process, for the input feature , multi-scale feature representation is obtained through three transformation functions of different scales:
[0050] #Small-scale branch (dimension is out_feature_size / 2), focusing on capturing local cell morphology and other microscopic features;
[0051] #Medium-scale branch (dimension is out_feature_size), focusing on extracting medium-range tissue structure features;
[0052] #Large-scale branch (dimension is out_feature_size*2), responsible for modeling a wider range of tissue distribution patterns and contextual information;
[0053] in , , They represent nonlinear transformation functions of different scales, that is, the above-mentioned preset transformation functions. Figure 4As shown in the performance of image feature extraction at different scales, in a specific implementation, the model is most active in local feature extraction, indicating that it is very sensitive to detail information, and medium-scale features show a relatively balanced feature extraction ability; global features are relatively conservative, which may indicate that the model is more cautious when integrating large-scale information; it should be noted that, Figure 5 As shown in the multi-scale feature activity in , local features tend to capture significant detail patterns (high activation values), medium-scale features maintain stable information encoding between different samples, and global features tend to extract common features (low variance).
[0054] In this embodiment, the process of weighted fusion of each initial feature information corresponding to the full slice image may specifically include: calculating the attention weight corresponding to each initial feature information to obtain each weighted initial feature information according to the corresponding calculation result; analyzing each weighted initial feature information using a preset attention mechanism to obtain the importance distribution result of each weighted initial feature information, and performing a weighted fusion operation on each initial feature information according to the importance distribution result; specifically, in order to make full use of the complementarity of multi-scale features, this embodiment uses a double-layer attention mechanism; at the feature level, this embodiment introduces an attention unit for each scale branch, and calculates the attention weight:
[0055] ;
[0056] in Represents the L2 norm. Then we get the weighted features:
[0057] ;
[0058] At the scale level, this embodiment designs a scale attention module by learning the importance distribution of three scale features:
[0059] , which is the above importance distribution result.
[0060] in Represents the importance weights of the three scales. This hierarchical attention design enables the model to flexibly focus on the most diagnostically valuable features according to the characteristics of different cases.
[0061] In the feature fusion stage, this embodiment adopts a progressive strategy, and the final output feature can be expressed as:
[0062] ;
[0063] in represents the nonlinear transformation of the feature integration module. Therefore, the forward propagation process of the entire module can be concisely expressed as:
[0064] ,in, .
[0065] In a specific embodiment, Figure 6 As shown in the scale preference analysis in the figure, the model obviously tends to use medium-scale features (60%), indicating that medium-range patterns and structures dominate the decision-making; small-scale and large-scale features receive equal attention allocation (20% each), indicating that the model relies on local details and global information to a similar degree, showing the balance of the model in processing information of different scales; taking medium-range feature patterns as the main decision-making basis means that the model is better at capturing medium-sized targets or features, paying equal attention to local details and global context, and forming complementary auxiliary information. This distribution shows that the model adopts an information processing strategy of "mainly in the middle and supplemented by both ends", which is particularly suitable for processing tasks that require comprehensive consideration of medium-range features, such as organ segmentation, lesion detection, etc., and has a good balance for tasks that require both details and overall considerations. This distribution shows that the model has a strong ability to extract mesoscopic features and achieves a certain balance between different scales, avoiding excessive reliance on one extreme, and may have better robustness when processing complex scenes.
[0066] It should be noted that in order to improve the generalization ability of the model, this embodiment introduces a dropout mechanism in each processing stage, and ensures the stability of the feature distribution through LayerNorm (layer normalization); and the above-mentioned multi-scale feature extraction module can effectively extract multi-level pathological features in WSI images, providing rich and reliable feature representations for subsequent multimodal fusion and disease diagnosis. By visually analyzing the scale attention weights, the model can adaptively adjust the degree of attention to features of different scales according to the characteristics of different cases, which is consistent with the diagnostic thinking process of clinicians. By introducing the dropout mechanism and using LayerNorm, the stability of the feature distribution is guaranteed; by extracting feature information of different scales of the full slice image, a rich and reliable feature representation is provided for subsequent multimodal fusion and disease diagnosis.
[0067] Step S13: dynamically adjust the weights of the clinical feature information and the image feature information based on a preset attention gating mechanism, and adaptively fuse the clinical feature information and the image feature information according to the corresponding adjustment results to obtain corresponding target multimodal features.
[0068] In the multimodal fusion diagnosis model, this embodiment designs an optimized gated fusion module (OptimizedGatedFusion), and the corresponding feature fusion process is as follows: Figure 7As shown in FIG. 1 , the module includes the above-mentioned attention gating mechanism, which is used to adaptively integrate WSI image features and clinical features. The module realizes the dynamic fusion of two modal information through mechanisms such as feature transformation, attention gating and residual learning. Accordingly, the process of dynamically adjusting the weights of clinical feature information and image feature information based on the preset attention gating mechanism may specifically include: obtaining target gating values corresponding to the clinical feature information and the image feature information according to the clinical feature information, the image feature information and the preset attention gating mechanism; dynamically adjusting the weights of the clinical feature information and the image feature information based on the target gating value; the above-mentioned process of dynamically adjusting the weights of the clinical feature information and the image feature information according to the clinical feature information, the image feature information and the preset attention gating mechanism. The process of obtaining target gating values corresponding to clinical feature information and image feature information by using a preset feature transformation network and a preset attention gating mechanism may specifically include: mapping the clinical feature information and the image feature information to a target feature space by using a preset feature transformation network, splicing the clinical feature information and the image feature information in the target feature space by using the preset attention gating mechanism to obtain corresponding spliced feature information, and obtaining the target gating value according to the spliced feature information; specifically, for the input image feature X_img and clinical feature X_cli, firstly performing modality-specific feature mapping through a feature transformation network:
[0069] ;
[0070] Where T_img and T_cli represent the transformation functions of image and clinical features, respectively, including linear mapping and LayerNorm normalization layer, which are used to map features of different modalities into the same feature space.
[0071] In order to achieve a dynamic balance between modalities, this embodiment introduces an attention gating mechanism. The gating value is calculated by concatenating the transformed bimodal features:
[0072] ;
[0073] in represents the sigmoid activation function, W_1 and W_2 are learnable weight matrices. It reflects the degree of model's reliance on the two modal information in the current sample.
[0074] Based on the calculated gating values, the model adaptively fuses the bimodal features:
[0075] ;
[0076] in This gated fusion mechanism enables the model to dynamically adjust the degree of attention paid to image and clinical information according to the characteristics of different cases.
[0077] The final fusion feature can be expressed as:
[0078] ;
[0079] It is worth noting that the module also reserves the design of residual learning:
[0080] ;
[0081] in Represents the residual mapping function. This design helps to alleviate the degradation problem in deep network training. It should be noted that through experiments, it can be found that directly using F_fused can obtain good performance, and using the residual learning mechanism can further alleviate the degradation problem in deep network training; it should be noted that in this embodiment, different attention mechanisms can be used, such as self-attention (Self-Attention) or graph attention network (Graph Attention Networks), to replace the existing attention-based gating unit; and in this embodiment, different feature fusion strategies can be adopted, such as tensor-level fusion or subspace-level fusion to replace the operation-level gating fusion mechanism.
[0082] Experimental results show that this feature fusion method based on the gating mechanism can effectively integrate the visual information and clinical phenotype information of WSI images. Figure 8 From the distribution of gating values shown, it can be observed that the model can adaptively adjust the degree of dependence on different modal information according to the characteristics of different cases. For example, when the WSI image contains obvious pathological features, the model will give higher weight to the image features; and when clinical indicators are more critical to the diagnosis, the model will rely more on clinical features. This dynamic adaptability is consistent with the diagnosis process of doctors comprehensively analyzing multi-source information in clinical practice. This embodiment effectively solves the problem of differences in feature representation between modalities through modality-specific feature transformation networks and residual learning structures, and improves the stability and generalization of the model; through visual analysis of the distribution of gating values, an intuitive explanation of the model decision process is provided, which enhances the credibility and usability of the model in clinical practice, and has important reference value for promoting the research and application of multimodal medical diagnosis; by analyzing the distribution of gating values, an intuitive explanation is provided for model decision-making. Doctors can understand the degree of dependence of the model on WSI images and clinical indicators in different cases by observing the gating values, which is of great significance to improving the credibility of clinical decision-making.
[0083] To verify the reliability of the multimodal data fusion method, this embodiment also evaluates the performance of the method on a large-scale open source tumor dataset, which includes core needle biopsy whole slice images (WSI) and corresponding clinical data of a total of 1,058 early breast cancer patients. The clinical data contains 5 features, including age (numeric value), tumor size (numeric value), ER (Estrogen Receptor) classification, PR classification (Progesterone Receptor) classification, and HER2 classification (Human epidermal growth factor receptor 2).
[0084] In addition, this example compares and analyzes the above multimodal data fusion method with the existing methods, and the results are shown in the following table:
[0085] Table 1 Model evaluation results
[0086]
[0087] Among them, GLAF-AB (global-and local-feature interaction with attention-based) is a global and local feature interaction network based on the attention mechanism, the DL-CNB+C model (DL core-needlebiopsy-clinical data) is a core needle biopsy-clinical data model based on deep learning, the DL-CNB (DLcore-needle biopsy) model is a core needle biopsy model based on deep learning, and the Clinical data only model is a clinical data model based on deep learning.
[0088] As can be seen from Table 1 above, this embodiment uses the following standards as evaluation indicators:
[0089] AUC (Area Under Curve, i.e. the area enclosed by a curve and the coordinate axis): reflects the overall discrimination ability of the model;
[0090] Accuracy: classification accuracy;
[0091] Sensitivity: sensitivity (true positive rate);
[0092] Specificity: specificity (true negative rate);
[0093] F1-score: A comprehensive indicator that balances precision and recall;
[0094] The calculation formulas for the above data are as follows:
[0095] ;
[0096] Among them, TP (True Positive) is the true positive example, TN (True Negative) is the true negative example, FP (False Positive) is the false positive example, FN (False Negative) is the false negative example, Prec is the precision rate, and Recall is the recall rate.
[0097] The experimental results show that compared with the GLAF-AB (global-and local-feature interaction with attention-based) model, the MSGF-DiagNet model in this embodiment, that is, the aforementioned target multimodal fusion diagnosis model, has significant improvements in AUC, sensitivity, and F1 index, with AUC increased by 3.7%, sensitivity increased by 1.14%, and F1 increased by 1.1%; compared with the DL-CNB+Cmodel, MSGF-DiagNet has significant improvements in AUC, sensitivity, and F1 index, with AUC increased by 4.2%, sensitivity increased by 2.02%, and F1 increased by 2.4%; compared with the Clinical data only (baseline) model, the comprehensive indicators of classification accuracy, sensitivity, specificity, and balanced precision and recall have been greatly improved. This result confirms the effectiveness of the dynamic fusion mechanism of this embodiment.
[0098] To verify the contribution of each module, this embodiment also conducts an ablation experiment:
[0099] Table 2 Comparison between MSGF-DiagNet model and GF-DiagNet model
[0100]
[0101] Table 3 Comparison between MSGF-DiagNet model and MS-DiagNet model
[0102]
[0103] Through ablation experiments, it can be found that the combination of MS and GF is better than using MS (Multi-Scale Diagnostic Network) or GF (Gated Fusion Diagnostic Network) alone, which proves the rationality of the model architecture design in this embodiment.
[0104] See also Fig. 9As shown, an embodiment of the present invention provides a multimodal data fusion device, including:
[0105] A data acquisition module 11 is used to acquire a target full-slice image and target clinical data corresponding to the target full-slice image from a target data source, and acquire corresponding clinical feature information based on the target clinical data;
[0106] The feature fusion module 12 is used to obtain each initial feature information of different scales corresponding to the target full-slice image by using the target multimodal fusion diagnosis model and each preset transformation function, and perform a weighted fusion operation on each initial feature information to obtain the image feature information corresponding to the target full-slice image;
[0107] The multimodal feature acquisition module 13 is used to dynamically adjust the weights of the clinical feature information and the image feature information based on a preset attention gating mechanism, and adaptively fuse the clinical feature information and the image feature information according to the corresponding adjustment results to obtain the corresponding target multimodal features.
[0108] It can be seen that this application avoids the problem of using only one modality of data to diagnose and analyze the disease by integrating data of two different modalities, namely images and clinical data, and improves the accuracy of the diagnostic analysis results; by using the attention gating mechanism to control the adaptive fusion of data of different modalities, the fusion weights of data of different modalities can be adaptively adjusted according to the case, avoiding the problem of inaccurate diagnostic results caused by the use of fixed weights, while improving the applicability of the model.
[0109] In some specific implementations, the data acquisition module 11 further includes:
[0110] A model acquisition unit is used to construct a target multimodal fusion framework according to the target analysis task corresponding to the target full-slice image and the target clinical data, and to acquire the target multimodal fusion diagnosis model based on the target multimodal fusion framework.
[0111] In some specific implementations, the feature fusion module 12 may specifically include:
[0112] The feature information calculation unit is used to use the preset transformation functions corresponding to the feature extraction layers in the target multimodal fusion diagnosis model to parallelly calculate the initial feature information of different scales corresponding to the target full slice image.
[0113] In some specific implementations, the feature fusion module 12 may specifically include:
[0114] A weight calculation unit, used to calculate the attention weight corresponding to each of the initial feature information, so as to obtain each weighted initial feature information according to the corresponding calculation result;
[0115] The feature fusion unit is used to analyze each of the weighted initial feature information using a preset attention mechanism to obtain an importance distribution result of each of the weighted initial feature information, and perform a weighted fusion operation on each of the initial feature information according to the importance distribution result.
[0116] In some specific implementations, the multimodal feature acquisition module 13 may specifically include:
[0117] A gating value acquisition submodule, used to acquire target gating values corresponding to the clinical feature information and the image feature information according to the clinical feature information, the image feature information and the preset attention gating mechanism;
[0118] A weight adjustment unit is used to dynamically adjust the weights of the clinical feature information and the image feature information based on the target gating value.
[0119] In some specific implementations, the gate value acquisition submodule may specifically include:
[0120] A gate control acquisition unit is used to map the clinical feature information and the image feature information to a target feature space using a preset feature transformation network, and to splice the clinical feature information and the image feature information in the target feature space using the preset attention gating mechanism to obtain corresponding spliced feature information, and to obtain the target gating value based on the spliced feature information.
[0121] Furthermore, the present application also discloses an electronic device. Fig.10 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0122] Fig.10 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the multimodal data fusion method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0123] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0124] In addition, the memory 22 as a carrier for storing resources may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0125] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, and may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the multimodal data fusion method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.
[0126] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the multimodal data fusion method disclosed above. The specific steps of the method can refer to the corresponding contents disclosed in the above embodiments, and will not be repeated here.
[0127] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0128] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0129] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0130] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0131] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A multimodal data fusion method, characterized in that: include: Acquire a target full-slice image and target clinical data corresponding to the target full-slice image from a target data source, and acquire corresponding clinical feature information based on the target clinical data; Using the target multimodal fusion diagnosis model and the preset transformation functions to obtain each initial feature information of different scales corresponding to the target full-slice image, and performing a weighted fusion operation on each of the initial feature information to obtain the image feature information corresponding to the target full-slice image; The weights of the clinical feature information and the image feature information are dynamically adjusted based on a preset attention gating mechanism, and the clinical feature information and the image feature information are adaptively fused according to the corresponding adjustment results to obtain corresponding target multimodal features.
2. The multimodal data fusion method according to claim 1, characterized in that: Before acquiring the target full-slice image and the target clinical data corresponding to the target full-slice image from the target data source, the method further includes: A target multimodal fusion framework is constructed according to the target analysis task corresponding to the target full-slice image and the target clinical data, and the target multimodal fusion diagnosis model is acquired based on the target multimodal fusion framework.
3. The multimodal data fusion method according to claim 1, characterized in that: The method of using the target multimodal fusion diagnosis model and the preset transformation functions to obtain initial feature information of different scales corresponding to the target full-slice image includes: The initial feature information of different scales corresponding to the target full slice image is calculated in parallel using the preset transformation functions corresponding to the feature extraction layers in the target multimodal fusion diagnosis model.
4. The multimodal data fusion method according to claim 1, characterized in that: The weighted fusion operation is performed on each of the initial feature information, including: Calculating the attention weight corresponding to each of the initial feature information to obtain each weighted initial feature information according to the corresponding calculation result; The preset attention mechanism is used to analyze each of the weighted initial feature information to obtain the importance distribution result of each of the weighted initial feature information, and a weighted fusion operation is performed on each of the initial feature information according to the importance distribution result.
5. The multimodal data fusion method according to claim 1, characterized in that: The dynamically adjusting the weights of the clinical feature information and the image feature information based on a preset attention gating mechanism includes: Acquire target gating values corresponding to the clinical feature information and the picture feature information according to the clinical feature information, the picture feature information and the preset attention gating mechanism; The weights of the clinical feature information and the image feature information are dynamically adjusted based on the target gating value.
6. The multimodal data fusion method according to claim 5, characterized in that: The obtaining, according to the clinical feature information, the image feature information and the preset attention gating mechanism, a target gating value corresponding to the clinical feature information and the image feature information comprises: The clinical feature information and the image feature information are mapped to a target feature space using a preset feature transformation network, and the clinical feature information and the image feature information are spliced in the target feature space using the preset attention gating mechanism to obtain corresponding spliced feature information, and the target gating value is obtained based on the spliced feature information.
7. The multimodal data fusion method according to any one of claims 1 to 6, characterized in that: The model structure of the target multimodal fusion diagnosis model includes a residual learning network, and the target multimodal fusion diagnosis model is pre-configured with a residual learning mechanism.
8. A multimodal data fusion device, characterized in that: include: A data acquisition module, used to acquire a target full-slice image and target clinical data corresponding to the target full-slice image from a target data source, and acquire corresponding clinical feature information based on the target clinical data; A feature fusion module is used to obtain each initial feature information of different scales corresponding to the target full-slice image by using the target multimodal fusion diagnosis model and each preset transformation function, and perform a weighted fusion operation on each initial feature information to obtain the image feature information corresponding to the target full-slice image; A multimodal feature acquisition module is used to dynamically adjust the weights of the clinical feature information and the image feature information based on a preset attention gating mechanism, and adaptively fuse the clinical feature information and the image feature information according to the corresponding adjustment results to obtain the corresponding target multimodal features.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the multimodal data fusion method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the multimodal data fusion method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal breast cancer classification training method and system based on graph attention network
CN114864076A
Multi-modal fusion survival prognosis method and device based on pathology and genes
CN117594225A
Multi-modal fusion survival prognosis method based on cross Transform and MLIF
CN117746201A
Breast cancer histological image segmentation and classification method based on multi-scale deep learning
CN119180838A
Cited By
Classification model, classification method and device based on multi-modal multi-scale fusion
CN120565112A
Facial paralysis grading method, system and equipment fusing multi-modal data and medium
CN120876479A
Multi-modal orthopedic infection diagnosis method and system based on double-flow encoder fusion
CN121122681A
A Multimodal Diagnostic Method and System for Orthopedic Infections Based on Dual-Stream Encoder Fusion
CN121122681B
Epilepsy prediction system based on multi-modal biological image and image data processing method
CN121565504A