Medical multi-modal data fusion method and system

Through the preprocessing of multimodal data, cross-attention mechanism and Q-Lora quantization technology, the problems of heterogeneity and computational complexity in medical multimodal data fusion are solved, efficient and accurate data fusion is achieved, and personalized diagnosis and treatment plan recommendations are supported.

CN120636847APending Publication Date: 2025-09-12YUNNAN UNITED VISION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510763782.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing medical multimodal data fusion technologies face many challenges in dealing with data heterogeneity, capturing potential relationships between modalities, inconsistent data quality, and computational complexity. It is difficult to fuse multimodal data efficiently and accurately to meet the needs of modern medicine.

Method used

A preprocessing step is used to convert multimodal data into high-dimensional feature vectors, and modal correlation calculation and alignment are performed through the cross-attention mechanism. The pre-trained large language model is fine-tuned by combining Q-Lora quantization technology and multi-head self-attention mechanism to ensure unified representation and efficient processing of data.

Benefits of technology

It achieves efficient and accurate fusion of multimodal data, improves the accuracy of diagnosis and treatment plans, and is suitable for personalized medicine and real-time health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636847A_ABST
    Figure CN120636847A_ABST
Patent Text Reader

Abstract

The invention relates to a medical multi-modal data fusion method and system, and belongs to the technical field of medical data processing. The method comprises the following steps: preprocessing multi-modal data, and converting the multi-modal data into a high-dimensional feature vector to obtain each modal feature; modal correlation calculation and alignment are carried out on the preprocessed multi-modal data, so that features of all modals are fused into a unified representation, and multi-modal fusion features are obtained; and performing fine tuning on the pre-trained large language model through Q-Lora, training the pre-trained large language model after fine tuning by using the multi-modal fusion features processed by the multi-head self-attention mechanism, and predicting input multi-modal data by using the trained pre-trained large language model. According to the method, different types of data can be seamlessly integrated, the consistency of data fusion is ensured, the accuracy of multi-modal data fusion is improved, and the training cost and computing resources of the model are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a medical multimodal data fusion method and system, and belongs to the technical field of medical data processing. Background Art

[0002] With the rapid development of modern medicine, the demand for data-driven personalized medicine in clinical practice continues to increase. The use of multimodal medical data is becoming increasingly widespread. This data comes from a variety of medical testing equipment and technologies, such as text data, tabular data, imaging data (X-rays, CT scans, MRIs), time series data (electrocardiograms, blood pressure monitoring, blood glucose monitoring), and biosignals (brain waves, electromyography). In addition, there is also unstructured text data such as electronic health records (EHRs) and physician diagnostic reports. Multimodal data can reflect a patient's health status from different perspectives. By integrating cross-modal data, a more comprehensive analysis of the patient's condition and the development of more accurate diagnosis and treatment plans can be achieved.

[0003] Despite the enormous potential of multimodal medical data, existing technologies face numerous challenges in processing and fusing this data. Current multimodal data fusion methods are mostly limited to single modalities or involve simple modality combinations, failing to fully exploit the potential connections and complementary information between data. Furthermore, due to the wide variety of data sources and their strong heterogeneity, the acquisition methods, formats, and timestamp characteristics of each modality vary significantly, posing significant challenges to the accuracy and efficiency of data fusion.

[0004] Existing medical multimodal data analysis technologies have made significant progress in some specific fields (such as deep learning processing of imaging data, modeling and prediction of time series data). These technologies can process different types of data separately and show strong effects in single-modal analysis. For example, convolutional neural networks (CNN) for imaging data and long short-term memory networks (LSTM) for time series data have been widely used in their respective fields and have achieved decent results. In the diagnosis and risk prediction of some diseases, artificial intelligence technologies based on single-modal data can also provide relatively accurate prediction and analysis results. Although existing technologies have made breakthroughs in single-modal data processing, there are still major deficiencies in the fusion of multimodal data:

[0005] Heterogeneous data fusion is difficult: Data from different modalities often come from different devices and detection methods, with varying spatiotemporal resolutions, acquisition frequencies, and data structures. For example, imaging data is high-resolution two- or three-dimensional images, while electrocardiogram (ECG) or blood pressure data is continuous time series data. This heterogeneity makes it difficult for traditional data fusion technologies to effectively integrate and unify data.

[0006] Capturing the correlations in multimodal data is difficult: Complex underlying relationships exist between different modal data, and traditional linear methods or those based on simple statistical features cannot accurately capture these complex cross-modal dependencies. For example, genomic data may have nonlinear interactions with imaging data or time series data, making it difficult for existing methods to reveal such correlations.

[0007] Inconsistent data quality: The acquisition process of data from various modalities can be affected by a variety of factors, including device performance, operator skill level, and patient compliance. This can lead to inconsistent data quality, including noise and outliers. Addressing these quality issues during the fusion process remains a technical challenge.

[0008] High computational complexity: Multimodal data fusion often involves large amounts of high-dimensional data, particularly genomic and imaging data. This data volume is enormous and consumes a lot of computing resources. Existing fusion methods often fail to meet real-time or high-efficiency requirements when processing such high-dimensional and complex data.

[0009] In summary, while existing medical multimodal data fusion technologies have made some progress in single-modal analysis and simple modality fusion, they still face numerous challenges in handling data heterogeneity, capturing potential relationships between modalities, data scarcity, and computational complexity. Therefore, there is an urgent need for new methods that can efficiently and accurately fuse multimodal data to meet the growing demands of modern medical practice. Summary of the Invention

[0010] In order to solve the above-mentioned problems, the present invention provides a medical multimodal data fusion method and system, including solving the problems of data heterogeneity, insufficient modal correlation mining, inconsistent data quality, and high data processing complexity in the medical multimodal data fusion process.

[0011] The technical solution of the present invention is: a medical multimodal data fusion method, the method comprising the following steps:

[0012] Step 1: Preprocess the multimodal data, convert the multimodal data into high-dimensional feature vectors, and obtain the features of each modality;

[0013] Step 2: Calculate and align the modal correlation of the preprocessed multimodal data to fuse the features of each modality into a unified representation to obtain the multimodal fusion feature;

[0014] Step 3: Fine-tune the pre-trained large language model through Q-Lora, and train the fine-tuned pre-trained large language model with the multimodal fusion features processed by the multi-head self-attention mechanism, and use the trained pre-trained large language model to predict the input multimodal data.

[0015] Furthermore, the step 1 includes: preprocessing each data type in the multimodal data, the multimodal data including text, image, table, and time series data;

[0016] Preprocessing text data: This process strictly adheres to the standard workflow of current mainstream large-scale language models (LLMs). This includes text normalization, word segmentation and tokenization, and vectorized representation. This preprocessing not only ensures high-quality representation but also effectively improves the overall performance and applicability of multimodal data fusion methods.

[0017] Preprocess the image data: extract features using the pre-trained ViT image model and convert CT and MRI image data into feature vectors;

[0018] Preprocessing of tabular data: For structured tabular data, Min-Max normalization is used to process the data of each field to eliminate the dimensional differences between different features, ensure a more balanced weight distribution of each feature in the model, and improve the data fusion effect; Min-Max normalization is expressed as:

[0019]

[0020] Among them, X i is the original eigenvalue, X max and X min are the minimum and maximum values ​​of the feature, X' i is the standardized value;

[0021] Preprocessing of time series data: For time series data with different sampling frequencies (such as blood pressure monitoring, electrocardiogram, etc.), use Z-Score normalization for processing. Specifically, the Z-Score normalization formula is as follows:

[0022]

[0023] Where X is the original data, μ is the mean of the time series data, σ is the standard deviation, and Z is the standardized value. Z represents the degree of deviation of the original data point from the mean μ, with the unit of standard deviation σ.

[0024] By preprocessing each data type as described above, the time series data of different sampling frequencies are normalized to the same numerical range to facilitate subsequent alignment and fusion. After preprocessing, the data of all modalities are converted into high-dimensional feature vectors.

[0025] Furthermore, to ensure effective alignment of data between different modalities, this solution introduces a cross-attention mechanism. Through this mechanism, the model can focus on the relevant features of one modality in another modality, thereby effectively capturing the potential dependencies between different modalities. Step 2 includes:

[0026] Step 2.1: Calculate the modal relevance of the preprocessed multimodal data through the cross-attention mechanism. Specifically, for any two modal data A and B, calculate the attention score of A to B through the cross-attention mechanism. The formula is as follows:

[0027]

[0028] Among them, Q A , K B , V B are the query, key, and value matrices of modal A and B, respectively, d k is the dimension scaling factor;

[0029] The cross-attention mechanism is used to transfer information between different modalities, eliminating the heterogeneity problem between modalities.

[0030] Step 2.2: Align modal data of different time scales and formats through attention scores. After completing the modal data alignment, the features of each modality are fused into a unified representation to obtain multimodal fusion features to ensure the uniformity of subsequent processing.

[0031] Furthermore, after the modal pair correlation calculation is completed, the features of each modality are fused into a unified representation; on this basis, the pre-trained large language model is efficiently fine-tuned using Q-Lora quantization technology; step 3 includes:

[0032] Step 3.1: Quantize the pre-trained large language model (such as Transformer or GPT) through Q-Lora to reduce the training overhead and memory usage of model parameters. Q-Lora uses a low-bitrate weight representation method to freeze some parameters of the pre-trained large language model and only fine-tune the quantization layer, making the fine-tuning process more efficient and reducing resource consumption.

[0033] Step 3.2: Apply a multi-head self-attention mechanism to the multimodal fusion features obtained in step 2, so that the quantized pre-trained large language model in subsequent processing can adaptively weigh the correlation between different modalities;

[0034] Each attention head calculates the attention score independently, and finally the results are spliced ​​and projected into the output space. The calculation formula is as follows:

[0035] MultiHead(Q,K,V)=Concat(head1,head2,...,head n )W o

[0036] Among them, Q, K, V represent the query, key and value of the multi-head self-attention mechanism, head n represents the nth attention head in the multi-head self-attention mechanism, W 0 It is a weight matrix used to concatenate the outputs of all attention heads and perform a linear transformation. MultiHead() represents the Multi-Head Self-Attention Mechanism, a key component in the Transformer architecture. The Multi-Head Self-Attention Mechanism captures diverse relationships between different positions in the input sequence by applying multiple attention heads in parallel, thereby enhancing the model's expressiveness and ability to capture complex patterns.

[0037] Step 3.3: Input the multimodal fusion features output in step 2 into the fine-tuned pre-trained large language model for training. In the encoder of the pre-trained model, the text data is directly processed together with the data of other modalities to ensure that the representation of each modal feature in the language model remains consistent.

[0038] Step 3.4: Use the pre-trained large language model trained in step 3.3 to predict the input multimodal data. The trained pre-trained large language model can make accurate diagnoses, predictions, or recommend personalized treatment plans based on the fused global representation.

[0039] The present invention also provides a medical multimodal data fusion system, comprising: a module for executing the above-mentioned medical multimodal data fusion method.

[0040] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned medical multimodal data fusion method when executing the program.

[0041] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned medical multimodal data fusion method when executed by a processor.

[0042] The present invention also provides a computer program product, comprising a computer program, which implements the above-mentioned medical multimodal data fusion method when executed by a processor.

[0043] The medical multimodal data fusion method of the present invention can be applied to multiple practical scenarios:

[0044] 1. Personalized disease auxiliary diagnosis: By integrating patients' imaging data, genetic data and clinical diagnostic data, it can provide accurate personalized auxiliary diagnosis support for complex diseases (such as cancer and rare diseases).

[0045] 2. Treatment plan recommendation: The integration of different modal data (such as treatment history, imaging results, and genetic information) can provide doctors with more personalized and effective treatment plan recommendations, especially in tumor treatment and precision medicine.

[0046] 3. Health monitoring and early warning: Utilizing the patient's real-time biological signals, time series data, and historical health data, the present invention can assist in the long-term monitoring of patients with chronic diseases, provide real-time early warnings of potential risks, and assist in providing personalized health management services.

[0047] 4. Telemedicine and mobile health assistance: This invention can be combined with wearable devices to realize assisted remote health monitoring. By fusing multimodal data, it provides doctors with real-time feedback on the patient's health status, and assists in supporting telemedicine and mobile health applications.

[0048] The beneficial effects of the present invention are:

[0049] (1) Efficiency: Q-Lora quantization technology greatly reduces the training cost and computing resources of the model, and is suitable for real-time analysis of large-scale medical data.

[0050] (2) Accuracy: The cross-attention mechanism effectively captures the dependencies between different modalities and improves the accuracy of multimodal data fusion.

[0051] (3) Unified representation: By mapping the data of each modality into a unified representation, different types of data can be seamlessly integrated, providing high-quality input for subsequent tasks.

[0052] (4) Z-Score standardization: It effectively solves the problem of different sampling frequencies of time series data and ensures the consistency of data fusion.

[0053] The present invention greatly improves the fusion accuracy and computational efficiency of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0055] Example 1: Figure 1As shown, a medical multimodal data fusion method is provided to solve the problems of data heterogeneity, insufficient modality correlation mining, inconsistent data quality, and high data processing complexity in the medical multimodal data fusion process:

[0056] (1) Data heterogeneity: Data sources of different modalities (such as images, time series, and text) vary significantly in terms of acquisition methods and structures, making it difficult to directly process them uniformly.

[0057] (2) Insufficient mining of modal correlation: There are often complex nonlinear dependencies between data of different modalities. Existing technologies cannot effectively capture and utilize these potential cross-modal correlations, which affects the accuracy of diagnosis and prediction.

[0058] (3) Inconsistent data quality: Since the acquisition process of each modal data is affected by various external factors, there are noise, outliers and inconsistent sampling frequencies in the data, which affects the accuracy of data fusion.

[0059] (4) High data processing complexity: Faced with multi-dimensional, high-frequency data streams, such as imaging and genomic data, existing fusion technologies consume large amounts of computing resources when processing large-scale complex data, making it difficult to meet the needs of real-time medical scenarios.

[0060] A medical multimodal data fusion method comprises the following steps:

[0061] Step 1: Preprocess the multimodal data, convert the multimodal data into high-dimensional feature vectors, and obtain the features of each modality;

[0062] Furthermore, the step 1 includes: preprocessing each data type in the multimodal data, the multimodal data including text, image, table, and time series data;

[0063] Preprocessing of text data: This preprocessing strictly follows the standard process of current mainstream large-scale language models (LLMs). The preprocessing of text data includes text normalization, word segmentation and tokenization, and vectorized representation. This preprocessing of text data not only ensures high-quality representation of text data, but also effectively improves the overall performance and applicability of multimodal data fusion methods. Preprocessing of image data: Features are extracted through the pre-trained ViT image model, and CT and MRI image data are converted into feature vectors.

[0064] Preprocessing of tabular data: For structured tabular data, Min-Max normalization is used to process the data of each field to eliminate the dimensional differences between different features, ensure a more balanced weight distribution of each feature in the model, and improve the data fusion effect; Min-Max normalization is expressed as:

[0065]

[0066] Among them, X i is the original eigenvalue, X max and X min are the minimum and maximum values ​​of the feature, X' i is the standardized value;

[0067] Preprocessing of time series data: For time series data with different sampling frequencies (such as blood pressure monitoring, electrocardiogram, etc.), use Z-Score normalization for processing. Specifically, the Z-Score normalization formula is as follows:

[0068]

[0069] Where X is the original data, μ is the mean of the time series data, σ is the standard deviation, and Z is the standardized value. Z represents the degree of deviation of the original data point from the mean μ, with the unit of standard deviation σ.

[0070] By preprocessing each data type as described above, the time series data of different sampling frequencies are normalized to the same numerical range to facilitate subsequent alignment and fusion. After preprocessing, the data of all modalities are converted into high-dimensional feature vectors.

[0071] Step 2: Calculate and align the modal correlation of the preprocessed multimodal data to fuse the features of each modality into a unified representation to obtain the multimodal fusion feature;

[0072] Furthermore, to ensure effective alignment of data between different modalities, this solution introduces a cross-attention mechanism. Through this mechanism, the model can focus on the relevant features of one modality in another modality, thereby effectively capturing the potential dependencies between different modalities. Step 2 includes:

[0073] Step 2.1: Calculate the modal relevance of the preprocessed multimodal data through the cross-attention mechanism. Specifically, for any two modal data A and B, calculate the attention score of A to B through the cross-attention mechanism. The formula is as follows:

[0074]

[0075] Among them, Q A , K B , V B are the query, key, and value matrices of modal A and B, respectively, d k is the dimension scaling factor;

[0076] The cross-attention mechanism is used to transfer information between different modalities, eliminating the heterogeneity problem between modalities.

[0077] Step 2.2: Align modal data of different time scales and formats through attention scores. After completing the modal data alignment, the features of each modality are fused into a unified representation to obtain multimodal fusion features to ensure the uniformity of subsequent processing.

[0078] Step 3: Fine-tune the pre-trained large language model through Q-Lora, and train the fine-tuned pre-trained large language model with the multimodal fusion features processed by the multi-head self-attention mechanism, and use the trained pre-trained large language model to predict the input multimodal data.

[0079] Furthermore, after the modal pair correlation calculation is completed, the features of each modality are fused into a unified representation; on this basis, the pre-trained large language model is efficiently fine-tuned using Q-Lora quantization technology; step 3 includes:

[0080] Step 3.1: Quantize the pre-trained large language model (such as Transformer or GPT) through Q-Lora to reduce the training overhead and memory usage of model parameters. Q-Lora uses a low-bitrate weight representation method to freeze some parameters of the pre-trained large language model and only fine-tune the quantization layer, making the fine-tuning process more efficient and reducing resource consumption.

[0081] Step 3.2: Apply a multi-head self-attention mechanism to the multimodal fusion features obtained in step 2, so that the quantized pre-trained large language model in subsequent processing can adaptively weigh the correlation between different modalities;

[0082] Each attention head calculates the attention score independently, and finally the results are spliced ​​and projected into the output space. The calculation formula is as follows:

[0083] MultiHead(Q,K,V)=Concat(head1,head2,...,head n )W o

[0084] Among them, Q, K, V represent the query, key and value of the multi-head self-attention mechanism, head n represents the nth attention head in the multi-head self-attention mechanism, W 0It is a weight matrix used to concatenate the outputs of all attention heads and perform a linear transformation. MultiHead() represents the Multi-Head Self-Attention Mechanism, a key component in the Transformer architecture. The Multi-Head Self-Attention Mechanism captures diverse relationships between different positions in the input sequence by applying multiple attention heads in parallel, thereby enhancing the model's expressiveness and ability to capture complex patterns.

[0085] Step 3.3: Input the multimodal fusion features output in step 2 into the fine-tuned pre-trained large language model for training. In the encoder of the pre-trained model, the text data is directly processed together with the data of other modalities to ensure that the representation of each modal feature in the language model remains consistent.

[0086] Step 3.4: Use the pre-trained large language model trained in step 3.3 to predict the input multimodal data. The trained pre-trained large language model can make accurate diagnoses, predictions, or recommend personalized treatment plans based on the fused global representation.

[0087] The present invention also provides a medical multimodal data fusion system, comprising:

[0088] The high-dimensional feature vector conversion module is used to preprocess the multimodal data, convert the multimodal data into high-dimensional feature vectors, and obtain the features of each modality;

[0089] The multimodal fusion module is used to calculate and align the modal correlation of the preprocessed multimodal data, so that the features of each modality are integrated into a unified representation to obtain the multimodal fusion feature;

[0090] The prediction module is used to fine-tune the pre-trained large language model through Q-Lora, train the fine-tuned pre-trained large language model with the multimodal fusion features processed by the multi-head self-attention mechanism, and use the trained pre-trained large language model to predict the input multimodal data.

[0091] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned medical multimodal data fusion method when executing the program.

[0092] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned medical multimodal data fusion method when executed by a processor.

[0093] The present invention also provides a computer program product, comprising a computer program, which implements the above-mentioned medical multimodal data fusion method when executed by a processor.

[0094] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.

Claims

1. A medical multimodal data fusion method, characterized by: The method comprises the following steps: Step 1: Preprocess the multimodal data, convert the multimodal data into high-dimensional feature vectors, and obtain the features of each modality; Step 2: Calculate and align the modal correlation of the preprocessed multimodal data to fuse the features of each modality into a unified representation to obtain the multimodal fusion feature; Step 3: Fine-tune the pre-trained large language model through Q-Lora, and train the fine-tuned pre-trained large language model with the multimodal fusion features processed by the multi-head self-attention mechanism, and use the trained pre-trained large language model to predict the input multimodal data.

2. The medical multimodal data fusion method according to claim 1, characterized in that: The step 1 includes: preprocessing each data type in the multimodal data, where the multimodal data includes text, images, tables, and time series data; Preprocessing of text data: Preprocessing of text data includes text normalization, word segmentation and tokenization, and vectorized representation; Preprocess the image data: extract features through a pre-trained image model and convert the image data into feature vectors; Preprocessing of tabular data: For structured tabular data, Min-Max normalization is used to process the data of each field to eliminate the dimensional differences between different features. Min-Max normalization is expressed as: Among them, X i is the original eigenvalue, X max and X min are the minimum and maximum values ​​of the feature, X' i is the standardized value; Preprocessing of time series data: For time series data with different sampling frequencies, Z-Score normalization is used for processing. Specifically, the Z-Score normalization formula is as follows: Where X is the original data, μ is the mean of the time series data, σ is the standard deviation, and Z is the standardized value. Z represents the degree of deviation of the original data point from the mean μ, with the unit of standard deviation σ. By preprocessing each data type as described above, the time series data of different sampling frequencies are normalized to the same value range.

3. The medical multimodal data fusion method according to claim 1, characterized in that: The step 2 includes: Step 2.1: Calculate the modal relevance of the preprocessed multimodal data using the cross-attention mechanism. Specifically, for any two modal data A and B, calculate the attention score of A to B using the cross-attention mechanism. The formula is as follows: Among them, Q A , K B , V B are the query, key, and value matrices of modal A and B, respectively, d k is the dimension scaling factor; The cross-attention mechanism is used to transfer information between different modalities, eliminating the heterogeneity problem between modalities. Step 2.2: Align modal data of different time scales and formats through attention scores. After completing the modal data alignment, the features of each modality are fused into a unified representation to obtain multimodal fusion features.

4. The medical multimodal data fusion method according to claim 1, characterized in that: The step 3 includes: Step 3.1: Quantize the pre-trained large language model using Q-Lora. Q-Lora uses a low-bitrate weight representation method to freeze some parameters of the pre-trained large language model and only fine-tune the quantization layer. Step 3.2: Apply a multi-head self-attention mechanism to the multimodal fusion features obtained in step 2, so that the quantized pre-trained large language model in subsequent processing can adaptively weigh the correlation between different modalities; Each attention head calculates the attention score independently, and finally the results are spliced ​​and projected into the output space. The calculation formula is as follows: MultiHead(Q,K,V)=Concat(head1,head2,...,head n )W o Among them, Q, K, V represent the query, key and value of the multi-head self-attention mechanism, head n represents the nth attention head in the multi-head self-attention mechanism, W 0 It is the weight matrix used to concatenate the outputs of all attention heads and perform linear transformation. MultiHead() represents the multi-head self-attention mechanism. Step 3.3: Input the multimodal fusion features output in step 2 into the fine-tuned pre-trained large language model for training; Step 3.4: Use the pre-trained large language model trained in step 3.3 to predict the input multimodal data.

5. A medical multimodal data fusion system, characterized in that: include: A module for executing a medical multimodal data fusion method as claimed in any one of claims 1 to 4.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, a medical multimodal data fusion method as described in any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the medical multimodal data fusion method according to any one of claims 1 to 4 is implemented.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the medical multimodal data fusion method according to any one of claims 1 to 4 is implemented.

Citation Information

Cited By

  • Multi-modal data fusion method, system and device for modal missing, processor and computer readable storage medium thereof

    CN121280839A

  • Methods, systems, devices, processors, and computer-readable storage media for multimodal data fusion addressing modal gaps

    CN121280839B