Data fusion method for multi-modal data

Through the multimodal data fusion method, using technologies such as EfficientNet, BioBERT, random forest, GBM, LSTM and large language model, the characteristics of multiple data types are extracted and fused, solving the problem of incomplete diagnosis of a single data source and improving the accuracy of diagnosis of infectious pneumonia.

CN120387129APending Publication Date: 2025-07-29MACAU UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342999.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the clinical diagnosis of infectious pneumonia relies on a single data source analysis, cannot fully reflect the patient's condition, and lacks a complete multimodal data fusion scheme.

Method used

Multimodal data fusion method is adopted, including feature extraction and fusion of image data, text data, structured data, multi-time point data, time series data and text information and experimental result data, and multi-modal fusion is used to feature extraction and multi-head attention mechanisms.

Benefits of technology

Through multimodal data fusion, the accuracy of diagnosis is improved and more comprehensive and reliable diagnostic support is provided to medical staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387129A_ABST
    Figure CN120387129A_ABST
Patent Text Reader

Abstract

The invention relates to a data fusion method for multi-modal data, and the method comprises the steps: collecting the multi-modal data, and carrying out the data feature extraction of image data, and obtaining a first data feature; performing data feature extraction on the text data to obtain a second data feature; performing data feature extraction on the structured data to obtain third data features; performing data feature extraction on the multi-time-point data to obtain fourth data features; performing data feature extraction on the time sequence data to obtain a fifth data feature; performing data feature extraction on the text information and the experimental result data to obtain a sixth data feature; and performing multi-modal fusion on the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature and the sixth data feature to obtain fused data. According to the invention, through multi-modal data fusion, different types of data are comprehensively considered, so that the diagnosis accuracy is improved, and great help is brought to the diagnosis of medical personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data fusion method for multi-modal data. Background Art

[0002] The clinical diagnosis of infectious pneumonia relies on data from multiple aspects, such as imaging data (e.g., CT, X-ray), biomarkers, laboratory test results, patient medical history, etc. Currently, data analysis often separately analyzes data of a single type and then aggregates the analysis results.

[0003] This approach may not comprehensively reflect the patient's condition due to considering a single data source. Through multi-modal data fusion, different types of data can be comprehensively considered, thereby improving the accuracy of diagnosis and bringing great help to the diagnosis of medical staff.

[0004] Currently, there is no perfect multi-modal data fusion solution in the market that can meet the industry business requirements. Summary of the Invention

[0005] The purpose of the present invention is to at least solve one of the deficiencies of the prior art and provide a data fusion method for multi-modal data.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] Specifically, a data fusion method for multi-modal data is proposed, including the following:

[0008] Obtain multi-modal data, where the multi-modal data includes image data, text data, structured data, multi-time point data, time series data, text information, and experimental result data;

[0009] Extract data features from the image data to obtain first data features;

[0010] Extract data features from the text data to obtain second data features;

[0011] Extract data features from the structured data to obtain third data features;

[0012] Extract data features from the multi-time point data to obtain fourth data features;

[0013] Extract data features from the time series data to obtain fifth data features;

[0014] Extract data features from the text information and experimental result data to obtain sixth data features;

[0015] Perform multimodal fusion on the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature to obtain the fused data.

[0016] Further, specifically, extract data features from the image data to obtain the first data feature, including,

[0017] Normalize the image data to adjust the pixel values to the standard range, as shown in the following formula,

[0018]

[0019] where I is the original image data, I min and I max are the minimum and maximum values of the original image data respectively;

[0020] Perform random angle adjustment on I′, that is, I rot = rotate(I norm , θ),

[0021] Perform horizontal flipping on I rot , that is, process I rot through the flip() function to obtain I rot ′,

[0022] Perform random cropping on I rot ′ to obtain I crop = crop(I rot ′, crop_size), where crop_size is the preset cropping size,

[0023] Based on the EfficientNet model pre-trained by transfer learning, perform feature extraction on I crop to obtain the high-dimensional image feature F img ,

[0024] F img = F EffNet (I crop ),

[0025] Perform principal component analysis (PCA) on F img to obtain F img-reduced ,

[0026] F img-reduced = PCA(F img ),

[0027] Reduce the dimension of F img-reduced by finding the eigenvectors of the feature covariance matrix,

[0028] F img-reduced ′ = WT F img-reduced ,

[0029] Among them, W T is the feature vector matrix, then F img-reduced ′ is the first data feature.

[0030] Furthermore, specifically, for text data, data feature extraction is performed to obtain the second data feature, including,

[0031] Preprocess the text data to remove stop words and punctuation marks to obtain the preprocessed text data Tokens;

[0032] Then use the BioBERT model to perform deep semantic encoding on the preprocessed text data Tokens to extract the deep semantic feature F txt ,

[0033] F txt = F BioBERT (Tokens),

[0034] Perform average pooling operation on F txt to obtain the second data feature F txt-fixed ,

[0035]

[0036] Furthermore, specifically, for structured data, data feature extraction is performed to obtain the third data feature, including,

[0037] Standardize the structured data to obtain X std ,

[0038] X std =(X - μ X ) / σ X ,

[0039] Among them, X represents the original structured data, μ X represents the mean of the original structured data, and σ X represents the standard deviation of the original structured data;

[0040] Extract features from X std through machine learning to obtain the feature F ml ,

[0041] F ml = MLFeatureExtraction(X std ),

[0042] Then input the feature F ml into the random forest model to extract the third data feature Fstr-selected ,

[0043]

[0044] Furthermore, specifically, data feature extraction is performed on multi-time point data to obtain a fourth data feature, including

[0045] The fourth data feature F GBM (X) is obtained by integrating multi-time point data through a GBM model trained based on minimizing a loss function. The GBM model is as follows

[0046]

[0047] where h m (X) is the m-th decision tree, and α m is the weight, and X represents the input multi-time point data

[0048] Furthermore, specifically, data feature extraction is performed on time series data to obtain a fifth data feature

[0049] Convert the structured data statistically calculated daily into a time series form

[0050] X time-series = DailyStatistics(X)

[0051] Calculate the residual signal of X time-series

[0052] r1(t) = f(t) - 〈f, g0〉g0(t)

[0053] Repeat the above formula to calculate a series of basis functions, and calculate the projection and residual for each until the preset condition is met

[0054] r k+1 (t) = r k (t) - 〈r k , g k 〉g k (t)

[0055] Finally, sum up the projection results of all basis functions to obtain the AFD reconstruction result of X time-series :

[0056]

[0057] Input the AFD reconstruction result of X time-series into the LSTM model, and output the fifth data feature F afd ,

[0058] F afd ​=AFD(X time-series )、F afd-lstm =LSTM(F afd ),

[0059] where F afd-lstm is the AFD reconstruction result of X time-series .

[0060] Furthermore, specifically, data feature extraction is performed on the text information and experimental result data to obtain the sixth data feature, including,

[0061] The first part is to extract the feature F LLM-text from the case text information based on the LLM model,

[0062] F LLM-text =LLM(T medical , K external ),

[0063] where T medical represents the medical record text data, and K external is the relevant information retrieved from the external medical knowledge base and the medical treatment experiences of multiple pre-selected famous doctors;

[0064] The second part is to extract the feature F LLM-lab from the experimental result data based on the LLM model,

[0065] F LLM-lab =LLM(F CRP , F PCT K external ),

[0066] where F CRP and F PCT represent the data features of CRP and PCT, and K external is the relevant information retrieved from the external medical knowledge base and the medical treatment experiences of multiple pre-selected famous doctors

[0067] Adding F LLM-text and F LLM-lab gives the sixth data feature F LLM ,

[0068] F LLM =F LLM-text +F LLM-lab .

[0069] Furthermore, specifically, multi-modal fusion is performed on the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature to obtain the fusion data, including,

[0070] F is obtained by fusing the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature into a feature matrix concat ,

[0071] F concat = [F img-reduced ', F txt-fixed , F str-selected , F GBM (X), F afd-lstm , F LLM ;

[0072] Multimodal fusion is performed based on the multi-head attention mechanism to obtain fused data. Specifically, F concat is superimposed with the self-attention output through residual connection,

[0073] H res = F concat + Z,

[0074] Then, the feature H res after residual connection is normalized through layer normalization to obtain the fused data H out ,

[0075] H out = LayerNorm(H res ),

[0076] where Z = AV, A is the attention weight, and V is the weighted value vector,

[0077]

[0078] Q = W Q F concat , K = W K F concat , V = W V F concat , W Q , W K , W V are preset weights respectively, k takes Q, K, V, and d k represents the dimension of each head respectively,

[0079] Q, K, V satisfy,

[0080] MultiHead(Q, K, V) = [Z1, Z2,..., Z h W O ,

[0081] Z i represents the output of the i-th attention head, i ∈ [1, h], and h is the number of attention heads, W Ois a linear transformation matrix.

[0082] The beneficial effects of the present invention are as follows:

[0083] The present invention provides a data fusion method for multi-modal data. By collecting multi-modal data, which includes image data, text data, structured data, multi-time point data, time series data, text information, and experimental result data, and then respectively extracting data features from the image data to obtain the first data feature; extracting data features from the text data to obtain the second data feature; extracting data features from the structured data to obtain the third data feature; extracting data features from the multi-time point data to obtain the fourth data feature; extracting data features from the time series data to obtain the fifth data feature; extracting data features from the text information and experimental result data to obtain the sixth data feature; and performing multi-modal fusion on the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature to obtain the fusion data. The present invention can comprehensively consider different types of data through multi-modal data fusion, thereby improving the accuracy of diagnosis and bringing great help to the diagnosis of medical staff. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] By describing the embodiments shown in the accompanying drawings in detail, the above and other features of the present disclosure will become more apparent. The same reference numerals in the drawings of the present disclosure denote the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0085] Figure 1 The flowchart of the data fusion method for multi-modal data according to the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] The following will clearly and completely describe the concept, specific structure, and technical effects of the present invention in combination with the embodiments and the drawings to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The same reference numerals used throughout the drawings indicate the same or similar parts.

[0087] Example 1, referring to Figure 1 , the present invention provides a data fusion method for multi-modal data, including the following:

[0088] Step 110: Obtain multimodal data, where the multimodal data includes image data, text data, structured data, multi-time point data, time series data, text information, and experimental result data;

[0089] Step 120: Extract data features from the image data to obtain the first data feature;

[0090] Step 130: Extract data features from the text data to obtain the second data feature;

[0091] Step 140: Extract data features from the structured data to obtain the third data feature;

[0092] Step 150: Extract data features from the multi-time point data to obtain the fourth data feature;

[0093] Step 160: Extract data features from the time series data to obtain the fifth data feature;

[0094] Step 170: Extract data features from the text information and experimental result data to obtain the sixth data feature;

[0095] Step 180: Perform multimodal fusion on the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature to obtain fused data.

[0096] In this Embodiment 1, by collecting multimodal data, where the multimodal data includes image data, text data, structured data, multi-time point data, time series data, text information, and experimental result data, and then respectively extracting data features from the image data to obtain the first data feature; extracting data features from the text data to obtain the second data feature; extracting data features from the structured data to obtain the third data feature; extracting data features from the multi-time point data to obtain the fourth data feature; extracting data features from the time series data to obtain the fifth data feature; extracting data features from the text information and experimental result data to obtain the sixth data feature; performing multimodal fusion on the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature to obtain fused data. The present invention can, through multimodal data fusion, comprehensively consider different types of data, thereby improving the accuracy of diagnosis and bringing great help to the diagnosis of medical staff.

[0097] As a preferred embodiment of the present invention, specifically, extracting data features from the image data to obtain the first data feature includes,

[0098] Imaging data played a crucial role in this study, mainly including CT and X-ray film imaging data, aiming to extract key imaging features of lung infections, such as alveolar infiltration, consolidation, ground-glass opacity, and tree-in-bud sign, etc. First of all, the image processing steps ensured the consistency of the data. By normalizing the pixel values to the standard range, the differences between images were reduced: normalizing the said imaging data adjusted the pixel values to the standard range, as shown in the following formula,

[0099]

[0100] where I is the original imaging data, I min and I max are the minimum and maximum values of the original imaging data respectively; the purpose of this operation is to provide stable input data for subsequent analysis. In addition, with the help of data augmentation techniques, we diversified the images, such as rotation, flipping, and cropping. These methods effectively expanded the dataset and improved the performance of the model on unknown data. For example, the image can be randomly rotated by a certain angle:

[0101] Randomly adjusting the angle of I′ means I rot = rotate(I norm , θ),

[0102] Horizontally flipping I rot means processing I rot through the flip() function to obtain I rot ′,

[0103] Randomly cropping I rot ′ gives I crop = crop(I rot ′, crop_size), where crop_size is the preset cropping size,

[0104] To capture deeper features from the preprocessed images, we adopted a pre-trained EfficientNet model and performed feature extraction through transfer learning: based on the pre-trained EfficientNet model by means of transfer learning, feature extraction was performed on I crop to obtain high-dimensional imaging features F img ,

[0105] F img = F EffNet (I crop ),

[0106] Performing principal component analysis PCA on F img gives F img-reduced, with the aim of reducing the feature dimension while retaining the key image information, thereby reducing the computational cost and improving the training efficiency of the model.

[0107] F img-reduced = PCA(F img ),

[0108] By finding the eigenvectors of the feature covariance matrix to reduce the dimension of F img-reduced dimensionality reduction,

[0109] F img-reduced ' = W T F img-reduced ,

[0110] where W T is the eigenvector matrix of, then F img-reduced ' is the first data feature. By extracting the eigenvectors of the feature matrix, the high-dimensional data is compressed into a lower-dimensional representation form.

[0111] As a preferred embodiment of the present invention, specifically, data feature extraction is performed on the text data to obtain the second data feature, including,

[0112] The feature extraction of text data covers medical record records (such as patient symptoms, medical history, previous medications) and laboratory test results (such as blood routine, C-reactive protein (CRP), procalcitonin (PCT), etc.). In the preprocessing step, we first removed stop words and punctuation marks. Then, the BioBERT model was used to perform deep semantic encoding on the processed text: BioBERT has been pre-trained on a dedicated biomedical corpus and can accurately capture the complex semantic structure and potential semantic associations in medical texts.

[0113] Preprocess the text data to remove stop words and punctuation marks to obtain the preprocessed text data Tokens;

[0114] Then use the BioBERT model to perform deep semantic encoding on the preprocessed text data Tokens to extract the deep semantic feature F txt ,

[0115] F txt = F BioBERT (Tokens),

[0116] In order to further unify the dimension of the text features and facilitate the fusion with other modal data, we introduced an average pooling operation to convert the extracted semantic information into a fixed-length vector representation: perform an average pooling operation on F txt to obtain the second data feature F txt-fixed ,

[0117]

[0118] As a preferred embodiment of the present invention, specifically, for structured data, data feature extraction is performed to obtain the third data feature, including

[0119] When processing structured data and feature extraction, we mainly process immune status data (such as inflammatory factor levels, flow cytometry analysis results) and microbial detection results (including metagenomic sequencing (mNGS), targeted gene sequencing (tNGS), microbial culture, and multiplex PCR, etc.). To ensure that data with different dimensions can be compared on the same scale, we standardized these data: standardized the structured data to obtain X std ,

[0120] X std =(X - μ X ) / σ X ,

[0121] where X represents the original structured data, μ X represents the mean of the original structured data, and σ X represents the standard deviation of the original structured data;

[0122] For the standardized data, we adopted a variety of advanced machine learning methods for feature extraction. Through algorithms such as the random forest model and support vector machine, we can identify the most important features: perform feature extraction on X std through machine learning to obtain the feature F ml ,

[0123] F ml = MLFeatureExtraction(X std ),

[0124] Then input the feature F ml into the random forest model to extract the third data feature F str-selected ,

[0125]

[0126] where represents the random forest model constructed based on the tree parameter k, and its input is the feature F ml .

[0127] As a preferred embodiment of the present invention, specifically, for multi-time point data, data feature extraction is performed to obtain the fourth data feature, including

[0128] Integrate multi-time point data from outpatient clinics, emergency departments, and ICUs. After standardization or normalization, extract trend changes and time series features, such as heart rate changes, white blood cell counts, etc. GBM is an ensemble learning method that gradually reduces the prediction error by iteratively training weak classifiers (usually decision trees). In each round of training, the prediction ability of the model is optimized by fitting the residuals of the previous round of the model. Integrate multi-time point data from outpatient clinics, emergency departments, and ICUs. After standardization or normalization, extract trend changes and time series features, such as heart rate changes, white blood cell counts, etc. The GBM model is trained by minimizing the loss function (such as logarithmic loss or mean squared error). Assume the training set is X, the target variable is F, and the prediction of the model is: The fourth data feature F is obtained by integrating multi-time point data through the GBM model trained based on minimizing the loss function GBM (X), the GBM model is as follows,

[0129]

[0130] where h m (X) is the m-th decision tree, α m is the weight, and X represents the input multi-time point data.

[0131] As a preferred embodiment of the present invention, specifically, data feature extraction is performed on time series data to obtain the fifth data feature,

[0132] For the structured data (time series data) statistically calculated daily, we apply adaptive Fourier decomposition (AFD) for feature extraction. AFD is a powerful feature extraction tool suitable for capturing the periodic and aperiodic components of complex biological signals and has an automatic denoising function. We convert the structured data into a time series form, and the daily statistical results are used as the input of the time series: Convert the structured data statistically calculated daily into a time series form,

[0133] X time-series = DailyStatistics(X),

[0134] AFD decomposes the signal into multiple orthogonal basis functions recursively. These basis functions are adaptively selected according to the spectral characteristics of the signal to ensure maximum extraction of the key features in the signal. Among them, each basis function is adaptively selected according to the specific spectral characteristics of the signal. Denote the standardized time series data as the signal f(t). First, select the initial basis function g0(t) that best matches the signal features according to the maximum selection principle of energy tracking matching. Subsequently, calculate the projection of the signal on the initial basis function

[0135] 〈f, g0〉, where the inner product here is the inner product in the Euclidean space. Thus, calculate the residual signal:

[0136] Calculate X time-series of the residual signal,

[0137] r1(t) = f(t) - 〈f,g0〉g0(t),

[0138] Repeat the above calculation to obtain a series of basis functions, and calculate the projection and residual for each one until the preset condition is met (for example, the energy of the residual signal is lower than the threshold).

[0139] r k+1 (t) = r k (t) - 〈r k ,g k 〉g k (t),

[0140] Finally, sum up the projection results of all basis functions to obtain the AFD reconstruction result of X time-series :

[0141]

[0142] Input the AFD reconstruction result of X time-series into the LSTM model to output the fifth data feature F afd , in this way, AFD extracts the main spectral features of the signal, reflecting the hidden periodic and aperiodic components in the signal. We further input the features extracted by AFD into the long short-term memory network (LSTM) model for learning and predicting time series features. The LSTM model can effectively handle long-term dependencies and capture the dynamic change features in the time series, thereby improving the prediction ability for complex time series data, especially having significant advantages in biomedical signal analysis.

[0143] F afd = AFD(X time-series ), F afd-lstm = LSTM(F afd ),

[0144] where F afd-lstm is the AFD reconstruction result of X time-series .

[0145] As a preferred embodiment of the present invention, specifically, data feature extraction is performed on the text information and experimental result data to obtain the sixth data feature, including,

[0146] To deeply mine the key information in medical data, this study uses large language model technology and focuses on the feature extraction of two types of data: one is the medical record text information, and the other is the laboratory test results. The large language model has powerful semantic understanding and reasoning capabilities, can automatically identify and extract the key features in different data sources, and associate this information to provide high-dimensional and semantically rich feature representations for model training.

[0147] a. Feature Extraction and Association of Medical Record Text Information

[0148] For medical record texts (such as symptoms, medical history, treatment plans, etc.), the large language model can deeply understand the medical entities and their relationships in the text. For example, the model can identify the patient's symptoms, such as "fever, cough", and infer its possible causes and development trends. By combining external medical knowledge bases, the model can automatically supplement the implicit association information in the text and construct a feature representation with a deeper semantic level.

[0149] The information extracted by the model from the medical records includes:

[0150] Symptoms: such as "fever, shortness of breath"

[0151] Medical history: such as "chronic obstructive pulmonary disease (COPD)"

[0152] Medication: such as "azithromycin"

[0153] Combined with external retrieval, the model can associate this information. For example, it can infer the potential impact of the medical history on the medication, thereby generating comprehensive features.

[0154] The output feature representation is:

[0155] The first part, the feature extraction of the case text information based on the LLM model to obtain feature F LLM-text ,

[0156] F LLM-text = LLM(T medical , K external ),

[0157] where T medical represents the medical record text data, and K external is the relevant information retrieved from external medical knowledge bases and the medical treatment experiences of multiple pre-selected famous doctors;

[0158] b. Feature Extraction and Association of Laboratory Test Results

[0159] Laboratory test results (such as C-reactive protein (CRP), procalcitonin (PCT), etc.) are important bases for diagnosing and evaluating the condition. The large language model standardizes these data and associates them with the information in the medical record text to generate a composite feature representation.

[0160] For example, an increase in CRP and PCT may indicate inflammation or infection, and this data can be combined with relevant symptoms (such as fever) in the medical record to form a more diagnostically valuable feature vector. By combining laboratory data and medical record text, the large language model can capture the trends and potential associations of pathological changes, providing support for subsequent disease prediction.

[0161] The output feature representation form is:

[0162] In the second part, the feature F is extracted from the experimental result data based on the LLM model LLM-lab ,

[0163] F LLM-lab = LLM(F CRP , F PCT K external ),

[0164] where F CRP and F PCT represent the data features of CRP and PCT, and K external is the relevant information retrieved from external medical knowledge bases and the medical treatment experiences of multiple pre-selected famous doctors.

[0165] Adding F LLM-text and F LLM-lab to obtain the sixth data feature F LLM ,

[0166] F LLM = F LLM-text + F LLM-lab

[0167] As a preferred embodiment of the present invention, specifically, the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature are subjected to multimodal fusion to obtain fusion data, including

[0168] Multimodal Fusion is to integrate data from different modalities (such as images, texts, time series, structured data, time series data, etc.) to make full use of the complementarity of multi-source information and improve the overall performance of the task. By fusing data of multiple modalities, the system can capture more context and features, improve the accuracy and robustness of tasks such as classification and detection, and thus provide more reliable decision support in complex scenarios. To achieve effective fusion of multimodal data, we adopt the self-attention mechanism (Transformer). First, the image features, text features, and structured data features, as well as the associated features extracted by the prediction model and the large language model, are integrated into a feature matrix:

[0169] Fuse the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature into a feature matrix to obtain F concat ,

[0170] F concat =[F img-reduced ′,F txt-fixed ,F str-selected ,F GBM (X),F afd-lstm ,F LLM ;

[0171] Perform multi-modal fusion based on the multi-head attention mechanism to obtain the fused data. Specifically, stack F concat with the self-attention output through residual connection,

[0172] H res =F concat +Z,

[0173] Then, normalize the feature H res after the residual connection through layer normalization to obtain the fused data H out ,

[0174] H out =LayerNorm(H res ),

[0175] where Z = AV, A is the attention weight, and V is the weighted value vector.

[0176]

[0177] Q = W Q F concat 、K = W K F concat 、V = W V F concat , W Q 、W K 、W V are preset weights respectively, k takes Q, K, V, and d k represents the dimension of each head respectively. The query vector is used to represent how the features of the current modality are related to other modalities, the key vector is used to capture the feature representations of other modalities, and the value vector represents the weights and information of the features themselves. Calculate the attention weight: The dot product operation between the query vector and the key vector reflects the correlation between the features of different modalities, and the dot product result will be normalized through the Softmax function;

[0178] Q, K, V satisfy,

[0179] MultiHead(Q,K,V) = [Z1,Z2,…,Z h W O ,

[0180] Z i represents the output of the i-th attention head, where i ∈ [1, h] and h is the number of attention heads, and W O is a linear transformation matrix.

[0181] This can effectively prevent the problem of gradient disappearance and accelerate model training. The decoder and the decision layer: After multi-modal fusion, the generated feature matrix will be input into the fully connected layer to output the task prediction result.

[0182] Although the description of the present invention has been quite detailed and particularly describes several of the described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but rather should be regarded as providing a broad interpretation of these claims in light of the prior art by reference to the appended claims, so as to effectively cover the intended scope of the present invention. In addition, the present invention is described above in terms of embodiments foreseeable by the inventor for the purpose of providing a useful description, and non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.

[0183] The above is only the preferred embodiment of the present invention. The present invention is not limited to the above-mentioned embodiments. As long as it achieves the technical effects of the present invention by the same means, it should fall within the protection scope of the present invention. Within the protection scope of the present invention, its technical solutions and / or embodiments may have various different modifications and changes.

Claims

1. A data fusion method for multimodal data, characterized in that Including the following: Obtain multimodal data, where the multimodal data includes image data, text data, structured data, multi-time point data, time series data, text information, and experimental result data; Extract data features from the image data to obtain first data features; Extract data features from the text data to obtain second data features; Extract data features from the structured data to obtain third data features; Extract data features from the multi-time point data to obtain fourth data features; Extract data features from the time series data to obtain fifth data features; Extract data features from the text information and experimental result data to obtain sixth data features; Perform multimodal fusion on the first data features, second data features, third data features, fourth data features, fifth data features, and sixth data features to obtain fused data.

2. The data fusion method for multimodal data according to claim 1, wherein Specifically, extracting data features from the image data to obtain first data features includes, Normalize the image data to adjust the pixel values to the standard range, as shown in the following formula, Where I is the original image data, I min and I max are the minimum value and the maximum value of the original image data, respectively; Randomly adjust the angle of I′ to obtain I rot = rotate(I norm , θ), For I rot perform a horizontal flip, that is, process I through the flip() function rot to obtain I rot ′ For I rot ′, perform random cropping to obtain I crop = crop(I rot ′, crop_size), where crop_size is the preset cropping size. Based on the EfficientNet model pre-trained by means of transfer learning, perform feature extraction on I crop to obtain high-dimensional image features F img , F img = F EffNet (I crop ) For F img Perform principal component analysis (PCA) processing on it to obtain F img-reduced , F img-reduced = PCA(F img ) Reduce the dimension of F by finding the eigenvectors of the characteristic covariance matrix img-reduced dimensionality reduction F img-reduced ′ = W T F img-reduced , Among them, W T is the eigenvector matrix, then F img-reduced ′ is the first data feature.

3. The data fusion method for multimodal data according to claim 2, wherein Specifically, extracting data features from the text data to obtain second data features includes, Preprocess the text data to remove stop words and punctuation marks to obtain preprocessed text data Tokens; Next, the BioBERT model is used to perform deep semantic encoding on the preprocessed text data Tokens to extract deep semantic features F txt , F txt = F BioBERT (Tokens) Average pool operation is performed on F txt to obtain the second data feature F txt-fixed , 4. The data fusion method for multimodal data according to claim 3, wherein Specifically, extracting data features from the structured data to obtain third data features includes, Standardize the structured data to obtain X std , X std = (X - μ X ) / σ X , Among them, X represents the original structured data, μ X represents the mean of the original structured data, and σ X represents the standard deviation of the original structured data; Feature extraction of X through machine learning std to obtain feature F ml , F ml = MLFeatureExtraction(X std ) Then input the feature F ml into the random forest model to extract the third data feature F str-selected , 5. The data fusion method for multimodal data according to claim 4, characterized in that, Specifically, extracting data features from the multi-time point data to obtain fourth data features includes, The fourth data feature F is obtained by integrating multi-time point data through a GBM model trained based on minimizing a loss function. GBM (X), and the GBM model is as follows. where h m (X) is the m-th decision tree, and α m is the weight, and X represents the input multi-timepoint data.

6. The data fusion method for multimodal data according to claim 5, wherein Specifically, extracting data features from the time series data to obtain fifth data features, Convert the structured data statistically calculated daily into a time series form, X time-series = DailyStatistics(X), Calculate X time-series of the residual signal, r1(t) = f(t) - 〈f,g0〉g0(t), Repeat the above formula calculation to obtain a series of basis functions, and calculate the projection and residual for each until the preset condition is met, r k+1 r(t) = r k r(t) - 〈r k , g k 〉g k r(t), Finally, the projection results of all basis functions are summed up as a series to obtain the AFD reconstruction result of X time-series : Input X time-series The AFD reconstruction result into the LSTM model, and output the fifth data feature F afd , F afd = AFD(X time-series ), F afd-lstm = LSTM(F afd ), Among which F afd-lstm is the AFD reconstruction result of X time-series .

7. The data fusion method for multimodal data according to claim 6, characterized in that Specifically, extracting data features from the text information and experimental result data to obtain sixth data features includes, The first part, feature extraction of case text information based on the LLM model to obtain feature F LLM-text , F LLM-text = LLM(T medical , K external ) Among them, T medical represents the medical record text data, and K external is the relevant information retrieved from the external medical knowledge base and the medical treatment experiences of multiple pre-selected famous doctors; Part Two: Feature F is obtained by extracting features from the experimental result data based on the LLM model LLM-lab , F LLM-lab = LLM(F CRP , F PCT K external ), Among them, F CRP and F PCT represent the data characteristics of CRP and PCT, and K external is the relevant information retrieved from external medical knowledge bases and the medical treatment experiences of multiple pre-selected famous doctors Add F LLM-text to F LLM-lab to obtain the sixth data feature F LLM , F LLM = F LLM-text + F LLM-lab 。 8. The data fusion method for multimodal data according to claim 7, wherein Specifically, performing multimodal fusion on the first data features, second data features, third data features, fourth data features, fifth data features, and sixth data features to obtain fused data includes, Fuse the first data feature, the second data feature, the third data feature, the fourth data feature, the fifth data feature, and the sixth data feature into a feature matrix to obtain F concat , F concat = [F img-reduced ′, F txt-fixed , F str-selected , F GBM (X), F afd-lstm , F LLM ; Multi-modal fusion is performed based on the multi-head attention mechanism to obtain fused data. Specifically, F is superimposed with the self-attention output through residual connection. concat ​ H res = F concat + Z, Then, the feature H after residual connection is normalized through layer normalization res to obtain the fused data H out , H out = LayerNorm(H res ) Where Z = AV, A is the attention weight, and V is the weighted value vector, Q = W Q F concat 、K = W K F concat 、V = W V F concat ,W Q 、W K 、W V are preset weights respectively, k takes Q, K, V, d k represents the dimension of each head respectively Q, K, V satisfy, MultiHead(Q,K,V)=[Z1,Z2,…,Z h W O , Z i represents the output of the i-th attention head, where i ∈ [1, h] and h is the number of attention heads, and W O is a linear transformation matrix.