A Multimodal Data Fusion Method, Device, Computer Device, and Storage Medium

Through special modal encoders and multi-head attention mechanisms and other technical means, advanced feature extraction and decoupling of multi-modal data is carried out, and feature fusion is combined with cosine alignment and orthogonal projection technology, which solves the problem of difficulty in utilizing heterogeneity and correlation in multi-modal data fusion, and efficient disease risk warning is achieved, and the accuracy and personalization of early warning are improved.

CN119691687BActive Publication Date: 2025-06-10SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510202089.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-10
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing multimodal data fusion methods are difficult to effectively process high heterogeneity-heterogeneous data, poor information extraction capabilities, unable to fully utilize the correlation between multimodal data, and lack medical interpretability, which affects the accuracy and reliability of cardiovascular event warnings.

Method used

Special modal encoder is used to perform high-order feature extraction, combine expert knowledge constraints and multi-head attention mechanisms to decouple features, use cosine alignment algorithms and orthogonal projection technology to process low-frequency and high-frequency features, and perform multi-level fusion through adaptive deep fusion networks. Finally, the fused features are input into the pre-trained risk prediction model to output the disease risk warning probability.

Benefits of technology

Effectively handle the heterogeneity and heterogeneity of multimodal data, improve signal-to-noise ratio, eliminate redundant information and noise, reduce the impact of high coupling between modes, realize spatio-temporal correlation modeling of data, provide transparent interpretation of data fusion, improve the credibility and availability of results, and significantly improve the accuracy and personalization of disease risk warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691687B_ABST
    Figure CN119691687B_ABST
Patent Text Reader

Abstract

The present application relates to a multi-modal data fusion method, apparatus, computer device, and storage medium. The method includes: using a dedicated modal encoder to perform high-order feature extraction on the multi-modal data to be fused, obtaining high-order embedded features of the multi-modal data; decoupling the high-order embedded features by combining expert knowledge constraints and a multi-head attention mechanism to respectively obtain high-frequency features representing local information and low-frequency features representing global information; using a cosine alignment algorithm to align the low-frequency features, using an orthogonal projection technique to perform orthogonal projection on the high-frequency features, and performing multi-level fusion on the aligned low-frequency features and the orthogonally projected high-frequency features through an adaptive deep fusion network to obtain fused multi-modal joint features. The present application solves the problems of data heterogeneity and heterogeneity, ensures effective fusion of different modal data, and can significantly improve the accuracy and personalization ability of disease risk early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical fields of data processing and medical health, and particularly relates to a multi-modal data fusion method, device, computer device, and storage medium. Background Art

[0002] With the rapid development of medical informatization, more and more medical data are collected and stored through different channels. Especially in the treatment process of multi-center perioperative patients, it involves various modal data such as structured electronic health records (EHRs), preoperative electrocardiograms, medical images (chest X-rays, abdominal CT scans), and intraoperative sequential physiological monitoring. By fusing multi-modal data, rich patient health information can be obtained to provide reference for clinical decision-making. In the prior art, multi-modal data fusion methods mostly rely on rule-based feature extraction and manually designed fusion strategies. Although data fusion can be achieved in some scenarios, due to significant differences in the structure, quality, and spatio-temporal scale of different modal data, many technical challenges are faced in multi-modal data fusion, especially there are great limitations in dealing with data heterogeneity, specifically manifested as follows:

[0003] (1) It is difficult to effectively process highly heterogeneous and heterogeneous data: The sources of multi-modal data are diverse. For the same modal data, due to differences in multi-center acquisition devices and environments, there are not only large differences in data formats such as quality, resolution, and scale, but also significant heterogeneity and heterogeneity in terms of data type, quality, and scale. Existing multi-modal data fusion methods usually cannot effectively process these differences, resulting in poor fusion effects and the inability to extract effective complementary information from each modality. Especially in the task of perioperative cardiovascular event warning, the data heterogeneity is more prominent, affecting the accuracy and reliability of the warning system.

[0004] (2) Poor information extraction ability: Existing technologies usually rely on traditional rule-based feature extraction methods and do not fully consider the noise and redundant information in the data. Especially in intraoperative sequential physiological monitoring and image data, noise and redundant information account for a large proportion, resulting in a low signal-to-noise ratio and making it difficult to effectively extract valuable medical information, thereby affecting the accuracy of cardiovascular event risk assessment and the accuracy of warning.

[0005] (3) Unable to fully utilize the correlation between multi-modal data: In multi-modal data, there are often complex horizontal coupling relationships between modalities. Especially the high coupling between image data and structured electronic health records. Existing data fusion methods often cannot effectively analyze the mutual relationships between different data modalities and are difficult to fully utilize the complementary information in multi-modal data for accurate prediction.

[0006] (4)Insufficient medical interpretability: Most of the existing multimodal data fusion methods rely on black-box models, lacking medical interpretability of the fusion process, which not only reduces the credibility of the model but also limits its popularization and application in clinical practice.

[0007] Therefore, in order to better address the heterogeneous and heterogeneous problems of multimodal data and improve the accuracy of perioperative cardiovascular event warning, there is an urgent need for a new multimodal data fusion algorithm to provide more accurate and personalized decision support for clinicians. Summary of the Invention

[0008] This application provides a multimodal data fusion method, device, computer device, and storage medium, aiming to at least partially solve one of the above technical problems in the prior art.

[0009] To solve the above problems, this application provides the following technical solutions:

[0010] A multimodal data fusion method includes:

[0011] Using a dedicated modal encoder to perform high-order feature extraction on the multimodal data to be fused, obtaining high-order embedded features of the multimodal data;

[0012] Combining expert knowledge constraints and multi-head attention mechanism to decouple the high-order embedded features, respectively obtaining high-frequency features representing local information and low-frequency features representing global information;

[0013] Using the cosine alignment algorithm to align the low-frequency features, using the orthogonal projection technique to perform orthogonal projection on the high-frequency features, and performing multi-level fusion on the aligned low-frequency features and the orthogonally projected high-frequency features through an adaptive depth fusion network to obtain fused multimodal joint features;

[0014] Inputting the fused multimodal features into a pre-trained risk prediction model, and outputting the disease risk warning probability through the risk prediction model.

[0015] Another technical solution adopted in the embodiments of this application is that the multimodal data includes structured electronic health records, preoperative electrocardiograms, medical image data, and intraoperative physiological monitoring. The step of using a dedicated modal encoder to perform high-order feature extraction on the multimodal data to be fused and obtaining high-order embedded features of the multimodal data is specifically as follows:

[0016] For the structured electronic health records, high-order feature extraction is performed through an encoder combining a multi-layer perceptron and a decision tree;

[0017] For the preoperative electrocardiogram and intraoperative physiological monitoring data, Fourier transform and frequency-domain attention mechanism are used to extract high-order features in the frequency domain;

[0018] For medical image data, extract the detailed features therein through a deep convolutional network.

[0019] The technical solution adopted in the embodiment of this application further includes: decoupling the high-order embedded features by combining expert knowledge constraints and the multi-head attention mechanism to respectively obtain high-frequency features representing local information and low-frequency features representing global information. Specifically:

[0020] Encode the expert knowledge related to disease risk warning into an embedding vector through an embedding layer, splice the embedding vector with the high-order embedded features, and input the result into a feature decoupling module;

[0021] The feature decoupling module processes the high-order embedded features through a normalization layer, a multi-head attention mechanism layer, and a feed-forward fully connected layer in sequence to capture the potential relationships and feature relationships between the multi-modal data;

[0022] Based on the potential relationships and feature relationships between the multi-modal data, analyze the high-order embedded features through a low-frequency feature common decoupling layer and a high-frequency feature specific decoupling layer respectively to extract the low-frequency features and high-frequency features of the multi-modal data;

[0023] Calculate the mean and variance of the low-frequency features and high-frequency features, and sample from the mean and variance through the Gaussian sampling method to obtain the final low-frequency features and high-frequency features.

[0024] The technical solution adopted in the embodiment of this application further includes: the feature decoupling module processes the high-order embedded features through a normalization layer, a multi-head attention mechanism layer, and a feed-forward fully connected layer in sequence to capture the potential relationships and feature relationships between the multi-modal data. Specifically:

[0025] Standardize the high-order embedded features of each modality data through the normalization layer; given the embedded feature , the normalization layer standardizes it through the following formula:

[0026]

[0027] where and are the mean and standard deviation of respectively;

[0028] Learn the potential relationships between the high-order embedded features of each modality data through the multi-head attention mechanism layer, and mine the semantic information of each modality data;

[0029] Introduce a non-linear transformation through the feed-forward fully connected layer to further extract the output features of the multi-head attention mechanism layer and capture the feature relationships of each modality data.

[0030] The technical solution adopted in the embodiment of the present application further includes: based on the potential relationship and feature relationship between the multimodal data, the high-order embedded features are analyzed through a low-frequency feature common decoupling layer and a high-frequency feature specific decoupling layer respectively, and the low-frequency features and high-frequency features of the multimodal data are extracted. Specifically:

[0031] The low-frequency feature common decoupling layer uses the spectral modulation attention mechanism of the weighted spectral matrix to adaptively adjust the low-frequency components in the high-order embedded features, enhancing the model's attention to and extraction of low-frequency features:

[0032]

[0033]

[0034] where is the absolute value of the spectrum of the th input feature, is the feature dimension, is the total number of features; represents the extracted low-frequency features, is obtained by three different linear transformations of the feature vector output from the previous layer, is the dimension of the key vector;

[0035] The high-frequency feature specific decoupling layer uses the spectral modulation attention mechanism to adaptively adjust the high-frequency components in the high-order embedded features, emphasizing the modality-specific detailed features:

[0036]

[0037]

[0038] where, is the square of the spectral intensity in the input high-order embedded features, representing the magnitude of the high-frequency components, represents the extracted high-frequency features.

[0039] The technical solution adopted in the embodiment of the present application further includes: using the cosine alignment algorithm to align the low-frequency features and using the orthogonal projection technique to perform orthogonal projection on the high-frequency features. Specifically:

[0040] Calculate the cosine similarity between the low-frequency features through the cosine alignment algorithm, and constrain the alignment of the low-frequency features by minimizing the cosine distance:

[0041]

[0042] where and respectively represent the and the low-frequency features of the modal data, is the Euclidean norm of the vector;

[0043] Combined with expert knowledge, the aligned low-frequency features are weighted by a multilinear adaptive weighting algorithm to obtain weighted low-frequency features ; the weighting formula is:

[0044]

[0045] where is the weighting coefficient adaptively calculated according to expert knowledge ;

[0046] All the weighted low-frequency features are fused to obtain the fused low-frequency feature representation as:

[0047]

[0048] where represents the weight coefficient of the low-frequency feature, represents the number of modalities;

[0049] The high-frequency features are standardized by an orthogonal projection algorithm:

[0050]

[0051] where is the unitized high-frequency feature vector of the th modal data, represents the Euclidean norm of the vector;

[0052] For the high-frequency feature , project it outside the space of other modal features:

[0053]

[0054] where, is the high-frequency feature representation of the th modal data after projection, represents the inner product operation;

[0055] The orthogonally projected high-frequency features are fused by linear weighting to obtain the final high-frequency feature representation :

[0056]

[0057] where is the weight coefficient of high-frequency features, indicating the number of modalities.

[0058] The technical solution adopted in the embodiment of the present application further includes: performing multi-level fusion on the aligned low-frequency features and orthogonally projected high-frequency features through an adaptive depth fusion network to obtain fused multi-modal joint features, specifically:

[0059] Performing one-dimensional temporal convolution on the low-frequency features and high-frequency features through a low-frequency-high-frequency spatio-temporal information collaboration sub-network and then performing weighted fusion to obtain temporal features ; calculating the spatial interaction relationship between the low-frequency features and high-frequency features through pointwise multiplication to obtain spatial features ;

[0060] Performing spatial consistency projection on the temporal features and spatial features and performing normalization processing to obtain normalized temporal features and spatial features and ;

[0061] Performing weighted fusion on the normalized temporal features and spatial features and to obtain unified multi-modal joint features :

[0062]

[0063] where is a learnable fusion weight coefficient.

[0064] Another technical solution adopted in the embodiment of the present application is: a multi-modal data fusion device, including:

[0065] A feature extraction module: used to extract high-order features of the multi-modal data to be fused by using a dedicated modal encoder to obtain high-order embedded features of the multi-modal data;

[0066] A feature decoupling module: used to decouple the high-order embedded features by combining expert knowledge constraints and a multi-head attention mechanism to respectively obtain high-frequency features representing local information and low-frequency features representing global information;

[0067] A feature fusion module: used to align the low-frequency features by using a cosine alignment algorithm, orthogonally project the high-frequency features by using an orthogonal projection technique, and perform multi-level fusion on the aligned low-frequency features and orthogonally projected high-frequency features through an adaptive depth fusion network to obtain fused multi-modal joint features;

[0068] Risk prediction module: configured to input the fused multi-modal features into a pre-trained risk prediction model, and output the disease risk warning probability through the risk prediction model.

[0069] Another technical solution adopted in the embodiments of the present application is: a computer device, which includes a processor and a memory coupled to the processor. Among them,

[0070] The memory stores program instructions for implementing the multi-modal data fusion method;

[0071] The processor is configured to execute the program instructions stored in the memory to control the multi-modal data fusion method.

[0072] Another technical solution adopted in the embodiments of the present application is: a storage medium storing program instructions executable by a processor, and the program instructions are used to execute the multi-modal data fusion method.

[0073] Compared with the prior art, the beneficial effects produced by the embodiments of the present application are as follows: The multi-modal data fusion method, device, computer device, and storage medium in the embodiments of the present application effectively process the differences of different modal data by designing dedicated encoders for each modal data, improve the signal-to-noise ratio of the data, solve the problems of data heterogeneity and isomerism, and ensure the effective fusion of different modal data; use expert knowledge to decompose the high-order features of each modal data into low-frequency features and high-frequency features, effectively eliminate redundant information and noise, reduce the influence of high coupling degree between modalities, and ensure the independence of each modal information; perform feature fusion on the low-frequency features and high-frequency features through a multi-level interpretable feature fusion module, realize spatio-temporal correlation modeling of the data, provide a transparent interpretation of the data fusion, and improve the credibility and usability of the results. The present application solves the problems of information sharing and differential extraction between multi-modal data through the decoupling of high and low frequency features, the extraction of commonalities and uniqueness of features between modalities, and multi-level interpretable feature fusion, and can significantly improve the accuracy and personalization ability of disease risk warning. Description of the Drawings

[0074] Figure 1 is a flowchart of the multi-modal data fusion method in the embodiments of the present application;

[0075] Figure 2 is a schematic structural diagram of a feature decoupling module based on expert knowledge constraints and multi-head attention mechanism in the embodiments of the present application;

[0076] Figure 3 is a schematic structural diagram of the multi-modal data fusion device in the embodiments of the present application;

[0077] Figure 4 is a schematic structural diagram of the computer device in the embodiments of the present application;

[0078] Figure 5 Schematic structural diagram of the storage medium according to the embodiment of the present application. Specific implementation manners

[0079] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0080] The terms "first", "second", and "third" in the present application are only for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or computer device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or computer devices.

[0081] Referring to herein "embodiment" means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0082] Specifically, please refer to Figure 1 , which is a flowchart of the multi-modal data fusion method according to the embodiment of the present application. The multi-modal data fusion method according to the embodiment of the present application includes the following steps:

[0083] S100: Obtain multi-modal data of a target patient;

[0084] In this step, the multimodal data obtained includes, but is not limited to, medical data such as the patient's structured electronic health record, preoperative electrocardiogram, medical imaging data, and intraoperative physiological monitoring. Among them, the medical imaging data includes, but is not limited to, chest X-ray plain film or abdominal CT plain film, etc. The intraoperative physiological monitoring data includes, but is not limited to, real-time physiological signals such as the patient's heart rate, blood oxygen saturation, and respiratory rate.

[0085] S110: Design corresponding dedicated modal encoders for each modal data, and use the dedicated modal encoders to perform high-order feature extraction on the multimodal data to be fused, obtaining high-order embedded features of the multimodal data;

[0086] In this step, in order to effectively learn the invariant features between different modal data and fully mine the key information in each modal data, this application designs corresponding dedicated modal encoders according to the properties of different modal data, and uses the dedicated modal encoders to extract high-order features from different modal data using different techniques. The output feature dimensions of each modal encoder are kept consistent, thereby providing a unified feature space for subsequent multimodal data fusion and avoiding the fusion difficulties caused by inconsistent feature dimensions.

[0087] Specifically, the high-order feature extraction process of multimodal data includes:

[0088] S111: For the structured electronic health record, perform high-order feature extraction through an encoder combining a multi-layer perceptron and a decision tree;

[0089] Among them, the data in the structured electronic health record usually appears in structured numerical and categorical forms, including important information such as the patient's medical history, diagnosis, and laboratory test results. In order to extract high-level semantic features from this information, this application designs an encoder that combines a multi-layer perceptron (MLP) and a decision tree model (LightGBM). Through the MLP, numerical data such as age, body mass index, and preoperative blood glucose are embedded and non-linearly feature-learned to capture univariate information and simple variable relationships; through LightGBM, categorical data such as gender, surgical site, and hypertension history are deeply learned to capture complex variable interactions and non-linear features, and then the features extracted by the MLP and LightGBM are concatenated to obtain the high-order features of the structured electronic health record, denoted as 。

[0090] S112: For preoperative electrocardiogram and intraoperative physiological monitoring data, use Fourier transform and frequency-domain attention mechanism to extract high-order features in the frequency domain;

[0091] Among them, the preoperative electrocardiogram and intraoperative physiological monitoring data are both time series data, with strong time dependence and dynamic changes. To better process time series data, this application designs a feature extraction network that combines the Fourier transform (FFT) and the frequency domain attention mechanism. First, the time domain signal is transformed into the frequency domain through the Fourier transform to extract frequency domain features such as power spectral density, enhancing the ability to capture key frequency components; then, based on the frequency domain attention mechanism, key frequency components are selected to further improve the expression ability of the frequency domain features, thereby extracting the high-order features of the preoperative electrocardiogram and intraoperative physiological monitoring data, denoted as , .

[0092] S113: For medical image data, extract the detailed features therein through a deep convolutional network;

[0093] Among them, medical image data provides a detailed view of the internal structure of the patient's body and can reflect the potential lesion areas of the disease. To extract effective features from medical image data, this application uses a DenseNet (Dense Convolutional Network) as the encoder of image features. Through the layer-by-layer dense connection design of DenseNet, multi-scale features in medical image data are efficiently captured, denoted as , which contains local and global information of the medical image data.

[0094] It can be understood that by using the above feature extraction method, this application can effectively extract high-order embedded features for disease risk warning from multi-modal data, which can improve the signal-to-noise ratio of the data, optimize the feature extraction ability, and provide rich basic information for subsequent multi-modal data fusion.

[0095] S120: Decouple the high-order embedded features by combining expert knowledge constraints and the multi-head attention mechanism to obtain high-frequency features representing local information and low-frequency features representing global information respectively;

[0096] In this step, to solve the problem of high horizontal structural coupling degree of multi-modal data, this application introduces a feature decoupling module based on expert knowledge constraints and the multi-head attention mechanism to analyze and separate the highly coupled parts between multi-modal data, realizing effective decoupling of the data, so as to better utilize the potential correlation between different modal data. Specifically, as Figure 2As shown in the figure, it is a schematic structural diagram of a feature decoupling module based on expert knowledge constraints and multi-head attention mechanism in an embodiment of the present application. The feature decoupling module encodes the high-order embedded features of a single modality and decomposes them into low-frequency features and high-frequency features. The low-frequency features are used to capture global patterns, reveal the common information of different modality data, and reflect the long-term patterns of disease risk early warning; the high-frequency features are used to extract local detailed information and reflect the sudden changes in disease risk early warning. Through this decoupling process, the complementary information of multi-modal data can be effectively separated, the influence of redundancy and noise can be removed, the expression ability of meaningful patterns in multi-modal data can be enhanced, and the efficient parsing and interpretable separation of shared information and unique information between multi-modalities are realized.

[0097] Further, as Figure 2 shown, the feature decoupling process of the feature decoupling module based on expert knowledge constraints and multi-head attention mechanism in an embodiment of the present application includes the following steps:

[0098] S121: Encode the expert knowledge related to disease risk early warning (such as disease characteristics, clinical experience, medical guidelines, etc.) into an embedding vector C through an embedding layer, and input the concatenated embedding vector C and high-order embedded features into the feature decoupling module;

[0099] S122: The feature decoupling module processes the high-order embedded features through a normalization layer, a multi-head attention mechanism layer, and a feed-forward fully connected layer in sequence to capture the potential relationships and feature relationships between multi-modal data;

[0100] Among them, the processing processes of the normalization layer, the multi-head attention mechanism layer, and the feed-forward fully connected layer on the high-order embedded features are specifically as follows:

[0101] Standardize the high-order embedded features of each modality through the normalization layer; given the embedded feature (the th feature), the normalization layer standardizes it through the following formula:

[0102]

[0103] where and are the mean and standard deviation of , respectively.

[0104] The multi-head attention mechanism layer parallelly learns the potential relationships between the features of multi-modal data from multiple different "perspectives", captures the long-range dependencies within and between modalities of multi-modal data, and thus fully excavates the semantic information of each modality data.

[0105] Introduce non - linear transformation through the feed - forward fully - connected layer to further extract the output features of the multi - head attention mechanism layer, enabling it to better capture the feature relationships of multi - modal data and enhancing the model's expression ability for disease risk warning.

[0106] S123: Based on the potential relationships and feature relationships among multi - modal data, respectively parse the high - order embedded features through the low - frequency feature common decoupling layer and the high - frequency feature specific decoupling layer, and extract the low - frequency features and high - frequency features of multi - modal data respectively;

[0107] Among them, the task of the low - frequency feature common decoupling layer is to extract the low - frequency features shared among multi - modal data by parsing the input high - order embedded features. The low - frequency features reflect the relatively stable common information in different modal data, usually representing the long - term and general patterns in disease risk warning, and can effectively reveal the potential risk patterns of diseases. Specifically, the low - frequency feature common decoupling layer adopts the spectral modulation attention mechanism of the weighted spectral matrix, and adaptively adjusts the low - frequency components in the input high - order embedded features to enhance the model's attention to and extraction of low - frequency features. The specific algorithm is as follows:

[0108]

[0109]

[0110] Among them is the spectral absolute value of the th input feature, reflecting the intensity of the low - frequency components in the spectrum, is the feature dimension, is the total number of features. represents the extracted low - frequency features, is obtained by three different linear transformations of the feature vector output from the previous layer, is the dimension of the key vector.

[0111] The high - frequency feature specific decoupling layer focuses on extracting the high - frequency features of multi - modal data, that is, the local details and sudden changes in the data, which can help distinguish the subtle differences between different modal data. Specifically, the high - frequency feature specific decoupling layer adopts the spectral modulation attention mechanism to adaptively adjust the high - frequency components in the input high - order embedded features, emphasizing the modality - specific detail features. The specific calculation formula is as follows:

[0112]

[0113]

[0114] Among them, is the square of the spectral intensity in the input high - order embedded features, representing the magnitude of the high - frequency components, Indicates the extracted high-frequency features.

[0115] S124: Calculate the mean ( ), and variance ( ) of the low-frequency features and high-frequency features, and sample from the mean ( ) and variance ( ) through the Gaussian sampling method to obtain the final low-frequency features and high-frequency features;

[0116] Among them, after passing through the decoder module, the low-frequency features and high-frequency features of different modality data can reconstruct the high-order embedded features of the original input ( ). This reconstruction process ensures the stability of model training and avoids the problem of unstable training caused by feature loss or deformation. Finally, the decoder module restores the embedded representation that matches the input high-order embedded features through the combined decoding of low-frequency features and high-frequency features.

[0117] It can be understood that through the above feature decoupling process, the model can efficiently separate the common information and unique information in multi-modal data, and through the alternating action of the encoder and decoder modules, ensure the stability and effectiveness of different modality features, providing a solid foundation for subsequent multi-modal data fusion, which is conducive to the model to more accurately identify and evaluate the potential risks of disease risk early warning, and improve the accuracy and personalization level of prediction.

[0118] S130: Align the decoupled low-frequency features using the cosine alignment algorithm, perform orthogonal projection on the high-frequency features using the orthogonal projection technique, and then perform multi-level fusion on the low-frequency features and high-frequency features through the adaptive deep fusion network to obtain the fused multi-modal joint features;

[0119] In this step, in the feature fusion stage, this application fuses the decoupled low-frequency features and high-frequency features through a multi-level interpretable feature fusion module to achieve an effective joint representation of multi-modal data and provide more accurate support for the disease risk early warning of the target patient. First, the interpretable feature fusion module captures the commonality between modalities by cosine-aligning the low-frequency features. Then, it performs orthogonal projection on the high-frequency features to maintain the uniqueness of different modalities. Finally, it fuses all modality features through the adaptive deep fusion network, which can clearly analyze the relationship between shared features and unique features and ensure that the model has good interpretability, thereby improving the accuracy and personalization ability of cardiovascular event risk assessment.

[0120] Specifically, the feature fusion process of the interpretable feature fusion module includes the following steps:

[0121] S131: Cosine-align the low-frequency features through the cosine alignment algorithm;

[0122] Among them, low-frequency features usually contain common information in multimodal data and can reflect the global pattern of cardiovascular events. In order to enhance the correlation between low-frequency features between multimodal data, this application first projects the low-frequency representation of each modal data to the same subspace through a cosine alignment algorithm. In this way, the shared information in different modal data can be uniformly represented to ensure that the low-frequency representations in different modal data remain consistent in space, thereby better revealing the potential risk patterns of the disease. Specifically, the cosine alignment algorithm calculates the cosine similarity between low-frequency features and constrains the alignment of low-frequency features through a strategy that minimizes the cosine distance, thereby further optimizing the collaborative representation of low-frequency features between different modal data. The cosine alignment formula is as follows:

[0123]

[0124] in and Respectively represent and The low-frequency characteristics of the modal data, is the Euclidean norm of the vector.

[0125] Define the loss function , minimizing the cosine distance between low-frequency features is:

[0126]

[0127] in, and Respectively represent the number of modes of multimodal data.

[0128] Through the above alignment operation, the low-frequency features of each modal data can work together better, providing a unified representation space for subsequent multimodal data fusion, helping to discover key risk factors for disease risk warning, and providing a quantitative basis for clinical decision-making.

[0129] S132: combining expert knowledge, weighting the aligned low-frequency features through a multilinear adaptive weighting algorithm to obtain weighted low-frequency features;

[0130] Among them, in order to further enhance the clinical applicability of low-frequency features, this application combines expert knowledge and uses a multilinear adaptive weighting algorithm to weight low-frequency features, giving higher weights to information related to disease risk warning, and obtaining weighted low-frequency features. , so that the model can more effectively focus on the characteristics that are closely related to the patient's perioperative risk. The specific weighting formula is as follows:

[0131]

[0132] Among them, is a weighted coefficient adaptively calculated based on expert knowledge and is calculated by the following formula:

[0133]

[0134] is a learnable weight vector.

[0135] S133: Fuse all weighted low-frequency features to obtain a fused low-frequency feature representation; among them, the finally fused low-frequency feature representation is:

[0136]

[0137] where represents the weight coefficient of the low-frequency feature, represents the number of modalities.

[0138] Through this process, the model can effectively integrate the low-frequency features in the multimodal data, thereby more accurately identifying the risk factors in the perioperative period.

[0139] S134: After performing redundancy removal processing on the high-frequency features of the multimodal data through the orthogonal projection algorithm, perform weighted fusion to obtain a fused high-frequency feature representation;

[0140] Among them, in the perioperative disease risk warning, the high-frequency features represent the local details and unique information in different modality data, and are of great significance for predicting sudden disease changes and their risks. However, due to the possible redundant information and noise in the high-frequency features of the multimodal data, in order to improve the robustness and accuracy of the model, it is necessary to effectively extract and fuse these high-frequency features. For this purpose, this application designs a high-frequency orthogonal projection layer to perform redundancy removal processing on the high-frequency features of the multimodal through the orthogonal projection algorithm, so as to maximize the retention of the modality-specific detailed information and reduce noise interference. Specifically, first normalize the high-frequency features to ensure that the lengths of the feature vectors are consistent:

[0141]

[0142] where is the unitized high-frequency feature vector of the th modality data unit, represents the Euclidean norm of the vector.

[0143] In order to further remove the redundant information between the high-frequency features of different modality data, enhance the independence of the high-frequency features of each modality data through orthogonal projection. For the high-frequency features of the , project it outside the space of other modal features, and the formula is as follows:

[0144]

[0145] Where is the high-frequency feature representation of the th modal data after projection, and represents the inner product operation.

[0146] Finally, linearly weight the high-frequency features after orthogonal projection for fusion to obtain the final high-frequency feature representation :

[0147]

[0148] Where is the weight coefficient of the high-frequency feature, represents the number of modalities, and the weight coefficient can be adjusted according to the importance of each modal data to ensure that the high-frequency features of each modal data receive appropriate attention during the final fusion.

[0149] Through the above process, the model can efficiently fuse the high-frequency features of each modal data, remove redundancy and noise, and enhance the accurate assessment of perioperative disease risk warning.

[0150] S135: Weightedly fuse the fused low-frequency feature representation and high-frequency feature representation through an adaptive deep fusion network to obtain a unified multi-modal joint feature;

[0151] Where, in order to further improve the fusion effect of low-frequency features and high-frequency features, this application performs feature fusion by designing an adaptive deep fusion network. This network performs spatio-temporal coordination and spatial consistency projection on low-frequency features and high-frequency features through two layers of sub-networks, realizing dynamic modeling of low-frequency features and high-frequency features in the time and space dimensions, so as to ensure that multi-modal data fully integrates low-frequency features and high-frequency features in the spatio-temporal dimension, enhancing the sensitivity and robustness of the model to disease risk warning. Map high-frequency features to a low-dimensional space through orthogonal projection technology to remove the interference of noise and redundant information, improve the signal-to-noise ratio of the data, and enhance the independence of modal features. At the same time, dynamically adjust the feature contribution weights of each modal data through a multi-linear adaptive weighting algorithm combined with clinical expert knowledge to improve the interpretability of the model. The specific fusion process includes:

[0152] S1351: Spatially and temporally synergize low-frequency features and high-frequency features through the low-frequency - high-frequency spatio-temporal information synergy sub-network to explore the dynamic relationship between low-frequency features and high-frequency features, ensuring that low-frequency features and high-frequency features can effectively synchronously interact and correlate in both the time dimension and the space dimension; the specific operation is as follows: perform one-dimensional temporal convolution on the low-frequency features and high-frequency features and then perform weighted fusion to obtain temporal features ; calculate the spatial interaction relationship between the low-frequency features and high-frequency features through element-wise multiplication to ensure the spatial synergy of the features, obtaining spatial features . Through spatio-temporal synergy operations, not only can the patterns of disease risk warnings in time changes be captured, but also the detailed information in spatial data is considered.

[0153] S1352: Project the temporal features and spatial features for spatial consistency and perform normalization;

[0154] Among them, by projecting the temporal features and spatial features for spatial consistency, the consistency and scale uniformity of low-frequency features and high-frequency features during fusion are ensured, avoiding fusion problems caused by inconsistent feature scales. The specific operation is as follows: First, perform linear projections on the temporal features and spatial features respectively to map them into a unified feature space:

[0155]

[0156] where is the projection weight matrix of the joint space, and are the dimensions of the original feature and the projection space respectively.

[0157] Then, normalize the projected temporal features and spatial features to obtain the normalized temporal features and spatial features .

[0158] Finally, perform weighted fusion on the normalized temporal features and spatial features to obtain the unified multi-modal joint feature :

[0159]

[0160] where is the learnable fusion weight coefficient.

[0161] Through the fusion process of the above-mentioned adaptive depth fusion network, effective multi-modal joint features can be obtained. The multi-modal joint features not only contain the common information in the multi-modal data, but also retain the unique information of each modal data, which is beneficial to helping clinicians make more accurate diagnoses and treatments.

[0162] S140: Input the fused multi-modal features into a pre-trained risk prediction model, and output the disease risk warning probability of the target patient through the risk prediction model;

[0163] In this step, in order to optimize the performance of the model in the risk prediction task, this application designs a high-low frequency collaborative contrast loss function , by jointly optimizing the similarity constraint between low-frequency features and the contrast between low-frequency and high-frequency features in two losses, while maintaining the shared information between low-frequency features, strengthening the uniqueness and time-varying information of high-frequency features, further enhancing the collaborative effect of low-frequency and high-frequency features, enhancing the model's ability to capture long-term patterns and sudden changes, and improving the robustness and accuracy of the model in perioperative disease risk warning.

[0164] Specifically, the similarity constraint between low-frequency features is: First, calculate the Euclidean distance between the low-frequency features of different modal data to ensure their similarity in the feature space and reduce the differences between low-frequency features. The calculation formula is as follows:

[0165]

[0166] where and are the low-frequency features of the rd and th modal data respectively.

[0167] To ensure the complementarity of low-frequency and high-frequency features in multi-modal data learning, this application designs a contrast mechanism between low-frequency and high-frequency features to encourage them to maintain differences, so that low-frequency and high-frequency features can complement each other in the feature space, avoiding information redundancy and insufficient feature decoupling. The contrast mechanism formula between low-frequency and high-frequency features is as follows:

[0168]

[0169] where, and represent the low-frequency and high-frequency features of the th modal data respectively, and represent the number of modalities of multi-modal data.

[0170] Combining the similarity constraint between low-frequency features and the contrast between low-frequency and high-frequency features is the final high-low frequency collaborative contrast loss :

[0171]

[0172] It can be understood that by jointly optimizing the similarity constraint between low-frequency features and the contrast between low-frequency and high-frequency features, this application can ensure that low-frequency features and high-frequency features can complement each other well in the feature space, avoid information redundancy, and at the same time improve the accuracy of feature fusion.

[0173] In addition, the risk prediction model optimizes the risk probability of the model output through cross-entropy loss. By calculating the above losses and performing backpropagation optimization, the model training reaches the optimal state. According to the risk warning probability output by the model, clinicians can timely understand the disease risks of patients during the perioperative period and take corresponding intervention measures, which can not only provide important risk warnings for surgeries, guide the smooth progress of surgeries, but also improve the success rate of surgeries and the survival rate of patients, and has important clinical application value.

[0174] It can be understood that in the above embodiments, this application takes the risk warning of perioperative cardiovascular events as an example for illustration. The applicant can also apply it to other types of disease risk warnings such as chronic disease monitoring. For different disease fields, this application does not need to adjust the model structure. Only by replacing the training samples with multi-modal data related to the target disease type during the training stage can efficient data fusion and feature extraction be achieved, and the corresponding disease risk warning probability can be output. In addition, the feature decoupling technology and adaptive deep fusion network with expert knowledge constraints designed in this application can also be applied to data fusion tasks in the industrial field, such as multi-sensor data anomaly detection, complex system fault prediction, etc. By adjusting the input data type and task objectives, this application can achieve a wider range of data fusion applications and has high generality and adaptability.

[0175] The multi-modal data fusion method of this application's embodiment effectively processes the differences of each modal data by designing corresponding dedicated encoders for each modal data, improves the signal-to-noise ratio of the data, solves the heterogeneity and isomerism problems of multi-modal data, and ensures the effective fusion of multi-modal data; uses expert knowledge to decompose the high-order features of each modal data into low-frequency features and high-frequency features, effectively eliminates redundant information and noise, reduces the influence of high coupling degree between modalities, and ensures the independence of each modal information; performs feature fusion on low-frequency features and high-frequency features through a multi-level interpretable feature fusion module, realizes the spatio-temporal correlation modeling of data, provides a transparent explanation for data fusion, and improves the credibility and usability of the results. This application solves the problems of information sharing and differential extraction between multi-modal data through the decoupling of high and low-frequency features, the extraction of commonalities and uniqueness between inter-modal features, and multi-level interpretable feature fusion, and can significantly improve the accuracy and personalization ability of disease risk warnings.

[0176] Please refer to Figure 3 , which is a schematic structural diagram of the multi-modal data fusion device according to the embodiment of the present application. The multi-modal data fusion method device 40 according to the embodiment of the present application includes:

[0177] Feature extraction module 41: used to perform high-order feature extraction on the multi-modal data to be fused by using a dedicated modal encoder, and obtain high-order embedded features of the multi-modal data;

[0178] Feature decoupling module 42: used to decouple the high-order embedded features by combining expert knowledge constraints and multi-head attention mechanisms, and respectively obtain high-frequency features representing local information and low-frequency features representing global information;

[0179] Feature fusion module 43: used to align the low-frequency features by using the cosine alignment algorithm, perform orthogonal projection on the high-frequency features by using the orthogonal projection technology, and perform multi-level fusion on the aligned low-frequency features and the orthogonally projected high-frequency features through an adaptive deep fusion network to obtain fused multi-modal joint features;

[0180] Risk prediction module 44: used to input the fused multi-modal features into a pre-trained risk prediction model, and output the disease risk warning probability through the risk prediction model.

[0181] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.

[0182] The device provided by the embodiment of the present application can be applied in the foregoing method embodiment, for details, please refer to the description of the foregoing method embodiment, and will not be elaborated here.

[0183] Please refer to Figure 4 , which is a schematic structural diagram of the computer device according to the embodiment of the present application. The computer device 50 includes:

[0184] A memory 51 storing executable program instructions;

[0185] A processor 52 connected to the memory 51;

[0186] The processor 52 is used to call the executable program instructions stored in the memory 51 and execute the following steps: performing high-order feature extraction on the multi-modal data to be fused by using a dedicated modal encoder to obtain high-order embedded features of the multi-modal data; decoupling the high-order embedded features by combining expert knowledge constraints and a multi-head attention mechanism to respectively obtain high-frequency features representing local information and low-frequency features representing global information; aligning the low-frequency features by using a cosine alignment algorithm, performing orthogonal projection on the high-frequency features by using an orthogonal projection technique, and performing multi-level fusion on the aligned low-frequency features and the orthogonally projected high-frequency features through an adaptive depth fusion network to obtain fused multi-modal joint features; inputting the fused multi-modal features into a pre-trained risk prediction model, and outputting a disease risk warning probability through the risk prediction model.

[0187] Among them, the processor 52 can also be called a CPU (Central Processing Unit, central processing unit). The processor 52 may be an integrated circuit chip with signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0188] Please refer to Figure 5, which is a schematic structural diagram of the storage medium according to an embodiment of the present application. The storage medium according to the embodiment of the present application stores program instructions 61 that can implement the following steps: using a dedicated modal encoder to perform high-order feature extraction on the multi-modal data to be fused to obtain high-order embedded features of the multi-modal data; combining expert knowledge constraints and a multi-head attention mechanism to decouple the high-order embedded features to respectively obtain high-frequency features representing local information and low-frequency features representing global information; using a cosine alignment algorithm to align the low-frequency features, using an orthogonal projection technique to perform orthogonal projection on the high-frequency features, and performing multi-level fusion on the aligned low-frequency features and the orthogonally projected high-frequency features through an adaptive depth fusion network to obtain fused multi-modal joint features; inputting the fused multi-modal features into a pre-trained risk prediction model, and outputting a disease risk warning probability through the risk prediction model. Among them, the program instructions 61 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network computer device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of the present application. And the foregoing storage medium includes: various media that can store program instructions such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, or a terminal computer device such as a computer, a server, a mobile phone, or a tablet. Among them, the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0189] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0190] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units. The above is only the implementation manner of the present application and does not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A multimodal data fusion method, characterized in that: include: Using a dedicated modality encoder to extract high-order features from the multimodal data to be fused, and obtaining high-order embedded features of the multimodal data; the multimodal data includes structured electronic health records, preoperative electrocardiograms, medical imaging data, and intraoperative physiological monitoring; Combining expert knowledge constraints and a multi-head attention mechanism to decouple the high-order embedded features, high-frequency features representing local information and low-frequency features representing global information are obtained respectively; The low-frequency features are aligned using a cosine alignment algorithm, the high-frequency features are orthogonally projected using an orthogonal projection technique, and the aligned low-frequency features and the orthogonally projected high-frequency features are multi-level fused using an adaptive deep fusion network to obtain fused multi-modal joint features; The fused multimodal features are input into the pre-trained risk prediction model, and the disease risk warning probability is output through the risk prediction model; wherein: The high-order embedding features are decoupled by combining expert knowledge constraints and multi-head attention mechanisms to obtain high-frequency features representing local information and low-frequency features representing global information, respectively. Specifically, The expert knowledge of the relevant disease risk warning is encoded into an embedding vector through an embedding layer, and the embedding vector is concatenated with the high-order embedding feature and then input into the feature decoupling module; The feature decoupling module processes the high-order embedded features through a normalization layer, a multi-head attention mechanism layer, and a feed-forward fully connected layer in sequence to capture the potential relationship and feature relationship between the multimodal data; Based on the potential relationship and feature relationship between the multimodal data, the high-order embedded features are parsed through a low-frequency feature common decoupling layer and a high-frequency feature specific decoupling layer to extract the low-frequency features and high-frequency features of the multimodal data; The mean and variance of the low-frequency features and the high-frequency features are calculated, and the mean and variance are sampled by a Gaussian sampling method to obtain the final low-frequency features and high-frequency features.

2. The multimodal data fusion method according to claim 1, characterized in that: The dedicated modal encoder is used to extract high-order features from the multimodal data to be fused, so as to obtain high-order embedding features of the multimodal data, specifically: For the structured electronic health record, high-order feature extraction is performed by an encoder based on a combination of a multi-layer perceptron and a decision tree; For the preoperative electrocardiogram and intraoperative physiological monitoring data, Fourier transform and frequency domain attention mechanism are used to extract high-order features in the frequency domain; For medical imaging data, detailed features are extracted through deep convolutional networks.

3. The multimodal data fusion method according to claim 2, characterized in that: The feature decoupling module processes the high-order embedded features through a normalization layer, a multi-head attention mechanism layer, and a feed-forward fully connected layer in sequence to capture the potential relationship and feature relationship between the multimodal data, specifically: The high-order embedding features of each modality data are standardized by the normalization layer; given the embedding features , the normalization layer normalizes it by the following formula: in and They are The mean and standard deviation of Through the multi-head attention mechanism layer, the potential relationship between the high-order embedded features of each modal data is learned to mine the semantic information of each modal data; The output features of the multi-head attention mechanism layer are further extracted by introducing nonlinear transformation through the feed-forward fully connected layer to capture the characteristic relationship of each modality data.

4. The multimodal data fusion method according to claim 3, characterized in that: Based on the potential relationship and feature relationship between the multimodal data, the high-order embedded features are parsed through the low-frequency feature common decoupling layer and the high-frequency feature specific decoupling layer respectively to extract the low-frequency features and high-frequency features of the multimodal data, specifically: The low-frequency feature commonality decoupling layer uses the spectrum modulation attention mechanism of the weighted spectrum matrix to adaptively adjust the low-frequency components in the high-order embedded features, thereby enhancing the model's attention to and extraction of low-frequency features: in It is The absolute value of the spectrum of the input features, is the feature dimension, is the total number of features; represents the extracted low-frequency features, The feature vector output by the previous layer is obtained through three different linear transformations. is the dimension of the key vector; The high-frequency feature-specific decoupling layer uses a spectral modulation attention mechanism to adaptively adjust the high-frequency components in the high-order embedded features, emphasizing the modality-specific detail features: in, is the square of the spectral intensity in the input high-order embedded feature, indicating the size of the high-frequency component. Represents the extracted high-frequency features.

5. The multimodal data fusion method according to any one of claims 1 to 4, characterized in that: The cosine alignment algorithm is used to align the low-frequency features, and the orthogonal projection technology is used to perform orthogonal projection on the high-frequency features, specifically: The cosine similarity between the low-frequency features is calculated by the cosine alignment algorithm, and the alignment of the low-frequency features is constrained by the strategy of minimizing the cosine distance: in and Respectively represent and The low-frequency characteristics of the modal data, is the Euclidean norm of the vector; Combined with expert knowledge, the aligned low-frequency features are weighted by a multilinear adaptive weighting algorithm to obtain weighted low-frequency features. ; The weighted formula is: in, Based on expert knowledge Adaptively calculated weighting coefficients; All weighted low-frequency features are fused to obtain the fused low-frequency feature representation for: in represents the weight coefficient of low-frequency features, Indicates the number of modes; The high-frequency features are To standardize: in It is The high-frequency eigenvector of the normalized modal data, Represents the Euclidean norm of a vector; For high frequency features , projecting it outside the space of other modal features: in, It is The high-frequency features of the modal data after projection are represented by represents the inner product operation; The high-frequency features after orthogonal projection are linearly weighted Fusion is performed to obtain the final high-frequency feature representation : in is the weight coefficient of high-frequency features, Indicates the number of modes.

6. The multimodal data fusion method according to claim 5, characterized in that: The adaptive deep fusion network is used to perform multi-level fusion of the aligned low-frequency features and the orthogonally projected high-frequency features to obtain the fused multi-modal joint features, specifically: The fused low-frequency features and the final high-frequency features are subjected to one-dimensional temporal convolution and weighted fusion through the low-frequency-high-frequency spatiotemporal information collaborative sub-network to obtain the temporal features. ; Calculate the spatial interaction relationship between the fused low-frequency features and the final high-frequency features by point-by-point multiplication to obtain the spatial features ; The time feature and spatial characteristics Perform spatial consistency projection and normalization to obtain normalized time features and spatial characteristics ; The normalized time feature and spatial characteristics Perform weighted fusion to obtain unified multi-modal joint features : in is the learnable fusion weight coefficient.

7. A multimodal data fusion device, used to execute the multimodal data fusion method according to claim 1, characterized in that: include: Feature extraction module: used to extract high-order features from the multimodal data to be fused using a dedicated modal encoder to obtain high-order embedded features of the multimodal data; Feature decoupling module: used to decouple the high-order embedded features by combining expert knowledge constraints and multi-head attention mechanism, and obtain high-frequency features representing local information and low-frequency features representing global information respectively; Feature fusion module: used to align the low-frequency features using the cosine alignment algorithm, perform orthogonal projection on the high-frequency features using the orthogonal projection technology, and perform multi-level fusion of the aligned low-frequency features and the orthogonally projected high-frequency features through an adaptive deep fusion network to obtain fused multi-modal joint features; Risk prediction module: used to input the fused multimodal features into a pre-trained risk prediction model, and output the disease risk warning probability through the risk prediction model.

8. A computer device, characterized in that: The computer device includes a processor and a memory coupled to the processor, wherein: The memory stores program instructions for implementing the multimodal data fusion method according to any one of claims 1 to 6; The processor is used to execute the program instructions stored in the memory to control a multimodal data fusion method.

9. A computer storage medium, characterized in that Program instructions executable by a processor are stored, and the program instructions are used to execute the multimodal data fusion method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image-text data identification method, electronic equipment and storage medium

    CN117953525A

  • Multi-modal content detection method and system based on multi-head attention mechanism

    CN118709022A