A Feature Fusion Method Based on Multimodal Medical Data
The hybrid GNN-Transformer model effectively integrates multi-modal medical data by capturing structural and contextual information, addressing heterogeneity and redundancy, and enhancing diagnostic and treatment accuracy.
Patent Information
- Application Number
- CN202510121709.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-26
AI Technical Summary
Existing methods struggle to effectively integrate and utilize the diverse and complex information from multiple modalities of medical data, such as imaging, clinical records, and genetic data, due to heterogeneity, noise, redundancy, and complexity, limiting accurate and personalized medical diagnostics and treatments.
A hybrid architecture combining Graph Neural Networks (GNNs) with Transformer models for feature extraction and fusion, utilizing self-attention mechanisms to capture long-range dependencies and adaptively weight features based on their importance, while constructing multi-layered graph structures to model relationships in medical data.
Enhances the integration of multi-modal medical data by capturing structural and contextual information, reducing noise and redundancy, and improving the accuracy of medical analysis, diagnosis, and personalized treatment planning.
Smart Images

Figure CN119557840B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data processing, and in particular to a feature fusion method based on multimodal medical data. Background Art
[0002] The multimodality of medical data refers to the fact that the data generated in the medical environment has many different forms and sources. These data forms include but are not limited to: medical images (such as computed tomography (CT) ), magnetic resonance imaging ( ), ultrasound), pathology images, genomic data (e.g., Sequencing), clinical records (such as patient history, signs and symptoms), and physiological signals collected by patients' wearable devices (such as heart rate, blood oxygen saturation). The diversity of this data brings unprecedented challenges and opportunities.
[0003] Traditional medical diagnosis and treatment often rely on a single data modality, which makes it difficult for the present invention to fully understand the patient's condition. Scans can provide detailed anatomical information but lack information about metabolic activity, which can be obtained through Scans are obtained. Clinical records provide the patient's medical history and symptoms, but lack direct observation of changes at the pathological level. Therefore, how to effectively integrate these multimodal data to achieve more comprehensive, accurate and personalized medical services has become an important research topic.
[0004] On the one hand, multimodal medical data can provide richer complementary information, enabling the present invention to provide a deeper understanding and more accurate diagnosis of diseases. Multimodal and genomic data can reveal the complex mechanisms of diseases, thus providing a basis for developing more targeted treatment plans. On the other hand, multimodal medical data also brings new challenges. These challenges include:
[0005] Data heterogeneity: Data of different modalities have different structures, scales, and semantic information, which makes direct data fusion very difficult. For example, medical images are high-dimensional, while clinical records are structured text data.
[0006] Data noise and missing information: Medical data often contains noise and missing information, which can reduce the reliability of fusion results. For example, a patient’s medical history may be incomplete or inaccurate, and medical images may contain artifacts.
[0007] Data redundancy: There may be information redundancy in multimodal data. If this redundant information cannot be effectively removed, the efficiency of the fused feature representation will be reduced.
[0008] Complexity of fusion methods: How to design an efficient feature fusion method to fully utilize the complementary information in multimodal data and avoid information loss and noise introduction is still an open research problem.
[0009] In order to solve these problems, researchers have tried a variety of different fusion methods, including statistical methods, machine learning methods, and deep learning methods. However, these methods often have certain limitations when dealing with complex medical data. Therefore, exploring more effective multimodal data fusion methods has important theoretical significance and application value.
[0010] In recent years, deep learning technology has made significant progress in the field of multimodal data fusion. Deep learning models, especially and , has shown powerful capabilities in image processing and sequence modeling. However, traditional and It is difficult to effectively process data with complex structural relationships, such as social networks in medicine, gene regulatory networks, and anatomical connections between organs.
[0011] As an emerging deep learning model, it is specifically used to process graph structured data. Node representations can be learned by aggregating the neighbor information of nodes, so it is very suitable for modeling the relational structure in medical data. For example, the similarities between patients can constitute a patient network, the interactions between genes can constitute a gene regulatory network, and the anatomical connections between organs can constitute an anatomical structure graph. There are also extensive applications in multimodal data fusion, such as fusing social networks and health records, combining gene expression data and clinical information, or integrating medical images and patient metadata. It may be insufficient when dealing with sequence data or long-distance dependencies.
[0012] Architecture, as a deep learning model based on self-attention mechanism, The self-attention mechanism can capture the long-range dependencies between different positions in the sequence, which makes It has significant advantages in processing time series data and global context information. It is also widely used in the fields of image processing and multimodal data fusion. For example, The image is divided into a series of patches and these patches are input as sequence data into In the model, the image classification task is achieved. In terms of multimodal data fusion, researchers use Capture the relationship between different modalities and perform effective feature fusion.
[0013] In response to the challenges of multimodal fusion of medical data, this paper proposes a and The main contributions of this invention include:
[0014] 1. New - Hybrid architecture: This paper proposes a novel hybrid model that utilizes Capture the relational structure in medical data and pass the extracted features to To capture global context information. This hybrid model takes full advantage of the advantages of both architectures and overcomes the limitations of a single model in processing complex medical data.
[0015] 2. Adaptive feature aggregation mechanism: This paper designs an adaptive feature aggregation mechanism that can dynamically adjust according to the data modality. and The output feature weights are adjusted to achieve more accurate feature fusion. This mechanism avoids the tedious process of manually setting weights and improves the adaptability of the model to different modalities. Summary of the invention
[0016] In order to solve the problems in the prior art, the present invention proposes a feature fusion method and system based on multimodal medical data. In the first aspect, Figure 1 As shown, the present invention provides a feature fusion method based on multimodal medical data, and the steps are as follows:
[0017] Step 1: First, perform systematic data preprocessing and embedding on multimodal medical data to uniformly map different types of medical data into a high-dimensional vector space, including:
[0018] Step 1.1. Embedding processing of medical images:
[0019] First, the medical images are preprocessed by standardization, including normalization to , pixel values are normalized to Interval, and Z-score standardization. For a given medical image ,in Represents the height of the image, Represents the width of the image, represents the number of channels, using the pre-trained on ImageNet As a feature extractor Specifically, retain The first four residual blocks of , remove the last fully connected layer, and add a new dimension The fully connected layer obtains the feature vector:
[0020] ;
[0021] in, represent The feature extraction part of GAP represents the global average pooling operation. Represents the newly added fully connected layer parameters, Represents the final image embedding vector.
[0022] Step 1.2. Embedding process of clinical records:
[0023] The clinical text is first preprocessed, including word segmentation, stop word removal, punctuation processing, etc. For a given clinical record text sequence ,in Representative words, Represents the length of the text. First, we use the Word2Vec model pre-trained in the medical field. Embed each word:
[0024] ;
[0025] in, Expressive words The embedding vector of Represents the dimension of word embedding. Considering the particularity of medical text, the present invention uses a bidirectional LSTM network to process word embedding sequences:
[0026] ;
[0027] ;
[0028] ;
[0029] in, and They represent the hidden states of the forward and backward LSTMs respectively, and finally the text representation is obtained through the attention mechanism:
[0030] ;
[0031] ;
[0032] in, , , is a learnable parameter, is the output dimension of text features.
[0033] Step 1.3. Embedding processing of genomic data:
[0034] The gene expression data were first logarithmically transformed and standardized. ,in Representative The expression value of a gene, Represents the total number of genes, and uses a multilayer perceptron for nonlinear transformation:
[0035] ;
[0036] ;
[0037] ;
[0038] in, , , is the transformation matrix, represents the final embedding dimension of gene expression data. Batch normalization and (ratio 0.2).
[0039] Step 1.4. Embedding processing of patient clinical attributes:
[0040] Different processing strategies are used for the clinical attributes of patients, including numerical and categorical data such as age, gender, and weight. , perform minimum-maximum normalization:
[0041] ;
[0042] For categorical attributes such as gender , using a learnable embedding layer:
[0043] ;
[0044] in, is the embedding matrix, is the category embedding dimension. All attribute features are concatenated and passed through a fully connected layer to obtain the final attribute representation:
[0045] ;
[0046] in, It is the unified dimension of attribute characteristics.
[0047] Step 2. This method is carefully designed Model complex relationships in multimodal medical data and extract node feature representations that incorporate local structural information:
[0048] Step 2.1. Construction of multi-level graph structure:
[0049] First, three different levels of graph structures are constructed: patient similarity graph , Organ Association Diagram and molecular interaction diagrams For the patient similarity graph, edge weights are calculated using cosine similarity:
[0050] ;
[0051] in, and is the patient node feature, is the similarity threshold. For the organ association graph, an adjacency matrix is constructed based on anatomical knowledge:
[0052] ;
[0053] For molecular interaction graphs, the edge weight matrix is constructed using protein-protein interaction information in the STRING database. .
[0054] Step 2.2. Hierarchical graph attention feature extraction:
[0055] The graph attention network is used to extract features for each graph structure. First, the attention coefficient between nodes is calculated:
[0056] ;
[0057] ;
[0058] in, are the query and key transformation matrices, is the attention weight vector, Represents a splicing operation, is the hidden dimension of the attention layer. Then perform feature aggregation:
[0059] ;
[0060] in, is the value transformation matrix, is the ELU activation function, Represents the number of layers. Use the multi-head attention mechanism to expand the above calculation:
[0061] ;
[0062] in, is the number of attention heads, Represents a concatenation operation.
[0063] Step 2.3. Cross-graph structural information fusion:
[0064] In order to fuse information from different graph structures, a cross-graph attention mechanism is designed:
[0065] ;
[0066] ;
[0067] in, Representation Node In the graph structure The feature representation in and is a learnable parameter matrix. Finally, we get a node representation that integrates multi-level graph structure information. .
[0068] Step 3. This method is designed based on the improved The feature encoding module of the architecture captures the long-range dependencies and global semantic information between multimodal medical data through a multi-level self-attention mechanism:
[0069] Step 3.1. Position-aware feature encoding:
[0070] In order to preserve the position information in the sequence, an improved position encoding scheme is designed. First, the absolute position encoding is calculated:
[0071] ;
[0072] ;
[0073] in, Indicates location, Represents the dimension, is the model dimension. At the same time, relative position encoding is introduced:
[0074] ;
[0075] ;
[0076] in, is the learnable relative position embedding matrix, is the maximum relative distance, The function limits the relative distance to a reasonable range. The final input is expressed as:
[0077] ;
[0078] in, For The node feature matrix of Normalizes operations for a layer.
[0079] Step 3.2. Multi-scale self-attention mechanism:
[0080] Design a multi-scale self-attention mechanism to capture dependencies of different scopes. First, calculate the query, key, and value matrix:
[0081] ;
[0082] in, is the learnable parameter matrix, is the dimension of the attention head. Then calculate the attention scores of different scales:
[0083] ;
[0084] in, is the scale mask matrix, which is used to control the receptive field range of attention:
[0085] ;
[0086] in, are window sizes of different scales. The output of multi-head attention is:
[0087] ;
[0088] in, is the number of attention heads, is the output transformation matrix.
[0089] Step 3.3. Feedforward network and normalization:
[0090] After the self-attention layer, a two-layer feed-forward network is used for feature transformation:
[0091] ;
[0092] in, , is the weight matrix, is the hidden dimension of the feedforward network, and GELU is the activation function. The calculation process of the layer is:
[0093] ;
[0094] ;
[0095] in, The ratio is set to 0.1, and residual connections and layer normalization are used to stabilize the training process. Repeat this process with 4 layers. Block, get the final feature representation .
[0096] Step 4. This method proposes a dynamic and adaptive feature fusion mechanism to achieve intelligent fusion of medical data by learning the importance weights of different modality features, while taking into account the complementarity and redundancy of features:
[0097] Step 4.1. Feature preprocessing and standardization:
[0098] First, the features from and Features To standardize:
[0099] ;
[0100] ;
[0101] in, The ratio is 0.2, Used to ensure the consistency of feature distribution. Then the features are mapped to the same dimensional space through nonlinear transformation:
[0102] ;
[0103] ;
[0104] in, is the transformation matrix, A unified fusion dimension.
[0105] Step 4.2. Adaptive weight calculation:
[0106] Design an attention mechanism to calculate feature fusion weights. First, calculate the feature correlation matrix:
[0107] ;
[0108] The modality-specific attention weights are then calculated based on the relevance matrix:
[0109] ;
[0110] ;
[0111] in, is a learnable parameter, is the sigmoid function. In order to enhance the sparsity of weights, the temperature parameter is introduced :
[0112] ;
[0113] ;
[0114] Step 4.3. Multi-level feature fusion:
[0115] A hierarchical fusion strategy is adopted, and fusion is first performed at the feature level:
[0116] ;
[0117] Then the fusion is performed at the semantic level through a gating mechanism:
[0118] ;
[0119] ;
[0120] The final fusion feature is obtained by weighted combination:
[0121] ;
[0122] in, is a learnable balance parameter, initialized to 0.5. Finally, the final fusion feature is obtained through residual connection and layer normalization:
[0123] ;
[0124] Step 5. Model training and optimization, including multi-stage training, adaptive learning rate adjustment, and regularization techniques for training and optimization.
[0125] Through the above steps, this method achieves efficient fusion of multimodal medical data.
[0126] On the other hand, Figure 2 As shown, the present invention also proposes a feature fusion system based on multimodal medical data, which aims to improve the accuracy of medical diagnosis, prognosis analysis and personalized treatment plan formulation by effectively integrating medical data from different modalities. The system includes multiple modules that cooperate with each other to perform steps such as data input, preprocessing, feature extraction, feature fusion and analysis.
[0127] The system includes: data input module, data preprocessing module (including medical image preprocessing submodule, clinical record preprocessing submodule, genome data preprocessing submodule and patient attribute data preprocessing submodule), feature embedding module, graph structure construction module, Modules, Module, adaptive feature fusion module, and output module. These modules work together to process multimodal medical data and finally output the analysis results.
[0128] The data input module is responsible for receiving input from a variety of medical data sources, including but not limited to medical imaging data (such as , ), clinical record data, genomic data, and patient attribute data. This module has a data format check unit to verify whether the format of the input data meets the system preset requirements, and contains a data routing unit to route different types of data to the corresponding data preprocessing module. The data input module outputs the original data stream to the data preprocessing module.
[0129] The data preprocessing module receives the raw data stream from the data input module and performs corresponding preprocessing operations according to the data type. The module includes a medical image preprocessing submodule, which is used to perform operations such as image standardization, image resizing, image enhancement and artifact removal, and outputs the preprocessed medical image data stream; a clinical record preprocessing submodule, which is used to perform operations such as text cleaning, word segmentation, stem extraction or word form restoration, stop word removal and numerical feature processing, and outputs the preprocessed clinical record data stream; a genomic data preprocessing submodule, which is used to perform operations such as data standardization, gene selection, data conversion and data integration, and outputs the preprocessed genomic data stream; and a patient attribute data preprocessing submodule, which is used to perform operations such as numerical feature processing, category feature encoding and missing value processing, and outputs the preprocessed patient attribute data stream.
[0130] The feature embedding module receives the preprocessed data stream (215, 226, 235, 244) from the data preprocessing module and converts the preprocessed data of different modalities into vector representations. For example, the preprocessed medical image data is converted into a medical image embedding vector Use the following formula: ,in For medical imaging data, A pre-trained convolutional neural network is used to convert the pre-processed clinical record data into a clinical record embedding vector Using the word vector average method, the formula is: ,in is the length of the text, Indicates word embedding vector; convert the preprocessed genomic data into genomic data embedding vector Using the formula ,in For gene expression data, is a linear transformation matrix; and converts the preprocessed patient attribute data into a patient attribute embedding vector The feature embedding module outputs the embedding vector flow to the graph structure building module and module.
[0131] The graph structure building module receives the embedding vector stream from the feature embedding module and builds a graph structure based on the intrinsic connections of the medical data. This module treats the embedding vectors as graph nodes and defines edges based on the characteristics of the data. For example, in the patient similarity network, the similarity is calculated using the following formula, and then the edges are built based on the similarity: in , For patients and The characteristic vector of is the similarity; or in the anatomical network, the edges are defined according to the adjacency relationship of organs. The graph structure building module outputs the graph structure to module.
[0132] The module receives the graph structure from the graph structure building module and the embedding vector stream from the feature embedding module, and adopts The model learns the features of graph nodes. This module uses Compute Node With its neighbors The attention coefficient between , the formula is as follows: ,in , Representative Node and The embedding vector of , is a learnable weight matrix, is the attention calculation function; then, the calculated attention coefficient is used to aggregate the information of neighboring nodes and update the node The feature representation , the formula is: in is a learnable weight matrix, is the activation function. Module output node feature representation flow arrive module and adaptive feature fusion module.
[0133] The module receives The node feature representation flow of the module , and learn global context information. This module will The features output by the module are considered as sequence data , and perform position encoding. The position encoding calculation formula is as follows: as well as ,in and For location The position encoding vector and elements, Represents the location, Represents the dimension of the model, Represents the dimension of the encoding vector. The sequence after adding the position encoding is This module uses a multi-head self-attention mechanism to learn global dependencies, and its core calculation formula is: and , and perform residual connection and layer standardization operations. The module outputs a global context feature representation stream to the adaptive feature fusion module.
[0134] The adaptive feature fusion module receives The node feature representation flow of the module and from The module's global context feature representation flow , and dynamically merges and This module calculates The output weight Use the formula: ,in, yes feature, yes feature, and is a learnable parameter, yes function; then use the calculated weights to fuse the and The characteristics of are: ,in, is the fused feature. The adaptive feature fusion module outputs the fused feature representation stream to the output module.
[0135] The output module receives the fused feature representation stream from the adaptive feature fusion module , and according to the specific medical analysis tasks, perform operations such as classification or regression, and output the final analysis results.
[0136] The system first receives multimodal medical data from the data input module and routes it to the data preprocessing module. The data preprocessing module performs corresponding preprocessing operations according to the data type. Then, the preprocessed data is converted into vector representation by the feature embedding module. Next, the graph structure construction module constructs a graph structure based on the embedded features and passes it to the module. The module learns the features of graph nodes and passes them to module to learn global context information. The adaptive feature fusion module dynamically fuses and Finally, the output module receives the fused features and outputs the final results according to the medical analysis task.
[0137] The system can effectively integrate information from multiple modalities of medical data and improve the accuracy and efficiency of medical analysis tasks.
[0138] The beneficial technical effects of the present invention are as follows: by effectively integrating multiple modal medical data, the accuracy and efficiency of medical analysis are significantly improved. Capturing structural information of medical data, The global context information is extracted, and the weights of different modal features are dynamically adjusted using an adaptive feature fusion mechanism, thereby achieving a more accurate integration of multimodal information. Compared with traditional methods, this system can more comprehensively utilize the complementary information contained in medical data, reduce information redundancy and noise interference, and make the feature representation learned by the model more discriminative, thereby improving the performance of disease diagnosis, prognosis prediction, and personalized treatment plan formulation. In addition, the modular design of the system improves its scalability and flexibility, enabling it to adapt to medical data of different types and sizes, providing an efficient and reliable solution for the field of medical artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0139] Figure 1 A method flow chart of a feature fusion method based on multimodal medical data provided by the present invention.
[0140] Figure 2 A system architecture diagram of a feature fusion system based on multimodal medical data provided by the present invention. DETAILED DESCRIPTION
[0141] The following is a detailed introduction to the proposed The core idea of this method is to first use Model the structural information contained in multimodal medical data, and then The extracted structural features are passed to , in order to learn the global context information, and finally realize the effective integration of multimodal data through the adaptive feature fusion mechanism.
[0142] The purpose of multimodal data embedding is to transform different types of medical data into a unified vector representation for subsequent processing by the model. Different data modalities require different embedding methods.
[0143] For medical imaging data, for example scanning, etc., the present invention uses a pre-trained convolutional neural network ( ) as a feature extractor. Specifically, the present invention uses a pre-trained Models such as , or , as a feature extractor. The weights of the pre-trained model will be fixed or only fine-tuned to avoid problems such as too many training parameters and overfitting.
[0144] For a given medical image ,in , , Represent the height, width and number of channels of the image respectively. Model The output feature vector is
[0145] ;
[0146] in Representative image The embedding vector of Represents the dimension of image embedding.
[0147] Clinical records are usually in the form of text, such as medical history, diagnosis report, treatment plan, etc. The model is converted into a vector representation. Specifically, the present invention can adopt , or The pre-trained word vector model maps each word to a vector representation. For a given clinical record text ,in Representative words, Represents the length of the text. Model , each word can be converted into a corresponding word vector :
[0148] ;
[0149] in, Expressive words The embedding vector of Represents the dimension of word embedding. The embedding vector of the text record This can be obtained in a variety of ways, such as averaging all word vectors:
[0150] ;
[0151] Or use or Equivalent sequence models to learn contextual information of text records:
[0152] ;
[0153] in represents a sequence model, Represents the dimension of the text embedding vector.
[0154] Genomic data, such as gene expression data or gene variation data, usually exists in the form of numerical values or sequences. Or pre-trained The model is converted into a vector representation. For gene expression data ,in Representative The expression value of a gene, Represents the total number of genes. Embedding vector of gene expression data Direct mapping or linear transformation can be used:
[0155] ;
[0156] in represents a linear transformation matrix, Dimensions of the embedding vector representing gene expression data.
[0157] For gene sequence data, the present invention can be processed in a manner similar to word vectors. Each base in the gene sequence (such as A, T, C, G) is mapped to a vector, and then the gene sequence is processed using a sequence model to learn the context information of the gene sequence and obtain the final gene sequence embedding vector.
[0158] In addition to medical images, clinical records, and genomic data, patient demographic data (such as age, gender, race) and other attribute data also need to be embedded. For numerical data, the original data can be used directly or standardized. For categorical data, this paper adopts one-hot encoding or Convert it to a vector representation. For example, for patient age It can be normalized and then expressed as For patient gender , if it exists Gender can be expressed as Only one position is 1, and the rest are 0. Finally, the vectors of different attributes are concatenated or fused to obtain the embedding vector of the patient attribute. .
[0159] In summary, the embedding process of multimodal data can be achieved using a mapping function express:
[0160] ;
[0161] in, represents data of any modality, represents the embedding vector of the modality, Represents the dimension of the embedding vector of this modality.
[0162] After obtaining the embedding vectors of the multimodal data, the present invention constructs a graph structure to represent the relationship of the medical data. This graph structure can be constructed based on prior knowledge, such as the similarity between patients, the anatomical structure between organs, or the interaction between proteins. Assume that the graph constructed by the present invention is ,in Represents a collection of nodes, Represents a collection of edges. Each node Represents a data sample or a data feature. For example, a node can represent a patient, an organ, or a gene. Each node Each corresponds to an embedding vector , which is the initial feature of the node. Each edge Representative Node and The construction of the graph structure depends on the specific problem. For example, in a patient similarity network, edges can be constructed based on the clinical characteristics of the patients; while in an organ anatomy graph, edges can be constructed based on the connection relationship between organs.
[0163] The present invention adopts ( ) as The core of the model. It can adaptively learn the importance of nodes through the self-attention mechanism, thereby improving the expressiveness of the graph model. The core operation of is to aggregate the information of neighboring nodes through attention weights to update the feature representation of the target node. , the set of its neighbor nodes is , The information aggregation process is as follows:
[0164] First, compute the nodes and its neighbor nodes The attention coefficient between :
[0165] ;
[0166] in, and Representative Node and The embedding vector of and represents the learnable weight matrix, Represents the attention calculation function, such as the inner product or a single-layer feedforward neural network, which is used to calculate the similarity between nodes. The output is a scalar value.
[0167] Then, using the calculated attention coefficient , weighted aggregation of neighbor node information is performed to obtain node Updated feature representation :
[0168] ;
[0169] in, is a learnable weight matrix, represents the activation function, for example or . is the updated node representation. Network, the present invention can aggregate the information of multi-hop neighbors, so as to better learn the feature representation of nodes. The output of the model is the feature representation of all nodes , is the number of nodes in the graph.
[0170] exist After the module extracts the graph structure information, the present invention will Output Input to In order to use Model, the present invention needs to Output Specifically, the present invention can directly convert the feature vector of each node into As an input of one time step, that is:
[0171] ;
[0172] in is the number of nodes, is the dimension of the node feature.
[0173] The present invention adopts the standard The encoder structure consists of multiple layers (Multi-Head Attention) and (feedforward neural network). Each layer of , first calculates the self-attention of the input sequence , and its calculation formula is as follows:
[0174] ;
[0175] in , , Represent the query vector, key vector and value vector respectively, which can be obtained by inputting the vector After different linear transformations, we get:
[0176] ;
[0177] ;
[0178] ;
[0179] in, , , is a learnable weight matrix, is the dimension of the key vector. Multi-head attention is obtained by concatenating the calculation results of multiple independent attention heads and linearly transforming them:
[0180] ;
[0181] in,
[0182] ;
[0183] in represents the number of attention heads, , , , All of them are learnable parameter matrices. After the multi-head attention layer, it passes through a feedforward neural network to further perform nonlinear transformation on the features, as well as a residual connection and layer normalization operation. The output is recorded as , Represents the encoded feature vector for each position.
[0184] In getting and After the output of , the present invention needs to effectively fuse the two features. Considering the requirements of different data modalities and different tasks, the present invention designs an adaptive feature aggregation module, which can dynamically adjust the features from and The weights of the features of the model.
[0185] For each node, the present invention obtains its Output and Output In order to achieve adaptive feature fusion, the present invention first and splice and input it into a simple and generate a fusion weight :
[0186] ;
[0187] in, represents the concatenation of vectors, and are learnable weights and biases, represent Activation function, which limits the output to between 0 and 1. express The output weight of The output weight of .
[0188] The final fusion feature It can be calculated as:
[0189] ;
[0190] Through the adaptive aggregation mechanism, the present invention can dynamically adjust the and The importance of features can be improved to achieve more accurate feature fusion.
[0191] Finally, the present invention can fuse features Input into the subsequent classification or regression module to perform specific medical analysis tasks.
[0192] As mentioned earlier, It is the core component of this method, and its main function is to use the graph structure of medical data to extract structural features. The construction method of the graph structure directly affects The performance of the model therefore needs to be carefully designed according to specific medical problems.
[0193] The graph structure in medical data can be constructed based on the intrinsic relationship of the data and medical expertise. The present invention can use the following graph network:
[0194] 1. Patient similarity network: In this network, each node represents a patient, and the edge represents the degree of similarity between patients. The similarity of patients can be calculated based on clinical characteristics, such as age, gender, medical history, symptoms, or laboratory test results. The similarity can be measured using Euclidean distance, cosine similarity, or other statistical distances. Specifically, if the clinical feature vector between two patients is and , then the similarity between them The calculation can be performed, for example, as follows:
[0195] ;
[0196] in, represents the dot product of vectors, Represents the Euclidean distance of vectors and measures the similarity between two vectors through cosine similarity.
[0197] When the similarity exceeds a certain threshold, the present invention can add an edge between two patients. The network can capture similar relationships between patient groups, for example, it can identify patient clusters with similar symptoms or medical histories to assist in diagnosis.
[0198] 2. Anatomical network: In the anatomical network, each node represents an organ or an anatomical region, and the edge represents the connection between organs. This connection relationship can be obtained from medical images, such as Scan or For example, the present invention may use the adjacency relationship of organs to define the edges of the graph, or use the vascular connections or neural connections between organs as the weights of the edges.
[0199] For example, if the adjacent relationship between the liver (L) and the gallbladder (G) is defined as 1, otherwise it is 0, it can be expressed as
[0200] ;
[0201] The relationship between anatomical structures is represented using an adjacency matrix.
[0202] This network can capture the interactions between different organs and assist in disease localization and functional analysis.
[0203] 3. Biological pathway network: In a biological pathway network, each node represents a gene, protein or metabolite, and the edges represent the interaction between biological molecules, such as gene regulation, protein-protein interaction and metabolic pathways. These interaction relationships can be obtained from biological databases, such as or This network can reveal the molecular mechanisms of diseases and provide a basis for drug discovery and personalized treatment.
[0204] For example, in a gene expression data, there is an interaction between gene A and gene B, then the interaction relationship between these two genes can be expressed as
[0205] ;
[0206] The adjacency matrix was used to represent the interaction relationship between genes.
[0207] In practical applications, the present invention can select a suitable graph structure according to specific research questions. In addition, different graph structures can be used simultaneously, and their outputs can be fused to more comprehensively describe the complexity of medical data.
[0208] After the graph structure is constructed, the present invention needs to select a suitable To learn node feature representation. In addition to the previously mentioned Model, the present invention can also consider other Models, for example:
[0209] 1. Graph Convolutional Network: It is a classic Model, which updates node features by weighted averaging the information of neighboring nodes. The information aggregation process can be expressed as:
[0210] ;
[0211] in, Represents neighbor node The embedding vector of represents the learnable weight matrix, and Represents nodes and The degree of is used for normalization, Representative Node Update representation. represents the activation function. compared to, Using fixed weights for information aggregation makes it impossible to adaptively adjust the weights according to the importance of neighbor nodes.
[0212] 2. Graph Attention Network: As mentioned above, The self-attention mechanism is used to learn the weights between nodes, allowing the model to pay more attention to important neighbor nodes. The information aggregation process is as follows:
[0213] First calculate the attention coefficient:
[0214] ;
[0215] in, and Representative Node and The embedding vector of and represents the learnable weight matrix, Represents the attention calculation function.
[0216] Then use the attention coefficient to aggregate the information of neighboring nodes:
[0217] ;
[0218] in, is a learnable weight matrix, represents the activation function, is the updated node representation.
[0219] 3. Graph Isomorphic Network: It is a powerful model, which is able to distinguish different graph structures. The information aggregation process is as follows:
[0220] ;
[0221] in, represents the feature vector of the current node, represents the feature vector of the neighboring nodes, and represents the linear transformation matrix, represents a multi-layer perceptron, is a learnable parameter.
[0222] By summing the information of neighbor nodes instead of averaging, and introducing a learnable parameter , making the model more expressive.
[0223] In practical applications, the present invention can often select appropriate For graph structures with complex node relationships, the present invention uses a more expressive model. or model, while for graph structures with simple node relationships, The present invention can also be used to combine different Models are used together, for example, first use Perform initial feature extraction and then use or Perform more advanced feature learning.
[0224] As mentioned before, the adaptive feature aggregation module aims to dynamically fuse and The characteristics of the model can fully utilize the complementary information of multimodal data.
[0225] For each node ,That The output features are expressed as , The output features are expressed as In order to achieve adaptive feature aggregation, the present invention first and concatenate, and pass a linear transformation and a Activation function to calculate feature fusion weights :
[0226] ;
[0227] in represents the concatenation of vectors, is a learnable weight matrix, is a learnable bias, and Respectively represent and Dimensions of the output feature vector. represent function, which compresses the output value to In the range, used to indicate The output weight of . Correspondingly, The output weight of .
[0228] Using the calculated weights , the present invention will Output and Output Perform weighted fusion to obtain the final fusion feature representation :
[0229] ;
[0230] The fusion process can adaptively adjust the and The weight of the output feature. If The value of is close to 1, indicating that the characteristics of the current node are more dependent on The output, on the contrary, if The value of is close to 0, which means that the characteristics of the current node are more dependent on If is 0.5, then and The advantage of this adaptive fusion is that it can better utilize the information of multimodal data, avoid the tediousness of manual selection of weights, and have better generalization ability.
[0231] The expressiveness of a model refers to the degree to which the model can learn complex functions. A model with strong expressiveness can better fit the training data and generalize to unseen data. For multimodal data fusion tasks, the present invention hopes that the model can learn the complex relationships between different modalities and extract useful features. The model of the present invention combines and Architecture, both models have powerful expressive power.
[0232] The model can learn the feature representation of graph structured data through a message passing mechanism. The model can approximate any continuous function under certain conditions. For example, (Graph Isomorphism Networks) can distinguish different graph structures under certain conditions. This means that the model of the present invention can capture complex structural relationships in medical data, such as patient networks, biological pathways, and anatomical structures. Specifically, Node representation learned by the model It can be regarded as a comprehensive representation of the information of a node and its neighboring nodes. As the number of layers increases, the node representation will contain information about more distant neighboring nodes, allowing the model to learn more global structural information.
[0233] The model can learn long-distance dependencies of sequence data through the self-attention mechanism. The expressive power of has been proven by a lot of research. In the multi-head attention mechanism, the model can learn information from multiple subspaces to capture more comprehensive features. The output is The model of the present invention can further capture the long-distance dependencies between structural features, thereby effectively combining local information and global information. The feedforward neural network layer in the model can learn more complex nonlinear transformations, thereby improving the overall expressiveness of the model.
[0234] In addition, the adaptive feature aggregation module proposed in the present invention also enhances the expressiveness of the model. and The model can adaptively adjust the importance of different features based on the characteristics of the input data by weighted combination of the output features. This adaptability makes the model more flexible in processing different data modalities.
[0235] In summary, the model of the present invention combines The ability to extract structural information and The global context learning capability of the proposed model and the use of an adaptive feature fusion mechanism ensure that the model has strong expressive power and can handle complex multimodal medical data.
[0236] The convergence of a model refers to whether the model can stably converge to a local optimal solution or a global optimal solution during the training process. A converged model can produce reliable results and can be generalized to unseen data. The model of the present invention uses stochastic gradient descent (Stochastic Gradient Descent, ) or other optimization algorithms.
[0237] The model is usually trained using a backpropagation method. The convergence of the model is affected by the network structure and parameter initialization. The convergence speed of the model can be improved by using some techniques, such as Activation function, using or Perform standardization operations or use residual connections.
[0238] The model also has convergence issues during training, especially when the model is deep. In order to achieve a high convergence speed, the present invention adopts techniques such as multi-head attention mechanism, residual connection and layer normalization. These techniques can effectively avoid the gradient vanishing and gradient exploding problems and accelerate the convergence of the model. In addition, a reasonable learning rate setting is also crucial for the convergence of the model.
[0239] The convergence of the adaptive feature aggregation module mainly depends on its internal linear transformation and Activation function. Since the structure of these modules is relatively simple, their convergence is relatively good. In addition, the present invention also adopts weight initialization strategy and gradient clipping techniques to ensure that the model can converge stably.
[0240] In general, the model of the present invention adopts a variety of techniques during training, thereby ensuring that the model has good convergence. Experimental results also prove that the model can stably converge to a good local optimal solution.
[0241] Feature fusion is a key step in multimodal data fusion, and its purpose is to integrate data from different modalities to obtain a more comprehensive feature representation. Effective feature fusion should be able to retain useful information from different modalities and eliminate noise in the information. The model of the present invention adopts an adaptive feature aggregation method to perform feature fusion.
[0242] As mentioned above, the present invention assumes represents the node features obtained from the graph structure information, Represents the node features obtained from the global context information. Adaptive feature fusion weight The calculation formula is:
[0243] ;
[0244] The final fusion feature is expressed as:
[0245] ;
[0246] The present invention analyzes this feature fusion process. In an ideal situation, and Complementary information can be captured, and all of them are task-related features. and are two random vectors. If the correlation between the two vectors is 0, that is, they are mutually orthogonal, then the fused features can contain complete information from two different vectors. However, if the two vectors are highly correlated, the fusion may not bring additional information gain. In actual medical data, and Usually not orthogonal, so the present invention hopes to adaptively fuse the weights Can be based on and The correlation between the two is adjusted adaptively. When the correlation between the two is low, The value of will be close to 0.5, thus balancing the two features. When there is a large correlation between the two features, It can automatically learn to make one of them have a high weight and the other one have a low weight, thus avoiding feature redundancy. This fusion strategy can effectively retain the complementary information of different modalities.
[0247] In addition, the present invention uses Function as activation function, so The value range is arrive This can be regarded as a gating mechanism, which makes the fusion feature more flexible to adapt to different data.
[0248] From the perspective of information theory, the goal of feature fusion is to reduce the mutual information redundancy between different modalities while retaining important information. The present invention hopes that the fused features can provide richer information than the features of a single modality. The model of the present invention can achieve this goal by learning adaptive fusion weights.
[0249] The robustness of a model refers to the stability of the model in the face of noisy data, missing data, or changes in data distribution. Medical data usually has noise and missing data, so the present invention hopes that the model has good robustness.
[0250] The model can learn node representation by aggregating information from neighboring nodes. This aggregation process can reduce the noise impact of a single node and improve the robustness of the model. In addition, the present invention can use some graph data enhancement methods, such as randomly deleting nodes or edges, to further improve Robustness of the model.
[0251] The model uses a self-attention mechanism to learn global context information. The self-attention mechanism can capture long-distance dependencies in the data, thus having a certain robustness to noisy data. In addition, the present invention can be used Technique, randomly masking the outputs of some neurons to further improve the robustness of the model.
[0252] The robustness of the adaptive feature aggregation module mainly depends on its linear transformation and Activation function. In order to improve the robustness of the fusion module, the present invention can use Regularization technology reduces the risk of overfitting of the model. In addition, the present invention can also improve the overall robustness of the model through data preprocessing steps, such as data standardization and data cleaning.
[0253] In summary, the model proposed in this paper adopts a variety of techniques to improve its robustness, including graph aggregation, self-attention mechanism, , regularization and data preprocessing. The experimental results also prove that the model of the present invention can have better performance when facing noisy data and missing data.
[0254] In order to evaluate the proposed The present invention conducts experiments on multiple datasets and compares them with some classic multimodal fusion methods.
[0255] Due to the sensitivity and privacy of medical data, the present invention uses some publicly available medical datasets and performs appropriate preprocessing and formatting on these data. The present invention considers the following representative multimodal medical datasets:
[0256] 1. Multimodal brain tumor dataset: This dataset contains patients’ Images (including , , and Four modalities), clinical information (such as patient age, gender, tumor type, etc.), and genomic data (such as gene expression profiles). The goal of this invention is to use these multimodal data to predict patient survival time or tumor histological grade. This dataset is mainly used to verify the performance of the model in integrating multiple imaging data, clinical data, and genomic data.
[0257] Multimodal Brain Tumor Dataset:
[0258] ;
[0259] 2. Multimodal heart disease dataset: This dataset contains patients’ Signal (ECG), images (echocardiograms), and clinical information (such as patient blood pressure, cholesterol levels, etc.). The goal of this invention is to predict whether a patient has heart disease and to predict the risk of a heart attack. This dataset is used to validate the model's performance in integrating time series signals, imaging data, and clinical data.
[0260] Multimodal Heart Disease Dataset:
[0261] ;
[0262] 3. Multimodal lung disease dataset: This dataset contains patients’ Scanned images, pathological images, and clinical records (such as patient smoking history, cough symptoms, etc.). The goal of this invention is to use these multimodal data to predict the probability of lung cancer and distinguish different types of lung diseases. This dataset is designed to test the model's ability to integrate image, pathology, and text data.
[0263] Multimodal Lung Disease Dataset:
[0264] ;
[0265] These datasets cover a variety of different medical data modalities and medical analysis tasks, and can comprehensively evaluate the performance of the multimodal fusion method of the present invention. It should be noted that these datasets are fictitious and are only used for experimental description.
[0266] For different medical analysis tasks, the present invention needs to use different evaluation indicators. The present invention mainly uses the following common evaluation indicators:
[0267] 1. Classification task: For the classification task, the present invention uses the following indicators for evaluation:
[0268] Accuracy: The ratio of correctly classified samples to the total number of samples.
[0269] ;
[0270] in, represents a true positive, represents true negative, represents a false positive, Represents a false negative.
[0271] Precision: The proportion of samples that are actually positive among all samples predicted to be positive.
[0272] ;
[0273] Recall: The proportion of samples that are correctly predicted as positive examples among all samples that are truly positive examples.
[0274] ;
[0275] F1-Score: The harmonic mean of precision and recall.
[0276] ;
[0277] Area Under the Receiver Operating Characteristic Curve ): The area under the curve is used to evaluate the overall performance of the binary classification model.
[0278] 2. Regression task: For regression tasks, the present invention uses the following indicators for evaluation:
[0279] Mean Squared Error ): The average of the squares of the errors between the predicted values and the true values.
[0280] ;
[0281] in, Representative The true value of the samples, Representative The predicted value of samples, is the total number of samples.
[0282] Root Mean Squared Error ): The square root of the mean square error.
[0283] ;
[0284] Mean Absolute Error ): The average of the absolute values of the errors between the predicted values and the true values.
[0285] ;
[0286] Coefficient of determination ( ): It is used to measure how well the regression model fits the data. The value range is 0 to 1. The larger the value, the better the model fits.
[0287] ;
[0288] in, Represents the average of all true values.
[0289] The present invention selects appropriate evaluation indicators according to specific medical analysis tasks. For example, for tumor grading tasks, the present invention uses , , ,and For the task of predicting patient survival time, the present invention uses , , as well as .
[0290] The present invention divides the data set into a training set, a validation set, and a test set in a ratio of 7:1.5:1.5. The present invention will use the validation set to adjust the model parameters and use the test set to evaluate the final performance of the model.
[0291] The configuration of the present invention is as follows:
[0292] Model: The present invention uses As . The number of layers is 2, the number of hidden units in each layer is 128, and the number of multi-head attention heads is 8. As an activation function. The present invention uses The optimizer performs parameter updates.
[0293] Model: The present invention uses Encoder as a sequence model. The number of layers is 2, the number of heads of multi-head attention is 8, and the dimension of the hidden layer is 256. The dimension of is 512. The present invention also uses As an activation function. The present invention uses The optimizer performs parameter updates.
[0294] Learning rate: The present invention uses 0.001 as the initial learning rate. In order to improve the stability of model training, the present invention uses a learning rate decay strategy, that is, when the loss on the validation set no longer decreases, the learning rate will gradually decrease.
[0295] Batch Size: The present invention uses 32 as the Batch Size.
[0296] Epoch: The model of the present invention is trained for 200 epochs.
[0297] Regularization: To prevent overfitting, this paper uses and L2 regularization techniques. The probability is set to 0.5. The weight decay coefficient of L2 regularization is 0.0001.
[0298] The parameters of the model will be fine-tuned according to the specific data set to achieve optimal performance.
[0299] In order to evaluate the proposed The performance of the model is compared with the following classic multimodal fusion models:
[0300] 1. Early Fusion: Directly concatenate the original data of different modalities, and then input the concatenated data into a traditional machine learning model or a deep learning model for training. This method is simple and direct, but it may lead to performance degradation due to differences in the scale and structure of data of different modalities.
[0301] 2. Late Fusion: First, extract features from different modal data separately, then concatenate or weightedly fuse the extracted features, and finally perform classification or regression. This method can retain the uniqueness of different modalities, but may not fully utilize the interaction between modalities.
[0302] 3. Based on Multimodal fusion: Using convolutional neural networks ( ) extracts features from different modalities and then fuses the features using concatenation or attention mechanisms. They excel in image processing tasks, but may have limitations when processing other types of data.
[0303] 4. Based on Multimodal fusion: Using recurrent neural networks ( ) processes sequence data, such as time series signals or text data, and then fuses features using concatenation or attention mechanisms. There is a gradient vanishing problem when processing long sequence data.
[0304] 5. Based on Multimodal fusion: Mapping data of different modalities to The self-attention mechanism is used in the model to fuse the long-range dependencies between different modalities.
[0305] 6. Based on Multimodal fusion: Convert multimodal data into a graph structure and use Perform feature extraction and fusion. This method can utilize the structural information in multimodal data, but may not fully utilize the global context information.
[0306] 7. : Based on the maximum average difference ( ) collaborative feature matching method, which maps the features of different modalities into the same latent space and minimizes To align the feature distributions of different modalities.
[0307] These models represent different multimodal fusion strategies and can be used to evaluate the proposed The performance of the model. The present invention will select representative models with similar model architectures to the present invention for comparison, so as to better illustrate the advantages of the model. In the experiment, the present invention will use the same data preprocessing steps and optimization strategies to train different models, so as to ensure the fairness of the experiment.
[0308] The present invention conducts experiments on three multimodal medical datasets and evaluates the models using different evaluation metrics. Table 1, Table 2, and Table 3 show the performance of different models on the three datasets.
[0309] Table 1 Performance comparison on multimodal brain tumor dataset:
[0310] ;
[0311] Table 1 shows the performance of different models on the multimodal brain tumor dataset. The model outperforms other comparison models in all evaluation indicators. In particular, , ,as well as In terms of indicators, the model of the present invention has significant advantages, indicating that the method proposed in the present invention can more effectively integrate medical data of multiple modalities, thereby improving the accuracy of prediction.
[0312] Table 2 Performance comparison on multimodal heart disease dataset:
[0313] ;
[0314] Table 2 shows the performance of different models on the multimodal heart disease dataset. The model of the present invention achieved the highest score in all indicators, which shows that the method of the present invention can effectively handle the fusion problem of time series signals, imaging data and clinical data. In terms of indicators, the model of the present invention has more significant advantages than other comparison models, which shows that the model of the present invention has stronger discrimination ability.
[0315] Table 3 Performance comparison on multimodal lung disease datasets:
[0316] ;
[0317] Table 3 shows the performance of different models on the multimodal lung disease dataset. Similar to the experimental results of the previous two datasets, the model of the present invention outperforms other models in all evaluation indicators. This further verifies the effectiveness and robustness of the method proposed in the present invention. It can be seen that , ,as well as The performance on the three datasets is better than that of the traditional early fusion and late fusion methods, proving that deep models have more advantages when processing multimodal medical data. and advantages, thus achieving the best performance.
[0318] The results show that the present invention is based on The multimodal fusion method of the present invention has achieved significant performance improvement on three multimodal medical data sets. Compared with the traditional early fusion and late fusion methods, the method of the present invention can more effectively utilize the complementary information in the multimodal data, thereby improving the accuracy of prediction. At the same time, the model performance of the present invention is also better than other fusion methods based on deep learning, such as , , as well as This illustrates the advantages of the hybrid model architecture proposed in this paper and the effectiveness of the adaptive feature fusion mechanism.
[0319] Specifically, The module is able to capture structural information in medical data, such as similarities between patients, anatomical connections between organs, and interactions between biomolecules. This structural information plays an important role in medical data analysis. The module can capture global context information, allowing the model to learn more comprehensive feature representations. Through the adaptive feature fusion module, the model can dynamically adjust and The weights of the output features are used to achieve more accurate feature fusion.
Claims
1. A feature fusion method based on multimodal medical data, characterized in that: The method comprises the following steps: Step 1. First, systematically preprocess and embed the multimodal medical data to uniformly map different types of medical data into a high-dimensional vector space. Step 2. Pass Model the relationships in multimodal medical data and extract node feature representations that incorporate local structural information; Step 3. Based on the improved The feature encoding module of the architecture captures the dependencies and global semantic information between multimodal medical data through a multi-level self-attention mechanism; Step 4: Through the dynamic adaptive feature fusion mechanism, learn the importance weights of different modality features to achieve intelligent fusion of medical data, specifically: Step 4.
1. Feature preprocessing and standardization: First, from Features and Features To standardize: ; ;in, The ratio is 0.2, Used to ensure the consistency of feature distribution; then the features are mapped to the same dimensional space through nonlinear transformation: ; ;in, is the transformation matrix, For the unified integration dimension; Step 4.
2. Adaptive weight calculation: Design an attention mechanism to calculate feature fusion weights; first calculate the feature correlation matrix: ; Then the modality-specific attention weights are calculated based on the correlation matrix: ; ;in, is a learnable parameter, is the sigmoid function; in order to enhance the sparsity of weights, the temperature parameter is introduced : ; ; Step 4.
3. Multi-level feature fusion: A hierarchical fusion strategy is adopted, and fusion is first performed at the feature level: ; Then fusion is performed at the semantic level through a gating mechanism: ; ; The final fusion feature is obtained by weighted combination: ;in, is a learnable balance parameter, initialized to 0.5; finally, the final fusion feature is obtained through residual connection and layer standardization: ; Step 5. Model training and optimization, including multi-stage training, adaptive learning rate adjustment, and regularization for training and optimization.
2. A feature fusion method based on multimodal medical data according to claim 1, characterized in that: Step 1 specifically includes: Step 1.
1. Embedding processing of medical images: First, the medical images are preprocessed by standardization, including normalization to , pixel values are normalized to Interval, and Z-score standardization; for a given medical image ,in Represents the height of the image, Represents the width of the image, Represents the number of channels, using the pre-trained on ImageNet As a feature extractor Specifically, retain The first four residual blocks of , remove the last fully connected layer, and add a new dimension The fully connected layer obtains the feature vector: ; in, represent The feature extraction part of GAP represents the global average pooling operation. Represents the newly added fully connected layer parameters, Represents the final image embedding vector; Step 1.
2. Embedding process of clinical records: The clinical text is first preprocessed, including word segmentation, stop word removal, punctuation processing, etc. For a given clinical record text sequence ,in Representative words, Represents the length of the text. First, we use the Word2Vec model pre-trained in the medical field. Embed each word: ; in, Expressive words The embedding vector of Represents the dimension of word embedding; considering the particularity of medical text, the present invention adopts a bidirectional LSTM network to process the word embedding sequence: ; ; ; in, and They represent the hidden states of the forward and backward LSTMs respectively, and finally the text representation is obtained through the attention mechanism: ; ;in, , , is a learnable parameter, is the output dimension of text features; Step 1.
3. Embedding processing of genomic data: The gene expression data were first logarithmically transformed and standardized; ,in Representative The expression value of a gene, Represents the total number of genes, and uses a multilayer perceptron for nonlinear transformation: ; ; ;in, , , is the transformation matrix, represents the final embedding dimension of gene expression data; batch normalization and (ratio 0.2); Step 1.
4. Embedding processing of patient clinical attributes: Different processing strategies are used for the clinical attributes of patients, including numerical and categorical data such as age, gender, and weight. , perform minimum-maximum normalization: ; For categorical attributes such as gender , using a learnable embedding layer: ;in, is the embedding matrix, is the category embedding dimension; all attribute features are concatenated and passed through a fully connected layer to obtain the final attribute representation: ;in, It is the unified dimension of attribute characteristics.
3. The feature fusion method based on multimodal medical data according to claim 1, characterized in that: Step 2 specifically includes: Step 2.
1. Construction of multi-level graph structure: First, three different levels of graph structures are constructed: patient similarity graph , Organ Association Diagram and molecular interaction diagrams ; For the patient similarity graph, edge weights are calculated using cosine similarity: ;in, and is the patient node feature, is the similarity threshold; for the organ association graph, the adjacency matrix is constructed based on anatomical knowledge: ; For molecular interaction graphs, the edge weight matrix is constructed using protein-protein interaction information in the STRING database ; Step 2.
2. Hierarchical graph attention feature extraction: For each graph structure, the graph attention network is used for feature extraction; first, the attention coefficient between nodes is calculated: ; ;in, are the query and key transformation matrices, is the attention weight vector, Represents a splicing operation, is the hidden dimension of the attention layer; then feature aggregation is performed: ;in, is the value transformation matrix, is the ELU activation function, Represents the number of layers; the above calculation is expanded using the multi-head attention mechanism: ;in, is the number of attention heads, Represents a splicing operation; Step 2.
3. Cross-graph structural information fusion: In order to fuse information from different graph structures, a cross-graph attention mechanism is designed: ; ;in, Representation Node In the graph structure The feature representation in and is a learnable parameter matrix; finally, we get a node representation that integrates multi-level graph structure information .
4. The feature fusion method based on multimodal medical data according to claim 1, characterized in that: Step 3 specifically includes: Step 3.
1. Position-aware feature encoding: In order to preserve the position information in the sequence, an improved position encoding scheme is designed; first, the absolute position encoding is calculated: ; ;in, Indicates location, Represents the dimension, is the model dimension; relative position encoding is introduced at the same time: ; ;in, is the learnable relative position embedding matrix, is the maximum relative distance, The function limits the relative distance to a reasonable range; the final input is expressed as: ;in, For The node feature matrix of Standardize operations for layers; Step 3.
2. Multi-scale self-attention mechanism: Design a multi-scale self-attention mechanism to capture dependencies at different scales; first calculate the query, key, and value matrix: ;in, is the learnable parameter matrix, is the attention head dimension; then calculate the attention scores of different scales: ;in, is the scale mask matrix, which is used to control the receptive field range of attention: ;in, are window sizes of different scales; the output of multi-head attention is: ;in, is the number of attention heads, is the output transformation matrix; Step 3.
3. Feedforward network and normalization: After the self-attention layer, a two-layer feed-forward network is used for feature transformation: ;in, , is the weight matrix, is the hidden dimension of the feedforward network, GELU is the activation function; The calculation process of the layer is: ; ;in, The ratio is set to 0.1, residual connections and layer normalization are used to stabilize the training process; 4 layers are repeated. Block, get the final feature representation .
5. A feature fusion system based on multimodal medical data, characterized in that: The system is used to execute the method described in any one of claims 1 to 4, and the system further includes: a data input module, a data preprocessing module, a feature embedding module, a graph structure building module, Modules, module, adaptive feature fusion module, and output module; these modules work together to process multimodal medical data and finally output the analysis results.
6. The system according to claim 5, characterized in that Said The module receives the graph structure from the graph structure building module and the embedding vector stream from the feature embedding module, and adopts The model learns graph node features; this module uses Compute Node With its neighbors The attention coefficient between , the formula is as follows: ,in , Representative Node and The embedding vector of , is a learnable weight matrix, is the attention calculation function; then, the calculated attention coefficient is used to aggregate the information of neighboring nodes and update the node The feature representation , the formula is: ,in is a learnable weight matrix, is the activation function; Module output node feature representation flow arrive module and adaptive feature fusion module.
7. The system according to claim 6, characterized in that The module receives The node feature representation flow of the module , and learn global context information; This module will The features output by the module are considered as sequence data , and perform position encoding. The position encoding calculation formula is as follows: as well as ,in and For location The position encoding vector and elements, Represents the location, Represents the dimension of the model, Represents the dimension of the encoding vector. The sequence after adding the position encoding is ; This module uses a multi-head self-attention mechanism to learn global dependencies, and its core calculation formula is: and , and perform residual connection and layer standardization operations; The module outputs a global context feature representation stream to the adaptive feature fusion module.
8. The system according to claim 7, characterized in that The adaptive feature fusion module receives The node feature representation flow of the module and from The module's global context feature representation flow , and dynamically merges and The module calculates the characteristics of The output weight Use the formula: ,in, yes feature, yes feature, and is a learnable parameter, yes function; then use the calculated weights to fuse the and The characteristics of are: ,in, The adaptive feature fusion module outputs the fused feature representation stream to the output module.
Citation Information
Patent Citations
Multi-modal medical image fusion based on expansion convolution and attention GCN
CN117392494A
Multi-modal medical image prediction method and device based on graph neural network, medium and product
CN118674701A