Method and system for progressive fusion of multimodal medical data
Through the multimodal medical data gradual fusion method, the cross-layer attention mechanism and dynamic graph learning are used to solve the problems of information abandonment and similarity calculation in the existing technology, and high-accuracy information transmission and adaptability enhancement are achieved.
Patent Information
- Application Number
- CN202411032143.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-07-30
AI Technical Summary
The existing multimodal medical data fusion methods tend to discard complementary information between modals during the fusion process, and the static graph representation and heuristic graph enhancement methods are unstable when calculating patient similarity.
The multimodal medical data gradual fusion method is adopted to maximize the hierarchical information and the accuracy of information transmission by extracting multi-layer features, cross-layer attention mechanism fusion and dynamic graph learning.
Through the progressive fusion method and cross-layer attention mechanism, the inclusion of hierarchical information is maximized, dynamic graph learning reduces the heterogeneity of similar matrix, improves the accuracy of information transmission, and enhances the adaptability to different downstream tasks.
Smart Images

Figure CN119004362B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of feature fusion, and particularly relates to a method and system for progressive fusion of multimodal medical data. Background Art
[0002] With the improvement of the technological level, we can obtain a vast amount of medical data of different modalities, such as electronic medical record data, imaging data, and text data, etc., so as to more comprehensively evaluate the health status of patients from different perspectives and help doctors and researchers more accurately understand the condition and predict its progression trend. Multimodal medical data fusion, by seeking complementary information between modalities and studying complex cross-modal interaction relationships, plays a huge role in improving the accuracy of disease diagnosis and promoting the early detection of diseases.
[0003] Multimodal medical data fusion strategies can be divided into early fusion, mid-term fusion, and late fusion according to the fusion stage. In the early fusion stage, after operations such as feature screening and feature mapping on data of different modalities, they are fused into a richer input feature and then input into the model; in the mid-term fusion stage, data of different modalities pass through the feature extraction modules of their respective modalities and then are fused; the late fusion stage is to fuse the prediction vectors of each modality for downstream tasks.
[0004] Currently, multimodal medical data fusion methods have the following deficiencies. First, as the multimodal fusion stage progresses, the features of each modality are continuously purified, filtering out the redundant information and noise interference therein. However, the information contained is more biased towards single modalities, and the complementary information between modalities may be discarded. Second, the graph representation structures constructed from non-graph data are often static graphs. When calculating patient similarity, all features are given the same importance and are not matched with downstream tasks. In addition, when removing low-correlation edges, it is impossible to perform adaptive screening for the patient himself. Third, in modality alignment, contrast learning plays an important role, but the dependence on high-quality data augmentation has always been its pain point, and graph augmentation is mostly based on heuristic methods, and the effects are unstable in different task scenarios. Summary of the Invention
[0005] The present invention aims to solve at least the above technical problem, and provides a method and system for progressive fusion of multimodal medical data.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] In the first aspect, the present invention provides a method for progressive fusion of multimodal medical data, comprising:
[0008] Extract the N - layer features of the patient's multimodal medical data, where N is an integer greater than or equal to 2. The multimodal medical data includes imaging data, structured electronic medical record data, and text data. The extracted N - layer features include N - layer imaging features, N - layer structured electronic medical record features, and N - layer text features;
[0009] Fuse all modal features in each layer of features to obtain N hierarchical fusion features;
[0010] Align the hierarchical fusion features of each layer respectively;
[0011] Based on the cross - layer attention mechanism, fuse the aligned hierarchical fusion features of the first layer and the second layer to obtain the first progressive fusion feature. Then, based on the cross - layer attention mechanism, fuse the first progressive fusion feature and the aligned hierarchical fusion feature of the third layer to obtain the second progressive fusion feature, and so on. Repeat this cross - layer attention mechanism fusion step until the fusion of the N - layer aligned hierarchical fusion features is completed.
[0012] Preferably, the method for fusing all modal features in each layer of features to obtain N hierarchical fusion features is as follows:
[0013] For each layer of features, map all modal features in this layer and splice them into a spliced feature;
[0014] Screen out the important features with high correlation with the downstream task in the spliced feature;
[0015] Calculate the similarity matrix of patients based on the screened important features;
[0016] After pruning the similarity matrix, perform information transfer with the important features to obtain the hierarchical fusion feature of each layer.
[0017] Preferably, through the loss function Make the feature vectors of different modalities of the same patient similar after mapping:
[0018]
[0019] where cos is the cosine similarity calculation, and the calculation formula is and are the imaging feature structured electronic medical record feature and text feature feature vectors after mapping respectively, and j = {1, 2,..., N}, representing the j - th layer.
[0020] Preferably, the pruning method is as follows:
[0021] Calculate the similarity threshold of each patient based on the screened important features;
[0022] Prune the similarity matrix according to the similarity threshold of the patient, remove the edges smaller than the similarity threshold, and retain the edges greater than or equal to the similarity threshold.
[0023] Preferably, through the loss function Measure the similarity between the important features reconstructed in each layer and the important features selected,
[0024] where, are the important features of the reconstruction, are the important features selected, j = {1, 2,..., N}, representing the j-th layer, cos is the cosine similarity calculation, and the calculation formula is
[0025] where, the important features of the reconstruction are obtained through a masked graph decoder, and the specific method is:
[0026] Perform path masking on the pruned similarity matrix, perform unsupervised feature masking on the hierarchical fusion features, and then input them into the graph convolutional decoder for decoding to obtain the important features of the reconstruction in each layer.
[0027] Preferably, the method of aligning the hierarchical fusion features of each layer is shown in the following formula:
[0028] where, FC gcn is a fully connected network, are the hierarchical fusion features are the feature vectors after mapping, BatchNormal is the operation of normalizing the feature vectors column by column, j = {1, 2,..., N}, representing the j-th layer.
[0029] Preferably, the formula for fusing based on the cross-layer attention mechanism is:
[0030]
[0031] where, is the progressive fusion feature after fusing the aligned hierarchical fusion features of the first j layers, softmax is the normalization exponential function, and its formula is are the aligned hierarchical fusion features is the dimension of, refers to is the transpose matrix of, is a fully connected network, j = {1, 2,..., N}, representing the j-th layer.
[0032] Preferably, the image data includes X-ray films, CT scan images, and MRI images; the structured electronic medical record data includes demographic characteristics and clinical test information; and the text data includes medical histories and treatment records.
[0033] In a second aspect, the present invention provides a system for progressive fusion of multimodal medical data, comprising:
[0034] An extraction module, configured to extract N-layer features of a patient's multimodal medical data, where N is an integer greater than or equal to 2, the multimodal medical data includes image data, structured electronic medical record data, and text data, and the extracted N-layer features include N-layer image features, N-layer structured electronic medical record features, and N-layer text features;
[0035] A first fusion module, configured to fuse all modal features in each layer of features to obtain N hierarchical fusion features;
[0036] An alignment module, configured to align the hierarchical fusion features of each layer respectively;
[0037] A second fusion module, configured to fuse the hierarchically fused features aligned in the first layer and the second layer respectively based on a cross-layer attention mechanism to obtain a first progressive fusion feature, and then fuse the first progressive fusion feature and the hierarchically fused features aligned in the third layer based on the cross-layer attention mechanism to obtain a second progressive fusion feature, and so on, repeating this cross-layer attention mechanism fusion step until the fusion of the N-layer aligned hierarchical fusion features is completed.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The present invention realizes the interaction of hierarchical fusion feature vectors based on a cross-layer attention mechanism through a progressive fusion method, and uses a residual network to alleviate the problem that shallow information is easily lost when the number of layers increases, so as to maximize the hierarchical information contained in the progressive fusion features in a progressive interaction manner.
[0040] The present invention provides a dynamic graph learning method, which captures the non-linear interaction relationship of features through graph adaptive learning, reduces the heterogeneity of the similarity matrix, and improves the accuracy of information transmission; in this adaptive learning process, a threshold is predicted for the features of each patient to prune the similarity matrix, making the pruned similarity matrix more accurate.
[0041] The present invention uses a masked graph autoencoder for modal alignment, which is more universal for different downstream tasks and can effectively capture the interaction relationship of cross-modal features. In this process, important features are decoded and reconstructed, and a feature reconstruction loss function is set, so that while encoding and compressing features, important information is learned to explore the complex interaction relationship between features. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Schematic flowchart of the method for progressive fusion of multimodal medical data provided by the embodiment of the present invention;
[0043] Figure 2 Schematic diagram of the masked graph autoencoder structure provided by an embodiment of the present invention;
[0044] Figure 3 Framework diagram of the progressive fusion of multimodal medical data provided by the embodiment of the present invention;
[0045] Figure 4 Schematic diagram of the cross-level attention mechanism structure provided by an embodiment of the present invention;
[0046] Figure 5 Schematic diagram of the structure of the system for progressive fusion of multimodal medical data provided by the embodiment of the present invention;
[0047] Figure 6 Schematic block diagram of the exemplary electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0049] The first embodiment of the present invention provides a method for progressive fusion of multimodal medical data, Figure 1 Schematic flowchart of the method for progressive fusion of multimodal medical data provided by the embodiment of the present invention.
[0050] As Figure 1 shown, the method for progressive fusion of multimodal medical data specifically includes the following steps:
[0051] S101, extract the N-layer features of the multimodal medical data of the patient, where N is an integer ≥ 2, the multimodal medical data includes image data, structured electronic medical record data, and text data, and the extracted N-layer features include N-layer image features, N-layer structured electronic medical record features, and N-layer text features.
[0052] It can be understood that N is a hyperparameter, and the user can determine it according to their own requirements for the accuracy and time cost of the downstream prediction task. For example, it can be 2, 3, 4, 7, 9, 10, 15, etc.
[0053] Specifically, the multimodal medical data of the patient includes image data (x image) Structured electronic medical record data (x mc ) and text data (x text ). Among them, the image data includes data forms such as X-ray films, CT scan images, MRI images, etc.; the structured electronic medical record data includes information such as demographic characteristics and clinical tests; the text data includes information such as medical history and treatment records. Let the symbol M represent the number of patients with complete multimodal medical data collected, and the entire multimodal medical data set is represented by P as:
[0054]
[0055] For the method of extracting the N-layer features of the multimodal medical data of patients, each modal medical data is input into its respective feature extraction module to obtain its respective feature vector. Specifically, the image data, structured electronic medical record data, and text data in the multimodal medical data of patients are respectively input into the image feature extraction module IFM, the structured electronic medical record feature extraction module MFM, and the text feature extraction module TFM to extract the feature vectors of the three modal data of each patient Among them, and are the image feature, electronic medical record feature, and text feature of the i-th patient respectively.
[0056] Since the image feature extraction module IFM is composed of N layers of image feature extraction functions IFE, the image data of the i-th patient can be obtained of multi-layer image features, that is Among them, is the feature of the image data of the i-th patient after passing through the first j layers of image extraction functions. For better description below, let Similarly, multi-layer structured electronic medical record features and text features can also be obtained, and finally N-layer image features, N-layer structured electronic medical record features, and N-layer text features are obtained.
[0057] S102, fuse all modal features in each layer of features to obtain N hierarchical fusion features.
[0058] It can be understood that the method of fusing features in each layer is the same, and finally the hierarchical fusion features of N layers of features are obtained respectively. Refer to Figure 2 , and the specific method includes the following steps:
[0059] S102-1, for each layer of features, map and splice all modal features in this layer of features into a spliced feature.
[0060] Exemplarily, for the j-th layer of features of a patient are respectively input into the j-th layer of image feature vector mapping function The mapping function of the electronic medical record feature vector of the j-th layer The mapping function of the text feature vector of the j-th layer And perform a concatenation operation on them. The specific formula is as follows:
[0061]
[0062] Among them, and are fully connected layer networks, and are respectively and The feature vectors after mapping. Concate is the operation of concatenating the feature vectors. is the feature vector after concatenating the j-th layer and j = {1, 2,..., N}, representing the N-th layer.
[0063] In some preferred embodiments, in order to reduce the heterogeneity between modalities, it is required that after mapping, the feature vectors of different modalities of the same patient should be similar. The following loss function is used to achieve this:
[0064]
[0065]
[0066] Among them, cos is the calculation of cosine similarity, and the calculation formula is is the modality mapping loss function of the j-th layer.
[0067] S102-2. Screen out the important features with high correlation with the downstream task in the concatenated features.
[0068] Exemplarily, input the concatenated features obtained in S102-1, that is, the j-th layer of the patient into the Gate feature screening module to screen out the important features with high correlation with the downstream task. The specific formula is as follows:
[0069]
[0070] Among them, FC Gate is a fully connected layer, Sigmoid and Relu are activation functions, AttentionSorce j is the importance of each patient's respective features in the j-th layer, mul is the operation of multiplying the corresponding elements of two vectors, is the screened important feature, j = {1, 2,..., N}.
[0071] In some preferred embodiments, in order to enable AttentionSorce j to recognize the importance of features, the following loss function is used to achieve this:
[0072]
[0073] where mean is to calculate the average value, is the loss function for the j-th layer Gate to screen features.
[0074] S102-3. Calculate the similarity matrix between patients based on the screened important features.
[0075] Exemplarily, input into the graph adaptive learning module to calculate the similarity between patients. The specific formula is as follows:
[0076]
[0077] where W is a learnable parameter matrix, Matrix j is the similarity matrix between patients in the j-th layer, is the transpose operation, and j = {1, 2,..., N}.
[0078] S102-4. After pruning the similarity matrix, perform information transfer with the important features to obtain the hierarchical fusion features of each layer.
[0079] In some preferred embodiments, the method for pruning the similarity matrix is as follows: Calculate the similarity threshold for each patient based on the screened important features; Prune the similarity matrix according to the similarity threshold of the patients, remove the edges smaller than the similarity threshold, and retain the edges greater than or equal to the similarity threshold.
[0080] The calculation formula for the similarity threshold is: where FC threshold is a fully connected network, Threshold j is the similarity threshold for each patient in the j-th layer, and j = {1, 2,..., N}.
[0081] The formula for pruning is:
[0082]
[0083] where Matrix j (m, n) represents the similarity between the m-th patient and the n-th patient in the j-th layer, Threshold j (m) represents the threshold of the m-th patient in the j-th layer, is The transposed matrix of represents the similarity matrix between patients after pruning, where j = {1, 2, ..., N}.
[0084] For information transmission, the pruned similarity matrix and the important features are input into the Graph Convolutional Network (GCN) module for information transmission. The specific formula is:
[0085]
[0086] where represents the feature vector after passing through the graph convolution module, which is the hierarchical fusion feature at the j-th layer and is called the hierarchical fusion feature. GCN encoder represents the graph convolution encoder module, where j = {1, 2, ..., N}.
[0087] It can be seen that S102-2 to S102-4 are the processes of dynamic graph learning, realizing the construction of a dynamic graph for non-graph data. Generally, there are often irrelevant and redundant information in the extracted features. Through dynamic graph learning, the features are screened to remove noise information and retain the features highly relevant to the downstream task. In the prior art, the graph construction method usually measures linear similarity and ignores the complex relationships between features. In the embodiments of the present invention, dynamic graph learning captures the non-linear interaction relationships of features through the graph adaptive learning module, reduces the heterogeneity of the similarity matrix, and improves the accuracy of information transmission. In addition, in the prior art, the similarity matrix pruning generally uses the KNN algorithm to retain k edges with high similarity, but this approach screens globally and ignores the local situation of each patient. In the embodiments of the present invention, dynamic graph learning predicts the threshold for the features of each patient, making the pruned similarity matrix more accurate.
[0088] In some more preferred embodiments, the similarity between the important features reconstructed in each layer and the selected important features is measured by the loss function to enhance the accuracy of the above adaptive graph construction:
[0089] where is the reconstructed important feature, is the selected important feature, j = {1, 2, ..., N} represents the j-th layer, and cos is the cosine similarity calculation. The calculation formula is
[0090] Among them, the important features of the reconstruction are obtained through a masked graph decoder. The specific method is as follows: The pruned similarity matrix is subjected to path masking, and the hierarchically fused features are subjected to unsupervised feature masking, and then input into the graph convolutional decoder for decoding to obtain the important features of the reconstruction.
[0091] Exemplarily, the pruned similarity matrix is input into the path masking module for path masking. In the path masking module, is sampled to obtain its subgraph. Whether each edge is retained or masked follows a Bernoulli distribution, and the specific formula is as follows:
[0092]
[0093] where is after path masking, where Bernouli is the Bernoulli distribution, Rate edge is the proportion of path masking, and ~ means follows the Bernoulli distribution, and j = {1, 2,..., N}.
[0094] Random masking is performed on the patients, that is, unsupervised feature masking is performed on the hierarchically fused features. The order of the patients is randomly shuffled, and masked patients and unchanged patients are selected according to a proportion. The specific formula is as follows:
[0095] NumMask = int(M * Rate Patient ),
[0096] MaskPatient = Shuffle(P)[:NumMask],
[0097]
[0098] where M is the number of patients, NumMask is the number of patients to be masked, * is the multiplication operation of two numbers, Rate Patient is the proportion of masking the patients, int is the rounding operation, Shuffle is the operation of shuffling the list to achieve the purpose of random sampling, and [:NumMask] is the operation of taking the first NumMask items of the list. is the feature vector after feature masking j = {1, 2,..., N}.
[0099] The hierarchically fused features after unsupervised feature masking and the similarity matrix after path masking are input into the graph convolutional decoder to decode the information. The specific formula is as follows:
[0100]
[0101] Among them, GCN decoder is a graph convolutional decoder module, is the feature vector after passing through the graph decoder, that is, the important reconstructed feature, j = {1, 2,..., N}.
[0102] By the above method of decoding and reconstructing important features, and setting the feature reconstruction loss function, while encoding and compressing features, important information is learned, and the complex interaction relationship between features is explored. The above masked encoder, by perturbing the features and the similarity matrix, makes neighbor nodes play a more important role during feature reconstruction, so as to improve the accuracy of the similarity matrix and maximize the extracted information.
[0103] S103, align the hierarchical fusion features of each layer respectively.
[0104] Since the hierarchical fusion feature vectors of different levels will be in different feature spaces and have heterogeneity, alignment operations are required. The specific method is shown in the following formula:
[0105]
[0106] Among them, FC gcn is a fully connected network, is the hierarchical fusion feature The feature vector after mapping, and BatchNormal is the operation of normalizing the feature vector column by column.
[0107] S104, after fusing the hierarchically fused features aligned by the first layer and the second layer respectively based on the cross-layer attention mechanism to obtain the first progressive fusion feature, then fuse the first progressive fusion feature and the hierarchically fused feature aligned by the third layer based on the cross-layer attention mechanism to obtain the second progressive fusion feature, and so on, repeating this cross-layer attention mechanism fusion step until the fusion of the hierarchically fused features aligned by N layers is completed.
[0108] Specifically, referring to Figures 3-4 , Figure 3 is a schematic diagram of the framework for progressive fusion of multimodal medical data, Figure 4It is a schematic diagram of the cross - level attention mechanism structure. This step explores the relationship of the fusion features between different levels through the self - attention mechanism, realizes the interaction of the hierarchical fusion feature vectors. And as the number of layers progresses, the influence of the early hierarchical fusion features on the progressive fusion features gradually weakens. By using the residual method, it alleviates the problem that the shallow - layer information is easily lost when the number of layers increases. Since there are N layers of aligned hierarchical fusion features in total, the cross - layer attention mechanism - based fusion needs to be performed N - 1 times. In this way, the hierarchical fusion features of different levels continuously interact. As the fusion stage progresses, progressive fusion is formed to obtain the final progressive fusion feature vector. By means of progressive interaction, the hierarchical information contained in the progressive fusion features is maximized.
[0109] The formula for the fusion process is:
[0110]
[0111] Among them, is the feature vector after fusing the hierarchical fusion features of the first j layers, called the progressive fusion feature. Softmax is the normalized exponential function, and its formula is is the dimension of refers to the transpose matrix of is the fully - connected network.
[0112] The second embodiment of the present invention provides a system for progressive fusion of multi - modal medical data.
[0113] Figure 5 It is a schematic diagram of the structure of the system for progressive fusion of multi - modal medical data provided by the embodiments of this application.
[0114] As Figure 5 shown, the system 500 for progressive fusion of multi - modal medical data includes an extraction module 501, a first fusion module 502, an alignment module 503, and a second fusion module 504. The specific functions of each module are described as follows:
[0115] The extraction module 501 is used to extract N - layer features of the multi - modal medical data of the patient, where N is an integer greater than or equal to 2. The multi - modal medical data includes image data, structured electronic medical record data, and text data. The N - layer features extracted include N - layer image features, N - layer structured electronic medical record features, and N - layer text features;
[0116] The first fusion module 502 is used to fuse all the modal features in each layer of features to obtain N hierarchical fusion features;
[0117] The alignment module 503 is used to align the hierarchical fusion features of each layer respectively;
[0118] The second fusion module 504 is configured to, after fusing the hierarchically fused features aligned at the first layer and the hierarchically fused features aligned at the second layer based on the cross-layer attention mechanism to obtain the first progressively fused features, then fuse the first progressively fused features and the hierarchically fused features aligned at the third layer based on the cross-layer attention mechanism to obtain the second progressively fused features, and repeat this cross-layer attention mechanism-based fusion step until the fusion of the hierarchically fused features aligned at N layers is completed.
[0119] Based on the above embodiments, the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the method for progressive fusion of multimodal medical data in the first aspect embodiments are implemented.
[0120] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0121] As Figure 6 shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0122] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0123] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as an omics data processing method or a model training method for omics data processing. For example, in some embodiments, the omics data processing method or the model training method for omics data processing can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the omics data processing method or the model training method for omics data processing described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the omics data processing method or the model training method for omics data processing by any other suitable means (e.g., by means of firmware).
[0124] Based on the above embodiments, the present invention also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method for progressive fusion of multimodal medical data disclosed in the embodiments of the present invention.
[0125] Based on the above embodiments, the present invention also provides a computer program product including a computer program, which implements the method for progressive fusion of multimodal medical data disclosed in the embodiments of the present invention when executed by a processor.
[0126] Wherein, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0127] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0128] In several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.
[0129] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0130] In addition, the functional units in the various embodiments of this application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0131] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application.
[0132] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A method for progressive fusion of multimodal medical data, characterized in that: Include: Extracting N layers of features from the patient's multimodal medical data, where N is an integer greater than 2, the multimodal medical data includes image data, structured electronic medical record data, and text data, and the extracted N layers of features include N layers of image features, N layers of structured electronic medical record features, and N layers of text features; All modal features in each layer of features are fused to obtain N levels of fused features; Align the hierarchical fusion features of each layer separately; Based on the cross-layer attention mechanism, the aligned hierarchical fusion features of the first layer and the second layer are fused to obtain the first progressive fusion feature, and then the first progressive fusion feature and the aligned hierarchical fusion features of the third layer are fused based on the cross-layer attention mechanism to obtain the second progressive fusion feature, and so on, repeating this fusion step based on the cross-layer attention mechanism until the fusion of the N layers of aligned hierarchical fusion features is completed; The method of fusing all modal features in each layer of features to obtain N-level fusion features is as follows: for each layer of features, all modal features in the layer of features are mapped and spliced into spliced features; Filter out important features in the splicing features that are highly relevant to downstream tasks; Calculate the patient-patient similarity matrix based on the important features obtained by screening; After pruning the similarity matrix, information is transferred with the important features to obtain the hierarchical fusion features of each layer.
2. The method for progressive fusion of multimodal medical data according to claim 1, characterized in that: Through the loss function After achieving the mapping, the feature vectors of different modalities for the same patient are similar: , where cos is the cosine similarity calculation, and the calculation formula is , , and Image features , Characteristics of structured electronic medical records and text features The mapped feature vector is , represents the jth layer.
3. The method for progressive fusion of multimodal medical data according to claim 1, characterized in that: The pruning method is: The similarity threshold of each patient was calculated based on the important features obtained by screening; The similarity matrix is pruned according to the patient's similarity threshold, edges less than the similarity threshold are removed, and edges greater than or equal to the similarity threshold are retained.
4. The method for progressive fusion of multimodal medical data according to claim 1, characterized in that: Through the loss function Measure the similarity between the important features reconstructed at each layer and the important features screened out, ,in, An important feature of reconstruction is is an important feature selected. , represents the jth layer, cos is the cosine similarity calculation, and the calculation formula is ; Among them, the important features of the reconstruction are obtained through the masked graph decoder, the specific method is as follows: The pruned similarity matrix is path-masked, the hierarchical fusion features are unsupervisedly masked, and then input into the graph convolution decoder for decoding to obtain the important features of each layer reconstruction.
5. The method for progressive fusion of multimodal medical data according to claim 1, characterized in that: The method of aligning the hierarchical fusion features of each layer is shown in the following formula: ,in, It is a fully connected network. It is a hierarchical fusion feature The mapped feature vector is It is the operation of normalizing the feature vector by column. , represents the jth layer.
6. The method for progressive fusion of multimodal medical data according to claim 1, characterized in that: The formula based on the cross-layer attention mechanism fusion is: ,in, It is the progressive fusion feature after fusing the hierarchical fusion features of the first j layers. Softmax is a normalized exponential function, and its formula is , is the aligned hierarchical fusion feature The dimension of means The transposed matrix of It is a fully connected network. , represents the jth layer.
7. The method for progressive fusion of multimodal medical data according to claim 1, characterized in that: The image data includes X-rays, CT scan images, and MRI images; the structured electronic medical record data includes demographic characteristics and clinical examination information; and the text data includes medical history and treatment records.
8. A system for progressive fusion of multimodal medical data, characterized in that: Include: An extraction module, used to extract N layers of features of the patient's multimodal medical data, where N is an integer greater than 2, the multimodal medical data includes image data, structured electronic medical record data, and text data, and the extracted N layers of features include N layers of image features, N layers of structured electronic medical record features, and N layers of text features; The first fusion module is used to fuse all modal features in each layer of features to obtain N hierarchical fusion features, specifically including, for each layer of features, mapping all modal features in the layer of features and splicing them into spliced features; screening out important features with high correlation with downstream tasks in the spliced features; calculating the similarity matrix between patients based on the screened important features; pruning the similarity matrix and transferring information with the important features to obtain hierarchical fusion features of each layer; Alignment module, used to align the hierarchical fusion features of each layer separately; The second fusion module is used to fuse the aligned hierarchical fusion features of the first layer and the second layer based on the cross-layer attention mechanism to obtain the first progressive fusion feature, and then fuse the first progressive fusion feature with the aligned hierarchical fusion features of the third layer based on the cross-layer attention mechanism to obtain the second progressive fusion feature, and so on, repeating this fusion step based on the cross-layer attention mechanism until the fusion of N layers of aligned hierarchical fusion features is completed.