Multi-modal pathological data fusion analysis system
Through the combination of adaptive modal interaction network and generative adversarial network, the problems of insufficient interaction between modals and poor information completion quality in traditional pathological data analysis are solved, and more efficient information interaction and data completion are achieved.
Patent Information
- Application Number
- CN202510453685.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
AI Technical Summary
Traditional pathological data analysis methods lack dynamic interactions between modes, resulting in information loss and difficulty in capturing potential correlations between cross-modal features. The features generated by the modal information completion method are prone to deviations, poor quality, and low information utilization.
Adaptive modal interaction network is used to improve cross-modal information interaction capabilities by calculating the importance of modality, and generate missing modal features using generative adversarial networks to ensure data integrity and complete quality.
Effectively avoid information loss, improve generalization ability, enhance the expression ability of fusion characteristics, ensure semantic consistency, and improve the ability to capture the correlation relationship between modals.
Smart Images

Figure CN119964837A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information technology, and in particular to a multimodal pathology data fusion analysis system. Background Art
[0002] With the rapid development of artificial intelligence and medical imaging technology, multimodal pathology data fusion analysis systems have shown great potential in disease diagnosis, treatment prediction and prognosis assessment. Traditional pathology data analysis methods lack dynamic interaction between modalities, and the imbalanced data processing leads to information loss, and it is difficult to capture the potential correlation between cross-modal features; traditional modality information completion methods usually use linear interpolation or mean filling, and the generated features are prone to deviation and poor quality, and direct splicing leads to low information utilization. Summary of the invention
[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a multimodal pathology data fusion analysis system. In view of the problems that traditional pathology data analysis methods lack dynamic interaction between modalities, uneven data processing leads to information loss, and it is difficult to capture the potential correlation between cross-modal features, this scheme uses an adaptive modality interaction network to improve the cross-modal information interaction capability by calculating the modality importance, dynamically adjusts the modality interaction weight, effectively avoids information loss, improves generalization ability, enhances the expression ability of fusion features, and ensures semantic consistency; in view of the problems that traditional modality information completion methods usually use linear interpolation or mean filling, the generated features are prone to deviations and poor quality, and direct splicing leads to low information utilization, this scheme uses a generative adversarial network to generate missing modality features to ensure data integrity, optimizes the generated features by constructing an adversarial loss function, ensures the true distribution of generated data, improves the completion quality, captures potential information relationships through feature propagation, and effectively enhances the correlation between modalities.
[0004] A multimodal pathology data fusion analysis system provided by the present invention includes a data acquisition module, a feature extraction module, a modality interaction module, a modality completion module, a multi-task optimization module and a result output module;
[0005] The data acquisition module acquires multimodal pathological data in real time, preprocesses the multimodal pathological data to obtain processed data, and sends the processed data to the feature extraction module;
[0006] The feature extraction module extracts high-dimensional features of each modality from the processed data, performs weighted optimization on the high-dimensional features of each modality, obtains optimized features, and sends the optimized features to the modality interaction module;
[0007] The modality interaction module dynamically models the correlation between modalities through an adaptive modality interaction network, constructs fusion features based on optimized features, and sends the fusion features to the modality completion module;
[0008] The modality completion module detects whether there is missing modality data, generates features of the missing modality using a generative adversarial network, constructs a complete fusion feature, and sends the complete fusion feature to the multi-task optimization module;
[0009] The multi-task optimization module constructs a multi-task collaborative optimization framework, uses the complete fusion features as shared underlying features, designs task-specific layers to achieve joint optimization among multiple tasks, generates optimization results and sends them to the result output module;
[0010] The result output module generates a final analysis report based on the optimization results and performs a visual display.
[0011] Furthermore, the modal interaction module constructs fusion features through an adaptive modal interaction network, specifically including the following steps:
[0012] Step S1: Matrix definition, initializing the feature matrix of multimodal pathology data based on the optimized features, and defining the weight matrix of the interaction relationship and interaction weight between modalities;
[0013] Step S2: Importance index calculation, calculate the correlation between modes, and calculate the importance index of each modal feature. The formula used is as follows: ;
[0014] In the formula, and Represents two different modal characteristics, Representing modality The importance index of the feature, Representing modality and modal The interaction weight between Representing modality and modal The correlation between
[0015] Step S3: Key feature screening, set the feature threshold according to the mean standard deviation. If the importance index is greater than the feature threshold, retain the key features of the modality and use the key features to define a new feature matrix. The formula used is as follows: ; ;
[0016] In the formula, represents the feature threshold, and Represent the mean and standard deviation of all modal feature importance indices, is the adjustment factor, Representing modality Key features of
[0017] Step S4: construct a modal interaction graph, taking each modality as a node, the interaction relationship between modalities as an edge, the correlation between modalities as the weight of the edge, and using a graph neural network to learn the semantic association between nodes;
[0018] Step S5: Weight adjustment: dynamically adjust the interaction weights between modalities according to the learned semantic associations. The formula used is as follows: ;
[0019] In the formula, represents the adjusted inter-modal interaction weight, Representing modality The distribution probability in the global feature, Representing modality Distribution probability among local features;
[0020] Step S6: Fusion feature generation: Use graph convolution method to propagate features between nodes in the modal interaction graph, aggregate the updated features into a unified high-dimensional feature vector as the fusion feature. The formula used is as follows: ;
[0021] In the formula, represents the fusion feature, Represents a splicing operation, They represent the updated features of modal A, B, and C respectively.
[0022] Furthermore, the modality completion module generates features of missing modalities using a generative adversarial network, specifically including the following steps:
[0023] Step A1: Network initialization, initializing the generator and discriminator. The generator is used to generate features of the missing modality, and the discriminator is used to distinguish between real features and generated features.
[0024] Step A2: Feature generation. The generator inputs the existing modal features and outputs the features of the missing modalities as generated features. The distribution difference between the generated features and the real features is calculated by KL divergence. The formula used is as follows: ;
[0025] In the formula, Represents real characteristics, represents the distribution of the real features, represents the distribution of generated features, Represents the distribution difference between generated features and real features;
[0026] Step A3: Construct an adversarial loss function and optimize the parameters of the generator and discriminator so that the generated feature distribution gradually approaches the real feature distribution;
[0027] Step A4: Introduce consistency constraints to ensure that the generated features are consistent with the existing modal features in the semantic space. The formula used is as follows: ;
[0028] In the formula, represents the consistency constraint, Indicates existing modal features, represents the generated features, Indicates the semantic similarity between the generated features and the existing modal features, and is the weight hyperparameter;
[0029] Step A5: Feature output, fusing the generated features with the existing modal features to form a complete fusion feature.
[0030] Furthermore, the multi-task optimization module constructs a multi-task collaborative optimization framework, takes the complete fusion features as shared underlying features, completes multiple analysis tasks at the same time, and designs task-specific layers to optimize the characteristics of each task; the loss function of the multi-task collaborative optimization framework is defined based on the joint loss of all tasks, and the parameters of the framework are optimized by minimizing the loss function until the loss function becomes stable.
[0031] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0032] (1) In view of the problems that traditional pathological data analysis methods lack dynamic interaction between modalities, uneven data processing leads to information loss, and it is difficult to capture the potential correlation between cross-modal features, this scheme uses an adaptive modality interaction network to improve the cross-modal information interaction capability by calculating the modality importance, dynamically adjust the modality interaction weight, effectively avoid information loss, improve generalization ability, enhance the expression ability of fusion features, and ensure semantic consistency.
[0033] (2) Traditional modal information completion methods usually use linear interpolation or mean filling, and the generated features are prone to deviation and poor quality. In addition, direct splicing leads to low information utilization. In this regard, this scheme uses a generative adversarial network to generate missing modal features to ensure data integrity. It optimizes the generated features by constructing an adversarial loss function to ensure the true distribution of the generated data and improve the completion quality. It captures potential information relationships through feature propagation and effectively enhances the correlation between modalities. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of a multimodal pathology data fusion analysis system proposed by the present invention;
[0035] Figure 2 Schematic diagram of the process used to construct fusion features for the modal interaction module;
[0036] Figure 3 Flowchart of the method used to generate features of missing modalities for the modality completion module.
[0037] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0039] Example 1, see Figure 1 , a multimodal pathology data fusion analysis system provided by the present invention includes a data acquisition module, a feature extraction module, a modality interaction module, a modality completion module, a multi-task optimization module and a result output module;
[0040] The data acquisition module acquires multimodal pathological data in real time, including pathological images, genomic data and clinical information, and performs preliminary cleaning and annotation on the multimodal pathological data to obtain processed data, and sends the processed data to the feature extraction module;
[0041] The feature extraction module extracts high-dimensional features of each modality from the processed data based on the deep learning model, uses the attention mechanism to perform weighted optimization on the high-dimensional features of each modality to obtain optimized features, and sends the optimized features to the modality interaction module;
[0042] The modality interaction module dynamically models the correlation between modalities through an adaptive modality interaction network, constructs fusion features based on optimized features, and sends the fusion features to the modality completion module;
[0043] The modality completion module detects whether there is missing modality data, generates features of the missing modality using a generative adversarial network, constructs a complete fusion feature, and sends the complete fusion feature to the multi-task optimization module;
[0044] The multi-task optimization module constructs a multi-task collaborative optimization framework, uses the complete fusion features as shared underlying features, designs task-specific layers to achieve joint optimization among multiple tasks, generates optimization results and sends them to the result output module;
[0045] The result output module generates a final analysis report based on the optimization results and performs a visual display.
[0046] Example 2, see Figure 1 This embodiment is based on the above embodiment. In the data acquisition module, the multimodal pathological data includes pathological images, genomic data and clinical information. In the tumor diagnosis scenario, the pathological image is a tissue section image taken by a microscope, the genomic data is the gene sequencing result of the patient, and the clinical information includes the patient's age, gender, and basic medical history information.
[0047] The initial cleaning process includes removing blurred areas in pathological images, correcting sequencing errors in genomic data, and supplementing missing clinical information fields; the annotation process is to classify and label multimodal pathological data based on known medical knowledge, and divide pathological images into normal tissue and diseased tissue.
[0048] Example 3, see Figure 1 This embodiment is based on the above embodiment. The feature extraction module extracts high-dimensional features of each modality based on a deep learning model: for pathological images, a convolutional neural network is used to extract local texture and global structural features; for genomic data, a long short-term memory network is used to capture the dependencies between sequences; for clinical information, a multi-layer perceptron is used to extract nonlinear features.
[0049] The attention mechanism is introduced to assign different weights to the high-dimensional features of each modality extracted above to highlight key information: in pathological images, the features of diseased tissues are assigned higher weights; in genomic data, disease-related mutation sites are assigned higher weights. Through weighted optimization, the optimized features are output.
[0050] Example 4, see Figure 1 and Figure 2 Based on the above embodiment, the modal interaction module constructs fusion features through an adaptive modal interaction network, which specifically includes the following steps:
[0051] Step S1: Matrix definition: Initialize the feature matrix of multimodal pathology data based on the optimized features, correspond the pathology image, genomic data and clinical information to modality A, modality B and modality C respectively, and define the weight matrix of the interaction relationship and interaction weight between modalities. The feature matrix and weight matrix are expressed as follows: F=[ F A , F B , F C ] ; W=[ W A,B , W A,C , W B,C ] ;
[0052] In the formula, represents the feature matrix, represents the weight matrix, Respectively represent the high-dimensional features of pathological images, genomic data and clinical information in the optimized features, Represents the interaction weight between each modality;
[0053] Step S2: Importance index calculation, using the Pearson correlation coefficient to calculate the correlation between the modes and the importance index of each modal feature. The formula used is as follows: ; ; ;
[0054] In the formula, Respectively represent the importance index of modal A, B, and C features, represents the Pearson correlation between the modalities;
[0055] Step S3: Key feature screening, set the feature threshold according to the mean standard deviation. If the importance index is greater than the feature threshold, retain the key features of the modality and use the key features to define a new feature matrix. The formula used is as follows: ; F * =[ F A * , F B * , F C * ] ;
[0056] In the formula, represents the feature threshold, and Represent the mean and standard deviation of all modal feature importance indices, is the adjustment factor, Represent the key features of modes A, B, and C respectively, Represents the new feature matrix;
[0057] Step S4: construct a modal interaction graph, taking each modality as a node, the interaction relationship between modalities as an edge, the correlation between modalities as the weight of the edge, and using a graph neural network to learn the semantic association between nodes;
[0058] Step S5: Weight adjustment: dynamically adjust the interaction weights between modalities according to the learned semantic associations. The formula used is as follows: ;
[0059] In the formula, represents the adjusted inter-modal interaction weight, Representing modality The distribution probability among all features, Representing modality Distribution probability among local features;
[0060] Step S6: Fusion feature generation: Use graph convolution method to propagate features between nodes in the modal interaction graph, aggregate the updated features into a unified high-dimensional feature vector as the fusion feature. The formula used is as follows: ;
[0061] In the formula, represents the fusion feature, Represents a splicing operation, They represent the updated features of modal A, B, and C respectively.
[0062] By performing the operations described above, in order to address the problems that traditional pathology data analysis methods lack dynamic interaction between modalities, suffer from information loss due to unbalanced data processing, and have difficulty in capturing potential correlations between cross-modal features, this solution uses an adaptive modality interaction network to improve cross-modal information interaction capabilities by calculating modality importance, dynamically adjusts modality interaction weights, effectively avoids information loss, improves generalization capabilities, enhances the expressiveness of fused features, and ensures semantic consistency.
[0063] Example 5, see Figure 1 and Figure 3 Based on the above embodiment, this embodiment uses a generative adversarial network to generate features of missing modalities, and specifically includes the following steps:
[0064] Step A1: Network initialization, initializing the generator and discriminator. The generator is used to generate features of the missing modality, and the discriminator is used to distinguish between real features and generated features.
[0065] Step A2: Feature generation. The generator inputs the existing modality features and outputs the features of the missing modality as generated features. The distribution difference between the generated features and the real features is calculated by KL divergence. Assuming that the pathological image is missing, the genomic data and clinical information are input into the generator to generate the initial features of modality A. The formula used is as follows: ;
[0066] In the formula, Represents real characteristics, represents the distribution of the true features of pathological images, represents the distribution of generated features, Represents the distribution difference between generated features and real features;
[0067] Step A3: Construct an adversarial loss function and optimize the parameters of the generator and discriminator so that the generated feature distribution gradually approaches the real feature distribution;
[0068] Step A4: Introduce consistency constraints to ensure that the generated features are consistent with the existing modal features in the semantic space. The formula used is as follows: ;
[0069] In the formula, represents the consistency constraint, Indicates existing modal features, represents the generated features, Indicates the semantic similarity between the generated features and the existing modal features, and is the weight hyperparameter;
[0070] Step A5: Feature output, fusing the generated features with the existing modal features to form a complete fusion feature.
[0071] By performing the operations, traditional modal information completion methods usually use linear interpolation or mean filling, and the generated features are prone to deviation and poor quality, and direct splicing leads to low information utilization. In this solution, a generative adversarial network is used to generate missing modal features to ensure data integrity, and the generated features are optimized by constructing an adversarial loss function to ensure the true distribution of generated data and improve the completion quality. The potential information relationship is captured through feature propagation, and the correlation between modalities is effectively enhanced.
[0072] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step A3, the construction of the adversarial loss function specifically includes the following steps:
[0073] Step A31: Define the generator loss function. The formula used is as follows: L G = -E[log(1-D(G(z)))] ;
[0074] In the formula, represents the generator loss, E[·] represents the cross entropy loss function, Represents the discriminator's judgment result on the generated features;
[0075] Step A32: Define the discriminator loss function and calculate the discriminator's ability to distinguish between generated features and real features. The formula used is as follows: L D =-E[ log D x +E[log(1-D(G(z)))] ;
[0076] In the formula, represents the discriminator loss, Represents the discriminator's discrimination result on the real feature;
[0077] Step A33: Network training, jointly optimizing the generator and the discriminator, minimizing the generator loss and maximizing the discriminator loss by alternating training until the Nash equilibrium state is reached.
[0078] Embodiment 7, see Figure 1 In this embodiment, based on the above embodiment, the multi-task optimization module constructs a multi-task collaborative optimization framework. In the tumor diagnosis scenario, the multi-task collaborative optimization framework uses the complete fusion features as shared underlying features, and simultaneously completes the three tasks of disease classification, survival analysis, and treatment effect prediction, and designs task-specific layers to optimize the characteristics of each task: in the disease classification task, a softmax classifier is designed to output the probability distribution of each category of disease; in the survival analysis task, the Cox proportional hazard model is used to predict the patient's survival time; in the treatment effect prediction task, a regression model is used to evaluate the effectiveness of the treatment plan.
[0079] The loss function of the multi-task collaborative optimization framework is defined based on the joint loss of all tasks. The parameters of the multi-task collaborative optimization framework are optimized by the gradient descent method until the loss function reaches stability. The formula used is as follows: ;
[0080] In the formula, represents the objective function, Represents the task index, Indicates The loss of a task, Indicates The weight of a task.
[0081] Embodiment 8, see Figure 1 This embodiment is based on the above embodiment. The result output module generates a final analysis report according to the optimization results, including the analysis results of disease classification, survival analysis and treatment effect prediction, and supports multiple visualization forms for display: for disease classification tasks, a detailed classification report is generated, listing the probability of patients belonging to various diseases; for survival analysis tasks, the patient's survival curve is drawn and the key time nodes are marked; for treatment effect prediction tasks, effect scores and recommended suggestions for different treatment plans are provided.
[0082] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0083] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
[0084] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A multimodal pathology data fusion analysis system, characterized by: Data acquisition module, feature extraction module, modality interaction module, modality completion module, multi-task optimization module and result output module; The data acquisition module acquires multimodal pathological data in real time, preprocesses the multimodal pathological data to obtain processed data, and sends the processed data to the feature extraction module; The feature extraction module extracts high-dimensional features of each modality from the processed data, performs weighted optimization on the high-dimensional features of each modality, obtains optimized features, and sends the optimized features to the modality interaction module; The modality interaction module dynamically models the correlation between modalities through an adaptive modality interaction network, constructs fusion features based on optimized features, and sends the fusion features to the modality completion module; The modality completion module detects whether there is missing modality data, generates features of the missing modality using a generative adversarial network, constructs a complete fusion feature, and sends the complete fusion feature to the multi-task optimization module; The multi-task optimization module constructs a multi-task collaborative optimization framework, uses the complete fusion features as shared underlying features, designs task-specific layers to achieve joint optimization among multiple tasks, generates optimization results and sends them to the result output module; The result output module generates a final analysis report based on the optimization results and performs a visual display.
2. A multimodal pathology data fusion analysis system according to claim 1, characterized in that: The modal interaction module constructs fusion features through an adaptive modal interaction network, specifically including the following steps: Step S1: Matrix definition, initializing the feature matrix of multimodal pathology data based on the optimized features, and defining the weight matrix of the interaction relationship and interaction weight between modalities; Step S2: importance index calculation, calculating the correlation between modes and calculating the importance index of each modal feature; Step S3: screening key features, setting feature thresholds, if the importance index is greater than the feature threshold, retaining the key features of the modality, and using the key features to define a new feature matrix; Step S4: construct a modal interaction graph, taking each modality as a node, the interaction relationship between modalities as an edge, the correlation between modalities as the weight of the edge, and using a graph neural network to learn the semantic association between nodes; Step S5: weight adjustment, dynamically adjusting the interaction weights between modalities according to the learned semantic associations; Step S6: Fusion feature generation: Use graph convolution method to propagate features between nodes in the modal interaction graph, and aggregate the updated features into a unified high-dimensional feature vector as the fusion feature.
3. The multimodal pathology data fusion analysis system according to claim 1, characterized in that: The modality completion module generates features of missing modalities using a generative adversarial network, specifically including the following steps: Step A1: Network initialization, initialization of the generator and discriminator; Step A2: Feature generation: the generator inputs the existing modal features, outputs the features of the missing modalities as generated features, and calculates the distribution difference between the generated features and the real features; Step A3: Construct an adversarial loss function and optimize the parameters of the generator and discriminator so that the generated feature distribution gradually approaches the real feature distribution; Step A4: Introduce consistency constraints to ensure that the generated features are consistent with the existing modal features in the semantic space; Step A5: Feature output, fusing the generated features with the existing modal features to form a complete fusion feature.
4. The multimodal pathology data fusion analysis system according to claim 1, characterized in that: The multi-task optimization module constructs a multi-task collaborative optimization framework, takes the complete fusion features as shared underlying features, completes multiple analysis tasks at the same time, and designs task-specific layers to optimize the characteristics of each task; the loss function of the multi-task collaborative optimization framework is defined based on the joint loss of all tasks, and the parameters of the framework are optimized by minimizing the loss function until the loss function reaches stability.
Citation Information
Patent Citations
Multi-modal target detection method and system suitable for modal deficiency
CN114359586A
Medical health management system based on big data
CN118471542A
Multi-mode lung disease diagnosis and screening system based on GAN algorithm
CN118824524A
Lung infectious disease prediction system based on multi-modal data fusion
CN119480124A
Multimodal machine learning based clinical predictor
US20200105413A1
Cited By
Traditional Chinese medicine tumor collaborative treatment clinical data analysis method
CN120853773A
A traditional Chinese medicine tumor synergistic treatment clinical data analysis method
CN120853773B
PCIS early risk prediction system driven by multi-modal data
CN121054265A
Alzheimer's disease data completion method based on multi-modal generation and fusion
CN121439269A
Medical multi-source cross-platform data real-time monitoring and intelligent risk early warning method
CN121983305A