Multi-modal medical image fusion diagnosis system based on artificial intelligence

Through the multimodal medical image fusion diagnosis system based on artificial intelligence, the problems of large registration errors, insufficient feature fusion and lack of modality in multimodal image fusion diagnosis are solved, and efficient and accurate diagnosis and data completion are achieved, which significantly improves diagnostic efficiency.

CN120219898APending Publication Date: 2025-06-27SHANXI MEDICAL UNIV
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510294731.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems in the fusion diagnosis of multimodal medical images that have large registration errors, insufficient feature fusion and modality loss.

Method used

A multimodal medical image fusion diagnosis system based on artificial intelligence is adopted, including a shared feature encoder module, a two-way deformation field prediction module, a multimodal feature fusion module, a cross-modal self-supervised pre-training module and a generative data completion network. Through these modules, unified feature extraction, cross-modal alignment, feature fusion, model generalization and data completion of multimodal images are realized.

Benefits of technology

It improves the alignment accuracy of multimodal images and the comprehensiveness and accuracy of feature fusion, enhances the generalization ability of the model, effectively solves the diagnostic problems when modal missing, and achieves a balance between efficient screening and fine diagnosis, significantly improving diagnostic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219898A_ABST
    Figure CN120219898A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal medical image fusion diagnosis system based on artificial intelligence, and the system comprises the following steps: extracting shared features of CT and MRI images through a convolutional neural network, and mapping the shared features to the same feature space; a two-way step-by-step alignment strategy is adopted, a three-dimensional deformation field matrix is generated, and cross-modal image anatomical structure alignment is achieved; calculating modal feature weights and eliminating distribution differences through an attention mechanism and an adversarial domain adaptation layer; constructing a CT-MRI image block contrast learning task, and optimizing a shared feature encoder; a conditional generative adversarial network is used for generating a false image of a missing mode according to the semantic segmentation map, and data distribution is constrained through a Wasserstein distance; uniform feature extraction of multi-modal medical images is realized through a shared feature encoder, the cross-modal image alignment accuracy is improved in combination with a bidirectional deformation field prediction module, and the comprehensiveness and accuracy of fusion features are enhanced by using a multi-modal feature fusion module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image fusion, and particularly to a multi-modal medical image fusion diagnosis system based on artificial intelligence. Background Art

[0002] In the field of medical image diagnosis, doctors often need to comprehensively analyze various modal image data (such as CT, MRI, etc.) to make accurate diagnoses. However, different modal image data have different imaging principles and characteristics. How to effectively fuse these multi-modal data to improve the accuracy and efficiency of diagnosis has always been a research hotspot in the field of medical image processing.

[0003] When traditional methods perform multi-modal image registration, they often rely on manual comparison by doctors, which is not only time-consuming and laborious, but also easily affected by subjective factors of doctors, resulting in relatively large registration errors.

[0004] Different modal image data have different feature representations. How to effectively fuse these features and extract information useful for diagnosis is a difficult problem faced by the existing technology. Traditional feature fusion methods are often too simple to fully utilize the complementarity of multi-modal data.

[0005] In actual clinical applications, there are sometimes cases where certain modal image data are missing. The existing technology often cannot effectively handle this modal missing problem, resulting in the accuracy of diagnosis being affected.

[0006] Therefore, there is an urgent need for a multi-modal medical image fusion diagnosis system based on artificial intelligence to solve the technical problems existing in the above-mentioned existing technology. Summary of the Invention

[0007] The present invention overcomes the deficiencies of the existing technology and provides a multi-modal medical image fusion diagnosis system based on artificial intelligence.

[0008] To achieve the above object, the technical solution adopted by the present invention is: A multi-modal medical image fusion diagnosis system based on artificial intelligence, comprising: a shared feature encoder module that simultaneously extracts shared features of CT and MRI multi-modal medical images through a convolutional neural network, and maps different modal data into the same feature space;

[0009] A bidirectional deformation field prediction module that adopts a bidirectional step-by-step alignment strategy based on the path independence of vector displacements, and generates a three-dimensional deformation field matrix through iterative optimization to achieve anatomical structure alignment of cross-modal images;

[0010] A multi-modal feature fusion module, including a cascaded attention mechanism layer and an adversarial domain adaptation layer, dynamically calculates different modal feature weights and eliminates the distribution differences between modalities;

[0011] Cross-modal self-supervised pre-training module, constructs a pre-training objective function including a CT-MRI image patch contrastive learning task, and optimizes the shared feature encoder module by maximizing the mutual information of positive sample pairs;

[0012] Generative data completion network, which consists of a conditional generative adversarial network, generates pseudo-image data of the missing modality according to the semantic segmentation map of the input modality, and constrains the generated data distribution through the gradient penalty Wasserstein distance.

[0013] In a preferred embodiment of the present invention, the bidirectional deformation field prediction module specifically includes:

[0014] Spatial transformation network, which uses a diffeomorphic transformation model to ensure the reversibility and smoothness of the deformation field;

[0015] Multi-scale registration sub-module, which parallelly calculates the local displacement field at multiple levels of the feature pyramid;

[0016] Deformation field fusion unit, which fuses the displacement field prediction results of different scales through a learnable gating mechanism.

[0017] In a preferred embodiment of the present invention, the multi-modal feature fusion module further includes:

[0018] Dynamic feature enhancement unit, which performs local feature enhancement in the feature space through deformable convolutional kernels;

[0019] Cross-modal interaction attention mechanism, which calculates the feature similarity matrix between modalities and generates an attention weight map;

[0020] Adversarial discriminator network, which uses a gradient reversal layer to achieve feature distribution alignment; wherein, the calculation expression of the loss function in the adversarial discriminator network is:

[0021] L adv =E[log(D(F s ))]+E[log(1-D(F t ))]

[0022] In the formula, L adv represents the adversarial loss; E[·] represents the expectation; D(·) represents the output of the adversarial discriminator, which is used to discriminate whether the input data is real or generated; F s represents the features or data of the source domain, that is, the real features or data; F t represents the features or data of the target domain, that is, the generated features or data.

[0023] Furthermore, E[log(D(F s))] represents the loss of the discriminator on the source domain data, where the discriminator hopes that the output of the source domain data (real data) is as close to 1 as possible; E[log(1 - D(F t ))] represents the loss of the discriminator on the target domain data, where the discriminator hopes that the output of the target domain data (generated data) is as close to 0 as possible; here, since log(1 - D(F t )) is used, it is actually encouraging the discriminator to wrongly think that the target domain data is real (i.e., the output is close to 1, but the form of the loss function is to estimate that its output is close to 0, so as to minimize D(F t )) by maximizing this loss; when training the generator, this part of the loss will be reversed (i.e., minimizing -E[log(1 - D(F t ))]) to encourage the generator to generate more real data to deceive the discriminator.

[0024] In a preferred embodiment of the present invention, the cross-modal self-supervised pre-training module includes: a data augmentation unit that applies joint augmentation of space and intensity including random elastic deformation and modality-specific noise injection to the input image; the calculation expression of the pre-training objective function of the contrastive learning task is:

[0025]

[0026] In the formula, L cont represents the contrastive loss; q represents the feature representation of the query sample; k + represents the feature representation of the positive sample similar to the query sample; k - represents the feature representation of the positive sample not similar to the query sample; sim(·,·) represents the similarity function; τ represents the temperature parameter, which is used to control the smoothness of the similarity distribution; log[·] is the logarithmic function, which is used to calculate the loss.

[0027] Further, sim(q, k + ) represents calculating the similarity between the query sample q and the positive sample k + ;

[0028] represents scaling and exponentiating the similarity; represents summing the exponentiated similarities of all negative samples k - , which represents the total similarity of negative samples; the entire denominator represents the total similarity of positive samples and all negative samples.

[0029] In a preferred embodiment of the present invention, the generative data completion network includes:

[0030] a U-Net structure generator and an embedded modality difference perception gating unit;

[0031] A multi-scale discriminator group, including three discriminators that respectively process features at the original resolution, 1 / 2 downsampling, and 1 / 4 downsampling;

[0032] An optimized loss function that combines adversarial loss, contrastive loss, and pixel-level L1 loss with weights; among them, the calculation expression of the optimized loss function is:

[0033] L total = λ pix L pix + λ cont L cont + λ adv L adv

[0034] In the formula, L total represents the total loss; λ pix , λ cont , λ adv represent learnable weight parameters used to control the contribution of each loss term to the total loss; L pix represents the pixel loss, which is used to measure the difference between the generated image and the real image at the pixel level; L cont represents the contrastive loss; L adv represents the adversarial loss.

[0035] In a preferred embodiment of the present invention, it further includes a cascaded processing architecture: a first-level lightweight screening network that uses MobileNetV3 after channel pruning to process single-modal data and outputs a heat map of suspicious regions; a second-level multi-modal fusion network that is only activated in regions where the heat map exceeds a threshold and performs the multi-modal fusion analysis described in claim 1; a dynamic computing resource allocation module that dynamically adjusts the batch size of each level of the network according to the GPU video memory occupancy rate.

[0036] In a preferred embodiment of the present invention, it further includes: a feature visualization interface that generates a cross-modal feature association map based on gradient class activation mapping; a diagnostic report generation module that integrates DICOM metadata and AI analysis results to automatically generate a structured diagnostic report.

[0037] In a preferred embodiment of the present invention, a multi-modal medical image diagnosis method is provided, which is applied to the multi-modal medical image fusion diagnosis system described above, and includes the following steps:

[0038] S1. Extract the shared feature representations of CT and MRI images through a shared encoder;

[0039] S2. Align the anatomical structures of multi-modal images using a bidirectional stepwise deformation field;

[0040] S3. Fuse cross-modal features through an attention mechanism and adversarial training;

[0041] S4. Enhance the generalization ability of the model through contrastive learning pre-training;

[0042] S5. When detecting the absence of a modality, generate pseudo-modal data to complete the input;

[0043] S6. Adopt a cascade architecture to achieve a balance between efficient screening and fine diagnosis.

[0044] In a preferred embodiment of the present invention, an electronic device includes: at least one processor; a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the electronic device executes the method.

[0045] In a preferred embodiment of the present invention, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.

[0046] The present invention solves the defects in the background art and has the following beneficial effects:

[0047] (1) Through the shared feature encoder, unified feature extraction of multi-modal medical images is achieved. Combining the bidirectional deformation field prediction module improves the accuracy of cross-modal image alignment, and the multi-modal feature fusion module enhances the comprehensiveness and accuracy of the fused features. The cross-modal self-supervised pre-training module improves the generalization ability of the model, enabling the model to better adapt to unseen data. The generative data completion network effectively solves the diagnostic problem when modalities are missing. At the same time, the cascade processing architecture balances efficient screening and fine diagnosis, and the dynamic computing resource allocation module optimizes the use of computing resources, significantly improving the diagnostic efficiency.

[0048] (2) Through the shared feature encoder module, the system can simultaneously extract the shared features of multi-modal medical images such as CT and MRI, map different modal data to the same feature space, and provide a unified basis for subsequent fusion analysis.

[0049] The bidirectional deformation field prediction module adopts a bidirectional step-by-step alignment strategy based on the path independence of vector displacements, effectively achieving the anatomical structure alignment of cross-modal images, reducing the registration error, and improving the accuracy of fusion.

[0050] The multi-modal feature fusion module dynamically calculates the weights of different modal features and eliminates the distribution differences between modalities through cascaded attention mechanism layers and adversarial domain adaptation layers, making the fused features more comprehensive and accurate.

[0051] (3) The cross-modal self-supervised pre-training module optimizes the shared feature encoder module and enhances the generalization ability of the model by constructing a pre-training objective function that includes a contrastive learning task for CT-MRI image patches, maximizing the mutual information of positive sample pairs. Through contrastive learning, the model can be pre-trained on unlabeled data, improving its adaptability to unseen data.

[0052] (4) The generative data completion network consists of a conditional generative adversarial network, which can generate pseudo-image data of the missing modality based on the semantic segmentation map of the input modality, and constrains the generated data distribution through the gradient penalty Wasserstein distance, effectively solving the diagnosis problem when there is modality missing.

[0053] (5) The cascade processing architecture quickly screens out suspicious regions through the first-level lightweight screening network, and the second-level multi-modal fusion network is only activated in regions where the heat map exceeds the threshold, achieving a balance between efficient screening and fine diagnosis. The dynamic computing resource allocation module dynamically adjusts the batch size of each level of the network according to the GPU video memory occupancy rate, optimizing the use of computing resources and improving the diagnosis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;

[0055] Figure 1 It is the system flowchart of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0057] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0058] Such as Figure 1As shown in the figure, the artificial intelligence-based multimodal medical image fusion diagnosis system includes: a shared feature encoder module that simultaneously extracts the shared features of CT and MRI multimodal medical images through a convolutional neural network, and maps data of different modalities to the same feature space;

[0059] A bidirectional deformation field prediction module that adopts a bidirectional step-by-step alignment strategy based on the path independence of vector displacement, and generates a three-dimensional deformation field matrix through iterative optimization to achieve anatomical structure alignment of cross-modal images;

[0060] Further, the bidirectional deformation field prediction module specifically includes:

[0061] A spatial transformation network that adopts a diffeomorphic transformation model to ensure the reversibility and smoothness of the deformation field;

[0062] A multi-scale registration sub-module that parallelly calculates local displacement fields at multiple levels of the feature pyramid;

[0063] A deformation field fusion unit that fuses the prediction results of displacement fields at different scales through a learnable gating mechanism.

[0064] A multimodal feature fusion module that includes a cascaded attention mechanism layer and an adversarial domain adaptation layer, dynamically calculates the weights of different modal features, and eliminates the distribution differences between modalities;

[0065] The multimodal feature fusion module further includes:

[0066] A dynamic feature enhancement unit that performs local feature enhancement in the feature space through a deformable convolution kernel;

[0067] A cross-modal interaction attention mechanism that calculates the feature similarity matrix between modalities and generates an attention weight map;

[0068] An adversarial discriminator network that adopts a gradient reversal layer to achieve feature distribution alignment; where the calculation expression of the loss function in the adversarial discriminator network is:

[0069] L adv =E[log(D(F s ))]+E[log(1-D(F t ))]

[0070] In the formula, L adv represents the adversarial loss; E[·] represents the expectation; D(·) represents the output of the adversarial discriminator, which is used to discriminate whether the input data is real or generated; F s represents the features or data of the source domain, that is, the real features or data; F t represents the features or data of the target domain, that is, the generated features or data.

[0071] Furthermore, E[log(D(F s ))] represents the loss of the discriminator on the source domain data, where the discriminator hopes that the output of the source domain data (real data) is as close to 1 as possible; E[log(1 - D(F t ))] represents the loss of the discriminator on the target domain data, where the discriminator hopes that the output of the target domain data (generated data) is as close to 0 as possible; here, since log(1 - D(F t )) is used, it is actually encouraging the discriminator to wrongly think that the target domain data is real (i.e., the output is close to 1, but the form of the loss function is to estimate that its output is close to 0, so as to minimize D(F t )) by maximizing this loss; when training the generator, this part of the loss will be reversed (i.e., minimizing -E[log(1 - D(F t ))]) to encourage the generator to generate more real data to deceive the discriminator.

[0072] The cross-modal self-supervised pre-training module constructs a pre-training objective function containing the CT-MRI image patch contrast learning task, and optimizes the shared feature encoder module by maximizing the mutual information of positive sample pairs;

[0073] Specifically, the cross-modal self-supervised pre-training module first calculates the similarity, including: calculating the similarity sim(q, k + ) between the query sample q and the positive sample k + ; calculating the similarity sim(q, k - ) between the query sample q and each negative sample k - .

[0074] Secondly, it calculates the exponential scaling, including: performing exponential scaling on the similarity of the positive sample performing exponential scaling on the similarity of each negative sample

[0075] After that, it normalizes the above data, including: calculating the sum of the exponential similarities of all negative samples calculating the proportion of the positive sample similarity in the total sum

[0076]

[0077] Finally, taking the negative logarithm of the above proportion to obtain the contrast loss, and its calculation expression is:

[0078]

[0079] In the formula, L cont represents the contrast loss; q represents the feature representation of the query sample; k +The feature representation of the positive sample similar to the query sample; k - The feature representation of the positive sample not similar to the query sample; sim(·,·) represents the similarity function; τ represents the temperature parameter, which is used to control the smoothness of the similarity distribution; log[·] is the logarithmic function, which is used to calculate the loss.

[0080] Further, sim(q, k + ) represents calculating the similarity between the query sample q and the positive sample k + ;

[0081] represents scaling and exponentiating the similarity; represents summing the exponentiated similarities of all negative samples k - , which represents the total similarity of the negative samples; the entire denominator represents the total similarity of the positive sample and all negative samples.

[0082] Furthermore, in the above formula, the negative logarithmic function is used to convert the similarity ratio into a loss value, so that the model maximizes the similarity of the positive sample pairs and minimizes the similarity of the negative sample pairs during training.

[0083] The generative data completion network, which is composed of a conditional generative adversarial network, generates pseudo-image data of the missing modality according to the semantic segmentation map of the input modality, and constrains the generated data distribution through the gradient penalty Wasserstein distance.

[0084] The generative data completion network includes:

[0085] A U-Net structure generator with an embedded modality difference perception gating unit;

[0086] A multi-scale discriminator group, which contains three discriminators that respectively process the features of the original resolution, 1 / 2 downsampling, and 1 / 4 downsampling;

[0087] An optimized loss function that combines the adversarial loss, contrastive loss, and pixel-level L1 loss by weighting; among them, the calculation expression of the optimized loss function is:

[0088] L total = λ pix L pix + λ cont L cont + λ adv L adv

[0089] In the formula, L total represents the total loss; λ pix , λ cont , λ advrepresents learnable weight parameters used to control the contribution of each loss term to the total loss; L pix represents the pixel loss, which is used to measure the difference between the generated image and the real image at the pixel level; L cont represents the contrastive loss; L adv represents the adversarial loss.

[0090] Specifically, for the pixel loss L pix the acquisition process is to use the mean squared error to calculate the pixel-level difference between the generated image G(x) and the real image y:

[0091]

[0092] In the formula, N represents the total number of pixels; i represents the order of the pixels; G(x) i represents the i-th pixel of the generated image; y i represents the i-th pixel of the real image.

[0093] Here, it needs to be further explained that the weights λ pix 、λ cont 、λ adv are used to balance the importance of different loss terms. According to the specific task and dataset, these weights may need to be adjusted to obtain the best performance.

[0094] Combined loss: By combining different loss terms, the model can optimize multiple objectives simultaneously, such as pixel-level accuracy, feature-level similarity, and adversarial realism of the generated image.

[0095] In a preferred embodiment, it further includes a cascaded processing architecture: a first-stage lightweight screening network that uses the MobileNetV3 after channel pruning to process single-modal data and outputs a heat map of suspicious regions; a second-stage multi-modal fusion network that is only activated in regions where the heat map exceeds a threshold and performs the multi-modal fusion analysis of claim 1; a dynamic computing resource allocation module that dynamically adjusts the batch size of each level of the network according to the GPU video memory occupancy rate. It also includes: a feature visualization interface that generates a cross-modal feature correlation map based on gradient class activation mapping; a diagnostic report generation module that integrates DICOM metadata and AI analysis results to automatically generate a structured diagnostic report.

[0096] A multi-modal medical image diagnosis method, applied to the above-mentioned multi-modal medical image fusion diagnosis system, includes the following steps:

[0097] S1. Extract the shared feature representations of CT and MRI images through a shared encoder; specifically including: by capturing the common features of each modal image, reducing the model parameters and the model complexity. The output of the shared encoder serves as the basis for subsequent alignment and fusion steps, ensuring that different modal features operate in the same feature space.

[0098] S2. Align the anatomical structures of multimodal images using a bidirectional progressive deformation field; specifically, include: adopting a bidirectional progressive alignment deformation field prediction strategy based on the path independence of vector displacement. This strategy realizes the precise alignment between different modality images by gradually adjusting the deformation field. The bidirectional alignment mechanism ensures the stability and accuracy of the alignment process, while considering the spatial correspondence relationship between different modalities.

[0099] S3. Fuse cross-modal features through the attention mechanism and adversarial training; specifically, include: introducing modality difference elimination methods, such as domain adaptation, adversarial training, etc., to reduce the negative impact of differences between different modalities on feature alignment and fusion. By learning the common representation or mapping relationship between modalities, the features of different modalities become more consistent and reliable after fusion.

[0100] S4. Enhance the model generalization ability using contrastive learning pre-training; specifically, include: using unlabeled data to construct cross-modal contrastive learning tasks, such as CT-MRI patch matching. Through contrastive learning, the model can learn the similarities and differences between different modalities, thereby improving the robustness of cross-modal representations. Self-supervised pre-training enables the model to also perform well on unseen data. By pre-training on a large-scale unlabeled dataset, the model can learn richer feature representations and improve the generalization ability in actual application scenarios.

[0101] S5. When detecting modality missing, generate pseudo-modal data to complete the input; specifically, include: when part of the modality data is missing, use cGAN to generate pseudo-modal data. cGAN can generate pseudo-modal data consistent with the real data distribution according to the existing modality data and conditional information (such as lesion location, type, etc.).

[0102] Introduce adversarial constraints to ensure that the generated pseudo-modal data has the same distribution as the real data in the feature space. Through continuous optimization of the discriminator, the generator can generate more realistic and reliable pseudo-modal data, improving the robustness of the system.

[0103] S6. Adopt a cascade architecture to achieve a balance between efficient screening and fine diagnosis; specifically, include: using a lightweight unimodal model (such as MobileNet for processing CT) in the preliminary screening stage. The lightweight model can quickly process a large amount of image data and screen out suspicious areas or lesion areas.

[0104] Only start the multi-modal fusion model in the suspicious area for more refined analysis and diagnosis. This cascade architecture takes into account both efficiency and accuracy, and can improve the processing speed of the system while ensuring the diagnostic accuracy.

[0105] Example 1

[0106] What is expected to be solved is that the traditional method relies on doctors to manually compare CT (showing calcified lesions) and MRI (showing edema zones) images, and there are problems of large registration errors and insufficient feature fusion.

[0107] Implementation steps: Input the patient's head CT (bone window) and T1-enhanced MRI (soft tissue window) data simultaneously, and automatically parse through the DICOM interface.

[0108] The shared encoder adopts an improved 3D ResNet-50 architecture, and branches out a modality-specific normalization layer after the first convolutional layer (CT data is normalized using Hounsfield units, and MRI is normalized using Z-score).

[0109] Construct a 4-level feature pyramid (resolution from 64×64×64 to 512×512×512), calculate the bidirectional deformation field on each level of the feature map, and use the SyN diffeomorphic model to ensure topological preservation.

[0110] Optimize the registration through cross-modal similarity measurement (NCC+MI hybrid loss), and reduce the CT→MRI registration error to 0.87±0.12mm (traditional method 1.53±0.21mm).

[0111] Apply the cross-modal attention mechanism in the tumor region (ROI), increase the weight of CT calcification features to 0.68, and the weight of MRI edema zone features to 0.72. The accuracy of the domain classifier drops from the initial 82% to 53%, proving the alignment of feature distributions.

[0112] When PET-CT is missing, cGAN generates pseudo-PET images based on MRI (SSIM reaches 0.89), and the cascaded architecture realizes real-time processing at 23fps on an RTX3090 graphics card, and the glioma grading accuracy reaches 94.7% (compared with the single-modal model 82.3%).

[0113] Example Two

[0114] For the rapid differential diagnosis of patients (pulmonary embolism, aortic dissection, coronary heart disease), multi-modal image analysis needs to be completed within 5 minutes.

[0115] Implementation steps: The first level uses the quantized MobileNetV3 to process CTA images and generate a heat map of suspected vascular stenosis.

[0116] Realize real-time processing at 0.2 seconds / frame on an NVIDIA Jetson edge device. When a suspected area of aortic dissection is detected (heat map value > 0.7), automatically call; the multi-modal fusion module: fuse echocardiogram and CTA features and generate pseudo-images when PET perfusion imaging is not available.

[0117] Cross-modal contrastive learning: During the pre-training stage, 2000 unlabeled CT-ultrasound image pairs were constructed. An improved MoCov3 framework was adopted, and the negative sample queue contained 5000 cross-modal samples. The pre-trained model improved the AUC of the downstream task by 12.3%.

[0118] The spatial association between the vascular wall stress distribution and the perfusion defect area was demonstrated through the feature visualization interface, and a structured report (including key sign descriptions and critical value alerts) was automatically generated.

[0119] Based on the ideal embodiments of the present invention as inspiration, through the above description, relevant personnel can make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and the technical scope must be determined according to the scope of the claims.

Claims

1. A multimodal medical image fusion diagnosis system based on artificial intelligence, characterized in that: include: The shared feature encoder module simultaneously extracts the shared features of CT and MRI multimodal medical images through convolutional neural networks, mapping different modality data into the same feature space; The bidirectional deformation field prediction module adopts a bidirectional stepwise alignment strategy based on the independence of vector displacement paths, generates a three-dimensional deformation field matrix through iterative optimization, and realizes the alignment of anatomical structures of cross-modal images; Multimodal feature fusion module, including cascaded attention mechanism layers and adversarial domain adaptation layers, dynamically calculates the weights of different modal features and eliminates the distribution differences between modalities; A cross-modal self-supervised pre-training module constructs a pre-training objective function including a CT-MRI image patch contrast learning task, and optimizes the shared feature encoder module by maximizing the mutual information of positive sample pairs; The generative data completion network, composed of a conditional generative adversarial network, generates pseudo image data of the missing modality based on the semantic segmentation map of the input modality, and generates data distribution through the gradient-penalized Wasserstein distance constraint.

2. The multimodal medical image fusion diagnosis system based on artificial intelligence according to claim 1, characterized in that: The bidirectional deformation field prediction module specifically includes: The spatial transformation network uses a differential homeomorphism transformation model to ensure the reversibility and smoothness of the deformation field; A multi-scale registration submodule that computes local displacement fields in parallel at multiple levels of the feature pyramid; The deformation field fusion unit fuses the displacement field prediction results of different scales through a learnable gating mechanism.

3. The multimodal medical image fusion diagnosis system based on artificial intelligence according to claim 1, characterized in that: The multimodal feature fusion module further comprises: Dynamic feature enhancement unit, which performs local feature enhancement in feature space through deformable convolution kernel; Cross-modal interactive attention mechanism, which calculates the feature similarity matrix between modalities and generates an attention weight map; The adversarial discriminator network uses a gradient reversal layer to achieve feature distribution alignment; the calculation expression of the loss function in the adversarial discriminator network is: L adv =E[log(D(F s ))]+E[log(1-D(F t ))] Where, L adv represents adversarial loss; E[·] represents expectation; D(·) represents the output of the adversarial discriminator, which is used to determine whether the input data is real or generated; F s The features or data representing the source domain are real features or data; t The features or data representing the target domain are generated features or data.

4. The multimodal medical image fusion diagnosis system based on artificial intelligence according to claim 1, characterized in that: The cross-modal self-supervised pre-training module includes: a data enhancement unit, which applies a joint enhancement of space and intensity including random elastic deformation and modality-specific noise injection to the input image; the calculation expression of the pre-training objective function of the contrastive learning task is: Where, L cont represents contrast loss; q represents the feature representation of the query sample; k + k represents the feature representation of positive samples similar to the query sample; - It represents the feature representation of positive samples that are dissimilar to the query samples; sim(·,·) represents the similarity function; τ represents the temperature parameter, which is used to control the smoothness of the similarity distribution; log[·] is the logarithmic function, which is used to calculate the loss.

5. The multimodal medical image fusion diagnosis system based on artificial intelligence according to claim 1, characterized in that: The generative data completion network includes: U-Net structure generator with embedded modality difference-aware gating unit; The multi-scale discriminator group includes three discriminators that process original resolution, 1 / 2 downsampled, and 1 / 4 downsampled features respectively; The optimization loss function is a weighted combination of adversarial loss, contrast loss and pixel-level L1 loss. The calculation expression of the optimization loss function is: L total =λ pix L pix +λ cont L cont +λ adv L adv Where, L total represents the total loss; λ pix , cont , adv represents a learnable weight parameter used to control the contribution of each loss term to the total loss; L pix represents pixel loss, which is used to measure the difference between the generated image and the real image at the pixel level; L cont represents contrast loss; L adv Represents resistance to loss.

6. The multimodal medical image fusion diagnosis system based on artificial intelligence according to claim 1, characterized in that: It also includes a cascade processing architecture: a first-level lightweight screening network, which uses MobileNetV3 after channel pruning to process single-modal data and outputs a heat map of suspicious areas; a second-level multimodal fusion network, which is only started in areas where the heat map exceeds a threshold and performs the multimodal fusion analysis described in claim 1; a dynamic computing resource allocation module, which dynamically adjusts the batch size of each level of the network according to the GPU memory occupancy rate.

7. The multimodal medical image fusion diagnosis system based on artificial intelligence according to any one of claims 1 to 6, characterized in that: Also includes: Feature visualization interface, generating cross-modal feature correlation maps based on gradient class activation mapping; The diagnostic report generation module integrates DICOM metadata and AI analysis results to automatically generate structured diagnostic reports.

8. A multimodal medical image diagnosis method, applied to a multimodal medical image fusion diagnosis system as claimed in any one of claims 1 to 7, characterized in that: The following steps are involved: S1, extract shared feature representations of CT and MRI images through a shared encoder; S2, anatomical structures of multimodal images are aligned using a bidirectional progressive deformation field; S3, integrating cross-modal features through attention mechanism and adversarial training; S4. Use contrastive learning pre-training to enhance model generalization ability; S5. When a modality is detected to be missing, generate pseudo-modality data to complete the input; S6. Use a cascade architecture to achieve a balance between efficient screening and precise diagnosis.

9. An electronic device, characterized in that: include: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method of claim 8.

10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the steps of the method according to claim 8 are implemented.

Citation Information

Cited By

  • Image lesion segmentation system based on convolutional neural network

    CN120599271A

  • Image lesion segmentation system based on convolutional neural network

    CN120599271B

  • Radiotherapy image information fusion method based on neural network model

    CN120725896A

  • Radiation therapy image information fusion method based on neural network model

    CN120725896B

  • Multi-modal image fusion identification method based on comparative learning

    CN120997635A