Multi-modal medical image intelligent fusion diagnosis process
By integrating cross-modal adaptive registration and fusion processing, dynamically weighted fusion feature extraction, and combining intelligent diagnostic analysis and interpretability enhancement, the problems of accuracy, adaptability, and interpretability in multimodal medical image fusion diagnosis are solved, achieving efficient and accurate multimodal image diagnosis.
Patent Information
- Application Number
- CN202610021329.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-05-29
AI Technical Summary
Existing multimodal medical image fusion diagnostic technologies suffer from problems such as cumbersome and error-prone registration and fusion processes, inability to adapt to different clinical scenarios, lack of interpretability, poor stability in cross-center applications, high model update costs, and insufficient personalized adaptation.
It adopts cross-modal adaptive registration and fusion integrated processing, dynamic weighted fusion feature extraction, combined with intelligent diagnostic analysis and interpretability enhancement, and optimizes the model through incremental training mechanism to achieve multi-level feature fusion and privacy protection.
It improves the accuracy and efficiency of image diagnosis, enhances the interpretability of diagnostic results, adapts to the needs of various clinical scenarios, reduces model update costs, and improves the stability and security of cross-center applications.
Smart Images

Figure CN122115324A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent medical imaging diagnosis technology, specifically involving a multimodal medical imaging intelligent fusion diagnosis process. Background Technology
[0002] With the deepening of the concept of precision medicine and the advancement of medical informatization, multimodal medical image fusion diagnosis has become a core supporting technology for clinical diagnosis and treatment. Different imaging modalities have complementary advantages: structural modalities (CT, MRI) can clearly present anatomical details, while functional modalities (PET, ultrasound) can reflect tissue metabolic activity or dynamic changes. Through multimodal data fusion, the information limitations of a single modality can be overcome, providing a more comprehensive basis for disease diagnosis. This technology has been widely recognized in the clinical diagnosis and treatment of various diseases.
[0003] However, existing multimodal medical image fusion diagnostic technologies still suffer from several key shortcomings, severely hindering their clinical translation and large-scale application: Traditional techniques separate cross-modal registration and feature fusion into independent stages, first achieving image alignment through registration and then performing feature fusion, a cumbersome process prone to error accumulation. In clinical scenarios with tissue deformation, traditional registration methods struggle to establish accurate global anatomical correspondences, resulting in blurred lesion boundaries after fusion, failing to meet the needs of accurate diagnosis; even with some deep learning methods attempting optimization, deep synergy between registration and fusion has not been achieved, and distortion easily occurs in the registration of subtle anatomical structures. Existing technologies often employ a unified architecture to extract features from different modalities, ignoring the differences in imaging characteristics among modalities—structural modalities require enhanced detail capture capabilities, while functional modalities need to balance noise suppression and functional information preservation. Simultaneously, the feature extraction process fails to effectively distinguish between common and private features between modalities, leading to information redundancy or loss of key features after fusion, failing to fully leverage the complementary advantages of multimodal approaches.
[0004] Most fusion methods employ fixed-weight linear fusion or simple attention mechanisms, failing to dynamically adjust the fusion strategy based on lesion type, image quality, and clinical scenario. Different clinical scenarios (such as early screening and efficacy evaluation) have varying requirements for each modality's features, but existing methods lack adaptive adjustment mechanisms, resulting in fusion effects that are difficult to adapt to diverse clinical needs and lacking versatility. Existing AI fusion diagnostic models are mostly "black box" architectures; while they can output diagnostic results, they cannot clearly explain the reasoning process, making it difficult to clarify the contribution of each modality's features to the diagnostic conclusion, thus failing to meet the requirements of traceability and interpretability in clinical diagnosis and treatment. Some models even exhibit situations where "the result is correct but the reasoning logic is flawed," increasing the risk of clinical application and reducing physicians' trust in the model.
[0005] Current technologies lack deep integration with clinical diagnostic processes, resulting in inconsistent diagnostic output formats and poor compatibility with hospital PACS systems and electronic medical records (EMR). In multi-center applications, factors such as data distribution differences and privacy protection requirements make it difficult for traditional training models to balance model performance and data security, leading to decreased stability when applied across centers. Furthermore, primary healthcare institutions lack specialized radiologists, and current technologies do not provide simplified procedures or tiered diagnostic guidelines adapted to primary care settings, failing to meet the needs of tiered healthcare policies. Existing models have fixed parameters after training and cannot be dynamically optimized based on new clinical data and treatment guidelines. As disease patterns change and imaging equipment is upgraded, model performance gradually declines, and retraining requires significant computing power and data resources, resulting in high maintenance costs. Moreover, the models do not consider the differences in diagnostic habits among different physicians, lacking sufficient personalized adaptation capabilities.
[0006] To address the shortcomings of existing technologies, this invention proposes a multimodal medical image intelligent fusion diagnostic process. Through innovation across the entire process, it solves the pain points of existing technologies in terms of accuracy, adaptability, and interpretability, thereby achieving more precise, efficient, and clinically adaptable multimodal image diagnosis. Summary of the Invention
[0007] In view of the problems raised in the background technology above, the purpose of this invention is to provide a multimodal medical image intelligent fusion diagnostic process, which solves the pain points of existing technologies in terms of accuracy, adaptability, and interpretability through full-process innovation, and realizes the precision, efficiency and clinical adaptability of multimodal image diagnosis.
[0008] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: The intelligent fusion diagnostic process for multimodal medical images includes the following core process steps: S1: Multimodal image data acquisition and preprocessing. Acquire raw data from at least two clinically commonly used image modalities, and perform format standardization, artifact suppression, grayscale normalization, and quality screening on the raw data to obtain standardized image data. S2: Cross-modal adaptive registration and fusion integrated processing, based on a cross-modal registration and fusion integrated network, simultaneously realizes spatial alignment and preliminary feature fusion of standardized image data, and generates aligned feature maps; S3: Hierarchical feature extraction, adopting a differentiated feature extraction architecture adapted to different modal imaging characteristics, extracting modal common features and modal private features respectively, forming a multi-level feature set; S4: Dynamic weighted fusion. Based on the lesion area identification results and clinical scenario requirements, a dynamic fusion weight allocation mechanism is constructed to adaptively weight and fuse multi-level feature sets to generate the globally optimal fusion feature map. S5: Intelligent diagnostic analysis, inputting the globally optimal fused feature map into the pre-trained diagnostic model, combining it with the clinical diagnostic rule base for multi-dimensional analysis, and outputting preliminary diagnostic results including disease type, lesion parameters, and risk level; S6: Enhanced interpretability and result output. Generate interpretable diagnostic reports through feature contribution visualization and diagnostic logic tracing, and simultaneously output standardized data interfaces to adapt to clinical systems. S7: Dynamic model optimization and privacy protection, building an incremental training mechanism based on clinical feedback data, and combining privacy computing technology to achieve secure iteration of multi-center data.
[0009] Further specifying, the preprocessing process of S1 specifically includes: S11: Format standardization, which converts the raw image data into a unified standardized format, extracts the relevant information in the image data and stores it in association; S12: Artifact suppression, using artifact suppression algorithms for specific correction of typical artifacts in different modalities; S13: Gray-level normalization, which uses a standardization method to map the gray-level values of different modal images to a unified numerical range, eliminating gray-level shifts caused by differences in equipment and scanning parameters; S14: Quality screening, automatically removes low-quality image data based on the image quality assessment index system, and triggers a re-acquisition command for data that does not meet the quality requirements.
[0010] Further specifying, the cross-modal adaptive registration and fusion integrated processing of S2 specifically includes: S21: Construct an integrated registration and fusion network that includes feature interaction units, multi-scale refinement units, and deformation field generation units; S22: Establish cross-modal global anatomical correspondence through feature interaction units to suppress feature confusion in large deformation regions; S23: The multi-scale refinement unit adopts a hierarchical refinement strategy to gradually optimize feature alignment accuracy from coarse to fine granular. S24: The deformation field generation unit achieves spatial alignment through multi-stage iterative optimization to ensure that the registration accuracy meets the requirements of clinical diagnosis. S25: During the registration process, preliminary feature fusion is completed simultaneously. Feature integration is achieved through cross-modal feature correlation calculation, and an aligned feature map is generated. The registration and fusion processing is timely and adapted to the needs of clinical applications.
[0011] Further specifying, the hierarchical feature extraction of S3 specifically includes: S31: The layered feature extraction adopts an asymmetric feature extraction architecture, and designs feature extraction paths with different depths for different modal imaging characteristics to adapt to the deep detail extraction of structural modalities and the noise suppression requirements of functional modalities. S32: Simultaneously captures global contour features and local fine structural features of organs through a global-local feature learning framework; S33: Introduce an attention mechanism to enhance the response weights of lesion-related feature channels and spatial regions; S34: Employ a feature decoupling strategy to separate modal common features from modal private features, avoiding information redundancy and loss of key features; S35: Perform dimensional unification and standardization on the extracted features to form a multi-level feature set containing a basic feature layer, a detailed feature layer, and a functional feature layer.
[0012] Further specifying, the dynamic weighted fusion of S4 specifically includes: S41: Determine the location, extent, and boundary information of lesions based on automatic lesion region segmentation technology, and calculate the salience index of lesion region; S42: The aforementioned clinical scenario recognition model is constructed to automatically identify clinical application scenarios such as early disease screening, lesion classification, efficacy evaluation, and prognosis prediction. S43: Design a dynamic weight allocation mechanism, whose weight allocation is based on modal feature contribution, lesion significance index, and clinical scenario requirements, to achieve adaptive adjustment of fusion weights; S44: Employs a multi-scale feature fusion strategy to decompose, weight, and reconstruct multi-level feature sets to generate the globally optimal fused feature map; S45: Introduce an uncertainty quantification mechanism to assess the reliability of fused features and mark high uncertainty regions.
[0013] Further specifying, the intelligent diagnostic analysis of S5 specifically includes: S51: The pre-trained diagnostic model is built on a deep learning architecture and contains multi-center labeled data, covering a variety of common disease types. S52: The clinical diagnostic rule base integrates the clinical experience and treatment standards of senior physicians, including general diagnostic rules and disease-specific diagnostic rules; S53: The multidimensional analysis includes lesion quantitative analysis, feature correlation analysis, and risk stratification analysis, and outputs lesion-related quantitative parameters and risk assessment results; S54: Based on the diagnostic results, a dual verification mechanism of model diagnosis and rule verification is adopted to trigger multi-model fusion decision for results with insufficient matching degree.
[0014] Further specifying, the interpretability enhancement and result output of S6 specifically include: S61: Generate a feature influence weight map based on feature contribution visualization technology to intuitively show the role of each modality feature in the diagnostic results; S62: Construct a diagnostic logic traceability chain to generate an interpretable diagnostic report, recording key nodes and reasoning basis throughout the entire process from feature extraction and fusion to diagnostic decision-making; S63: The diagnostic report includes an image fusion diagram, lesion annotation diagram, feature contribution diagram, diagnostic conclusion, and clinical recommendations. The format of the diagnostic report is compatible with general clinical standards.
[0015] Further specifying, the model dynamic optimization and privacy protection of S7 specifically include: S71: Construct an incremental training dataset, collect clinical feedback correction data and new case data, and adopt a dynamic weight adjustment mechanism to improve the training priority of high-value cases; S72: Employs a distributed collaborative training framework to achieve multi-center data collaborative training. Data is stored locally and not transmitted externally; only model parameter update information is shared. S73: Based on privacy protection technologies, ensure that the risk of leakage of patient privacy information meets compliance requirements; S74: Establish a model performance monitoring system to continuously track core diagnostic indicators. When an indicator drops below a preset threshold, the incremental training process is automatically triggered.
[0016] Furthermore, the imaging modality includes any combination of CT, MRI, PET, and ultrasound.
[0017] Furthermore, the diagnostic process is applicable to a variety of clinical scenarios: In tumor diagnosis, structural and functional modal data are integrated to achieve tumor staging and differentiation between benign and malignant tumors; in neurological disease diagnosis, multi-parameter neuroimaging data are integrated to assess the extent of lesions and the integrity of neural structures; and in cardiovascular disease diagnosis, cardiovascular-related imaging data are integrated to quantify the degree of vascular lesions and myocardial function. The beneficial effects of this invention are: This invention reduces process connection errors through an integrated registration-fusion design, and ensures registration accuracy for large deformation scenes and fine structures through multi-scale refinement and multi-stage optimization. Simultaneous feature fusion further improves information utilization efficiency, significantly improving the quality and diagnostic usability of the fused image. By visualizing feature contribution and tracing diagnostic logic, the invention achieves transparency and traceability of the diagnostic process, allowing clinicians to clearly understand the diagnostic basis, meeting the interpretability requirements of clinical diagnosis and treatment, significantly increasing physicians' trust in diagnostic results, and lowering the threshold for clinical application. This invention achieves seamless integration with existing clinical systems and adapts to clinical diagnosis and treatment processes through standardized diagnostic reports and diverse data interfaces. The simplified process design for primary care settings enhances the applicability of the technology in grassroots medical institutions, facilitating the implementation of hierarchical medical services. A comprehensive quality control and anomaly handling mechanism ensures the stability and continuity of the process operation. The diagnostic process outputs lesion-related quantitative parameters, risk assessment results, and clinical recommendations, providing comprehensive support for disease diagnosis and treatment and efficacy evaluation, thus improving clinical efficiency and accuracy. Interpretable design and optimized clinical adaptability accelerate the clinical translation and large-scale application of the technology. Standardized data interfaces in telemedicine scenarios facilitate the downward flow of high-quality medical resources and improve the accessibility of medical services. Attached Figure Description
[0018] The present invention can be further illustrated by the non-limiting embodiments given in the accompanying drawings; Figure 1 This is a flowchart illustrating the steps of an embodiment of the multimodal medical image intelligent fusion diagnostic process of the present invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments. The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. The technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0020] like Figure 1 As shown, the multimodal medical image intelligent fusion diagnostic process of the present invention includes the following core process steps: S1: Multimodal image data acquisition and preprocessing. Acquire raw data from at least two clinically commonly used image modalities, and perform format standardization, artifact suppression, grayscale normalization, and quality screening on the raw data to obtain standardized image data. S2: Cross-modal adaptive registration and fusion integrated processing, based on a cross-modal registration and fusion integrated network, simultaneously realizes spatial alignment and preliminary feature fusion of standardized image data, and generates aligned feature maps; S3: Hierarchical feature extraction, adopting a differentiated feature extraction architecture adapted to different modal imaging characteristics, extracting modal common features and modal private features respectively, forming a multi-level feature set; S4: Dynamic weighted fusion. Based on the lesion area identification results and clinical scenario requirements, a dynamic fusion weight allocation mechanism is constructed to adaptively weight and fuse multi-level feature sets to generate the globally optimal fusion feature map. S5: Intelligent diagnostic analysis, inputting the globally optimal fused feature map into the pre-trained diagnostic model, combining it with the clinical diagnostic rule base for multi-dimensional analysis, and outputting preliminary diagnostic results including disease type, lesion parameters, and risk level; S6: Enhanced interpretability and result output. Generate interpretable diagnostic reports through feature contribution visualization and diagnostic logic tracing, and simultaneously output standardized data interfaces to adapt to clinical systems. S7: Dynamic model optimization and privacy protection, building an incremental training mechanism based on clinical feedback data, and combining privacy computing technology to achieve secure iteration of multi-center data.
[0021] Specifically, the acquired image modalities are selected as a combination of CT (structural modality), MRI (soft tissue structural modality), and PET (functional metabolic modality), commonly used in lung cancer diagnosis; preprocessing employs DICOM to NIfTI format standardization, CT partial volume effect SIRT algorithm correction, z-score grayscale normalization, and SNR≥20dB quality screening; the registration-fusion integrated network adopts an improved PCRNet architecture; the hierarchical feature extraction uses an asymmetric architecture of 6-layer MRI Restormer + 3-layer PET Restormer; the dynamic weighted fusion allocates PET weights to 45%, CT weights to 35%, and MRI weights to 20% according to the early screening scenario; and the diagnostic model is SwingTransformer.
[0022] In practical applications of this embodiment, the preprocessing process of S1 specifically includes: S11: Format standardization, which converts the raw image data into a unified standardized format, extracts the relevant information in the image data and stores it in association; Specifically, the raw DICOM format data output from equipment from different manufacturers such as Siemens and GE is converted to NIfTI format using Python's pydicom library. Information such as patient ID, scan slice thickness, tube voltage, and scan time in the metadata is extracted and stored in the database in association with the image data. S12: Artifact suppression, using artifact suppression algorithms for specific correction of typical artifacts in different modalities; Specifically, the partial volume effect of CT is corrected using the SIRT iterative reconstruction algorithm (20 iterations), the magnetic susceptibility artifacts of MRI are processed using multi-echo signal fusion correction technology, and the speckle noise of ultrasound is suppressed using the wavelet threshold denoising algorithm. S13: Gray-level normalization, which uses a standardization method to map the gray-level values of different modal images to a unified numerical range, eliminating gray-level shifts caused by differences in equipment and scanning parameters; Specifically, the z-score standardization formula (x'=(x-μ) / σ) is used to uniformly map the gray values of CT, MRI, and PET to the [0,1] interval, where μ is the mean gray value of the image and σ is the standard deviation. S14: Quality screening, automatically removes low-quality image data based on the image quality assessment index system, and triggers a re-acquisition command for data that does not meet the quality requirements.
[0023] Specifically, the quality assessment index system includes three core indicators: signal-to-noise ratio (SNR), image entropy, and edge sharpness. Screening thresholds are set for SNR ≥ 20dB, image entropy ≥ 6.5 bits / pixel, and edge sharpness ≥ 0.8 to automatically remove low-quality data.
[0024] In practical applications of this embodiment, the cross-modal adaptive registration and fusion integration processing of S2 specifically includes: S21: Construct an integrated registration and fusion network that includes feature interaction units, multi-scale refinement units, and deformation field generation units; Specifically, the registration and fusion integrated network includes a cross-attention transformer (feature interaction unit), a pyramid-shaped attention module (multi-scale refinement unit), and a three-stage deformation field generator (deformation field generation unit), which is built on the PyTorch framework. S22: Establish cross-modal global anatomical correspondence through feature interaction units to suppress feature confusion in large deformation regions; Specifically, the feature interaction unit establishes a global anatomical correspondence between CT and MRI through bidirectional attention weight calculation. For the abdominal liver respiratory deformation area, the attention weight enhancement coefficient is set to 1.2 to suppress feature confusion. S23: The multi-scale refinement unit adopts a hierarchical refinement strategy to gradually optimize feature alignment accuracy from coarse to fine granular. Specifically, the multi-scale refinement unit is divided into three levels: 128×128×128 (coarse-grained), 192×192×192 (medium-grained), and 256×256×256 (fine-grained), with each level using a 3×3×3 convolution kernel to refine the features; S24: The deformation field generation unit achieves spatial alignment through multi-stage iterative optimization to ensure that the registration accuracy meets the requirements of clinical diagnosis. Specifically, the deformation field generation unit adopts a three-stage iteration of "coarse correction - intermediate optimization - fine alignment". The coarse stage corrects the proportional distortion (5 iterations), the intermediate stage corrects the local deformation (10 iterations), and the fine stage achieves pixel-level alignment (15 iterations). S25: During the registration process, preliminary feature fusion is completed simultaneously. Feature integration is achieved through cross-modal feature correlation calculation, and an aligned feature map is generated. The registration and fusion processing is timely and adapted to the needs of clinical applications.
[0025] In the practical application of this embodiment, the hierarchical feature extraction of S3 specifically includes: S31: The layered feature extraction adopts an asymmetric feature extraction architecture, and designs feature extraction paths with different depths for different modal imaging characteristics to adapt to the deep detail extraction of structural modalities and the noise suppression requirements of functional modalities. Specifically, in the asymmetric feature extraction architecture, MRI (structural modality) uses a 6-layer Restormer stack, and PET (functional modality) uses a 3-layer Restormer stack; S32: Simultaneously captures global contour features and local fine structural features of organs through a global-local feature learning framework; Specifically, in the global-local feature learning framework, global features are extracted using a 7×7×7 large convolutional kernel (such as the overall outline of the liver), and local features are captured using a 1×1×1 small convolutional kernel (such as hepatic vein branches and tumor microinfiltration areas). S33: Introduce an attention mechanism to enhance the response weights of lesion-related feature channels and spatial regions; Specifically, the attention mechanism uses the SE module (squeezing coefficient r=16) to enhance the weight of lesion-related feature channels, and the spatial region uses the CBAM module to focus on the lesion spatial region through a 2×2×2 pooling kernel; S34: Employ a feature decoupling strategy to separate modal common features from modal private features, avoiding information redundancy and loss of key features; Specifically, the feature decoupling strategy employs an orthogonal constraint loss function to separate modal common features (lesion location, size, and shape) from private features (T2-weighted signal intensity of MRI and metabolic activity distribution of PET). S35: Perform dimensional unification and standardization on the extracted features to form a multi-level feature set containing a basic feature layer, a detailed feature layer, and a functional feature layer.
[0026] In the practical application of this embodiment, the dynamic weighted fusion of S4 specifically includes: S41: Determine the location, extent, and boundary information of lesions based on automatic lesion region segmentation technology, and calculate the salience index of lesion region; Specifically, the U-Net++ network is used for automatic segmentation of lesion regions (4 layers of encoder and 4 layers of decoder), and the significance index of lesion regions (SI = number of lesion pixels / total number of pixels in the image) is calculated. A lesion with SI ≥ 0.05 is considered a highly significant lesion. S42: The aforementioned clinical scenario recognition model is constructed to automatically identify clinical application scenarios such as early disease screening, lesion classification, efficacy evaluation, and prognosis prediction. Specifically, the clinical scene recognition model is built based on the LightGBM algorithm. The input features are three types: image type, examination purpose, and patient medical history. The output features are four types: early disease screening, lesion classification, efficacy evaluation, and prognosis prediction. S43: Design a dynamic weight allocation mechanism, whose weight allocation is based on modal feature contribution, lesion significance index, and clinical scenario requirements, to achieve adaptive adjustment of fusion weights; Specifically, the dynamic weight allocation mechanism adopts a linear weighting function W=α×F+β×L+γ×S, (α+β+γ=1), where F is the modal feature contribution, L is the lesion significance index, and S is the scenario weight. For example, in the early disease screening scenario: α=0.4, β=0.3, γ=0.3. S44: Employs a multi-scale feature fusion strategy to decompose, weight, and reconstruct multi-level feature sets to generate the globally optimal fused feature map; S45: Introduce an uncertainty quantification mechanism to assess the reliability of fused features and mark high uncertainty regions.
[0027] In practical applications of this embodiment, the intelligent diagnostic analysis of S5 specifically includes: S51: The pre-trained diagnostic model is built on a deep learning architecture and contains multi-center labeled data, covering a variety of common disease types. Specifically, the pre-trained diagnostic model is based on the Swing Transformer V2 architecture, and the pre-training dataset contains 5,000 multi-center labeled data (covering 10 types of diseases such as lung cancer, brain tumors, and liver cancer). S52: The clinical diagnostic rule base integrates the clinical experience and treatment standards of senior physicians, including general diagnostic rules and disease-specific diagnostic rules; Specifically, the clinical diagnostic rule base is constructed using the Delphi method and includes 30 general diagnostic rules (such as "lesion diameter > 3cm and irregular borders suggest malignancy") and 50 disease-specific rules (lung cancer-specific rule: "lobulation sign + pleural traction sign + SUV > 2.5 suggest malignancy"). The rule base is stored in XML format and supports dynamic updates. S53: The multidimensional analysis includes lesion quantitative analysis, feature correlation analysis, and risk stratification analysis, and outputs lesion-related quantitative parameters and risk assessment results; S54: Based on the diagnostic results, a dual verification mechanism of model diagnosis and rule verification is adopted to trigger multi-model fusion decision for results with insufficient matching degree.
[0028] In practical applications of this embodiment, the interpretability enhancement and result output of S6 specifically include: S61: Generate a feature influence weight map based on feature contribution visualization technology to intuitively show the role of each modality feature in the diagnostic results; Specifically, the feature contribution visualization uses SHAP value calculation (kernel function type RBF) to generate a two-dimensional heat map and a three-dimensional lesion overlay map, which intuitively shows the contribution ratio of each modality feature to the diagnostic results; S62: Construct a diagnostic logic traceability chain to generate an interpretable diagnostic report, recording key nodes and reasoning basis throughout the entire process from feature extraction and fusion to diagnostic decision-making; S63: The diagnostic report includes an image fusion diagram, lesion annotation diagram, feature contribution diagram, diagnostic conclusion, and clinical recommendations. The format of the diagnostic report is compatible with general clinical standards.
[0029] In the practical application of this embodiment, the model dynamic optimization and privacy protection of S7 specifically include: S71: Construct an incremental training dataset, collect clinical feedback correction data and new case data, and adopt a dynamic weight adjustment mechanism to improve the training priority of high-value cases; Specifically, the incremental training dataset includes 1,000 misdiagnosed / missed diagnoses corrected by clinicians and 500 new rare disease cases. The dynamic weight adjustment uses the Focal Loss function to increase the training weights for rare and difficult cases. S72: Employs a distributed collaborative training framework to achieve multi-center data collaborative training. Data is stored locally and not transmitted externally; only model parameter update information is shared. Specifically, the distributed collaborative training adopts the federated learning framework (FedAvg algorithm), with participating nodes including at least 5 tertiary hospitals. The data of each node is stored locally on the hospital's intranet server, and the encrypted model parameters are shared only through the HTTPS protocol. S73: Based on privacy protection technologies, ensure that the risk of leakage of patient privacy information meets compliance requirements; Specifically, the privacy protection technology uses homomorphic encryption (Paillier algorithm) to encrypt model parameters and differential privacy technology (noise intensity ε=1.5) to protect training data; S74: Establish a model performance monitoring system to continuously track core diagnostic indicators. When an indicator drops below a preset threshold, the incremental training process is automatically triggered.
[0030] Specifically, the model performance monitoring system includes four core indicators: diagnostic accuracy, sensitivity, specificity, and AUC value. A trigger threshold is set when the indicator decreases by more than 5% for three consecutive months, and the incremental training process is automatically started.
[0031] In practical applications of this embodiment, the imaging modality includes any combination of CT, MRI, PET, and ultrasound.
[0032] In practical applications of this embodiment, the diagnostic process is applicable to a variety of clinical scenarios: In tumor diagnosis, structural and functional modal data are integrated to achieve tumor staging and differentiation between benign and malignant tumors; in neurological disease diagnosis, multi-parameter neuroimaging data are integrated to assess the extent of lesions and the integrity of neural structures; and in cardiovascular disease diagnosis, cardiovascular-related imaging data are integrated to quantify the degree of vascular lesions and myocardial function.
[0033] The working principle of this embodiment is as follows: By standardizing preprocessing, raw images from different modalities and devices are converted into standardized data in a unified format, eliminating data heterogeneity and providing a high-quality data foundation for subsequent registration, fusion, and feature extraction. Then, based on the registration-fusion integrated network, spatial alignment and feature fusion of cross-modal images are realized simultaneously, reducing errors caused by process separation, ensuring the accuracy and integrity of fused features, and improving processing efficiency. By adopting a differentiated feature extraction architecture and attention mechanism, common and private features of each modality are extracted in a targeted manner, maximizing the complementary advantages of multimodal approaches and providing high-quality feature support for accurate diagnosis. By employing a dual verification mechanism of "model diagnosis + rule validation" combined with multi-dimensional analysis, accurate diagnostic results are achieved. Simultaneously, feature contribution visualization and diagnostic logic traceability enhance the interpretability of the diagnostic process and increase clinical trust. Based on clinical feedback data and multi-center collaborative training, incremental training continuously optimizes model performance while introducing privacy protection technologies to ensure data security and compliance in multi-center applications. This forms a closed-loop operation mechanism of "data-model-diagnosis-feedback-optimization," guaranteeing the long-term applicability and clinical value of the process.
[0034] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A multimodal medical image intelligent fusion diagnostic process, characterized in that, The core process steps include the following: S1: Multimodal image data acquisition and preprocessing. Acquire raw data from at least two clinically commonly used image modalities, and perform format standardization, artifact suppression, grayscale normalization, and quality screening on the raw data to obtain standardized image data. S2: Cross-modal adaptive registration and fusion integrated processing, based on a cross-modal registration and fusion integrated network, simultaneously realizes spatial alignment and preliminary feature fusion of standardized image data, and generates aligned feature maps; S3: Hierarchical feature extraction, adopting a differentiated feature extraction architecture adapted to different modal imaging characteristics, extracting modal common features and modal private features respectively, forming a multi-level feature set; S4: Dynamic weighted fusion. Based on the lesion area identification results and clinical scenario requirements, a dynamic fusion weight allocation mechanism is constructed to adaptively weight and fuse multi-level feature sets to generate the globally optimal fusion feature map. S5: Intelligent diagnostic analysis, inputting the globally optimal fused feature map into the pre-trained diagnostic model, combining it with the clinical diagnostic rule base for multi-dimensional analysis, and outputting preliminary diagnostic results including disease type, lesion parameters, and risk level; S6: Enhanced interpretability and result output. Generate interpretable diagnostic reports through feature contribution visualization and diagnostic logic tracing, and simultaneously output standardized data interfaces to adapt to clinical systems. S7: Dynamic model optimization and privacy protection, building an incremental training mechanism based on clinical feedback data, and combining privacy computing technology to achieve secure iteration of multi-center data.
2. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The preprocessing process of S1 specifically includes: S11: Format standardization, which converts the raw image data into a unified standardized format, extracts the relevant information in the image data and stores it in association; S12: Artifact suppression, using artifact suppression algorithms for specific correction of typical artifacts in different modalities; S13: Gray-level normalization, which uses a standardization method to map the gray-level values of different modal images to a unified numerical range, eliminating gray-level shifts caused by differences in equipment and scanning parameters; S14: Quality screening, automatically removes low-quality image data based on the image quality assessment index system, and triggers a re-acquisition command for data that does not meet the quality requirements.
3. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that, The cross-modal adaptive registration and fusion integrated processing of S2 specifically includes: S21: Construct an integrated registration and fusion network that includes feature interaction units, multi-scale refinement units, and deformation field generation units; S22: Establish cross-modal global anatomical correspondence through feature interaction units to suppress feature confusion in large deformation regions; S23: The multi-scale refinement unit adopts a hierarchical refinement strategy to gradually optimize feature alignment accuracy from coarse to fine granular. S24: The deformation field generation unit achieves spatial alignment through multi-stage iterative optimization to ensure that the registration accuracy meets the requirements of clinical diagnosis. S25: During the registration process, preliminary feature fusion is completed simultaneously. Feature integration is achieved through cross-modal feature correlation calculation, and an aligned feature map is generated. The registration and fusion processing is timely and adapted to the needs of clinical applications.
4. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The hierarchical feature extraction of S3 specifically includes: S31: The layered feature extraction adopts an asymmetric feature extraction architecture, and designs feature extraction paths with different depths for different modal imaging characteristics to adapt to the deep detail extraction of structural modalities and the noise suppression requirements of functional modalities. S32: Simultaneously captures global contour features and local fine structural features of organs through a global-local feature learning framework; S33: Introduce an attention mechanism to enhance the response weights of lesion-related feature channels and spatial regions; S34: Employ a feature decoupling strategy to separate modal common features from modal private features, avoiding information redundancy and loss of key features; S35: Perform dimensional unification and standardization on the extracted features to form a multi-level feature set containing a basic feature layer, a detailed feature layer, and a functional feature layer.
5. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The dynamic weighted fusion of S4 specifically includes: S41: Determine the location, extent, and boundary information of lesions based on automatic lesion region segmentation technology, and calculate the salience index of lesion region; S42: The aforementioned clinical scenario recognition model is constructed to automatically identify clinical application scenarios such as early disease screening, lesion classification, efficacy evaluation, and prognosis prediction. S43: Design a dynamic weight allocation mechanism, whose weight allocation is based on modal feature contribution, lesion significance index, and clinical scenario requirements, to achieve adaptive adjustment of fusion weights; S44: Employs a multi-scale feature fusion strategy to decompose, weight, and reconstruct multi-level feature sets to generate the globally optimal fused feature map; S45: Introduce an uncertainty quantification mechanism to assess the reliability of fused features and mark high uncertainty regions.
6. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The intelligent diagnostic analysis of S5 specifically includes: S51: The pre-trained diagnostic model is built on a deep learning architecture and contains multi-center labeled data, covering a variety of common disease types. S52: The clinical diagnostic rule base integrates the clinical experience and treatment standards of senior physicians, including general diagnostic rules and disease-specific diagnostic rules; S53: The multidimensional analysis includes lesion quantitative analysis, feature correlation analysis, and risk stratification analysis, and outputs lesion-related quantitative parameters and risk assessment results; S54: Based on the diagnostic results, a dual verification mechanism of model diagnosis and rule verification is adopted to trigger multi-model fusion decision for results with insufficient matching degree.
7. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The enhanced interpretability and output of S6 specifically include: S61: Generate a feature influence weight map based on feature contribution visualization technology to intuitively show the role of each modality feature in the diagnostic results; S62: Construct a diagnostic logic traceability chain to generate an interpretable diagnostic report, recording key nodes and reasoning basis throughout the entire process from feature extraction and fusion to diagnostic decision-making; S63: The diagnostic report includes an image fusion diagram, lesion annotation diagram, feature contribution diagram, diagnostic conclusion, and clinical recommendations. The format of the diagnostic report is compatible with general clinical standards.
8. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The specific aspects of model dynamic optimization and privacy protection in S7 include: S71: Construct an incremental training dataset, collect clinical feedback correction data and new case data, and adopt a dynamic weight adjustment mechanism to improve the training priority of high-value cases; S72: Employs a distributed collaborative training framework to achieve multi-center data collaborative training. Data is stored locally and not transmitted externally; only model parameter update information is shared. S73: Based on privacy protection technologies, ensure that the risk of leakage of patient privacy information meets compliance requirements; S74: Establish a model performance monitoring system to continuously track core diagnostic indicators. When an indicator drops below a preset threshold, the incremental training process is automatically triggered.
9. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The imaging modalities include any combination of CT, MRI, PET, and ultrasound.
10. The multimodal medical image intelligent fusion diagnostic process according to claim 1, characterized in that: The diagnostic process is applicable to a variety of clinical scenarios: In tumor diagnosis, structural and functional modal data are integrated to achieve tumor staging and differentiation between benign and malignant tumors; in neurological disease diagnosis, multi-parameter neuroimaging data are integrated to assess the extent of lesions and the integrity of neural structures; and in cardiovascular disease diagnosis, cardiovascular-related imaging data are integrated to quantify the degree of vascular lesions and myocardial function.