Medical image processing method and system, computer equipment and storage medium

By combining cross-modal image generation with deep learning segmentation networks, the problem of low diagnostic accuracy caused by missing key sequences in MRI scans is solved, enabling high-precision automated identification and quantitative analysis of MS lesions, supporting early screening and treatment of MS.

CN121661002APending Publication Date: 2026-03-13FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing deep learning models have low diagnostic accuracy when key sequences are missing in MRI scans, making it difficult to meet the needs for accurate identification and assessment of multiple sclerosis (MS) lesions.

Method used

By combining cross-modal image generation with segmentation that incorporates prior knowledge, a multi-scale feature extraction and generation network is constructed to generate high-fidelity images of brain and neck targets. Using a three-level generative adversarial network and a deep learning segmentation network, the automated identification and quantitative analysis of brain and cervical spinal cord lesions are achieved.

Benefits of technology

In cases where key sequences are missing, the identification rate and segmentation accuracy of MS lesions are improved, and comprehensive and reliable quantitative information on lesions is provided, supporting early screening and treatment of MS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661002A_ABST
    Figure CN121661002A_ABST
Patent Text Reader

Abstract

The invention provides a medical image processing method and system, computer equipment and a storage medium, and belongs to the field of image processing.The medical image processing method comprises the steps that T1 weighted imaging and T2 weighted imaging in brain and neck medical images are collected, multi-scale feature extraction is conducted on the T1 weighted imaging and the T2 weighted imaging, and brain and neck region features are obtained; generating a first target image according to the brain and neck region features; generating a second target image from the first target image based on a dense connection mechanism; optimizing a signal of a focus area in the second target image through a focus perception loss function fused with the focus automatic segmentation result to obtain an optimized target image; and synthesizing a cross-modal brain target image and a cross-modal neck target image according to the first target image, the second target image and the optimized target image. According to the method, the problem of key sequence deletion in multiple sclerosis diagnosis is solved, the retention rate of small focuses is increased, and reliable technical support is provided for MS early screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing, and specifically relates to a medical image processing method, system, computer equipment, and storage medium. Background Technology

[0002] Multiple sclerosis (MS) is a chronic inflammatory demyelinating disease of the central nervous system, and its diagnosis and assessment are highly dependent on magnetic resonance imaging (MRI). Accurate MRI assessment requires comprehensive analysis of multimodal sequences: T1-weighted imaging (T1WI) and T2-weighted imaging (T2WI) are routinely used as the basic scanning sequences in clinical practice; in addition, fluid attenuated inversion recovery sequence (FLAIR) has high sensitivity for detecting MS lesions in the brain, especially periventricular and subcortical lesions; while phase-sensitive inversion recovery sequence (PSIR) and its magnetic moment map (PSIRMAG) can significantly improve the contrast between cervical spinal cord lesions and surrounding cerebrospinal fluid, and have unique advantages in the detection and identification of cervical MS lesions. By comprehensively using the above multimodal sequences, MS lesions distributed in the brain and cervical spinal cord can be identified comprehensively and accurately. However, in actual clinical scanning, due to factors such as scanning time limitations, equipment differences, or patient intolerance, key sequences (such as FLAIR, PSIR, and PSIRMAG) are often missing or incomplete, which seriously affects the accuracy and efficiency of diagnosis.

[0003] To address the shortcomings of manual evaluation of MS images, researchers have employed two main approaches: firstly, traditional image processing methods or early machine learning techniques for sequence generation. However, the resulting images are often of low quality, exhibiting issues such as blurred details and distorted contrast, failing to meet clinical diagnostic needs. Secondly, deep learning segmentation models have been used for lesion segmentation. However, these models typically rely on complete sequence input, and their performance deteriorates sharply when key sequences are missing. In summary, existing techniques offer limited accuracy when key sequences are lacking in the image. Summary of the Invention

[0004] To address the problem that existing deep learning models cannot handle missing sequences and have low evaluation accuracy, this invention provides a medical image processing method, system, computer device, and storage medium.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A medical image processing method, comprising: T1-weighted and T2-weighted imaging data were acquired from brain medical images and neck medical images. Multi-scale feature extraction was performed on the T1-weighted and T2-weighted imaging data of the brain medical images to obtain brain region features; multi-scale feature extraction was performed on the T1-weighted and T2-weighted imaging data of the neck medical images to obtain neck region features. First brain target images and first neck target images are constructed based on the features of the brain and neck regions, respectively. Based on a dense connection mechanism, a second brain target image is generated from the first brain target image, and a second neck target image is generated from the first neck target image. The resolution of the second brain target image is higher than that of the first brain target image, and the resolution of the second neck target image is higher than that of the first neck target image. Using a pre-set lesion perception loss function, the signals of the lesion regions in the second brain target image and the second neck target image are optimized, respectively, to obtain optimized brain target images and optimized neck target images. Cross-modal brain target images and cross-modal neck target images are generated based on the first brain target image and the first neck target image, the second brain target image and the second neck target image, and the optimized brain target image and the optimized neck target image, respectively.

[0006] Optionally, the medical image processing method provided by the present invention processes brain region features and neck region features through a three-level generative adversarial network to obtain cross-modal brain target images and cross-modal neck target images; The three-level generative adversarial network consists of a feature extraction layer, a sequence generation layer, and a discriminator connected in sequence. The feature extraction layer is composed of a ResNet-50 network. The sequence generation layer includes a first-level generator, a second-level generator, and a third-level generator set in parallel. The first-level generator is based on the U-net architecture, the second-level generator has densely connected blocks, and the third-level generator has a lesion-sensing loss function.

[0007] Optionally, the medical image processing method provided by the present invention further includes: The structural similarity index and peak signal-to-noise ratio of the first brain target image, the first neck target image, the second brain target image, the second neck target image, the optimized brain target image, and the optimized neck target image are calculated by the discriminator respectively. When the structural similarity index is less than the structural similarity index threshold, new first brain target image and first neck target image, second brain target image and second neck target image, and optimized brain target image and optimized neck target image are regenerated from the texture features of the brain and neck regions through the sequence generation layer. When the peak signal-to-noise ratio is less than the signal-to-noise ratio threshold, new first brain target image and first neck target image, second brain target image and second neck target image, and optimized brain target image and optimized neck target image are regenerated from the texture features of the brain and neck regions through the sequence generation layer. When the structural similarity index is greater than the structural similarity index threshold and the peak signal-to-noise ratio is greater than the signal-to-noise ratio threshold, the first brain target image and the first neck target image, the second brain target image and the second neck target image, and the optimized brain target image and the optimized neck target image are merged to obtain the cross-modal brain target image and the cross-modal neck target image.

[0008] Optionally, in the medical image processing method provided by the present invention, the cross-modal brain target image is a brain FLAIR image, and the cross-modal neck target image includes a neck PSIR image and a PSIRMAG image.

[0009] Optionally, the medical image processing method provided by the present invention further includes: By processing cross-modal brain target images and cross-modal neck target images through a pre-trained deep learning segmentation network, the lesion segmentation results of brain lesions and cervical spinal cord lesions are obtained. The quantitative information and spatial distribution characteristics of the lesions are obtained based on the lesion segmentation results.

[0010] Optionally, in a medical image processing method provided by the present invention, the deep learning segmentation network specifically includes an encoding module and a feature fusion and reconstruction module connected in sequence, and the encoding module includes a parallel SAM image encoder and an nnUNet encoder. Visual features were extracted from cross-modal brain and neck target images using the SAM image encoder, and medical-specific features were extracted from cross-modal brain and neck target images using the nnUNet encoder. The feature fusion and reconstruction module integrates visual features and medical-specific features to obtain lesion segmentation results. Based on the lesion segmentation results, the quantitative information and spatial distribution characteristics of brain lesions and cervical spinal cord lesions are determined.

[0011] Optionally, in the medical image processing method provided by the present invention, the quantitative information includes the number, volume, and major and minor axes of the lesions; the spatial distribution characteristics include the location distribution of the lesions around the ventricles, subcortical, infratentorial, and cervical spinal cord.

[0012] The present invention also provides a medical image processing system, comprising: The image acquisition and feature extraction module is used to acquire T1-weighted and T2-weighted images from brain medical images and neck medical images; to perform multi-scale feature extraction on T1-weighted and T2-weighted images from brain medical images to obtain brain region features; and to perform multi-scale feature extraction on T1-weighted and T2-weighted images from neck medical images to obtain neck region features. The image generation and optimization module is used to construct a first brain target image and a first neck target image based on brain region features and neck region features, respectively; based on a dense connection mechanism, a second brain target image is generated from the first brain target image, and a second neck target image is generated from the first neck target image, wherein the resolution of the second brain target image is higher than that of the first brain target image, and the resolution of the second neck target image is higher than that of the first neck target image; through a pre-set lesion perception loss function, the signals of lesion regions in the second brain target image and the second neck target image are optimized, respectively, to obtain optimized brain target images and optimized neck target images; A cross-modal image generation module is used to synthesize cross-modal brain target images and cross-modal neck target images based on a first brain target image and a first neck target image, a second brain target image and a second neck target image, and optimized brain target images and optimized neck target images, respectively.

[0013] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any of the steps of a medical image processing method.

[0014] The present invention also provides a computer-readable storage medium storing a computer program that, when loaded by a processor, can execute any step of a medical image processing method.

[0015] The medical image processing method provided by this invention has the following beneficial effects: Because the medical image processing method provided by this invention can combine prior knowledge of lesions to generate high-fidelity brain and neck target images from multi-scale features of T1 and T2 imaging sequences, it can improve the retention rate of smaller MS lesions while maintaining the accuracy of anatomical structures. This effectively solves the problem of low image evaluation accuracy caused by missing key sequences and provides reliable technical support for doctors' early MS screening. Attached Figure Description

[0016] To more clearly illustrate the embodiments and design schemes of the present invention, the accompanying drawings required for this embodiment will be briefly described below. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of a medical image processing method provided in an embodiment of the present invention; Figure 2 Flowchart for generating brain FLAIR and neck PSIR, PSIRMAG target images from brain and neck T1WI / T2WI sequences provided in embodiments of the present invention; Figure 3 This is a schematic diagram of the MS lesion segmentation model provided in an embodiment of the present invention; Figure 4 This is an example of a complete medical image processing workflow provided for embodiments of the present invention. Detailed Implementation

[0018] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.

[0019] To address the need for lesion identification in multiple sclerosis (MS), existing technologies suffer from two main challenges. Firstly, current models rely excessively on key sequences, which can significantly reduce lesion segmentation performance when these sequences are missing. Secondly, existing methods often focus on single sites, lacking a collaborative analysis and unified diagnostic framework for lesions in both the brain and cervical spinal cord. This results in limited lesion segmentation effectiveness and hinders physicians' ability to comprehensively assess disease progression. Therefore, there is an urgent need for a technological solution that can overcome the limitations of missing key sequences and simultaneously provide accurate, automated identification and quantitative analysis of MS lesions in both the brain and cervical spinal cord.

[0020] This invention provides a medical image processing method that addresses the problem of missing key diagnostic sequences, such as FLAIR in the brain, PSIR in the neck, and PSIRMAG, in MRI scans, hindering diagnosis. Its core lies in combining cross-modal image generation with segmentation incorporating prior knowledge to construct a comprehensive solution from conventional sequence input to multi-site quantitative analysis. Specifically, it includes acquiring T1-weighted and T2-weighted images from brain medical images, obtaining brain region features through multi-scale feature extraction, generating brain FLAIR images based on these features using a pre-trained cross-modal image generation model, similarly acquiring T1-weighted and T2-weighted images from neck medical images, obtaining neck region features through multi-scale feature extraction, and generating neck PSIR and PSIRMAG images based on these features using a pre-trained cross-modal image generation model; annotating MS lesions in multi-site, multi-modal MRI images, training a segmentation model using a network framework fusing visual priors and a large-scale basic segmentation model; and inputting the MRI images into the trained segmentation model to obtain quantitative information and spatial distribution characteristics of MS lesions in the brain and cervical spinal cord.

[0021] First, the medical image processing method provided by this invention can generate high-quality key sequences such as FLAIR, PSIR, and PSIRMAG required for clinical diagnosis from conventional sequences such as T1WI and T2WI using advanced cross-modal image generation models, such as GAN models or diffusion models. This effectively overcomes the diagnostic difficulties caused by sequence loss, ensures the integrity and reliability of the diagnostic process, and solves the problem of key sequence dependence.

[0022] Secondly, the medical image processing method provided by this invention is the first to collaboratively process brain and neck images in an automated process, and uses a uniformly trained segmentation model to analyze lesions, providing more comprehensive quantitative information and spatial distribution characteristics of lesions, which meets the clinical practice requirements of MS diagnosis, realizes integrated analysis of multiple sites, and improves the diagnostic efficiency of doctors.

[0023] Furthermore, the medical image processing method provided by this invention innovatively integrates the general prior knowledge of basic large-scale visual models such as SAM with the medical domain adaptive capabilities of nnUNet using the nnSAM segmentation framework. This integration makes the model highly robust to changes in image quality, contrast, and lesion morphology diversity, significantly improving the recognition rate and segmentation accuracy of MS lesions in the brain and neck, especially small and blurred lesions, thereby enhancing the segmentation accuracy and robustness of the segmentation model.

[0024] Finally, the medical image processing method provided by this invention can also output quantitative data and distribution maps of lesions, providing objective and accurate data support for doctors to assess the condition, monitor the efficacy and predict the prognosis, which is helpful for the early screening and precise diagnosis and treatment of MS.

[0025] Example 1 This invention provides a medical image processing method, specifically as follows: Figure 2 As shown, it includes the following steps: Step 11: Acquire T1-weighted and T2-weighted imaging from brain medical images and T1-weighted and T2-weighted imaging from neck medical images.

[0026] Step 12: Preprocess the T1-weighted and T2-weighted images in brain and neck medical images to obtain standardized T1-weighted and T2-weighted images.

[0027] Step 13: Perform multi-scale feature extraction on T1-weighted and T2-weighted imaging of brain medical images to obtain brain region features; perform multi-scale feature extraction on T1-weighted and T2-weighted imaging of neck medical images to obtain neck region features.

[0028] Step 14: Construct a first brain target image and a first neck target image based on the features of the brain and neck regions, respectively; Based on a dense connection mechanism, generate a second brain target image from the first brain target image and a second neck target image from the first neck target image, wherein the resolution of the second brain target image is higher than that of the first brain target image, and the resolution of the second neck target image is higher than that of the first neck target image; By using a lesion perception loss function that integrates the results of automatic lesion segmentation, optimize the signals of the lesion regions in the second brain target image and the second neck target image, respectively, to obtain optimized brain target images and optimized neck target images.

[0029] Step 15: Synthesize the final cross-modal brain target image and cross-modal neck target image based on the first brain target image, the first neck target image, the second brain target image, the second neck target image, the optimized brain target image, and the optimized neck target image, respectively. For example, the cross-modal brain target image is a brain FLAIR image, and the cross-modal neck target image includes a neck PSIR image and a PSIRMAG image.

[0030] Specifically, a three-tiered generative adversarial network architecture constructed using generative adversarial networks (GANs) or diffusion models can be used to process brain and neck region features, resulting in cross-modal brain and neck target images. The three-tiered GAN consists of a feature extraction layer, a sequence generation layer, and a discriminator connected in sequence. The feature extraction layer is composed of a ResNet-50 network. The sequence generation layer includes a first-level generator, a second-level generator, and a third-level generator set in parallel. The first-level generator uses a U-net architecture, the second-level generator has densely connected blocks, and the third-level generator uses a lesion perception loss function.

[0031] The three-stage generative adversarial network architecture consists of three levels. The first-stage generator receives brain or neck region features and initially generates low-resolution images, such as the first brain FLAIR and neck PSIR / PSIR / PSIRMAG target images generated based on brain and neck region features. The second-stage generator introduces densely connected modules on top of the U-Net to enhance details and improve resolution of the first brain FLAIR and neck PSIR / PSIR / PSIRMAG target images, outputting high-resolution images, such as the second brain FLAIR and neck PSIR / PSIR / PSIRMAG target images. The third-stage generator, based on the first two stages, integrates the lesion perception loss function from the automatic lesion segmentation results, focusing on optimizing the signal accuracy and structural consistency of the lesion region to obtain optimized brain FLAIR and neck PSIR / PSIRMAG target images. Furthermore, the three-stage generative adversarial network architecture includes a discriminator for quality assessment, calculating structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) metrics for the generated images, and triggering a regeneration mechanism for images that do not meet a preset quality threshold.

[0032] Specifically, after generating the first brain FLAIR and neck PSIR, PSIRMAG target images, the second brain FLAIR and neck PSIR, PSIRMAG target images, and optimizing the brain FLAIR and neck PSIR, PSIRMAG target images through the sequence generation layer, the image quality can be determined through the following steps: Step 151: Calculate the structural similarity index and peak signal-to-noise ratio of the first brain target image, the first neck target image, the second brain target image, the second neck target image, the optimized brain target image, and the optimized neck target image using a discriminator.

[0033] Step 152: When the structural similarity index is less than the structural similarity index threshold, new first brain target image and first neck target image, second brain target image and second neck target image, and optimized brain target image and optimized neck target image are regenerated from the texture features of the brain and neck regions through the sequence generation layer.

[0034] Step 153: When the peak signal-to-noise ratio is less than the signal-to-noise ratio threshold, new first brain target image and first neck target image, second brain target image and second neck target image, and optimized brain target image and optimized neck target image are regenerated from the texture features of the brain and neck regions through the sequence generation layer.

[0035] Step 154: When the structural similarity index is greater than the structural similarity index threshold and the peak signal-to-noise ratio is greater than the signal-to-noise ratio threshold, the first brain target image and the first neck target image, the second brain target image and the second neck target image, and the optimized brain target image and the optimized neck target image are merged to obtain the final cross-modal brain target image and cross-modal neck target image.

[0036] Specifically, when liquid attenuation inversion recovery (FLAIR) sequences, phase-sensitive inversion recovery (PSIR) sequences, or susceptibility-weighted imaging (PSIRMAG) sequences are missing, images can be generated from T1-weighted imaging and T2-weighted imaging sequences.

[0037] like Figure 2 As shown, firstly, a pre-trained ResNet-50 network is used to extract multi-scale features from T1-weighted and T2-weighted imaging sequences, focusing on capturing texture features of MS-prone areas such as periventricular white matter, corpus callosum, and cervical spinal cord. Then, a three-level generator is used to generate target images of brain FLAIR and neck PSIR and PSIRMAG at different resolutions with different emphases, thereby determining the target images of brain FLAIR and neck PSIR and PSIRMAG.

[0038] The three-level generator can employ deep learning models such as GAN, Diffusion Model, and CNN. For example, the first-level generator G1, based on the U-net architecture, performs preliminary image generation based on texture features extracted from T1-weighted and T2-weighted imaging sequences, yielding 256×256 resolution brain FLAIR and neck PSIR and PSIRMAG target images. The second-level generator G2 adds densely connected blocks to the deep learning model, enabling the generation of 512×512 high-resolution brain FLAIR and neck PSIR and PSIRMAG target images. The third-level generator G3, based on pre-set prior knowledge constraints on lesions, automatically segments lesion regions in the high-resolution brain FLAIR and neck PSIR and PSIRMAG target images and focuses on optimizing the signal accuracy of lesion regions, resulting in optimized brain FLAIR and neck PSIR and PSIRMAG target images.

[0039] After the three types of brain FLAIR and neck PSIR and PSIRMAG target images are generated, the structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR) of these three types of images are calculated by the discriminator network. When both meet the preset threshold, the generated three types of images are judged to be qualified and the brain FLAIR and neck PSIR and PSIRMAG target images are output. When either of them fails to meet the standard, a new three types of brain FLAIR and neck PSIR and PSIRMAG target images are reconstructed by the three-level generator, and the SSIM and PSNR are recalculated until the quality requirements are met.

[0040] Step 16: Process cross-modal brain target images and cross-modal neck target images using a pre-trained deep learning segmentation network to obtain the lesion segmentation results of brain lesions and cervical spinal cord lesions.

[0041] Step 17: Obtain the quantitative information and spatial distribution characteristics of the lesions based on the lesion segmentation results.

[0042] The quantitative information on brain lesions and cervical spinal cord lesions includes the number, volume, and length and short diameter of the lesions; the spatial distribution characteristics include the location of the lesions around the ventricles, in the subcortex, infratentorial region, and in the cervical spinal cord.

[0043] The deep learning segmentation network specifically includes an encoding module and a feature fusion and reconstruction module connected in sequence. The encoding module includes a parallel SAM image encoder and an nnUNet encoder. Visual features were extracted from FLAIR target images, neck PSIR images, neck PSIMAG images, brain medical images, and neck medical images using the SAM image encoder, and medical-specific features were extracted from FLAIR target images, neck PSIR images, neck PSIMAG images, brain medical images, and neck medical images using the nnUNet encoder. The feature fusion and reconstruction module fuses visual features and medical-specific features to obtain fused features. Based on the fused features, brain lesions and cervical spinal cord lesions are automatically segmented to obtain lesion segmentation results. Based on the lesion segmentation results, the quantitative information and spatial distribution characteristics of brain lesions and cervical spinal cord lesions are determined.

[0044] Specifically, such as Figure 3As shown, after generating supplementary brain FLAIR and neck PSIR and PSIRMAG target images through cross-modal generation, pre-trained deep learning segmentation networks such as U-Net and nnU-Net are used for automatic segmentation of brain MS lesions and neck MS lesions, respectively, obtaining the spatial distribution characteristics of lesions in multiple brain regions and multiple cervical spinal cord segments. Then, based on the spatial distribution characteristics of lesions in multiple brain regions, the affected distribution of each brain region, as well as lesion data such as the number, size, and proportion of lesions, are calculated. Similarly, based on the spatial distribution characteristics of lesions in multiple cervical spinal cord segments, the affected distribution of the C1-C7 segments, as well as lesion data such as the number, size, and proportion of lesions, are calculated. Finally, a spatial correlation index is calculated based on the lesion data from the brain and neck regions.

[0045] For example, the Segment Anything Model (SAM), pre-trained on a large-scale natural image dataset, is used as the basic feature extractor to encode a general and robust visual representation. The nnUNet framework is integrated to automatically optimize the network structure and parameters based on a multi-site, multi-modal magnetic resonance imaging training set to adapt to the domain specificity of medical images. Then, the general visual features extracted by the SAM image encoder are deeply fused with the medical domain specific features learned by nnUNet to jointly guide the segmentation model in the segmentation and recognition process of MS lesions, thus combining general visual understanding capabilities with professional medical image segmentation.

[0046] The deep learning segmentation network can be trained using a training set containing multimodal magnetic resonance imaging (MRI) images of the brain and neck, along with annotations of multiple sclerosis (MS) lesions in the images, and a network framework combining visual priors and a large-scale basic segmentation model.

[0047] Specifically, after obtaining the spatial distribution features of lesions through deep learning segmentation networks, a multi-channel 3D ResNet network is used to process multimodal brain and neck data separately, incorporating attention guidance for lesion regions to extract higher-order features. Furthermore, the calculated lesion quantification indicators and higher-order features are input into a pre-trained evaluation decision tree model to calculate a 0-100% McDonald's standard compliance score and acute, subacute, or chronic lesion activity grading. This allows doctors or existing algorithms to match the McDonald's standard compliance score and lesion activity grading with treatment recommendations, such as matching hormone therapy or immunomodulatory therapy.

[0048] Furthermore, the results obtained from the above processing can generate structured reports conforming to DICOM standards, and output MPR, three-dimensional lesion distribution display, and brain-neck correlation analysis diagrams. It can also combine historical data to output dynamic follow-up comparison views, allowing doctors to determine the presence of new or enhancing lesions, such as T1 low signal areas (black holes), atrophy, or spreading lesions indicating tissue destruction. Moreover, the above output results can be used by existing machine learning classifiers to classify lesion subtypes such as "Relapsing-Relieving MS (RRMS)," "Secondary Progressive MS (SPMS)," and "Primary Progressive MS (PPMS)" based on imaging burden such as lesion volume or number, and quantitative changes in lesion characteristics, and to perform disease activity grading, such as the MSSS level of the MS severity score. Alternatively, it can be extended to the EDSS level of the disability status grading, allowing doctors to recommend medication regimens such as immunomodulators and hormone pulse therapy based on lesion distribution and burden, patient subtype, and activity, or to guide the frequency of follow-up MRI, monitor new lesions, assess drug efficacy and resistance, and assist in decisions regarding medication changes / combination therapy.

[0049] Existing single-modal deep learning models not only suffer from drawbacks such as difficulty in effectively handling missing sequences, lack of collaborative analysis capabilities for lesions in multiple locations, low accuracy in identifying small lesions, and low automation of quantitative analysis, but also often develop for specific devices or data sources, resulting in limited generalization ability and difficulty in adapting to the actual needs of different medical institutions.

[0050] To address the aforementioned shortcomings, this invention provides a medical image processing method that applies cross-modal image generation to the accurate assessment of multiple sclerosis (MS). Targeting the unique signal characteristics of MS lesions, such as ring enhancement in demyelinating lesions and small spinal cord lesions, a lesion-aware generative network architecture was developed. This generative network architecture can integrate multi-scale features from T1 / T2 weighted images and incorporate prior lesion knowledge constraints to achieve high-fidelity generation of key sequences such as FLAIR and PSIR. While maintaining anatomical structural accuracy, it improves the retention rate of smaller MS lesions, effectively solving the problem of sequence loss due to equipment limitations in primary hospitals, and providing reliable technical support for early MS screening.

[0051] Furthermore, in response to the clinical characteristics of MS being multifocal and involving multiple sites, this invention uses deep learning automatic segmentation technology to extract the spatial distribution characteristics of brain lesions and cervical spinal cord lesions, such as the pattern of involvement of the C2-C4 segments, and further analyzes and evaluates them based on the spatial distribution characteristics. It outputs a structured report integrating key indicators of the evaluation criteria for doctors to refer to in order to determine treatment recommendations, which significantly improves the efficiency of clinical decision-making.

[0052] In summary, the medical image processing method provided by this invention not only fills the technological gap in the field of intelligent assessment of MS, but also creates a new closed-loop diagnosis and treatment model from image analysis to clinical decision-making, significantly reducing the dependence of assessment results on equipment and physician experience, and providing a reliable tool for the standardized assessment of MS.

[0053] Example 2 Based on Example 1, the present invention also provides a complete example of a medical image processing workflow: First, a brain and neck MRI dataset was used, containing complete imaging data of confirmed or suspected MS patients. Each case included T1WI, T2WI, and FLAIR sequences of the brain region, and T1WI, T2WI, PSIR, and corresponding PSIRMAG images of the neck region. Next, the voxel intensity of all images was Z-score normalized to a mean of 0 and a standard deviation of 1 to eliminate differences caused by different scanners and protocols. All brain sequences were resampled, and T2WI and FLAIR images were rigidly registered to the T1WI space of the same patient. Similarly, all neck sequences were resampled and registered to the T1WI space. Then, three radiologists with more than 5 years of experience in neuroimaging independently annotated MS lesions on ITK-SNAP software. Lesions in the periventricular, subcortical, infratentorial, and intracervical spinal cord regions were delineated at the pixel level. The final annotation results were resolved through majority voting or expert arbitration to establish the gold standard label.

[0054] Subsequently, a hybrid training strategy combining phased and end-to-end approaches was adopted, in which labeled datasets were input into the three-level generative network for training.

[0055] The three-tiered generative network comprises three cascaded generators (G1, G2, and G3) and a quality evaluation module. In the G1 generator, a U-Net structure is used as the basic architecture. It takes registered and channel-sequentially stitched brain and neck T1WI and T2WI sequences as input and outputs brain FLAIR and neck PSIR and PSIRMAG images at half the original resolution, thereby learning the overall mapping relationship and basic structural information from the source modality to the target modality. In the G2 generator, an enhanced U-Net structure with densely connected blocks is used. Taking the output of G1 and upsampled versions of the T1WI and T2WI source images as input, multi-scale feature reuse is achieved through dense connections to enhance details and restore resolution in the images generated by G1, with the output being the same as the original. Figure 1The G3 generator produces images with higher resolution. Its structure is similar to the G2 generator, but it incorporates a lesion-aware loss function. This loss function, based on the standard L1 loss, assigns higher weights to lesion regions using a saliency map generated by a pre-trained lesion segmentation network, ensuring high accuracy in signal intensity and texture patterns of MS lesion regions in the generated images.

[0056] Furthermore, a quality assessment module is integrated at the end of the three-tiered generative network. This module consists of a pre-trained discriminator and an image quality calculation unit. The discriminator determines the authenticity of the generated images, while the quality calculation unit simultaneously calculates the structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) between the generated brain FLAIR and neck PSIR, PSIRMAG images and the real brain FLAIR and neck PSIR, PSIRMAG images. When the SSIM is below 0.90 or the PSNR is below 28 dB, the system automatically triggers a regeneration process, feeding the sample back to G1 for a new round of generation until the output image meets the quality threshold or the maximum number of retries is reached.

[0057] The specific hybrid training strategy combining phased and end-to-end approaches is as follows: G1 is trained separately until convergence, and the weights of G1 are fixed. G2 is then connected and trained. After that, G1 and G2 are fixed, and G3 and the quality assessment module are connected for end-to-end fine-tuning. The loss functions during training include generative adversarial loss, pixel-level L1 reconstruction loss, and lesion perception loss specific to the third level, as shown in formula (1):

[0058] (1) in, To combat the losses, For L1 reconstruction loss, Loss of perception of the lesion , and These represent the weights for the three types of losses. After the model training is complete, the three-level generator and the quality evaluation module are encapsulated into an independent inference pipeline and deployed in the system's image generation module. For new inputs, the inference process is fully automated, requiring no manual intervention.

[0059] The deep network segmentation model adopts the nnSAM (nnUNet-powered Segment Anything Model) architecture, which trains independent segmentation models for brain and neck MS lesions to make full use of the image features of different sites, including general feature extraction branches and medical-specific branches.

[0060] In the general feature extraction branch, the ViT-H image encoder, pre-trained on the SA-1B dataset from the SAM model, is used for feature extraction. The input consists of generated images uniformly scaled to a fixed resolution, such as brain FLAIR, neck PSIR, and neck PSIRMAG images, as well as lesion region images from the original T1WI and T2WI. The encoder's weights are frozen, serving as a fixed general visual feature extractor that provides rich low-level edge and high-level semantic features. In the medical-specific branch, the nnUNetv2 framework is used as a learnable branch. It automatically analyzes attributes such as image size and voxel spacing in the training set, dynamically configures optimal hyperparameters such as network depth and number of convolutional kernels, and formulates corresponding training strategies.

[0061] Subsequently, the multi-scale general features extracted by the SAM encoder are fused with the medical features extracted by the nnUNet encoder. Specifically, a combination of element-wise addition and channel concatenation is used to integrate these fused features at the corresponding scale of the nnUNet decoder, thereby balancing general visual representation with medical specificity.

[0062] In the deep network segmentation model, the brain model and the neck model are trained separately and independently. The brain model receives three channels of data: the original T1WI, T2WI, and the generated FLAIR images of the brain. The neck model receives four channels of data: the original T1WI, T2WI, and the generated PSIR and PSIRMAG images. After training, the brain and neck segmentation models are encapsulated and deployed separately in the system's lesion analysis module. During actual inference, the system automatically calls the corresponding lesion segmentation model for processing and outputs the fused quantitative analysis results based on the lesion segmentation results.

[0063] Once the model training is complete, such as Figure 4 As shown, the image acquisition module automatically receives and retrieves the patient's raw DICOM image data from the brain and neck. The minimum input for the raw DICOM image data is brain T1WI and T2WI sequences, and neck T1WI and T2WI sequences. Then, in the brain pathway of the image generation module, the patient's brain T1WI and T2WI are input into a pre-trained brain cascade generation model to automatically generate the missing high-quality FLAIR images. In the neck pathway of the image generation module, the patient's neck T1WI and T2WI are input into a pre-trained neck cascade generation model to automatically generate the missing PSIR and PSIRMAG images. The quality assessment module calculates the SSIM and PSNR indices of the generated images. If the quality is substandard, the system triggers an internal regeneration process without manual intervention.

[0064] After multimodal image preprocessing, the images are fed into two independent segmentation models in parallel. In the brain lesion segmentation model, the original T1WI, T2WI, and generated FLAIR images of the brain are input into the brain nnSAM segmentation model to obtain a binary segmentation mask for lesions within the brain parenchyma. In the neck lesion segmentation model, the original T1WI, T2WI, and generated PSIR and PSIRMAG images of the neck are input into the neck nnSAM segmentation model to obtain a binary segmentation mask for lesions within the cervical spinal cord. Subsequently, in the lesion analysis module, the segmentation results are post-processed and quantitatively analyzed, and key indicators are automatically extracted, such as the total number and volume of lesions in each location, the major and minor axes of the largest lesion, and other quantitative information. Spatial distribution features of lesions in the periventricular, subcortical, infratentorial, and specific cervical spinal cord segments are also extracted.

[0065] In addition, once the quantitative information and spatial distribution characteristics are determined, further automatic verification can be performed based on spatial diffusion. For example, it can be verified whether there are lesions in at least two of the three locations: the periventricular, infratentorial, and spinal cord areas. The temporal diffusion can be assessed in conjunction with clinical history or previous imaging, and the assessment results can be output and displayed together with the quantitative information and spatial distribution characteristics.

[0066] Example 3 The present invention also provides a medical image processing system, comprising: The image acquisition and feature extraction module is used to acquire T1-weighted and T2-weighted images from brain medical images and neck medical images; to perform multi-scale feature extraction on T1-weighted and T2-weighted images from brain medical images to obtain brain region features; and to perform multi-scale feature extraction on T1-weighted and T2-weighted images from neck medical images to obtain neck region features. The image generation and optimization module is used to construct a first brain target image and a first neck target image based on brain region features and neck region features, respectively; based on a dense connection mechanism, a second brain target image is generated from the first brain target image, and a second neck target image is generated from the first neck target image, wherein the resolution of the second brain target image is higher than that of the first brain target image, and the resolution of the second neck target image is higher than that of the first neck target image; through a pre-set lesion perception loss function, the signals of lesion regions in the second brain target image and the second neck target image are optimized, respectively, to obtain optimized brain target images and optimized neck target images; A cross-modal image generation module is used to synthesize cross-modal brain target images and cross-modal neck target images based on a first brain target image and a first neck target image, a second brain target image and a second neck target image, and optimized brain target images and optimized neck target images, respectively.

[0067] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps in an embodiment of a medical image processing method. Specific implementation methods can be found in the method embodiments, and will not be repeated here.

[0068] Furthermore, the present invention also provides a non-transitory computer-readable storage medium containing instructions on which a computer program is stored. For example, a memory containing instructions that can be executed by a processor of a computer device to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the computer program is executed by the processor, it can implement the steps in an embodiment of a medical image processing method. Specific implementation methods can be found in the method embodiments, which will not be repeated here.

[0069] Those skilled in the art will understand that embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0070] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0071] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0072] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0073] It should be noted that the specific embodiments described above enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present invention has been described in detail in this specification and embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered within the protection scope of the present invention patent. No reference numerals in the claims should be construed as limiting the scope of the claims. Any simple variations or equivalent substitutions of technical solutions that can be readily obtained by those skilled in the art within the scope of the technology disclosed in the present invention are within the protection scope of the present invention.

Claims

1. A medical image processing method, characterized in that, include: T1-weighted and T2-weighted imaging data were acquired from brain medical images and neck medical images. Multi-scale feature extraction was performed on the T1-weighted and T2-weighted imaging data of the brain medical images to obtain brain region features. Multi-scale feature extraction was also performed on the T1-weighted and T2-weighted imaging data of the neck medical images to obtain neck region features. A first brain target image and a first neck target image are constructed based on the brain region features and neck region features, respectively. A second brain target image is generated from the first brain target image and a second neck target image is generated from the first neck target image based on a dense connection mechanism. The resolution of the second brain target image is higher than that of the first brain target image, and the resolution of the second neck target image is higher than that of the first neck target image. The signals of the lesion regions in the second brain target image and the second neck target image are optimized using a pre-set lesion perception loss function, respectively, to obtain optimized brain target images and optimized neck target images. A cross-modal brain target image and a cross-modal neck target image are synthesized based on the first brain target image and the first neck target image, the second brain target image and the second neck target image, and the optimized brain target image and the optimized neck target image, respectively.

2. The medical image processing method according to claim 1, characterized in that, The brain region features and neck region features are processed by a three-level generative adversarial network to obtain cross-modal brain target images and cross-modal neck target images; The three-tiered generative adversarial network comprises a feature extraction layer, a sequence generation layer, and a discriminator connected in sequence. The feature extraction layer is composed of a ResNet-50 network. The sequence generation layer comprises a first-level generator, a second-level generator, and a third-level generator set in parallel. The first-level generator is a U-net architecture, the second-level generator has densely connected blocks, and the third-level generator has a lesion perception loss function.

3. The medical image processing method according to claim 1, characterized in that, After obtaining optimized brain target images and optimized neck target images, the process also includes: The structural similarity index and peak signal-to-noise ratio of the first brain target image, the first neck target image, the second brain target image, the second neck target image, the optimized brain target image, and the optimized neck target image are calculated by the discriminator respectively. When the structural similarity index is less than the structural similarity index threshold, new first brain target image and first neck target image, second brain target image and second neck target image, and optimized brain target image and optimized neck target image are regenerated from the texture features of the brain and neck regions through the sequence generation layer. When the peak signal-to-noise ratio is less than the signal-to-noise ratio threshold, new first brain target image and first neck target image, second brain target image and second neck target image, and optimized brain target image and optimized neck target image are regenerated from the texture features of the brain and neck regions through the sequence generation layer. When the structural similarity index is greater than the structural similarity index threshold and the peak signal-to-noise ratio is greater than the signal-to-noise ratio threshold, the first brain target image and the first neck target image, the second brain target image and the second neck target image, and the optimized brain target image and the optimized neck target image are merged to obtain the cross-modal brain target image and the cross-modal neck target image.

4. The medical image processing method according to claim 1, characterized in that, The cross-modal brain target image is a brain FLAIR image, and the cross-modal neck target image includes a neck PSIR image and a PSIRMAG image.

5. The medical image processing method according to claim 1, characterized in that, After synthesizing cross-modal brain target images and cross-modal neck target images, the process also includes: The cross-modal brain target image and cross-modal neck target image are processed by a pre-trained deep learning segmentation network to obtain the lesion segmentation results of brain lesions and cervical spinal cord lesions; The quantitative information and spatial distribution characteristics of the lesions are obtained based on the lesion segmentation results.

6. A medical image processing method according to claim 5, characterized in that, The deep learning segmentation network specifically includes an encoding module and a feature fusion and reconstruction module connected in sequence. The encoding module includes a parallel SAM image encoder and an nnUNet encoder. Visual features are extracted from the cross-modal brain target image and the cross-modal neck target image using the SAM image encoder, and medical-specific features are extracted from the cross-modal brain target image and the cross-modal neck target image using the nnUNet encoder; The visual features and medical-specific features are fused by the feature fusion and reconstruction module to obtain lesion segmentation results. Based on the lesion segmentation results, the quantitative information and spatial distribution characteristics of the brain lesions and cervical spinal cord lesions are determined.

7. A medical image processing method according to claim 5, characterized in that, The quantitative information includes the number, volume, and major and minor diameters of the lesions; the spatial distribution characteristics include the location and distribution of the lesions around the ventricles, subcortical, infratentorial, and cervical spinal cord.

8. A medical image processing system, characterized in that, include: The image acquisition and feature extraction module is used to acquire T1-weighted and T2-weighted images from brain medical images and T1-weighted and T2-weighted images from neck medical images; to perform multi-scale feature extraction on the T1-weighted and T2-weighted images from the brain medical images to obtain brain region features; and to perform multi-scale feature extraction on the T1-weighted and T2-weighted images from the neck medical images to obtain neck region features. The image generation and optimization module is used to construct a first brain target image and a first neck target image based on the brain region features and neck region features, respectively; Based on a dense connection mechanism, a second brain target image is generated from a first brain target image, and a second neck target image is generated from a first neck target image. The resolution of the second brain target image is higher than that of the first brain target image, and the resolution of the second neck target image is higher than that of the first neck target image. The signals of the lesion regions in the second brain target image and the second neck target image are optimized by using a pre-set lesion perception loss function, respectively, to obtain optimized brain target images and optimized neck target images. A cross-modal image generation module is used to synthesize a cross-modal brain target image and a cross-modal neck target image based on the first brain target image and the first neck target image, the second brain target image and the second neck target image, and the optimized brain target image and the optimized neck target image, respectively.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the medical image processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is loaded by the processor, it is able to execute the steps of the medical image processing method according to any one of claims 1 to 7.