Dl-based multi-temporal brain ncct image processing method, device, medium and program product
By using a multi-temporal brain NCCT image processing method based on deep learning, the problems of insufficient spatiotemporal registration accuracy and feature extraction in traditional methods are solved, enabling high-precision dynamic analysis and personalized treatment support for cerebrovascular diseases, and improving clinical prediction capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE SECOND AFFILIATED HOSPITAL TO NANCHANG UNIV
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies suffer from insufficient spatiotemporal registration accuracy, inadequate feature extraction, and poor clinical interpretability in the prognostic assessment of cerebrovascular diseases, especially aSAH, making it difficult to achieve high-precision dynamic analysis and individualized treatment decisions.
We employ a multi-temporal brain NCCT image processing method based on deep learning. Through CDGM image enhancement, ANTs frame deformation registration, improved 3D Transformer-UNet feature extraction, and Grad-CAM++ visualization module, we construct spatiotemporally consistent four-dimensional image data, extract dynamic feature embedding vectors, and achieve high-precision analysis of the dynamic evolution of lesions.
It enables precise diagnosis and treatment of cerebrovascular diseases, provides high-precision dynamic feature extraction and clinical interpretability, and supports early prediction of complications and individualized treatment decisions.
Smart Images

Figure CN121687412B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and artificial intelligence technology, and in particular to a multi-temporal brain NCCT (non-contrast Computed Tomography) image processing method, device, medium and program product based on DL (Deep Learning), which is used to construct discrete multi-time point brain NCCT images into spatiotemporally consistent four-dimensional image data and extract dynamic feature embedding vectors that characterize the dynamic evolution of lesions. Background Technology
[0002] Among the spectrum of cerebrovascular diseases, hemorrhagic lesions are a key focus and challenge in clinical diagnosis and treatment due to their rapid onset and high mortality rate. Aneurysmal subarachnoid hemorrhage (aSAH), caused by the rupture of an intracranial aneurysm, is a particularly life-threatening cerebrovascular disease. The amount of blood leaking into the subarachnoid space can lead to a series of clinical symptoms, such as headache, coma, altered consciousness, seizures, and even death. Despite continuous advancements in diagnostic and treatment techniques, nearly half of the survivors are unable to return to a normal life due to neurological dysfunction, and their clinical prognosis remains unsatisfactory.
[0003] Currently, prognostic assessment of aSAH heavily relies on single-timepoint NCCT brain imaging at admission and clinical scoring scales such as the WFNS (World Federation of Neurosurgical Societies Grade) and Hunt-Hess classification. However, these methods have significant limitations: clinical scoring is highly subjective and difficult to personalize; and single-timepoint NCCT imaging only provides a static snapshot of the disease, failing to capture the dynamic evolution of secondary brain injuries such as hematoma absorption, cerebral edema, and delayed cerebral ischemia after aSAH. The outcome of the disease is essentially determined by the dynamic changes following acute hemorrhage; therefore, relying on static assessment strategies severely restricts the early and accurate prediction of prognosis and the formulation of individualized dynamic treatment decisions for aSAH patients.
[0004] In recent years, deep learning (DL) technology has revolutionized medical image analysis, demonstrating high accuracy in the automated detection, segmentation, and quantification of aSAH (aortic submucosal hemorrhage). However, most existing AI-based research remains limited to analyzing brain NCCT images at a single time point (usually at admission), failing to fully utilize the dynamic information contained in brain NCCT images at different time points. Even though some studies have attempted to integrate brain NCCT images from different time points, their methods suffer from fundamental limitations and technical deficiencies.
[0005] Specifically, in the process of implementing the technical solution of this invention, at least the following technical problems were found in the traditional single-time-point brain NCCT image evaluation method:
[0006] First, due to the spatial differences between brain NCCT images at different time points, and the interference of lesions such as intracranial hemorrhage with the registration results, it is difficult to align brain NCCT images at different time points with high precision, making it impossible to construct spatiotemporally consistent four-dimensional image data for subsequent analysis.
[0007] Second, existing technologies mostly use simple feature concatenation or traditional machine learning models for feature fusion, lacking effective feature extraction methods. Even if the brain NCCT images at different time points are aligned, it is difficult to accurately extract and quantify the spatiotemporal evolution of lesions (such as hematoma volume and edema range) from the aligned images.
[0008] Third, most existing deep learning models are "black box" structures, with opaque decision-making processes and a general lack of systematic interpretable analytical frameworks. Therefore, it is difficult to quantify the dynamic contribution of image features at different time points to the prediction results, and thus cannot provide explanations related to clinical pathophysiological mechanisms.
[0009] In summary, traditional methods suffer from technical bottlenecks in spatiotemporal registration accuracy, feature extraction, and clinical interpretability, resulting in insufficient prognostic performance and low clinical reliability for aSAH, thus hindering their implementation and application in clinical practice. There is an urgent need in this field for a technical solution that can automatically and accurately process brain NCCT images at different time points, construct spatiotemporally consistent four-dimensional image data, and extract clinically interpretable dynamic feature embedding vectors from it, in order to break through the current bottlenecks in aSAH prognostic prediction. Summary of the Invention
[0010] This invention provides a multi-temporal brain NCCT image processing method, device, medium, and program product based on deep learning, which effectively breaks through the technical bottlenecks of traditional methods in terms of spatiotemporal registration accuracy, feature extraction, and clinical interpretability, and provides reliable technical support for the precise diagnosis and treatment of cerebrovascular diseases.
[0011] In a first aspect, the present invention provides a multi-temporal brain NCCT image processing method based on deep learning (DL), applied to a multi-temporal brain NCCT image processing system based on DL, comprising:
[0012] After image enhancement using CDGM, a deformation registration process based on the ANTs framework and fused with a cost function mask is initiated to complete the spatial alignment of multi-temporal brain NCCT images.
[0013] The improved 3D Transformer-UNet uses the encoder of the SwinUnetR network, embeds the self-attention mechanism of local window and shift window, and uses spatial prior probability graphs for spatial constraints to complete the extraction of location-specific depth spatial features.
[0014] Based on the Transformer encoder and multi-head self-attention mechanism, the aggregation of temporal features is completed;
[0015] Using downstream tasks as constraints, a joint multi-task optimization strategy is adopted to obtain dynamic feature embedding vectors that characterize the dynamic evolution of lesions. The three types of sub-modules of the downstream tasks include a three-dimensional reconstruction decoder module, a prognostic prediction module, and a Grad-CAM++ visualization module.
[0016] Optionally, the steps for image enhancement using CDGM include:
[0017] During the forward noise diffusion process, the input unenhanced image is gradually mapped to the noisy state image, and an intermediate random state is formed through multi-step diffusion.
[0018] In the conditional denoising process, an I2SB model based on the unenhanced image is introduced to perform progressive reverse denoising on the intermediate random state, and finally reconstruct the enhanced image.
[0019] Optionally, a deformation registration process based on the ANTs framework and incorporating a cost function mask includes:
[0020] Using unenhanced images as the main input, enhanced images as auxiliary input, and baseline images as a reference, the SyN symmetric normalized deformation registration algorithm based on the ANTs framework is used for registration. The cross-correlation coefficient is used as the similarity measure, and the lesion segmentation mask output by the upstream segmentation network is used as the cost function mask. At the same time, a four-level hierarchical optimization strategy from 1 / 8 resolution to full resolution is implemented to achieve sub-millimeter-level spatiotemporal alignment of images at each time point.
[0021] The SSIM of normal brain tissue regions in the registered images was calculated for quality assessment. Only high-quality registration results with SSIM > 0.80 were retained, and finally, spatiotemporally aligned four-dimensional image data were generated.
[0022] Optionally, the upstream segmentation network uses SwinUnetR as the backbone network and embeds a self-attention mechanism of local window and shift window, and loads a spatial prior probability map as a spatial constraint. At the same time, the upstream segmentation network is optimized through supervised training, adopts the joint loss function of cross-entropy and Dice, and improves robustness through data augmentation and spatial prior regularization, and finally outputs the lesion segmentation mask at each time point.
[0023] Optionally, the steps for deep spatial feature extraction specifically include:
[0024] The encoder part of the already trained SwinUnetR network in the upstream segmentation network is extracted and used as the encoder of the improved 3D Transformer-UNet. The encoder retains the self-attention mechanism of local window and shift window and uses spatial prior probability graph for spatial constraints.
[0025] While maintaining the upstream segmentation performance, the pre-trained weights of the encoder are fine-tuned;
[0026] The spatiotemporally aligned four-dimensional image data is used as the input to the encoder, which outputs the depth spatial features at each time point.
[0027] Optionally, the aggregation steps for time-series features specifically include:
[0028] Design a temporal feature aggregation module based on Transformer encoder and multi-head self-attention mechanism. Input the deep space feature sequence into Transformer encoder with multiple attention heads, calculate the dependencies in the time dimension through multi-head self-attention mechanism, generate attention weight matrix and adaptively assign weights to features at different time points, perform weighted fusion of features according to weights to achieve intelligent fusion of deep space features, and output aggregated features.
[0029] Optionally, the steps to obtain the dynamic feature embedding vector specifically include:
[0030] The aggregated features are optimized by using regression / classification loss and spatial gradient constraints to guide the encoder and temporal feature aggregation module of the improved 3DTransformer-UNet to learn iteratively, ultimately obtaining dynamic feature embedding vectors with discriminative, temporal sensitivity and clinical relevance.
[0031] In a second aspect, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the aforementioned multi-temporal brain NCCT image processing method based on DL.
[0032] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned multi-temporal brain NCCT image processing method based on DL.
[0033] Fourthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned multi-temporal brain NCCT image processing method based on deep learning.
[0034] One or more technical solutions provided by this invention have at least the following technical effects or advantages:
[0035] In the field of medical image processing, traditional methods have technical bottlenecks in terms of spatiotemporal registration accuracy, feature extraction, and clinical interpretability. This invention proposes an innovative solution that effectively overcomes these bottlenecks and lays a solid foundation for the precise diagnosis and treatment of cerebrovascular diseases.
[0036] To address the technical bottlenecks in spatiotemporal registration accuracy of traditional methods, this invention innovatively employs a two-step strategy to achieve high-precision registration. First, CDGM is used to enhance the images, particularly low-contrast or noisy regions. This enhanced image serves as an auxiliary input for registration, effectively overcoming the deformation field calculation bias caused by abnormal lesion density or blurred boundaries in traditional registration methods. This improves the spatial alignment accuracy of multi-temporal brain NCCT images, providing a reliable four-dimensional data foundation for subsequent multi-temporal dynamic analysis. Subsequently, a deformation registration process using the ANTs framework and cost function masking is initiated. This process constructs spatiotemporally aligned four-dimensional image data from multi-timepoint brain NCCT images. By masking the influence of lesion regions on similarity metrics, a more reliable deformation field is obtained, achieving sub-millimeter-level spatiotemporal alignment of images at each time point.
[0037] To address the technical bottlenecks in feature extraction using traditional methods, this invention constructs a complete four-dimensional image data processing workflow based on an improved 3DTransformer-UNet and a temporal feature aggregation module. The 3DTransformer-UNet structure effectively extracts spatial context information, while the local window and shift window self-attention mechanisms enhance the localization ability of key pathological regions. The temporal feature aggregation module fully utilizes the dynamic information of multi-time-point data. Specifically, in the deep spatial feature extraction stage, the improved 3D Transformer-UNet employs a spatially prior-constrained local window and shift window self-attention mechanism, using the spatial prior probability map as a constraint condition for the self-attention mechanism. This guides the network to prioritize key pathological regions such as the hemorrhage core area and surrounding edema zone when extracting features. This mechanism not only preserves the sensitivity of the self-attention mechanism to spatial location information but also enhances the network's ability to identify the spatial distribution of specific pathological features, thereby enabling the accurate extraction of pathology-specific deep spatial features from four-dimensional image data. This stage, while ensuring feature extraction capabilities, improves the identification accuracy of various lesion types, enabling precise capture of early imaging changes related to complications such as intracranial aneurysm re-rupture and hemorrhage, and cerebral vasospasm, providing crucial evidence for prognostic prediction. In the temporal aggregation stage, a multi-head self-attention-based temporal feature aggregation module automatically learns the dependencies between features at different time points and assigns differentiated importance weights to features at each time point, achieving adaptive modeling of key disease progression dynamics such as hemorrhage absorption rate and edema expansion trend. This temporal modeling approach accurately identifies the evolutionary patterns of complications such as delayed cerebral ischemia, significantly improving the ability of dynamic features to represent disease outcomes and providing data support for individualized treatment plans. This hybrid architecture ensures both the model's feature extraction capabilities for single-time-point images and enhances its cross-time-point dynamic analysis capabilities, accurately quantifying the changes in key indicators such as hemorrhage volume and edema extent, achieving precise quantitative analysis of the dynamic evolution of cerebral hemorrhage, particularly effectively tracking the temporal changes of various hematomas such as subarachnoid hemorrhage and parenchymal hemorrhage.
[0038] To address the technical bottlenecks in the clinical interpretability of traditional methods, this invention constructs a clinically interpretable Grad-CAM++ visualization module. This module calculates the spatial region that contributes most to the prediction results and then visualizes and transforms this region, clearly showing the contribution of key pathological regions to the prediction results. This will greatly improve the ability to assess the disease progression of aSAH patients and has wide applicability in primary hospitals or medical environments with a shortage of professional personnel. It will play an important clinical role in assisting clinicians in judging disease outcomes, predicting early complications such as post-hemorrhagic hydrocephalus, re-rupture of intracranial aneurysms, cerebral vasospasm and delayed cerebral ischemia, and making prognostic assessments and treatment decisions.
[0039] In summary, this invention enables automated analysis of multi-temporal medical images, providing reliable technical support for the precise diagnosis and treatment of cerebrovascular diseases, and has good prospects for widespread application. Attached Figure Description
[0040] Figure 1 A flowchart of a multi-temporal brain NCCT image processing method based on deep learning provided by the present invention;
[0041] Figure 2 A schematic diagram of the structure of the backbone network SwinUnetR of the upstream segmentation network;
[0042] Figure 3 This is a comparison of the results of gold standard segmentation and upstream lesion segmentation on unenhanced images;
[0043] Figure 4 The image enhancement flowchart for the I2SB model;
[0044] Figure 5 A comparison image before and after the image enhancement process;
[0045] Figure 6 Comparison images before and after high-precision registration across multiple time series (3 time points);
[0046] Figure 7 The processing logic diagram of the encoder for the improved 3D Transformer-UNet;
[0047] Figure 8 This is the processing logic diagram for the time-series feature aggregation module;
[0048] Figure 9 This is a processing logic diagram for the downstream multi-task optimization module. Detailed Implementation
[0049] This invention provides a multi-temporal brain NCCT image processing method, device, medium, and program product based on deep learning, which effectively breaks through the technical bottlenecks of traditional methods in terms of spatiotemporal registration accuracy, feature extraction, and clinical interpretability, and provides reliable technical support for the precise diagnosis and treatment of cerebrovascular diseases.
[0050] First, the terms appearing in the instruction manual will be explained.
[0051] ANTs (Advanced Normalization Tools) is a top-tier open-source framework in the field of medical image processing. It is a high-level toolset based on multidimensional image registration, template construction, and segmentation. It is particularly adept at spatial alignment, deformation correction, structural segmentation, and quantitative analysis of brain images (such as brain NCCT images of aSAH patients). It is one of the most commonly used tools in neuroimaging and radiological imaging research.
[0052] CDGM (Conditional Diffusion Generative Model) is a type of network that introduces conditional constraints into the diffusion process of "stepwise noise addition and reverse noise reduction" to achieve controllable generation of target data, rather than randomly generating data from pure noise. It has great application value in the precise generation, repair, and segmentation of medical images.
[0053] The I2SB (Image-to-Image Schrödinger Bridge) model is a type of conditional diffusion model based on the Schrödinger bridge. It performs well in tasks such as image restoration and image conversion, and can also be used in medical imaging-related scenarios such as cone-beam X-ray fluoroscopy distortion correction.
[0054] GradCAM++ (Gradient-Weighted Class Activation Mapping++) is an improved interpretable AI algorithm of GradCAM, specifically designed to explain the decision-making process of convolutional neural networks. It has significant advantages in solving problems such as inaccurate multi-target localization, and is particularly suitable for the field of medical imaging.
[0055] Grad-CAM++ heatmaps are visual heatmaps output by the Grad-CAM++ algorithm. Their core function is to intuitively show which regions of an image the convolutional neural network focuses on when making decisions. They quantify the contribution of pixels to the prediction results through color intensity, making them highly valuable in medical imaging AI.
[0056] I. Multi-temporal Brain NCCT Image Processing Method Based on DL
[0057] The inspiration for this invention stems from the clinical need to monitor the dynamic evolution of cerebrovascular diseases. In cerebrovascular diseases, secondary brain injury following aSAH (atypical autoimmune hemorrhage), such as hematoma absorption, cerebral edema, and delayed cerebral ischemia, constitutes a complex dynamic process. Traditional single-timepoint brain NCCT imaging assessment methods (hereinafter referred to as traditional methods) cannot capture the dynamic evolution of secondary brain injury. Clinically, there is an urgent need for an objective assessment method that can quantify the dynamic evolution of the disease. Based on this, this invention combines deep learning (DL) with multi-temporal medical image analysis, proposing a DL-based multi-temporal brain NCCT image processing method. This method is applied to a DL-based multi-temporal brain NCCT image processing system. The system's main interface includes modules for data management, preprocessing settings, registration parameters, feature extraction, temporal analysis, and result display. After importing brain NCCT images from multiple key time points during a patient's hospitalization into the system, the system can automatically process the data to obtain spatiotemporally aligned four-dimensional image data, dynamic feature embedding vectors characterizing the dynamic evolution of the lesion, and analysis reports.
[0058] The purpose of this invention is to overcome the technical bottlenecks of traditional methods in terms of spatiotemporal registration accuracy, feature extraction, and clinical interpretability. It constructs spatiotemporally consistent four-dimensional image data and accurately quantifies and extracts features from the dynamic evolution of diseases. This helps clinicians to accurately assess disease progression and predict prognosis, provides more comprehensive decision-making basis, significantly improves the predictive ability of disease outcome, and has important clinical value in early warning of complications such as hydrocephalus and delayed cerebral ischemia after cerebral hemorrhage, in developing individualized treatment plans, accurately assessing treatment effects, and developing rehabilitation plans.
[0059] This invention enables the complete transformation from raw multi-temporal brain NCCT image data to dynamic feature embedding vectors, providing reliable quantitative evidence for clinical decision-making. Its basic principle relies on four core steps: First, a deformation registration process based on the ANTs framework and incorporating a cost function mask is used to complete the spatial alignment of multi-temporal brain NCCT images; second, an improved 3D Transformer-UNet uses the encoder of the SwinUnetR network, embedding a self-attention mechanism with local and shift windows, and uses a spatial prior probability graph for spatial constraints to complete depth spatial feature extraction; third, based on the Transformer encoder and multi-head self-attention mechanism, temporal features are aggregated; finally, using downstream tasks as constraints, a joint multi-task optimization strategy is employed to complete the dynamic feature embedding vector representing the dynamic evolution of lesions.
[0060] The multi-temporal brain NCCT image processing method based on deep learning proposed in this invention can be widely applied to the monitoring of the course and prognostic assessment of various cerebrovascular diseases such as aSAH, cerebral hemorrhage, and cerebral infarction. For better understanding, aSAH is used as an example of a cerebrovascular disease, and a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the embodiments described in this invention are only some, not all, of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0061] like Figure 1 As shown, steps 1) to 5) need to be executed first to prepare for the subsequent execution of the four core steps.
[0062] 1) Data Acquisition: Brain NCCT images of aSAH patients at multiple time points are acquired from the PACS (Picture Archiving and Communication System) of medical institutions. These multiple time points include at least three key time points: admission, early postoperative period, and subacute phase. The acquired brain NCCT images can be in standard medical image formats such as Dicom, Nrrd, and Mha; this invention does not impose any restrictions on this. Brain NCCT images from multiple time points obtained from multiple medical institutions constitute multi-center, multi-time series brain NCCT data; brain NCCT images from multiple time points obtained from a single medical institution constitute single-center, multi-time series brain NCCT data. Hereinafter, multi-center, multi-time series brain NCCT data and single-center, multi-time series brain NCCT data will be referred to simply as multi-time series brain NCCT data.
[0063] 2) Image Loading and Verification: When the DL-based multi-temporal brain NCCT image processing system (hereinafter referred to as the system) of this invention is started, the system first loads and verifies the multi-temporal brain NCCT data. The system supports data import in standard medical image formats such as Dicom, Nrrd, and Mha. During the data loading process, the system automatically performs data validity verification, specifically checking data integrity, spatial resolution consistency, and image quality indicators. If the system detects problems such as missing slices, abnormal spatial resolution, or severe artifacts, it will refuse to load the data and display an error message; qualified data that passes the verification will be automatically displayed on the interactive terminal for physician confirmation.
[0064] 3) Standardized processing: After receiving confirmation feedback from the physician, the system establishes a standardized processing pipeline based on the ANTs framework. It sequentially performs FOV adaptive adjustment to accurately define the region of interest in the brain, applies the N4 off-field correction algorithm of the ANTs framework to eliminate intensity inhomogeneity caused by different scanning devices, implements brain stripping processing based on template matching to remove interference from the skull and extracranial tissues, and uses anisotropic resampling method (hematoma size 0.5×0.5×5mm³) to balance spatial resolution and computational efficiency. Finally, it uses grayscale standardization to uniformly map CT values (unit: HU) to the sensitive range of 0-100, forming unenhanced images at multiple time points before registration.
[0065] After the unenhanced image is formed, continue to steps 4) to 6) to construct spatiotemporally consistent four-dimensional image data.
[0066] 4) Upstream Lesion Segmentation: Using the unenhanced images at each time point before registration as input, the upstream segmentation network obtains an automatic segmentation mask (i.e., lesion segmentation mask) for aSAH-related lesions at each time point. Specifically, the upstream segmentation network uses SwinUnetR as the backbone network and embeds a self-attention mechanism of local windows and shift windows. It also loads a spatial prior probability map of subarachnoid hemorrhage (hereinafter referred to as the spatial prior probability map) obtained from a large amount of clinical data as a spatial constraint to improve the ability to identify the core hemorrhage area and surrounding edema zone. At the same time, the upstream segmentation network is optimized through supervised training, using a cross-entropy and Dice joint loss function, and improving robustness through data augmentation and spatial prior regularization. Finally, it outputs a binary segmentation mask or a probabilistic segmentation mask at each time point, i.e., the lesion segmentation mask.
[0067] Specifically, the structure of the backbone network SwinUnetR of the upstream segmentation network is as follows: Figure 2 As shown. Figure 2 This demonstrates a self-attention mechanism that combines local windows and shifted windows (corresponding to...) Figure 2 The encoder and decoder structure of the self-attention module in the model emphasizes the spatial prior probability map as a constraint.
[0068] Specifically, the spatial prior probability map is the prior modeling part for each lesion subtype. Based on the assumption that different subtypes of hemorrhage have relatively fixed spatial distribution patterns in the brain, a statistical model of three types of hemorrhage (parenchymal hemorrhage, intraventricular hemorrhage, and subarachnoid hemorrhage) is performed using a multi-center dataset of 2800 labeled lesions. The specific process includes: First, spatial standardization is performed on all brain NCCT images. Rigid and affine registration is used to align the images to the standard MNI space, and the corresponding lesion labels are mapped to this space. Then, the frequency of occurrence of each type of hemorrhage is statistically analyzed on a voxel-by-voxel basis, thereby generating a spatial prior probability map reflecting the spatial distribution characteristics of different hemorrhage subtypes, providing structured spatial prior information for subsequent networks.
[0069] The encoder consists of multi-level residual convolutional modules and downsampling modules. It extracts local spatial features through residual convolutional modules and embeds self-attention modules that combine local windows and shift windows in each level to perform block modeling and cross-window information interaction on features. At the same time, it integrates spatial prior probability maps as explicit constraints to guide the network to focus on spatial regions with high prior probabilities during the feature encoding stage. The decoder is symmetrically set with the encoder. It restores spatial resolution through progressive upsampling modules and residual convolutional modules, and uses skip connections to introduce multi-scale features from the encoding stage into the decoding stage. This achieves refined spatial reconstruction while maintaining global context consistency, and finally outputs the target prediction result, i.e., the lesion label.
[0070] To visually demonstrate the segmentation effect of the upstream segmentation network, Figure 3 Example output results are given on unenhanced images at three time points. Figure 3 The left image in the image is the unenhanced image. Figure 3 The middle image in the image shows the manually annotated gold standard segmentation result. Figure 3 The right image shows the lesion segmentation result automatically generated by the upstream segmentation network. Figure 3 It can be seen that the upstream segmentation network can distinguish and label the hemorrhage core area and its surrounding edema zone. The lesion segmentation results automatically generated by the network are consistent with the corresponding gold standard results in terms of spatial distribution, indicating that the upstream segmentation network has the ability to extract spatial features related to lesion categories.
[0071] 5) Image Enhancement: To address the interference of lesion areas on registration results, the system uses a pre-trained CDGM to enhance the contrast and refine the structure of low-contrast or noisy areas in the unenhanced image, outputting an enhanced image. The enhanced image serves as an auxiliary input in the subsequent high-precision registration step, with the core purpose of improving the robustness and accuracy of deformation estimation during the registration process.
[0072] Specifically, such as Figure 4As shown, Figure 4 This demonstration shows the overall process of inputting an unenhanced image into an I2SB model and outputting an enhanced image through an image-to-image diffusion modeling process. Specifically, the input unenhanced image is progressively mapped to a noisy image X during the forward noise diffusion process. t The intermediate random state is formed through multi-step diffusion. Subsequently, in the conditional denoising process, an I2SB model conditioned on the unenhanced image is introduced to progressively reverse denoise the intermediate random state, finally reconstructing the enhanced image. By introducing conditional constraints in the diffusion and denoising stages, conditional generation and enhancement of cross-modal image information are achieved.
[0073] To demonstrate the difference before and after image enhancement, please refer to [reference needed]. Figure 5 , Figure 5 The first row shows the unenhanced image, and the second row shows the enhanced image generated by the I2SB model. Figure 5 The comparison results show that the enhanced images exhibit a clearer contrast between different brain tissues, and some local brain tissue structural details are further revealed. From Figure 5 It is evident that CDGM can enhance brain NCCT images across modalities while maintaining the consistency of the original anatomical structures. Using enhanced images to replace the original brain NCCT images for subsequent high-precision registration steps helps provide downstream multi-task optimization modules with more discriminative representations of image tissue features.
[0074] 6) High-precision registration: Using the baseline image (i.e. the unenhanced image at the first time point) as a reference, and combined with the lesion segmentation mask generated in step 4), the deformation registration process based on the ANTs framework and fused with the cost function mask is started for the unenhanced images and their corresponding enhanced images at subsequent time points (other than the first time point). Specifically, the deformation registration process at subsequent time points uses unenhanced images as the main input and enhanced images as auxiliary input. During the registration process, the SyN symmetric normalized deformation registration algorithm based on the ANTs framework is used for high-precision registration, aligning the images at subsequent time points with the reference image. The cross-correlation coefficient (CC) is used as the similarity measure. At the same time, the lesion segmentation mask output by the upstream segmentation network in step 4) is used as the cost function mask to shield the influence of the lesion region on the similarity measure. The deformation field is gradually optimized using a coarse-to-fine multi-resolution strategy to obtain a more reliable deformation field. Simultaneously, a stepwise multi-resolution optimization strategy from 1 / 8 resolution to full resolution (1 / 8 resolution, 1 / 4 resolution, 1 / 2 resolution, full resolution) is executed, i.e., a four-level hierarchical optimization strategy, to achieve sub-millimeter-level spatiotemporal alignment of images at each time point. Then, the quality is evaluated by calculating the structural similarity coefficient (SSIM) of normal brain tissue regions in the registered images. Only high-quality registration results with SSIM > 0.80 are retained, and finally, spatiotemporally aligned four-dimensional image data are generated.
[0075] Please refer to Figure 6 Taking three time points as examples, the results before and after high-precision registration are compared. Figure 6 The left half shows the unenhanced image and lesion segmentation mask before registration. Figure 6 (seg). By Figure 6 It can be seen that there are global rotation and translation differences between images at different time points. At the same time, due to the evolution of brain lesions and surrounding tissues over time, there is also a certain degree of nonlinear deformation, which causes the lesion segmentation masks corresponding to each time point to be misaligned in spatial position. Figure 6 The right half shows the registered image and the lesion segmentation mask. Figure 6 It can be seen that after high-precision registration processing, the global structure, local brain tissue and corresponding lesion labels at different time points are all consistent in spatial position, indicating that the present invention can simultaneously handle global displacement and local nonlinear deformation, thereby achieving high-precision spatial alignment of multi-temporal images.
[0076] To ensure the quality of data registration, the system establishes a fully automated quality control system. It uses customized scripts to detect and screen anomalies in the raw data (i.e., multi-temporal brain NCCT data), visualizes the brain region segmentation results and provides them for expert review, and uses multi-dimensional quantitative indicators such as normalized cross-correlation coefficient (NCC) and structural similarity coefficient (SSIM) to evaluate the registration quality. It also automatically triggers a manual review mechanism for suspected abnormal results, and finally outputs four-dimensional image data that meets the quality standards.
[0077] 7) Deep Spatial Feature Extraction: Based on the high-quality four-dimensional image data, a downstream feature extraction module is constructed. The encoder part of the SwinUnetR network, which has been trained in the upstream lesion segmentation step (step 4), is extracted and used as the encoder of the improved 3D Transformer-UNet. This encoder retains the self-attention mechanism of local windows and shift windows to achieve hierarchical feature extraction from micro-texture to macro-morphology. It continues to use the spatial prior probability map obtained from the upstream lesion segmentation step (step 4) for spatial constraints. While maintaining the upstream segmentation performance, the pre-trained weights of the encoder are fine-tuned to enhance its feature stability in multi-temporal tasks. The spatiotemporally aligned four-dimensional image data (organized using channel stacking) is used as the input to the encoder. The encoder outputs 1024-dimensional spatial feature latent space codes (i.e., location-specific deep spatial features) for each time point. The 1024-dimensional spatial feature latent space codes for multiple time points constitute a 1024-dimensional spatial feature latent space code sequence (i.e., a deep spatial feature sequence).
[0078] Specifically, the encoder structure and its input-output relationship of the improved 3D Transformer-UNet are as follows: Figure 7 As shown, Figure 7 The upper part of the SwinUnetR encoder module (upstream encoder module) is based on Figure 2 The network structure shown is a pre-trained encoder subnetwork, whose parameter weights are inherited and used for feature extraction in this module. This upstream encoder takes a spatial prior probability map and four-dimensional image data as input, and outputs a multi-scale spatial feature representation through multi-layer feature encoding based on local and shifted window self-attention mechanisms. Figure 7 (256-dimensional time-point t-space feature latent space encoding). Figure 7The lower half encoder (referred to as the downstream encoder) inherits the above encoding results and further encodes the four-dimensional image data at time point t to generate the corresponding spatial feature latent space code (i.e., the depth spatial features at time point t). Finally, the downstream encoder outputs a set of 1024-dimensional spatial feature latent space codes at time point t (the depth spatial feature sequence at time point t). Its spatial resolution is determined by the feature scale of the lowest layer of UNet and is used for subsequent temporal modeling or downstream task processing.
[0079] 8) Temporal Feature Aggregation: A temporal feature aggregation module based on a Transformer encoder and a multi-head self-attention mechanism is designed. The 1024-dimensional spatial feature latent space encoding sequence (i.e., deep spatial feature sequence) is input into an encoder with 4 attention heads. The multi-head self-attention mechanism is used to calculate the dependencies in the temporal dimension (i.e., calculate the correlation of features across time phases), generate an attention weight matrix, and adaptively allocate the weights of features at different time points. The features are weighted and fused according to the weights to realize the intelligent fusion of deep spatial features and output a 512-dimensional spatiotemporal feature latent space encoding (i.e., aggregated features).
[0080] Specifically, such as Figure 8 As shown, 1024-dimensional spatial feature latent space codes (i.e., deep spatial features) from three time points are input into a Transformer encoder module with shared parameters to obtain spatial feature latent space codes for the corresponding time points. The encoded feature dimension for each time point is 256×20×20×20. These multi-time-point spatial feature latent space codes are then stacked along the time dimension to form a four-dimensional multi-time-point latent space code, which is then input into a multi-head self-attention module for temporal feature modeling and information interaction. Through the multi-head self-attention mechanism, the correlation between different time points is adaptively modeled, ultimately outputting a unified 512-dimensional spatiotemporal feature latent space code (i.e., aggregated features). Figure 8 It can be seen that this temporal feature aggregation module can achieve the fusion expression of temporal features while maintaining the deep spatial features at each time point, providing deep spatiotemporal feature representation for subsequent reconstruction, prediction or discrimination tasks.
[0081] The system adopts an end-to-end modular architecture. The training process is optimized through an adaptive weighted multi-task loss function and a temporal consistency regularization method. The upstream segmentation network and the downstream feature extraction module are pre-trained first. Then, the 512-dimensional spatiotemporal feature latent space encoding (i.e., the aggregated features) is used as the basic feature input. Subsequently, it will be optimized through downstream tasks to further improve its discriminative power, temporal sensitivity and clinical relevance. Therefore, step 9 is executed.
[0082] 9) Downstream Multi-Task Optimization: A downstream multi-task optimization module is constructed. Using downstream tasks as constraints, a joint multi-task optimization strategy is adopted. The core optimization objective is to optimize the 512-dimensional spatiotemporal feature latent space encoding (i.e., the aggregated features) output from step 8). Through regression / classification loss and spatial gradient constraints, the encoder and temporal feature aggregation module of the improved 3D Transformer-UNet are guided to iteratively learn, ultimately obtaining a dynamic feature embedding vector (i.e., an optimized version of the 512-dimensional spatiotemporal feature latent space encoding) with discriminative, temporal sensitivity, and clinical relevance. The dynamic feature embedding vector can characterize the dynamic evolution of lesions, including key pathological processes such as hematoma clearance dynamics and edema expansion trajectory. Thus, the downstream multi-task optimization module constitutes the final utilization and output form of spatiotemporal feature information of the system, taking into account both prediction result generation and feature interpretability presentation.
[0083] Specifically, such as Figure 9 As shown, the 512-dimensional spatiotemporal feature latent space encoding (i.e., the aggregated features) output by the temporal feature aggregation module is used as input. After feature transformation and compression by a multi-layer downsampling encoding module, the features are connected to the classification and regression branches respectively to output the corresponding class labels and risk prediction results, thereby realizing multi-task joint modeling based on the same spatiotemporal feature representation. The training parameters are set as follows: the initial learning rate is 0.0001, the Adam optimizer is used, and five-fold cross-validation is used to ensure the robustness of the model.
[0084] It should be noted that the three types of sub-modules of the downstream tasks are joint training constraint tasks to achieve the above optimization objectives, specifically including:
[0085] 9.1) 3D Reconstruction Decoder Module: The 1024-dimensional spatial feature latent space encoding (i.e., depth spatial features) extracted by the encoder of the improved 3D Transformer-UNet, or the 512-dimensional spatiotemporal feature latent space encoding (i.e., aggregated features) after temporal aggregation, is input into the decoder to restore the feature map consistent with the original 3D brain NCCT image space, so as to realize the 3D spatial explicit expression of depth spatial features and provide spatial dimension constraints for feature optimization.
[0086] 9.2) Prognostic prediction module: The 512-dimensional spatiotemporal feature latent space encoding (i.e., the aggregated features) is input into a lightweight multilayer perceptron (MLP) to output the prognostic risk score of aSAH patients; the temporal features are supervised by regression / classification loss to guide the temporal aggregation module to learn the time patterns and dynamic features of the disease course related to the real clinical outcome, providing clinical prediction dimension constraints for feature optimization, so that the final optimized dynamic feature embedding vector has clear clinical prediction significance.
[0087] 9.3) Grad-CAM++ Visualization Module: To enhance clinical usability, a clinically interpretable Grad-CAM++ visualization module was constructed. Its core logic involves locating and calculating the spatial region that contributes most to the prediction result from the gradient information of downstream prognostic prediction branches. This region is then transformed into a visual heatmap, with the heatmap's color intensity corresponding to the magnitude of contribution; darker colors represent higher contributions, thus highlighting the contribution of key pathological regions to the prediction result. For more details, please refer to [link / reference]. Figure 9 In the feature encoding process, gradient calculation of feature maps and weighted summation mechanisms are introduced to generate feature response heatmaps based on the Grad-CAM++ algorithm. Figure 9 (Grad-CAM++ heatmap).
[0088] The system integrates the Grad-CAM++ heatmap, enabling the back-mapping of key information related to model decisions back to the original unenhanced image space. This achieves interpretable visualization of spatiotemporal features, intuitively presenting the correlation between key pathological regions and prediction results, while providing interpretable dimensional constraints for optimizing dynamic feature embedding vectors. Furthermore, the importance weights of different monitoring time windows calculated through the multi-head self-attention mechanism in step 8), along with the heatmap, provide physicians with intuitive decision-making support, helping them quickly locate key time points and core pathological regions. To adapt to clinical application needs, the system features a user-friendly graphical interface, including modules for data management, processing monitoring, and result display, supporting real-time monitoring and interactive adjustments of the processing workflow by physicians.
[0089] II. Computer Equipment
[0090] The present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of a multi-temporal brain NCCT image processing method based on deep learning.
[0091] The memory can be volatile or non-volatile, or a combination of both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory of this invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0092] The processor can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the aforementioned multi-temporal brain NCCT image processing method based on deep learning can be completed through integrated logic circuits in the processor's hardware or through software instructions. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the steps and logic block diagrams disclosed in this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps disclosed in this invention can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the field. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the aforementioned multi-temporal brain NCCT image processing method based on deep learning.
[0093] The steps of this invention can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof.
[0094] Software implementation can be achieved by executing functional modules (such as procedures, functions, etc.). Software code can be stored in memory and executed by the processor. Memory can be implemented in the processor or outside the processor.
[0095] III. Computer-readable storage media
[0096] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the aforementioned multi-temporal brain NCCT image processing method based on DL.
[0097] Computer storage media can include various media that can store program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] IV. Computer Program Products
[0099] The present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned multi-temporal brain NCCT image processing method based on deep learning.
[0100] Specifically, computer program products include: data signals, data signals embodied in a carrier wave, or computer-readable storage media.
[0101] It should be noted that the technical solutions described in this invention can be combined arbitrarily without conflict.
[0102] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention also includes these modifications and variations.
Claims
1. A multi-temporal brain NCCT image processing method based on deep learning, applied to a multi-temporal brain NCCT image processing system based on deep learning, characterized in that, include: After image enhancement using CDGM, a deformation registration process based on the ANTs framework and fused with a cost function mask is initiated to complete the spatial alignment of multi-temporal brain NCCT images. The improved 3D Transformer-UNet uses the encoder of the SwinUnetR network, embeds the self-attention mechanism of local window and shift window, and uses spatial prior probability graphs for spatial constraints to complete the extraction of location-specific depth spatial features. Based on the Transformer encoder and multi-head self-attention mechanism, the aggregation of temporal features is completed; Using downstream tasks as constraints, a joint multi-task optimization strategy is adopted to obtain dynamic feature embedding vectors that characterize the dynamic evolution of lesions. The three types of sub-modules of the downstream tasks include a three-dimensional reconstruction decoder module, a prognostic prediction module, and a Grad-CAM++ visualization module. The steps for image enhancement using CDGM include: During the forward noise diffusion process, the input unenhanced image is gradually mapped to the noisy state image, and an intermediate random state is formed through multi-step diffusion. In the conditional denoising process, CDGM is introduced with the unenhanced image as the condition, and the intermediate random state is gradually reversed to denoise, and finally the enhanced image is reconstructed. The deformation registration process based on the ANTs framework and incorporating a cost function mask specifically includes: Using unenhanced images as the main input, enhanced images as auxiliary input, and baseline images as a reference, the SyN symmetric normalized deformation registration algorithm based on the ANTs framework is used for registration. The cross-correlation coefficient is used as the similarity measure, and the lesion segmentation mask output by the upstream segmentation network is used as the cost function mask. At the same time, a four-level hierarchical optimization strategy from 1 / 8 resolution to full resolution is implemented to achieve sub-millimeter-level spatiotemporal alignment of images at each time point. The SSIM of normal brain tissue regions in the registered images was calculated for quality assessment. Only high-quality registration results with SSIM > 0.80 were retained, and finally, spatiotemporally aligned four-dimensional image data were generated. The upstream segmentation network uses SwinUnetR as its backbone network and embeds a self-attention mechanism of local window and shift window, and loads a spatial prior probability map as a spatial constraint. Meanwhile, the upstream segmentation network is optimized through supervised training, adopts a joint loss function of cross-entropy and Dice, and improves robustness through data augmentation and spatial prior regularization, and finally outputs the lesion segmentation mask at each time point. The specific steps of deep spatial feature extraction include: The encoder part of the already trained SwinUnetR network in the upstream segmentation network is extracted and used as the encoder of the improved 3D Transformer-UNet. The encoder retains the self-attention mechanism of local window and shift window and uses spatial prior probability graph for spatial constraints. While maintaining the upstream segmentation performance, the pre-trained weights of the encoder are fine-tuned; The spatiotemporally aligned four-dimensional image data is used as the input to the encoder, which outputs the depth spatial features at each time point.
2. The multi-temporal brain NCCT image processing method based on deep learning as described in claim 1, characterized in that, The specific steps for aggregating time-series features include: Design a temporal feature aggregation module based on Transformer encoder and multi-head self-attention mechanism. Input the deep space feature sequence into Transformer encoder with multiple attention heads, calculate the dependencies in the time dimension through multi-head self-attention mechanism, generate attention weight matrix and adaptively assign weights to features at different time points, perform weighted fusion of features according to weights to achieve intelligent fusion of deep space features, and output aggregated features.
3. The multi-temporal brain NCCT image processing method based on DL as described in claim 1, characterized in that, The specific steps to obtain the dynamic feature embedding vector include: The aggregated features are optimized by using regression / classification loss and spatial gradient constraints to guide the encoder and temporal feature aggregation module of the improved 3DTransformer-UNet to learn iteratively, ultimately obtaining dynamic feature embedding vectors with discriminative, temporal sensitivity and clinical relevance.
4. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the multi-temporal brain NCCT image processing method based on DL as described in any one of claims 1-3.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-temporal brain NCCT image processing method based on DL as described in any one of claims 1-3.
6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-temporal brain NCCT image processing method based on DL as described in any one of claims 1-3.
Citation Information
Patent Citations
Improved M-Net-based RGB color remote sensing image cloud detection method and system
CN109934200A
Brain tissue hemorrhage location classification and hemorrhage quantification method, device, medium and program
CN116245951A