Image recognition and enhanced CT image generation method based on plain-scan CT
By introducing hemodynamic parameters and deep learning models into NCE-CT images to generate CE-CT images, the accuracy problem of AD diagnosis in NCE-CT is solved, achieving efficient AD detection and assessment and reducing reliance on CE-CT.
Patent Information
- Application Number
- CN202511598397.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies struggle to accurately diagnose aortic dissection (AD) in non-contrast enhanced CT (NCE-CT) images. The low contrast between the aorta and the intimal flap increases the likelihood of misdiagnosis, and traditional pixel-level methods cannot fully capture hemodynamic factors, affecting diagnostic accuracy.
An image recognition and enhanced CT image generation method based on plain CT images is adopted. Hemodynamic parameters are calculated by CFD, and a PIAD deep learning model is constructed by combining cross-attention mechanism and Transformer global information extraction module to realize the generation of NCE-CT images to CE-CT images and aortic dissection classification.
It improves the accuracy of AD detection, reduces reliance on expensive CE-CT equipment, and enhances diagnostic reliability by integrating physical prior information to provide a comprehensive AD assessment.
Smart Images

Figure CN121074019A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image processing, and particularly relates to an image recognition and enhanced CT image generation method based on a plain CT image. BACKGROUND
[0002] Aortic dissection (AD) is a relatively rare but potentially fatal disease, which results in the aorta being divided into true and false lumens due to a tear in the aortic wall. The rupture of these lumens often leads to fatal outcomes. Due to the complexity of its symptoms, accurate diagnosis is crucial to avoid serious consequences. Contrast-enhanced computed tomography (CE-CT) is the gold standard for AD diagnosis, renowned for its over 95% sensitivity and specificity. This outstanding diagnostic performance is attributed to the significantly higher X-ray attenuation coefficient of the contrast agent than the vessel wall, enabling CE-CT to clearly delineate the structure. However, CE-CT is not a routine examination in clinical screening, and the use of contrast agents also poses potential risks, especially for patients with allergies or acute kidney failure. Therefore, many patients, especially those exhibiting symptoms similar to AD and overlapping with other cardiovascular disease symptoms, are usually first evaluated with non-contrast-enhanced CT (NCE-CT). However, the diagnostic dilemma of NCE-CT lies in the small contrast between the aorta and the intimal flap, which obscures the diagnosis of AD and increases the likelihood of misdiagnosis.
[0003] Currently, there are many methods that utilize 2D axial non-contrast-enhanced CT (NCE-CT) images, 3D NCE-CT volumes, and morphological features, or synthetic contrast-enhanced CT (CE-CT) images from NCE-CT scans to distinguish AD and non-AD cases. These methods are based on CT image pixels, employing convolutional feature extraction to detect AD. However, given the inherent complex biomechanical and hemodynamic properties of the cardiovascular system, traditional pixel-level methods may not fully capture all factors involved in AD. Studies have shown that by utilizing computational fluid dynamics (CFD) and 4D flow to acquire hemodynamic parameters in aortic dissection and comparing them with a control group, it was found that pulse pressure and blood flow velocity were significantly increased in AD patients. The physical conditions of blood flow have a crucial impact on the dynamics of the aortic wall, which also affects the performance of AD in imaging modalities. Therefore, it is essential to incorporate physical priors that reflect actual environmental conditions into the diagnostic paradigm. This will provide a more comprehensive understanding of the causes of AD, thereby improving the accuracy and reliability of AD diagnosis. SUMMARY
[0004] The application provides a kind of image recognition and enhanced CT image generation method based on plain CT image at least solves the technical problem mentioned above, specifically adopts the following technical scheme:
[0005] An image recognition and enhanced CT image generation method based on plain CT images, comprising:
[0006] A NCE-CT image and CE-CT image paired data set is obtained, which is divided into a training set, a validation set and a test set in proportion;
[0007] Based on the paired data, the blood flow parameters of the true lumen of the aorta are calculated by CFD, and the physical prior features are obtained through feature selection;
[0008] A PIAD deep learning model is constructed and trained, which takes NCE-CT images as input and jointly performs CE-CT image generation, true lumen and false lumen segmentation and aortic dissection classification;
[0009] The PIAD deep learning model includes an encoder guided by the physical prior features through a cross-attention mechanism, a Transformer global information extraction module, and three functional heads that output classification results, CE-CT images and segmentation masks respectively;
[0010] Hyperparameter search is used to train different hyperparameter combinations on the training set and verify their performance on the validation set, and the best hyperparameter combination is selected to test the PIAD deep learning model with the best hyperparameters on the test set;
[0011] The trained PIAD deep learning model is used in the inference stage of inputting only NCE-CT images to simultaneously obtain the corresponding CE-CT images, true lumen and false lumen segmentation results and aortic dissection identification results.
[0012] Further, the blood flow parameters of the true lumen of the aorta calculated by CFD are as follows:
[0013] The true lumen of the aorta is segmented in Mimics-19;
[0014] The three-dimensional model is optimized in Geomagic Studio;
[0015] The mesh is generated in ICEM CFD;
[0016] Blood flow simulation is performed in ANSYS Fluent to output three-dimensional blood flow parameters.
[0017] Further, the XGBoost is used to sort the importance of all blood flow parameters output by CFD, select the top n parameters, and calculate the ratio of the maximum value to the minimum value of each parameter as the physical prior feature.
[0018] Further, n is 3.
[0019] Further, the encoder of the PIAD deep learning model is a 3D UNet structure, and the down-sampling path and the up-sampling path thereof are fused with high and low layer features through a jump connection.
[0020] The cross attention mechanism takes the physical prior feature as a query vector and takes the features at each level of the encoder as key and value vectors, so as to inject physical information into the latent space.
[0021] Further, the three functional heads include:
[0022] A classification head is configured to output a binary classification result of whether the aortic dissection exists or not.
[0023] A generative decoder is configured to generate a CE-CT image corresponding to the input NCE-CT image.
[0024] A segmentation decoder is configured to segment the true lumen and the false lumen of the aorta.
[0025] Further, the PIAD deep learning model further includes:
[0026] A discriminator is connected to the generative decoder and is configured to distinguish between a real CE-CT image and a generated CE-CT image.
[0027] Further, a total loss function is defined, and the total loss function is minimized during the training process to achieve the training optimization of the model.
[0028] The classification loss in the generator is defined as:
[0029] .
[0030] where t is the output of the Transformer encoder, is the classification head, is the label of AD.
[0031] The segmentation loss in the generator is defined as:
[0032] .
[0033] where is the Dice loss of segmentation, is a real segmentation image, is a predicted segmentation image.
[0034] The perception loss in the generator is defined as:
[0035] .
[0036] where the perceptual loss measured using the pre-trained VGG model, the real CE-CT image, the predicted CE-CT image;
[0037] The adversarial definition in the generator is:
[0038]
[0039] wherein is the discriminator with parameters , is the output NCE-CT, is the predicted CE-CT image;
[0040] The generated pixel-level loss in the generator is:
[0041] .
[0042] wherein is the real CE-CT image, is the predicted CE-CT image, and || ||1represents the L1 loss;
[0043] The total loss of the generator is:
[0044] + + + + .
[0045] wherein , , , is the weight coefficient of the corresponding loss;
[0046] The loss of the discriminator is:
[0047] .
[0048] wherein is the discriminator with parameters , is the output NCE-CT, is the predicted CE-CT image, is the real CE-CT image;
[0049] The total loss of the whole model is:
[0050] .
[0051] Further, after obtaining the NCE-CT image and the CE-CT image paired data set, the paired data set is further preprocessed as follows:
[0052] The NCE-CT image and the CE-CT image are registered using B-spline non-rigid registration;
[0053] The NCE-CT image and the CE-CT image are resampled to 1.3x1.3x5mm 3 ;
[0054] The image is cropped to the size of 128x128x64 around the thoracic aorta;
[0055] The CT value of the NCE-CT image is clipped to [0, 200], and the CT value of the CE-CT image is clipped to [0, 800];
[0056] The clipped CT value is normalized to the range of [-1, 1].
[0057] Further, the deep learning model is trained by minimizing the total loss function through the Adam optimizer, and the hyperparameters are determined through the grid search method.
[0058] The present application has the advantages that the provided image recognition and enhanced CT image generation method based on plain CT image can provide comprehensive aortic dissection (AD) detection using only non-enhanced CT (NCE-CT) data. By integrating physical prior information, combining the segmentation tasks of true lumen and false lumen and the CE-CT generation task, the complementary information of NCE-CT and CE-CT is effectively utilized, the accuracy of AD detection is improved, and in practical application, comprehensive AD evaluation can be performed only through NCE-CT, reducing the dependence on expensive and complex CE-CT equipment.
[0059] The present application also has the advantages that the provided aortic dissection detection method based on physical information introduces a novel physical prior knowledge, which combines hemodynamic parameters in the training process, promotes seamless and efficient combination of physical information and CT data in the latent space, and solves the consistency problem of different modal data. Experimental results show that integrating physical prior information significantly improves the efficiency of AD detection, proving the potential value of the method in practical clinical applications. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0061] Figure 1 The overall schematic diagram of the image recognition and enhanced CT image generation method based on plain CT image provided by the present application is shown. DETAILED DESCRIPTION
[0062] The embodiments of the present application will be described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.
[0063] As shown in Figure 1 , the present application discloses an image recognition and enhanced CT image generation method based on plain CT image, comprising the following steps:
[0064] S1: Obtain NCE-CT (non-enhanced CT) image, CE-CT (enhanced CT) image paired dataset, and divide it into training set, validation set and test set according to the proportion.
[0065] Specifically, the NCE-CT, CE-CT dataset is , and each sample , wherein, NCE-CT, CE-CT and cardiovascular label data matched respectively, denotes the total sequence length.
[0066] As a preferred embodiment, after obtaining the NCE-CT image, CE-CT image paired dataset, the paired dataset is also preprocessed as follows:
[0067] Registration: NCE-CT image and CE-CT image are registered using B-spline non-rigid registration.
[0068] Resampling: NCE-CT image and CE-CT image are resampled to 1.3x1.3x5mm 3 .
[0069] Crop: Crop the image to the size of 128x128x64 around the thoracic aorta.
[0070] CT value clipping:
[0071] The CT values of the NCE-CT images are clipped to [0, 200], and the CT values of the CE-CT images are clipped to [0, 800].
[0072] Normalization: The clipped CT values are normalized to the range [-1, 1].
[0073] S2: Based on the paired data, the blood flow dynamics parameters of the true lumen of the aorta are calculated using CFD, and physical prior features are obtained through feature selection.
[0074] In order to integrate the blood flow dynamics parameters into the model, the flow field in the true lumen of the patient's aorta is calculated using computational fluid dynamics (CFD), and a neural network is used to predict the CFD results, as there is no label for the true lumen in the inference phase.
[0075] In embodiments of the present application, the blood flow dynamics parameters of the true lumen of the aorta calculated using CFD are:
[0076] The true lumen of the aorta is segmented in Mimics-19, and the blood vessel geometry is simplified by removing the aortic arch branches and smoothing the surface. Mimics-19 is a medical image processing software used to segment and reconstruct three-dimensional anatomical models from CT / MRI scan data, and can export STL and other formats for subsequent CFD or finite element analysis.
[0077] The three-dimensional model is optimized in Geomagic Studio. Geomagic Studio is a reverse engineering / surface optimization software that converts STL, point cloud, and other polygonal data into smooth NURBS surfaces, and performs geometry repair, hole filling, surface fitting, etc. to improve the quality of CFD mesh and calculation accuracy. The STL file generated by Mimics is refined in Geomagic Studio, the blood vessel is segmented into an inlet and an outlet, and the mesh size is optimized to adapt to fluid dynamics simulation.
[0078] The mesh is generated in ICEM CFD. ICEM CFD is a high-level pre-processing software in the ANSYS series, designed specifically for CFD, supporting tetrahedral, hexahedral, and hybrid mesh generation for complex geometries.
[0079] Specifically, the mesh is generated in ICEM CFD, and the mesh is imported into CFX-Pre to set blood flow parameters such as steady-state, density, viscosity, relative pressure, inlet velocity, and outlet pressure, so that the simulation can be performed in ANSYS Fluent to generate detailed blood flow dynamics data.
[0080] Steady-state: (Steady-state).
[0081] Density: 1066 kg·m−3.
[0082] Viscosity: 0.0035 Pa·s.
[0083] Relative pressure: 9490 Pa.
[0084] Inlet velocity: 0.5 m·s−1.
[0085] Outlet pressure: 0 Pa.
[0086] Blood flow simulation is performed in ANSYS Fluent to output three-dimensional blood flow hemodynamic parameters, generating comprehensive blood flow hemodynamic data. These data contain physical parameters of each point in the blood vessel.
[0087] Next, considering the diverse and potentially redundant blood flow parameters obtained from ANSYS Fluent, XGBoost is employed to rank the importance of all blood flow parameters output by CFD, and to reduce the complexity of the subsequent prediction network, the top n parameters are selected, and the ratio of the maximum value to the minimum value of each parameter is calculated as the physical prior feature. In this application, the top 3 parameters are selected. Without direct access to the true lumen, 3D UNet and Transformer encoder networks are used to predict these parameters on NCE-CT images and are trained through L1 loss. This method enhances the practicality of the model in real-world applications by estimating key hemodynamic parameters without direct measurement.
[0088] S3: Construct and train the PIAD deep learning model, which takes NCE-CT images as input and jointly performs CE-CT image generation, true lumen false lumen segmentation, and aortic dissection classification.
[0089] In the embodiments of the present application, the PIAD deep learning model includes: an encoder guided by physical prior features through cross-attention mechanism, a Transformer global information extraction module, and three functional heads that output classification results, CE-CT images, and segmentation masks, respectively. The encoder of the PIAD deep learning model is a 3D UNet structure, and its down-sampling path and up-sampling path fuse high and low-level features through skip connection. The cross-attention mechanism takes the physical prior features as the query vector and the encoder features at each level as the key-value vector, achieving the injection of physical information into the latent space. The three functional heads include: a classification head, a generation decoder, and a segmentation decoder. The classification head is used to output a binary classification result of the existence of aortic dissection. The generation decoder is used to generate the CE-CT image corresponding to the input NCE-CT image. The segmentation decoder is used to segment the true lumen and false lumen of the aorta. Further, the PIAD deep learning model also includes: a discriminator. The discriminator is connected to the generation decoder and is used to distinguish between real CE-CT images and generated CE-CT images.
[0090] In particular, the original NCE-CT data First, it is encoded by a Unet CNN encoder guided by cross-attention from predicted physical information parameters. The output is passed through a Transformer encoder to capture global information of the NCE-CT data. Next, the output of the Transformer encoder is input to a classification head for classification, to a decoder to generate CE-CT data, and to another decoder to segment the true and false lumen. The classification loss is where is the label of AD. Each layer of the two decoders also receives data from each layer of the preceding UNet encoder, forming a skip connection. For the segmentation and generation tasks, the present invention defines as the real CE-CT image, as the real segmentation image, as the predicted CE-CT image, as the predicted segmentation image, as the entire encoder and two decoders, as the discriminator of GANs. The loss of the generator is where is the perceptual loss obtained by a pre-trained VGG network, is the classification loss after a fixed classifier. The loss of the discriminator is .
[0091] The total loss function is defined, and the total loss function is minimized during the training process to achieve the training optimization of the model;
[0092] The classification loss in the generator is defined as:
[0093] .
[0094] where t is the output of the Transformer encoder, is the classification head, is the label of AD.
[0095] The segmentation loss in the generator is defined as:
[0096] = .
[0097] where is the Dice loss of segmentation, is the real segmentation image, is the predicted segmentation image
[0098] The perceptual loss in the generator is defined as:
[0099] .
[0100] where is the perceptual loss measured using a pre-trained VGG model, is the real CE-CT image, is the predicted CE-CT image.
[0101] The adversarial definition in the generator is:
[0102] .
[0103] where is the discriminator with parameters , is the output NCE-CT, is the predicted CE-CT image.
[0104] The generated pixel-level loss in the generator is:
[0105] .
[0106] where is the real CE-CT image, is the predicted CE-CT image, || ||1denotes the L1 loss.
[0107] The total loss of the generator is:
[0108] + + + + .
[0109] where , , , is the weight coefficient of the corresponding loss.
[0110] The loss of the discriminator is:
[0111] .
[0112] where is the discriminator with parameters , is the output NCE-CT, is the predicted CE-CT image, is the real CE-CT image.
[0113] Therefore, the total loss of the entire model is:
[0114] .
[0115] S4: Using hyperparameter search, training on the training set with different hyperparameter combinations, and verifying its performance on the validation set, selecting the best hyperparameter combination, testing the PIAD deep learning model with the best hyperparameter on the test set to obtain the predicted intervention effect.
[0116] Specifically, the data set is divided into training set, test set and validation set according to a certain proportion. Different hyperparameter combinations are used to construct the model, and the training set samples are input. The model is trained by minimizing the loss function through the Adam optimizer, and the trained model is verified on the validation set. After selecting the optimal hyperparameter combination on the validation set, the corresponding model is tested on the test set to verify the performance, and finally the performance of the method is obtained.
[0117] S5: The trained PIAD deep learning model is used in the inference stage of inputting only NCE-CT images, and the corresponding CE-CT images, true cavity and false cavity segmentation results and aortic dissection identification results are obtained synchronously.
[0118] In the embodiments of the present application, the deep learning model is trained by minimizing the total loss function through the Adam optimizer, and the hyperparameters are determined by the grid search method.
[0119] Specifically, taking the experimental data set as an example (200 for training set and validation set, and 50 for test set). As shown in Table 1, the model of the present application has better performance.
[0120] Table 1: Performance of the prediction method of the present application and the current most advanced model on the data set
[0121]
[0122] The basic principles, main features and advantages of the present application are shown and described above. Those skilled in the art should understand that the above examples do not limit the present application in any form, and any technical solution obtained by equivalent replacement or equivalent transformation falls within the scope of the present application.
Claims
1. A method for image recognition and enhanced CT image generation based on plain CT images, characterized in that, Include: Obtain paired datasets of NCE-CT images and CE-CT images, and divide them into training, validation, and test sets according to proportions. Based on the paired data, CFD is used to calculate the hemodynamic parameters of the aortic true lumen, and physical prior features are obtained through feature selection. A PIAD deep learning model is constructed and trained. The PIAD deep learning model takes NCE-CT images as input and jointly performs CE-CT image generation, true lumen and false lumen segmentation and aortic dissection classification. The PIAD deep learning model includes: an encoder guided by the physical prior features through a cross-attention mechanism, a Transformer global information extraction module, and three functional heads that output classification results, CE-CT images, and segmentation masks, respectively. Using hyperparameter search, we train on the training set with different combinations of hyperparameters, validate their performance on the validation set, select the best combination of hyperparameters, and test the PIAD deep learning model with the best hyperparameters on the test set. The trained PIAD deep learning model is used in the inference stage with only NCE-CT images as input, and the corresponding CE-CT images, true and false lumen segmentation results and aortic dissection identification results are obtained simultaneously.
2. The image recognition and enhanced CT image generation method based on plain CT images according to claim 1, characterized in that, The specific steps for calculating the hemodynamic parameters of the aortic true lumen using CFD are as follows: Dissecting the true lumen of the aorta in Mimics-19; Optimize 3D models in Geomagic Studio; Generate the mesh in ICEM CFD; Perform blood flow simulations in ANSYS Fluent to output three-dimensional hemodynamic parameters.
3. The image recognition and enhanced CT image generation method based on plain CT images according to claim 2, characterized in that, The feature selection step includes: using XGBoost to sort all blood flow parameters output by CFD by importance, selecting the top n parameters, and calculating the ratio of the maximum value to the minimum value of each parameter as the physical prior feature.
4. The image recognition and enhanced CT image generation method based on plain CT images according to claim 3, characterized in that, n is 3.
5. The method for image recognition and enhanced CT image generation based on plain CT images according to claim 1, characterized in that, The encoder of the PIAD deep learning model is a 3D UNet structure, and its downsampling path and upsampling path fuse high and low layer features through skip connections. The cross-attention mechanism uses the physical prior features as query vectors and the features at each level of the encoder as key vectors to inject physical information into the latent space.
6. The method for image recognition and enhanced CT image generation based on plain CT images according to claim 1, characterized in that, The three functional headers include: The classification head is used to output a binary classification result indicating whether aortic dissection exists; A generator decoder is used to generate a CE-CT image corresponding to the input NCE-CT image; A separator decoder is used to separate the true lumen from the false lumen of the aorta.
7. The method for image recognition and enhanced CT image generation based on plain CT images according to claim 6, characterized in that, The PIAD deep learning model also includes: A discriminator, connected to the generator decoder, is used to distinguish between real CE-CT images and generated CE-CT images.
8. The method for image recognition and enhanced CT image generation based on plain CT images according to claim 7, characterized in that, Define a total loss function, and minimize the total loss function during training to optimize model training; The classification loss in the generator is defined as: , Where t is the output of the Transformer encoder, For classification header, Tags for AD; The segmentation loss in the generator is defined as: = , in For the segmentation of Dice loss, For real segmented images, For the predicted segmented image; The perceptual loss in the generator is defined as: , in The perceptual loss is measured using a pre-trained VGG model. For real CE-CT images, For the predicted CE-CT image; Adversarial relationships in generators are defined as follows: , in For parameters The discriminator, For the output NCE-CT, For the predicted CE-CT image; The pixel-level loss in the generator is: , in For real CE-CT images, For the predicted CE-CT image, || ||1 represents L1 loss; The total loss of the generator is then: + + + + , in , , , These are the weighting coefficients for the corresponding losses; The loss of the discriminator is; , in For parameters The discriminator, For the output NCE-CT, For the predicted CE-CT image, These are real CE-CT images; The total loss of the entire model is: 。 9. The method for image recognition and enhanced CT image generation based on plain CT images according to claim 1, characterized in that, After acquiring the paired datasets of NCE-CT images and CE-CT images, the paired datasets were preprocessed as follows: The NCE-CT image was registered with the CE-CT image using B-spline non-rigid registration. The NCE-CT and CE-CT images were resampled to 1.3×1.3×5mm. 3 ; Crops the image around the thoracic aorta to a size of 128×128×64; The CT values of the NCE-CT images are cropped to [0, 200], and the CT values of the CE-CT images are cropped to [0, 800]. The clipped CT values are normalized to the range of [-1, 1].
10. The method for image recognition and enhanced CT image generation based on plain CT images according to claim 1, characterized in that, The deep learning model is trained by minimizing the total loss function using the Adam optimizer, and the hyperparameters are determined using a grid search method.
Citation Information
Patent Citations
Multi-temporal liver tumor segmentation method based on multi-head cross attention conversion network
CN115330816A
Method and equipment for constructing aortic dissection prediction model based on multi-modal features
CN118887401A
Cited By
Flat scanning and enhanced CT (Computed Tomography) automatic discrimination method and system based on aorta semantic segmentation and blood vessel region intensity statistics
CN122289275A