CT radiography intelligent imaging method and system based on DCT-GAN multi-scale fusion
By adopting the DCT-GAN multi-scale fusion method in CTA technology, discrete cosine transformation and multi-scale feature processing, the problem of insufficient clarity and artifacts of CTA image details is solved, and higher image conversion accuracy and authenticity are achieved.
Patent Information
- Application Number
- CN202510066197.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing CTA technology has insufficient image feature extraction, resulting in insufficient clarity in the details of the generated CTA images, which has artifacts, affecting the accuracy of the diagnosis.
The intelligent CT contrast imaging method based on DCT-GAN multi-scale fusion is adopted to process images in the frequency domain through discrete cosine transformation, enhance the loss calculation ability of the model in the frequency domain, and improve the accuracy and authenticity of image conversion through multi-scale feature extraction, feature alignment and fusion.
It improves the accuracy and authenticity of image conversion, reduces the problems of artifacts and blurred details, and makes the generated virtual CTA images clearer and more accurate.
Smart Images

Figure CN120013880A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing technology, and in particular to a CT angiography intelligent imaging method and system based on DCT-GAN multi-scale fusion. Background Art
[0002] CT angiography (CTA) is a non-invasive medical imaging technique that combines computed tomography (NCCT) and angiography. By injecting contrast agents into the veins, the CT device can capture the flow of contrast agents in the blood vessels, thereby generating a three-dimensional image of the blood vessels. CTA has high precision and three-dimensional imaging capabilities, and can provide millimeter-level vascular images to help doctors accurately identify lesions such as vascular stenosis, aneurysms, and vascular malformations. However, the injection of contrast agents can bring about many problems: (1) There are certain adverse reactions, such as urticaria, bronchospasm, laryngeal edema, and cardiac arrest; (2) Some patients are not suitable for contrast agent examinations, such as those with heart, liver, and kidney dysfunction. (3) It increases the examination time. After the contrast agent is injected, it takes a certain amount of time for the contrast agent to enter the corresponding organ before the examination can begin. The extension of the examination time will lead to delayed treatment of diseases such as acute stroke, thereby affecting the patient's prognosis.
[0003] With the development of deep learning technology, especially the introduction of the generative adversarial network (GAN) framework, some researchers have realized the image conversion from non-contrast CT (NCCT) to CTA by constructing an adversarial network model based on focused learning. This conversion technology can simulate the flow effect of contrast agents in blood vessels, thereby generating images close to real CTA. Although this technology can significantly shorten the inspection time and reduce costs, there are still some problems. For example, due to insufficient image feature extraction, the generated CTA images are not clear enough in details, or artifacts appear in certain areas, affecting the accuracy of diagnosis. Therefore, in order to improve the quality, realism and stability of the generated images, it is necessary to apply a deep learning model that combines advanced image feature extraction with multi-scale fusion, by enhancing the model's loss calculation ability in the frequency domain and deep fusion of multi-scale features, to improve the accuracy and authenticity of image conversion and reduce artifacts and blurred details. Summary of the invention
[0004] In order to improve the accuracy and authenticity of image conversion and reduce the problems of artifacts and blurred details, the purpose of the present invention is to provide a CT angiography intelligent imaging method and system based on DCT-GAN multi-scale fusion. The technical solutions adopted are as follows:
[0005] In a first aspect, the present application discloses a CT angiography intelligent imaging method based on DCT-GAN multi-scale fusion, wherein the method comprises:
[0006] S1, obtaining a preprocessed target NCCT-CTA image dataset, wherein the target NCCT-CTA image dataset includes a plurality of target NCCT images and target CTA images registered therewith;
[0007] S2, dividing the target NCCT-CTA image dataset according to a preset ratio to obtain a training set and a validation set;
[0008] S3, inputting the training set and the validation set into the initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the generator generates a virtual CTA image, and the discriminator discriminates the input original CTA image and the virtual CTA image in the image domain and the frequency domain processed by discrete cosine transform, and determines the overall loss function of the network based on the similarity difference in the image domain and the frequency domain, and the discrimination difference of the discriminator in the two domains;
[0009] S4. The preprocessed real-time NCCT image is input into the trained target generative adversarial network, and the generator processes it based on the DCT-GAN multi-scale fusion strategy to generate high-quality and high-resolution virtual CTA images.
[0010] Furthermore, in step S1, each registered image pair in the target NCCT-CTA image dataset is obtained by processing the following steps:
[0011] S11, acquiring an initial NCCT image and an initial CTA image registered therewith;
[0012] S12, operating the initial NCCT image and the initial CTA image registered therewith according to a preprocessing process based on the adaptive image analysis technology to obtain a preprocessed target NCCT image and a registered target CTA image therewith.
[0013] Furthermore, in step S12, the initial NCCT image and the initial CTA image registered therewith are operated according to the preprocessing process based on the adaptive image analysis technology to obtain the preprocessed target NCCT image and the target CTA image registered therewith, including:
[0014] S121, using an adaptive threshold segmentation algorithm to perform binarization processing on the initial NCCT image and the initial CTA image respectively, to obtain a binarized intermediate NCCT image and an intermediate CTA image;
[0015] S122, performing pixel-by-pixel intersection operation on the intermediate NCCT image and the intermediate CTA image to obtain an initial mask image;
[0016] S123, removing discrete regions in the initial mask image by maximum connected domain analysis to obtain a target mask image;
[0017] S124, extracting a region of interest excluding irrelevant background and noise regions from the initial NCCT image and the initial CTA image based on the target mask image, to obtain a preprocessed target NCCT image and a target CTA image registered therewith.
[0018] Furthermore, in step S2, the target NCCT-CTA image dataset is divided according to a preset ratio to obtain a training set, a validation set and a test set, including:
[0019] S21, performing image standardization processing on the target NCCT-CTA image dataset according to preset imaging characteristic difference adjustment rules, and slicing and screening standards, to obtain a standardized NCCT-CTA two-dimensional slice image dataset;
[0020] S22, dividing the standardized NCCT-CTA two-dimensional slice image data set according to a preset ratio to obtain a training set, a validation set, and a test set.
[0021] Furthermore, in step S21, the target NCCT-CTA image dataset is subjected to image standardization processing according to the preset imaging characteristic difference adjustment rules and the slice and screening criteria to obtain a standardized NCCT-CTA two-dimensional slice image dataset, including:
[0022] S211, adjusting the window level and window width of the target NCCT image and the target CTA image according to a preset imaging characteristic difference adjustment rule to obtain an adjusted target NCCT image and a target CTA image;
[0023] S212, performing a slicing operation on the adjusted target NCCT image and the target CTA image according to preset slicing parameters to obtain a preliminary NCCT-CTA two-dimensional slicing image data set;
[0024] S213. When it is determined that the proportion of the brain area in the corresponding two-dimensional slice image in the data set is less than a preset threshold, the two-dimensional slice image is deleted from the data set to obtain a final standardized NCCT-CTA two-dimensional slice image data set.
[0025] Furthermore, in step S3, during the training process, the overall network loss function is determined by the following steps:
[0026] S31, based on the generator in the initial generative adversarial network, perform multi-scale feature extraction, feature alignment and feature fusion on the input original NCCT image through multi-scale processing to obtain a fused feature map;
[0027] S32, inputting the fused feature map into a corresponding decoder for reconstruction to obtain a virtual CTA image;
[0028] S33, performing discrete cosine transform processing on the input original NCCT image, original CTA image, and virtual CTA image according to the following formula to obtain corresponding frequency domain images:
[0029]
[0030] Where f represents the image to be processed, DCT(f) u,v represents the frequency domain image obtained after the discrete cosine transform operation of the image f at the frequency coordinate (u, v), α(u) and α(v) represent the preset normalization coefficients, N represents the signal sampling length, and f(x, y) represents the pixel value of the image f at the spatial coordinate (x, y);
[0031] S34, determining a total loss function of the generator based on differences in the image domain and the frequency domain between the original CTA image and the virtual CTA image generated by the generator;
[0032] S35, determining an image domain adversarial loss function of the discriminator and a frequency domain adversarial loss function of the discriminator based on the discrimination results of the discriminator on the original NCCT image and the original CTA image, and the original NCCT image and the virtual CTA image in the image domain and the frequency domain;
[0033] S36: Perform weighted summation based on the determined loss functions to obtain the overall network loss function.
[0034] Further, in step S31, the generator based on the initial generative adversarial network performs multi-scale feature extraction, feature alignment and feature fusion on the input original NCCT image through multi-scale processing to obtain a fused feature map, including:
[0035] S311, based on the feature extraction layers of different scales included in the generator, extracting feature information of different levels from the input original NCCT image to obtain a multi-level feature map;
[0036] S312, aligning all feature maps except the last level according to the number of channels and sizes, so that these feature maps match the number of channels of the feature map of the last level and have consistent spatial resolution;
[0037] S313, performing feature fusion on each feature map after feature alignment and the feature map of the last level by adding elements one by one to obtain a fused feature map.
[0038] Furthermore, in step S34, the total loss function of the generator is determined based on the difference in the image domain and the frequency domain between the original CTA image and the virtual CTA image generated by the generator, including:
[0039] S341, constructing an image domain similarity loss function of a generator based on the difference between the original CTA image and the virtual CTA image in the image domain;
[0040] S342, constructing a frequency domain similarity loss function of a generator based on the difference between the original CTA image and the virtual CTA image in the frequency domain;
[0041] S343, constructing an adversarial loss function of the generator based on the difference in discrimination between the original CTA image and the virtual CTA image in the image domain and the frequency domain by the discriminator;
[0042] S344, performing weighted summation on the image domain similarity loss function, frequency domain similarity loss function, and adversarial loss function of the generator to obtain the total loss function of the generator.
[0043] Furthermore, in step S35, the image domain adversarial loss function of the discriminator is as follows:
[0044]
[0045] The frequency domain adversarial loss function of the discriminator is as follows:
[0046]
[0047] In step S36, the overall network loss function is as follows:
[0048]
[0049] Among them, NCCT original Represents the original NCCT image, CTA original Represents the original CTA image, CTA virtual represents a virtual CTA image, D represents the discriminator, and L BCE represents the binary cross entropy loss function, It represents the frequency domain image obtained after the original NCCT image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the original CTA image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the virtual CTA image is subjected to discrete cosine transform operation, Lgenerator Represents the total loss function of the generator.
[0050] In the second aspect, the present application discloses a CT angiography intelligent imaging system based on DCT-GAN multi-scale fusion, the system comprising a data acquisition module, a data division module, a GAN network training module and a virtual CTA image generation module, wherein:
[0051] The data acquisition module is used to acquire a preprocessed target NCCT-CTA image data set, wherein the target NCCT-CTA image data set includes a plurality of target NCCT images and target CTA images registered therewith;
[0052] The data partitioning module is used to partition the target NCCT-CTA image data set according to a preset ratio to obtain a training set and a validation set;
[0053] The GAN network training module is used to input the training set and the validation set into an initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the training set and the validation set are input into an initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the generator generates a virtual CTA image, and the discriminator discriminates the input original CTA image and the virtual CTA image in the image domain and the frequency domain processed by discrete cosine transform, and determines the overall network loss function based on the similarity difference in the image domain and the frequency domain, and the discrimination difference of the discriminator in the two domains;
[0054] The virtual CTA image generation module is used to input the preprocessed real-time NCCT image into the trained target generative adversarial network, and the generator processes it based on the DCT-GAN multi-scale fusion strategy to generate high-quality and high-resolution virtual CTA images.
[0055] The present invention has the following beneficial effects:
[0056] 1) Since the processing of discrete cosine transform in the frequency domain helps the model capture the high-frequency components in the image, these high-frequency components are crucial for the detailed representation of the image. Through the DCT-GAN multi-scale fusion strategy, the model can retain more detailed information when generating virtual CTA images, thereby reducing the problem of blurred details;
[0057] 2) By utilizing the synergy between the generator and the discriminator in the image domain and the frequency domain, the model can capture more detailed information and further improve the accuracy of image conversion;
[0058] 3) Based on the discriminant function in the frequency domain, the discriminator can also effectively suppress artifacts in the generated image, because artifacts usually appear as high-frequency noise or abnormal structures in the image. By optimizing the loss function, the model can gradually reduce these artifacts, making the generated virtual CTA images clearer and more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0060] Figure 1 A method flow chart of a CT angiography intelligent imaging method based on DCT-GAN multi-scale fusion provided by an embodiment of the present invention;
[0061] Figure 2 Comparison chart of CTA image generation effect compared with Pix2Pix method;
[0062] Figure 3 This is a model framework diagram for generating CTA images from NCCT images;
[0063] Figure 4 A system structure diagram of a CT angiography intelligent imaging system based on DCT-GAN multi-scale fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a CT angiography intelligent imaging method and system based on DCT-GAN multi-scale fusion proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0065] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0066] The following is a specific solution of a CT angiography intelligent imaging method and system based on DCT-GAN multi-scale fusion provided by the present invention, which is described in detail with reference to the accompanying drawings.
[0067] See also Figure 1, which shows a method flow chart of a CT angiography intelligent imaging method based on DCT-GAN multi-scale fusion provided by an embodiment of the present invention. The method can be applied to the scene of acute ischemic stroke, and specifically includes the following steps:
[0068] Step S1, obtaining a preprocessed target NCCT-CTA image dataset, wherein the target NCCT-CTA image dataset includes a plurality of target NCCT images and target CTA images registered therewith.
[0069] Specifically, during the preprocessing process, the application will perform binarization processing on the acquired initial NCCT and CTA images based on the adaptive threshold segmentation algorithm. After that, the target mask image is obtained by taking the intersection of each pixel and the maximum connected domain analysis to extract the region of interest. As a result, the common region of interest in both imaging modes is retained in the cropped NCCT image and CTA image, and irrelevant background and noise areas are removed, thereby ensuring the accuracy and efficiency of subsequent image processing, while reducing the consumption of computing resources.
[0070] Step S2, dividing the target NCCT-CTA image dataset according to a preset ratio to obtain a training set and a validation set.
[0071] Specifically, this application will divide the target NCCT-CTA image dataset into a training set and a validation set in a preset ratio according to the size and distribution characteristics of the dataset. The training set is used to train the generative adversarial network model so that the generator can learn the mapping relationship from NCCT images to CTA images. The validation set is used to evaluate the performance of the model, including the quality of the generated virtual CTA images, the similarity with the real CTA images, etc., to ensure that the model can perform well on unseen data.
[0072] In one of the embodiments, in the application scenario of acute ischemic stroke, in order to evaluate the generation performance of the model, the present application uses structural similarity as an evaluation indicator, and its calculation formula is:
[0073]
[0074] Among them, X is the virtual CTA image generated by the generative adversarial network, Y is the input original CTA image, μ X and μ Y represents the mean of the generated virtual CTA image and the original CTA image, δ X and δ Y represents the variance of the generated virtual CTA image and the original CTA image, and c1 and c2 are two constants introduced to avoid the denominator being zero.
[0075] Further, please refer to the following Table 1, it can be seen that the method of the present application performs better than the Pix2Pix method in the SSIM evaluation index, with an average value of 0.810, while the average value of Pix2Pix is 0.790. Therefore, it can be inferred that in the application scenario of acute ischemic stroke, the method of the present application can more accurately retain the structural information of the original image when generating a virtual CTA image, so that the visual effect is closer to the real CTA image.
[0076] Table 1
[0077] method SSIM Pix2Pix 0.790 This application 0.810
[0078] For further information, please refer to Figure 2 ,After comparative analysis of the quality of the virtual CTA images generated, especially the clarity and continuity of the vascular structure and the accuracy of the identification of the lesion area, it can be considered that the visual effect of the virtual CTA images generated by the present application is better than that of the Pix2Pix method, and the generation of blood vessels is more accurate.
[0079] Step S3, input the training set and the validation set into the initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the generator generates a virtual CTA image, and the discriminator discriminates the input original CTA image and the virtual CTA image in the image domain and the frequency domain processed by discrete cosine transform, and determines the overall loss function of the network based on the similarity difference in the image domain and the frequency domain, as well as the discrimination difference of the discriminator in these two domains.
[0080] For details, please refer to Figure 3 During the training process, the generator will try to generate more and more realistic virtual CTA images to deceive the discriminator. The discriminator needs to continuously improve its discrimination ability to accurately distinguish whether the input image is an original CTA image or a virtual CTA image generated by the generator. At the same time, in order to more comprehensively evaluate the image quality, the training process of this application not only focuses on the similarity in the image domain, but also introduces frequency domain discrimination processed by discrete cosine transform. Among them, frequency domain discrimination can capture high-frequency and low-frequency information in the image, so as to more accurately evaluate the texture, edges and other features of the image. Ultimately, the overall loss function of the network is mainly determined based on the difference in similarity between the image domain and the frequency domain, as well as the difference in discrimination of the discriminator in these two domains. By minimizing this loss function, the parameters of the generator and the discriminator can be further optimized, so that the generator can generate higher quality virtual CTA images, while enabling the discriminator to make more accurate discriminations.
[0081] In step S4, the preprocessed real-time NCCT image is input into the trained target generative adversarial network, and the generator processes it based on the DCT-GAN multi-scale fusion strategy to generate a high-quality, high-resolution virtual CTA image.
[0082] Specifically, when the preprocessed real-time NCCT images are input into the trained target generative adversarial network, the generator will process these images using the DCT-GAN multi-scale fusion strategy. It should be noted that the present application not only generates and adversarially trains images at a single scale, but also performs it at multiple scales simultaneously. In this way, the generator can better capture the detailed information in the image. During the forward propagation process, the generator will try to generate a virtual CTA image that is as visually close as possible to the real CTA image. At the same time, due to the introduction of discrete cosine transform processing technology, the generator can also learn the characteristics of the image in the frequency domain. Finally, after processing by the generator, high-quality, high-resolution virtual CTA images can be obtained to provide users with more accurate and reliable medical diagnostic information.
[0083] As can be seen from the above, the present application discloses a CT angiography intelligent imaging method based on DCT-GAN multi-scale fusion. Since the processing of discrete cosine transform in the frequency domain helps the model capture the high-frequency components in the image, these high-frequency components are crucial to the detailed expression of the image. Through the DCT-GAN multi-scale fusion strategy, the model can retain more detail information when generating virtual CTA images, thereby reducing the problem of blurred details; by utilizing the synergistic effect of the generator and the discriminator in the image domain and the frequency domain, the model can capture more detail information and further improve the accuracy of image conversion; based on the discriminator in the frequency domain, the discriminant effect can also effectively suppress artifacts in the generated image, because artifacts are usually manifested as high-frequency noise or abnormal structures in the image. By optimizing the loss function, the model can gradually reduce these artifacts, making the generated virtual CTA images clearer and more accurate.
[0084] In one embodiment, in step S1, each registered image pair in the target NCCT-CTA image dataset is obtained by processing the following steps:
[0085] Step S11, acquiring an initial NCCT image and an initial CTA image registered therewith.
[0086] Step S12, operating the initial NCCT image and the initial CTA image registered therewith according to a preprocessing process based on the adaptive image analysis technology to obtain a preprocessed target NCCT image and a registered target CTA image therewith.
[0087] In one embodiment, in step S12, the initial NCCT image and the initial CTA image registered therewith are operated according to a preprocessing process based on adaptive image analysis technology to obtain a preprocessed target NCCT image and a target CTA image registered therewith, including:
[0088] Step S121 , using an adaptive threshold segmentation algorithm to perform binarization processing on the initial NCCT image and the initial CTA image respectively, to obtain a binarized intermediate NCCT image and an intermediate CTA image.
[0089] Specifically, the present application uses the Otsu threshold segmentation algorithm to perform binarization processing on the initial NCCT image and the initial CTA image, respectively, wherein the algorithm can automatically determine a threshold value so that the intra-class variance between the segmented foreground and background is minimized, thereby maximizing the inter-class variance. In this way, each pixel in the initial NCCT image and the initial CTA image is divided into two categories: one category represents the region of interest, and the other category represents the background or noise. After such processing, in the intermediate NCCT image and the intermediate CTA image obtained after binarization, the region of interest will be clearly represented as white, and the background will be represented as black.
[0090] Step S122: performing pixel-by-pixel intersection operation on the intermediate NCCT image and the intermediate CTA image to obtain an initial mask image.
[0091] Specifically, the pixel-by-pixel intersection operation is performed on the intermediate NCCT image and the intermediate CTA image, which means that for each pixel position in the image, only when the pixel values of the two images at this position are both represented as the target area (i.e., both are the highlight values after binarization), the pixel value of the result image (i.e., the initial mask image) at this position is set to the value of the target area (usually 1 or white), otherwise it is set to the background value (usually 0 or black). This pixel-by-pixel intersection operation makes it possible to effectively retain those anatomical structures or features that are clearly visible in both the NCCT image and the CTA image, such as blood vessels, bones, etc.
[0092] Step S123, removing discrete regions in the initial mask image through maximum connected component analysis to obtain a target mask image.
[0093] Specifically, removing discrete regions in the initial mask image through maximum connected domain analysis means that the present application will identify and analyze connected regions in the initial mask image.
[0094] It should be noted that a connected region refers to a set of pixels that are connected to each other (i.e., adjacent and have the same pixel value) in an image. Maximum connected region analysis usually involves determining the largest connected region in an image and considering this region to be the most interesting part, while smaller discrete regions are considered noise or irrelevant background.
[0095] In the current embodiment, the present application will only retain the largest connected area in the initial mask image, and set all other discrete areas as background values, so as to obtain a clearer and more coherent target mask image.
[0096] In one embodiment, the present application removes discrete areas with an area less than 100 pixels in the initial mask image through maximum connected domain analysis, and by setting 100 pixels as the area threshold, these unimportant areas can be effectively removed while retaining connected areas with larger areas that are more likely to be real anatomical structures, thereby reducing the computational burden in subsequent image processing steps.
[0097] Step S124: extracting a region of interest excluding irrelevant background and noise regions from the initial NCCT image and the initial CTA image based on the target mask image, and obtaining a preprocessed target NCCT image and a target CTA image registered therewith.
[0098] Specifically, the present application will use the target mask image as a template to selectively extract information from the initial NCCT and CTA images. Specifically, the present application will traverse each pixel of the initial NCCT image and the CTA image, and only when the pixel value of the target mask image at the corresponding position represents the target area (such as 1 or white), the pixel value of the position will be retained in the preprocessed target NCCT and target CTA images; otherwise, the pixel value of the position is set to the background value. This process will generate two new images, namely the preprocessed target NCCT image and the target CTA image aligned with it, both of which only contain the region of interest corresponding to the target mask image, and irrelevant background and noise areas have been effectively removed.
[0099] In the above embodiment, through adaptive threshold segmentation, pixel-by-pixel intersection, maximum connected domain analysis and region of interest extraction, irrelevant background and noise areas in the initial NCCT and CTA images are effectively removed, the target mask image is accurately generated, and a high-quality pre-processed target image is extracted based on it, providing clearer and more accurate input data for subsequent image processing, thereby improving the overall processing efficiency and accuracy.
[0100] In one embodiment, in step S2, the target NCCT-CTA image dataset is divided according to a preset ratio to obtain a training set, a validation set, and a test set, including:
[0101] Step S21 , performing image standardization processing on the target NCCT-CTA image dataset according to preset imaging characteristic difference adjustment rules, and slicing and screening criteria, to obtain a standardized NCCT-CTA two-dimensional slice image dataset.
[0102] Step S22, dividing the standardized NCCT-CTA two-dimensional slice image data set according to a preset ratio to obtain a training set, a validation set and a test set.
[0103] In one embodiment, in step S21, the target NCCT-CTA image dataset is subjected to image standardization processing according to preset imaging characteristic difference adjustment rules and slice and screening criteria to obtain a standardized NCCT-CTA two-dimensional slice image dataset, including:
[0104] Step S211 , adjusting the window level and window width of the target NCCT image and the target CTA image according to a preset imaging characteristic difference adjustment rule, to obtain an adjusted target NCCT image and target CTA image.
[0105] It should be noted that due to the differences in imaging principles and tissue displays between NCCT images and CTA images, different window widths and window positions need to be set for them in order to highlight their respective anatomical structural features.
[0106] Specifically, the present application sets the window level and window width of the target NCCT image to 40 and 80, respectively, and sets the window level and window width of the target CTA image to 100 and 200, respectively, so that the soft tissue structure in the NCCT image and the vascular structure in the CTA image can be highlighted when the image is displayed.
[0107] Step S212: Slice the adjusted target NCCT image and target CTA image according to preset slice parameters to obtain a preliminary NCCT-CTA two-dimensional slice image data set.
[0108] Specifically, the present application performs a slice operation on the target NCCT image and the target CTA image with adjusted window level and window width according to preset slice thickness and slice interval parameters to obtain corresponding two-dimensional slice images.
[0109] Step S213, when it is determined that the proportion of the brain area in the corresponding two-dimensional slice image in the data set is less than a preset threshold, the two-dimensional slice image is deleted from the data set to obtain a final standardized NCCT-CTA two-dimensional slice image data set.
[0110] Specifically, in order to improve the image quality of training, when it is determined that the brain area accounts for less than 10% of the corresponding two-dimensional slice image, it is considered that the slice image contributes little to the training of the generative adversarial network to generate high-quality virtual CTA images, and may introduce unnecessary noise and interference. At this time, the two-dimensional slice image will be deleted from the dataset to streamline the dataset and reduce the computational burden during model training.
[0111] In the above embodiment, by adaptively adjusting the window level and window width of NCCT and CTA images, performing precise slicing operations according to preset parameters, and eliminating slice images with substandard brain area proportions, the standardization and quality of the data set are effectively improved, providing a more accurate and efficient data foundation for subsequent image processing and virtual CTA image generation.
[0112] In one embodiment, in step S3, during the training process, the overall network loss function is determined by the following steps:
[0113] Step S31, based on the generator in the initial generative adversarial network, multi-scale feature extraction, feature alignment and feature fusion are performed on the input original NCCT image through multi-scale processing to obtain a fused feature map.
[0114] Specifically, the present application extracts multi-level feature maps from the input original NCCT image through the multi-scale feature extraction layer of the generator. Subsequently, the number of channels and sizes of these feature maps are aligned to match the last level. Finally, all feature maps are fused by element-by-element addition to generate a fused feature map containing comprehensive information.
[0115] Step S32: input the fused feature map into a corresponding decoder for reconstruction to obtain a virtual CTA image.
[0116] Step S33, discrete cosine transform is performed on the input original NCCT image, original CTA image, and virtual CTA image according to the following formula to obtain corresponding frequency domain images:
[0117]
[0118] Where f represents the image to be processed, DCT(f) u,vIt represents the frequency domain image obtained after the discrete cosine transform operation of the image f at the frequency coordinate (u, v), α(u) and α(v) represent the preset normalization coefficients, N represents the signal sampling length, and f(x, y) represents the pixel value of the image f at the spatial coordinate (x, y).
[0119] Step S34, determining the total loss function of the generator based on the differences between the original CTA image and the virtual CTA image generated by the generator in the image domain and the frequency domain.
[0120] Specifically, the present application determines the difference in the image domain and the frequency domain between the original CTA image and the virtual CTA image generated by the generator, and comprehensively calculates the loss values of the two differences, such as by weighted sum, to obtain the total loss function of the generator. In the current embodiment, by comprehensively considering the information in the image domain and the frequency domain, the virtual CTA image generated by the generator can maintain consistency with the original CTA image in the frequency domain while maintaining the image details, thereby improving the quality and accuracy of the generated image.
[0121] Step S35, based on the discrimination results of the original NCCT image and the original CTA image, and the original NCCT image and the virtual CTA image in the image domain and the frequency domain, the image domain adversarial loss function of the discriminator and the frequency domain adversarial loss function of the discriminator are determined.
[0122] For details, please refer to Figure 3 , the role of the discriminator is to discriminate the input image pairs to distinguish whether they are from real data pairs or generated data pairs. Based on the discriminant results in the image domain and frequency domain, the image domain adversarial loss function and frequency domain adversarial loss function of the discriminator can be determined respectively. Among them, in the image domain, the discriminator will try to maximize the discrimination accuracy of the real data pairs (i.e., the original NCCT image and the original CTA image), while minimizing the discrimination accuracy of the generated data pairs (i.e., the original NCCT image and the virtual CTA image). This adversarial training process is specifically measured by defining a cross entropy loss function. In the frequency domain, the discriminator will also perform similar discrimination tasks, but the input at this time is image data processed by discrete cosine transform. In the process, the discriminator will try to distinguish the real data pairs and the generated data pairs in the frequency domain, and define a frequency domain adversarial loss function based on the discrimination results. This function can also use the cross entropy loss function, which calculates the discrimination error of the discriminator for the real data pairs and the generated data pairs in the frequency domain. Ultimately, the total loss function of the discriminator can be expressed as a weighted sum of the image domain adversarial loss function and the frequency domain adversarial loss function to guide the parameter optimization process of the discriminator.
[0123] Step S36, performing weighted summation based on the determined loss functions to obtain the overall network loss function.
[0124] Specifically, after determining the image domain loss function and frequency domain loss function of the generator, as well as the image domain adversarial loss function and frequency domain adversarial loss function of the discriminator, these loss functions need to be weighted and summed according to certain weights to obtain a unified overall network loss function.
[0125] In the above embodiment, through multi-scale feature extraction, feature alignment and fusion, and comprehensive loss optimization in the frequency domain and image domain, not only the detail fidelity and structural accuracy of the virtual CTA image are enhanced, but also the spectral consistency of the image is improved by introducing frequency domain processing, so that a more realistic and high-quality virtual CTA image can be generated.
[0126] In one embodiment, in step S31, the generator in the initial generative adversarial network performs multi-scale feature extraction, feature alignment and feature fusion on the input original NCCT image through multi-scale processing to obtain a fused feature map, including:
[0127] Step S311, based on the feature extraction layers of different scales contained in the generator, feature information of different levels is extracted from the input original NCCT image to obtain a multi-level feature map.
[0128] Specifically, the generator contains feature extraction layers of different scales, which are assumed to be E1, E2, E3, and E4. These feature extraction layers of different scales will gradually extract feature information of different levels from the input image. For details, please refer to the formulas: f1=E1(x), f2=E2(f1), f3=E3(f2), and f4=E4(f3). Among them, the formula f1=E1(x) means that the input original NCCT image x obtains the first layer feature f1 through the first layer feature extraction layer E1. Then, f1 is used as the input of the second layer feature extraction layer E2, and the second layer feature f2 is obtained after feature extraction, that is, f2=E2(f1). This process continues until it is processed by the fourth layer feature extraction layer E4, and the final feature f4 is obtained, that is, f4=E4(f3). In other words, each layer of feature extraction is performed on the basis of the previous layer of features, so the obtained features f1, f2, f3, and f4 represent image features of different scales and levels respectively.
[0129] In step S312, all feature maps except the last level are aligned according to the number of channels and size, so that these feature maps match the number of channels of the feature map of the last level and have consistent spatial resolution.
[0130] Specifically, for each feature map to be aligned, the present application performs a 3*3 convolution operation to adjust the number of channels of the feature map and obtain a feature map with the adjusted number of channels. Afterwards, in order to keep the spatial resolution of the feature map consistent with the last level, the present application further performs a global flat pooling operation on the feature map with the adjusted number of channels to adjust its size to be the same as the feature map of the last level.
[0131] In one embodiment, assuming that after multi-scale feature extraction, four levels of feature maps are obtained, namely f1, f2, f3, and f4, then the feature alignment processing method for all feature maps f1, f2, and f3 except the last level can refer to the following formula: ′ =Avgpool(Con 3×3 (f1)), f2 ′ =Avgpool(Con 3×3 (f2)), f3 ′ =Avgpool(Con 3×3 (f3)). Among them, Con 3×3 It indicates a 3*3 convolution operation on the input feature map, and Avgpool(*) indicates a global average pooling operation.
[0132] In step S313, each feature map after feature alignment is fused with the feature map of the last level by adding elements one by one to obtain a fused feature map.
[0133] Specifically, when it is determined that the aligned feature map is completely consistent with the last-level feature map in terms of the number of channels and size, the present application uses an element-by-element addition method to add each aligned feature map to the elements of the last-level feature map at the corresponding position to generate a new fused feature map. It should be noted that this fused feature map contains information from feature maps of different levels, which realizes cross-level interaction and integration during the fusion process, thereby enhancing the expressiveness and robustness of the feature map.
[0134] In the above embodiment, through multi-scale feature extraction, feature alignment and fusion, the feature information of different levels in the NCCT image is effectively integrated, the expressiveness and robustness of the feature map are enhanced, and a richer and more comprehensive feature representation is provided for subsequent image analysis.
[0135] In one embodiment, in step S34, determining the total loss function of the generator based on the difference between the original CTA image and the virtual CTA image generated by the generator in the image domain and the frequency domain includes:
[0136] Step S341: constructing an image domain similarity loss function of a generator based on the difference between the original CTA image and the virtual CTA image in the image domain.
[0137] Specifically, this application will construct the image domain similarity loss function L of the generator based on this formula mage :L mage =|CTA original -CTA virtual |, among which, CTA original Represents the original CTA image, CTA virtual Represents a virtual CTA image generated by the generator.
[0138] Step S342: constructing a frequency domain similarity loss function of a generator based on the difference between the original CTA image and the virtual CTA image in the frequency domain.
[0139] Specifically, this application will construct the frequency domain similarity loss function L of the generator based on this formula frequency : in, It represents the frequency domain image obtained after the original CTA image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the virtual CTA image is subjected to discrete cosine transform operation.
[0140] Step S343: constructing an adversarial loss function of the generator based on the difference in the discriminant between the original CTA image and the virtual CTA image in the image domain and the frequency domain.
[0141] Specifically, this application will construct the adversarial loss function L of the generator based on the following formula: gan :
[0142]
[0143] Where D represents the discriminator, L BCE Represents the binary cross entropy loss function used to measure the discrimination error of the discriminator for real data pairs and generated data pairs.
[0144] Step S344, based on the image domain similarity loss function, frequency domain similarity loss function, and adversarial loss function of the generator, a weighted sum is performed to obtain the total loss function of the generator.
[0145] Specifically, this application will perform weighted summation based on the following formula to obtain the total loss function L of the generator: generator :
[0146] L generator =αL image +βL frequency+γL gan ;
[0147] Among them, α, β and γ are all preset weight coefficients, which are set to 10, 5 and 1 respectively in this embodiment.
[0148] In one embodiment, in step S35, the image domain adversarial loss function of the discriminator is as follows:
[0150]
[0151] The frequency domain adversarial loss function of the discriminator is as follows:
[0152]
[0153] In step S36, the overall network loss function is as follows:
[0154]
[0155] Among them, NCCT original Represents the original NCCT image, CTA original Represents the original CTA image, CTA virtual represents a virtual CTA image, D represents the discriminator, and L BCE represents the binary cross entropy loss function, It represents the frequency domain image obtained after the original NCCT image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the original CTA image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the virtual CTA image is subjected to discrete cosine transform operation, L generator Represents the total loss function of the generator.
[0156] Please refer to Figure 4 The present application discloses a CT angiography intelligent imaging system based on DCT-GAN multi-scale fusion, which can be applied to the application scenario of acute ischemic stroke. The system includes a data acquisition module, a data division module, a GAN network training module and a virtual CTA image generation module, wherein:
[0157] The data acquisition module is used to acquire a preprocessed target NCCT-CTA image data set, wherein the target NCCT-CTA image data set includes a plurality of target NCCT images and target CTA images registered therewith.
[0158] The data partitioning module is used to partition the target NCCT-CTA image data set according to a preset ratio to obtain a training set and a validation set.
[0159] The GAN network training module is used to input the training set and the validation set into the initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the training set and the validation set are input into the initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the generator generates a virtual CTA image, and the discriminator discriminates the input original CTA image and the virtual CTA image in the image domain and the frequency domain processed by discrete cosine transform, and determines the overall network loss function based on the similarity difference in the image domain and the frequency domain, and the discrimination difference of the discriminator in these two domains.
[0160] The virtual CTA image generation module is used to input the preprocessed real-time NCCT image into the trained target generative adversarial network, and the generator processes it based on the DCT-GAN multi-scale fusion strategy to generate high-quality and high-resolution virtual CTA images.
[0161] In one embodiment, the above modules are also used to implement the CT angiography intelligent imaging method based on DCT-GAN multi-scale fusion as described in any of the above method embodiments, which is not limited in this application.
[0162] As can be seen from the above, the present application discloses a CT angiography intelligent imaging system based on DCT-GAN multi-scale fusion. Since the processing of discrete cosine transform in the frequency domain helps the model capture the high-frequency components in the image, these high-frequency components are crucial to the detailed expression of the image. Through the DCT-GAN multi-scale fusion strategy, the model can retain more detail information when generating virtual CTA images, thereby reducing the problem of blurred details; by utilizing the synergistic effect of the generator and the discriminator in the image domain and the frequency domain, the model can capture more detail information and further improve the accuracy of image conversion; based on the discriminator in the frequency domain, the discriminant effect can also effectively suppress artifacts in the generated image, because artifacts usually appear as high-frequency noise or abnormal structures in the image. By optimizing the loss function, the model can gradually reduce these artifacts, making the generated virtual CTA images clearer and more accurate.
[0163] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0164] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
Claims
1. A CT angiography intelligent imaging method based on DCT-GAN multi-scale fusion, characterized in that: The method comprises: S1, obtaining a preprocessed target NCCT-CTA image dataset, wherein the target NCCT-CTA image dataset includes a plurality of target NCCT images and target CTA images registered therewith; S2, dividing the target NCCT-CTA image dataset according to a preset ratio to obtain a training set and a validation set; S3, inputting the training set and the validation set into the initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the generator generates a virtual CTA image, and the discriminator discriminates the input original CTA image and the virtual CTA image in the image domain and the frequency domain processed by discrete cosine transform, and determines the overall loss function of the network based on the similarity difference in the image domain and the frequency domain, and the discrimination difference of the discriminator in the two domains; S4. The preprocessed real-time NCCT image is input into the trained target generative adversarial network, and the generator processes it based on the DCT-GAN multi-scale fusion strategy to generate high-quality and high-resolution virtual CTA images.
2. The method according to claim 1, characterized in that In step S1, each registered image pair in the target NCCT-CTA image dataset is obtained by processing the following steps: S11, acquiring an initial NCCT image and an initial CTA image registered therewith; S12, operating the initial NCCT image and the initial CTA image registered therewith according to a preprocessing process based on the adaptive image analysis technology to obtain a preprocessed target NCCT image and a registered target CTA image therewith.
3. The method according to claim 2, characterized in that In step S12, the initial NCCT image and the initial CTA image registered therewith are operated according to the preprocessing process based on the adaptive image analysis technology to obtain the preprocessed target NCCT image and the target CTA image registered therewith, including: S121, using an adaptive threshold segmentation algorithm to perform binarization processing on the initial NCCT image and the initial CTA image respectively, to obtain a binarized intermediate NCCT image and an intermediate CTA image; S122, performing pixel-by-pixel intersection operation on the intermediate NCCT image and the intermediate CTA image to obtain an initial mask image; S123, removing discrete regions in the initial mask image by maximum connected domain analysis to obtain a target mask image; S124, extracting a region of interest excluding irrelevant background and noise regions from the initial NCCT image and the initial CTA image based on the target mask image, to obtain a preprocessed target NCCT image and a target CTA image registered therewith.
4. The method according to claim 1, characterized in that: In step S2, the target NCCT-CTA image dataset is divided according to a preset ratio to obtain a training set, a validation set, and a test set, including: S21, performing image standardization processing on the target NCCT-CTA image dataset according to preset imaging characteristic difference adjustment rules, and slicing and screening standards, to obtain a standardized NCCT-CTA two-dimensional slice image dataset; S22, dividing the standardized NCCT-CTA two-dimensional slice image data set according to a preset ratio to obtain a training set, a validation set, and a test set.
5. The method according to claim 4, characterized in that In step S21, the target NCCT-CTA image dataset is subjected to image standardization processing according to the preset imaging characteristic difference adjustment rules and the slice and screening criteria to obtain a standardized NCCT-CTA two-dimensional slice image dataset, including: S211, adjusting the window level and window width of the target NCCT image and the target CTA image according to a preset imaging characteristic difference adjustment rule to obtain an adjusted target NCCT image and a target CTA image; S212, performing a slicing operation on the adjusted target NCCT image and target CTA image according to preset slicing parameters to obtain a preliminary NCCT-CTA two-dimensional slicing image data set; S213. When it is determined that the proportion of the brain area in the corresponding two-dimensional slice image in the data set is less than a preset threshold, the two-dimensional slice image is deleted from the data set to obtain a final standardized NCCT-CTA two-dimensional slice image data set.
6. The method according to claim 1, characterized in that In step S3, during the training process, the overall network loss function is determined by the following steps: S31, based on the generator in the initial generative adversarial network, perform multi-scale feature extraction, feature alignment and feature fusion on the input original NCCT image through multi-scale processing to obtain a fused feature map; S32, inputting the fused feature map into a corresponding decoder for reconstruction to obtain a virtual CTA image; S33, performing discrete cosine transform processing on the input original NCCT image, original CTA image, and virtual CTA image according to the following formula to obtain corresponding frequency domain images: Where f represents the image to be processed, DCT(f) u,v represents the frequency domain image obtained after the discrete cosine transform operation of the image f at the frequency coordinate (u, v), α(u) and α(v) represent the preset normalization coefficients, N represents the signal sampling length, and f(x, y) represents the pixel value of the image f at the spatial coordinate (x, y); S34, determining a total loss function of the generator based on differences in the image domain and the frequency domain between the original CTA image and the virtual CTA image generated by the generator; S35, determining an image domain adversarial loss function of the discriminator and a frequency domain adversarial loss function of the discriminator based on the discrimination results of the discriminator on the original NCCT image and the original CTA image, and the original NCCT image and the virtual CTA image in the image domain and the frequency domain; S36: Perform weighted summation based on the determined loss functions to obtain the overall network loss function.
7. The method according to claim 6, characterized in that In step S31, the generator based on the initial generative adversarial network performs multi-scale feature extraction, feature alignment and feature fusion on the input original NCCT image through multi-scale processing to obtain a fused feature map, including: S311, based on the feature extraction layers of different scales included in the generator, extracting feature information of different levels from the input original NCCT image to obtain a multi-level feature map; S312, aligning all feature maps except the last level according to the number of channels and sizes, so that these feature maps match the number of channels of the feature map of the last level and have consistent spatial resolution; S313, performing feature fusion on each feature map after feature alignment and the feature map of the last level by adding elements one by one to obtain a fused feature map.
8. The method according to claim 6, characterized in that In step S34, the total loss function of the generator is determined based on the difference in the image domain and the frequency domain between the original CTA image and the virtual CTA image generated by the generator, including: S341, constructing an image domain similarity loss function of a generator based on the difference between the original CTA image and the virtual CTA image in the image domain; S342, constructing a frequency domain similarity loss function of a generator based on the difference between the original CTA image and the virtual CTA image in the frequency domain; S343, constructing an adversarial loss function of the generator based on the difference in discrimination between the original CTA image and the virtual CTA image in the image domain and the frequency domain by the discriminator; S344, performing weighted summation on the image domain similarity loss function, frequency domain similarity loss function, and adversarial loss function of the generator to obtain the total loss function of the generator.
9. The method according to claim 6, characterized in that In step S35, the image domain adversarial loss function of the discriminator is as follows: The frequency domain adversarial loss function of the discriminator is as follows: In step S36, the overall network loss function is as follows: Among them, NCCT original Represents the original NCCT image, CTA original Represents the original CTA image, CTA virtual represents a virtual CTA image, D represents the discriminator, and L BCE represents the binary cross entropy loss function, It represents the frequency domain image obtained after the original NCCT image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the original CTA image is subjected to discrete cosine transform operation. It represents the frequency domain image obtained after the virtual CTA image is subjected to discrete cosine transform operation, L generator Represents the total loss function of the generator.
10. A CT angiography intelligent imaging system based on DCT-GAN multi-scale fusion, characterized in that: The system includes a data acquisition module, a data partitioning module, a GAN network training module and a virtual CTA image generation module, wherein: The data acquisition module is used to acquire a preprocessed target NCCT-CTA image data set, wherein the target NCCT-CTA image data set includes a plurality of target NCCT images and target CTA images registered therewith; The data partitioning module is used to partition the target NCCT-CTA image data set according to a preset ratio to obtain a training set and a validation set; The GAN network training module is used to input the training set and the validation set into an initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the training set and the validation set are input into an initial generative adversarial network based on DCT-GAN multi-scale fusion for model training. During the training process, the generator generates a virtual CTA image, and the discriminator discriminates the input original CTA image and the virtual CTA image in the image domain and the frequency domain processed by discrete cosine transform, and determines the overall network loss function based on the similarity difference in the image domain and the frequency domain, and the discrimination difference of the discriminator in the two domains; The virtual CTA image generation module is used to input the preprocessed real-time NCCT image into the trained target generative adversarial network, and the generator processes it based on the DCT-GAN multi-scale fusion strategy to generate high-quality and high-resolution virtual CTA images.
Citation Information
Cited By
Automatic blood vessel extraction method and system based on CT (Computed Tomography) image
CN120746986A
Automatic Vascular Extraction Method and System Based on CT Images
CN120746986B