Low-dose CT imaging method and device based on semantic priori knowledge guidance

By introducing semantic prior knowledge and joint network structures in low-dose CT imaging, the difficulties of denoising and detail retention in the prior art are solved, and more accurate noise processing and anatomical structure retention are achieved, image quality and diagnostic accuracy are improved.

CN119919519AActive Publication Date: 2025-05-02WUHAN UNIV

Patent Information

Application Number
CN202411990690.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing low-dose CT imaging methods are difficult to take into account both noise removal and image detail retention during the denoising process, and ignore the differences in different anatomical areas, resulting in blurring of details in areas with lower noise levels.

Method used

Using a low-dose CT imaging method based on semantic prior knowledge, semantic features are extracted through pre-trained visual models, joint network structures are designed, including denoising networks and semantic segmentation networks, and detailed features and semantic features are fused using semantic perception modules, and joint optimization of noise reduction and semantic segmentation is achieved through multi-task learning strategies.

Benefits of technology

Effectively distinguish different tissue types, capture specific noise patterns in different regions, improve the perception of noise levels in different anatomical structures by the denoising network, preserve key anatomical details, and improve image quality and diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919519A_ABST
    Figure CN119919519A_ABST
Patent Text Reader

Abstract

The invention discloses a low-dose CT (Computed Tomography) imaging method and equipment based on semantic priori knowledge guidance. In order to solve the problems of excessive smoothness of images, detail loss and the like in low-dose CT imaging, semantic prior information is extracted through a pre-trained visual model by utilizing significant differences of different anatomical structures (such as bones, soft tissues, gas regions and the like) on noise levels, so that different tissue types are effectively distinguished, and noise modes of specific regions are captured. Meanwhile, a semantic perception module is provided, detail features concerned by the noise reduction network and high-level anatomical semantic features captured by the semantic network are aligned and fused, and it is ensured that the anatomical structure is perceived more accurately in the denoising process. Further, a segmentation network is introduced behind the noise reduction network, and a multi-task learning strategy is adopted to realize joint optimization of noise reduction and semantic segmentation. According to the design, the denoising performance of the image is improved, key anatomical details are effectively reserved, and clearer image support is provided for clinical diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the intersection of medical image processing and computer vision, and specifically refers to a low-dose CT imaging method and device guided by semantic prior knowledge. Background Art

[0002] CT imaging plays an irreplaceable role in medical diagnosis, providing accurate anatomical and pathological information for tumor screening, cardiovascular disease, and lung disease. However, the radiation dose generated by CT scanning has long raised concerns about patient safety, especially when multiple scans are required, and the cumulative effect of radiation may increase the risk of cancer. In order to reduce radiation hazards, low-dose CT imaging is gradually adopted in clinical practice, for example, by reducing tube current and tube voltage, but this method inevitably increases image noise, which in turn affects the accuracy of diagnosis. Although existing denoising methods can improve image quality to a certain extent, they often cannot balance noise removal and image detail preservation, and they strongly rely on manual regularization and hyperparameters. The direct mapping relationship between normal-dose CT (NDCT) and low-dose CT (LDCT) does not follow clear rules, such as sparse sampling leading to stripe artifacts and low tube current leading to mixed Gaussian noise and Poisson noise.

[0003] In recent years, low-dose CT image post-processing algorithms based on deep learning have been widely studied and have made significant progress. The earliest studies mostly optimized the encoder-decoder network by minimizing the pixel-level loss between the denoised image and the normal-dose CT image. Typical models include residual encoder-decoder convolutional neural network (RED-CNN). Although these methods have outstanding denoising performance, they are prone to produce over-smoothed images due to the use of pixel-level loss functions (such as L1 or L2 loss), resulting in the loss of some details. To address this problem, researchers introduced generative adversarial networks (GANs) as well as structural similarity loss and perceptual loss to generate more realistic denoised images. Typical methods include the WGAN-VGG model. However, due to the characteristics of adversarial training, GANs usually require careful design of the optimization process and network architecture to ensure the convergence and stability of the model. In addition, some model-based methods attempt to combine iterative reconstruction methods and neural networks to improve the interpretation of deep learning models.

[0004] Although deep learning-based methods have greatly improved the image quality of low-dose CT, most of these methods ignore the differences between different anatomical regions, resulting in blurred details in areas with low noise levels. In addition, previous studies have shown that the noise level in CT images varies significantly depending on the tissue type. Therefore, in recent years, some studies have begun to explore the use of anatomical semantics to guide CT image denoising. For example, by using the absorption of X-rays in different tissues to generate anatomical feature domains (such as subcutaneous fat, muscle, and visceral fat), and combining semantic features and image features in the semantic fusion module, a structural semantic loss function is introduced to measure segmentation differences. However, the semantics obtained by these methods are too coarse, resulting in unsatisfactory denoising effects. Summary of the invention

[0005] In view of the deficiencies of the prior art, the present invention provides a low-dose CT imaging method and device based on prior knowledge guidance. It aims to introduce accurate semantic priors through a large visual model to distinguish different anatomical structures, so that the denoising network can implicitly distinguish the noise levels of different regions. At the same time, a semantic perception module is proposed to align and fuse the detail features that the denoising network focuses on with the high-level anatomical semantic features captured by the semantic network, ensuring more accurate perception of the anatomical structure during the denoising process. Finally, a segmentation network is introduced after the denoising network, and a multi-task learning strategy is adopted to achieve joint optimization of denoising and semantic segmentation, thereby obtaining better denoising effects.

[0006] The low-dose CT imaging method designed by the present invention based on the guidance of semantic prior knowledge comprises the following steps:

[0007] The semantic features of low-dose CT images are obtained using large models as prior knowledge of anatomical structures;

[0008] Designing a joint network structure, including a denoising network and a semantic segmentation network, wherein the denoising network is used to reduce noise in the CT image, and the semantic segmentation network uses the extracted semantic features to identify the anatomical structure;

[0009] The mean square error is used as the loss in the denoising process, and the Dice coefficient and cross entropy are used as the loss in the segmentation process;

[0010] Using the above loss, the denoising network and the segmentation network are trained using an alternating optimization method;

[0011] The trained joint network structure is used for low-dose CT imaging.

[0012] Furthermore, the visual branch in the pre-trained CLIP-Driven general model is used to extract semantic features.

[0013] Furthermore, the training dataset uses two public low-dose CT datasets released by the "NIHAAPM-Mayo Clinic Low-Dose CT GrandChallenge" in 2016 and 2020. Datasets can also be constructed for network training based on actual conditions.

[0014] Furthermore, the denoising network includes a shallow feature extraction layer, a deep feature extraction layer, and an output layer in sequence; the shallow feature extraction layer and the output layer are both composed of convolutions; the deep feature extraction layer is composed of three semantic perception modules, and the semantic perception module is composed of cross attention.

[0015] Preferably, the denoising process is as follows:

[0016] Shallow feature extraction is expressed as:

[0017] f l =conv(X)

[0018] Where conv represents the convolutional layer, X represents the low-dose CT image,

[0019] The semantic features are first used as query keys for the cross-attention module through the self-attention module, and the formula is expressed as:

[0020] Q=LN(SA(LN(GN(S))))

[0021] Among them, LN, GN, SA, and S represent layer normalization, group normalization, self-attention module, and semantic features. Then the low-dose CT features are convolved as K and V values. Finally, the semantic features transformed by cross attention and the original features f are concatenated as the final features:

[0022]

[0023] where f s and f represent the image features after fusion of semantic features and the output of the previous layer, respectively. k and GELU represent the scale factor and Gaussian error linear unit respectively, and concat represents the concatenation operation;

[0024] Semantic features can be expressed as:

[0025] S = conv(GELU(conv(Y s )))

[0026] where Y s represents the semantic labels extracted by the pre-trained visual model on NDCT,

[0027] The output layer can be expressed as:

[0028]

[0029] in and f d They represent the denoised low-dose CT and the output of the deep network feature extraction layer, respectively. Skip connections are used to accelerate the learning process of the network.

[0030] Furthermore, the segmentation network includes an encoder-decoder structure, both the encoder and the decoder adopt residual modules, and the encoder and the decoder use skip connections to retain low-level features of the image.

[0031] Furthermore, in each training cycle, the alternating optimization training method first uses the output of the pre-trained large model as prior knowledge, and fuses it into the denoising network through the semantic perception module to guide the optimization of the parameters of the denoising network; then uses the denoised image as input and the semantic prior knowledge as the prediction target to optimize the parameters of the segmentation network.

[0032] Furthermore, after the joint network structure training is completed, an independent validation dataset is used for performance evaluation, and the denoising performance is quantitatively evaluated using image quality evaluation indicators, including peak signal-to-noise ratio, structural similarity index, and root mean square error, to select the best network parameters for testing.

[0033] Furthermore, the training process of the alternating optimization is specifically as follows:

[0034] First, pairs of low-dose CT images and normal-dose CT images are selected from the dataset. Then, the normal-dose CT images are output to the pre-trained visual model to obtain semantic labels. The semantic labels and low-dose CT images are output to the denoising network together to obtain denoised low-dose CT images. The loss function is calculated according to the following formula:

[0035] Where Y represents the standard dose CT image, represents the denoised low-dose CT image; then back-propagation is performed to optimize the parameters of the denoising network, and then the denoised normal-dose CT image is output to the segmentation network to obtain the semantic label after segmentation. The loss function is calculated according to the following formula

[0036] L s =L Dice +αL CE

[0037] where p CE Cross entropy objective function, L Dice represents the Dice coefficient, and α represents the weight factor for balancing different loss functions; then back propagation is performed to optimize the parameters of the denoising network and the segmentation network; this cycle is repeated until the loss function converges.

[0038] Based on the same inventive concept, the present invention also discloses an electronic device, comprising:

[0039] one or more processors;

[0040] A storage device for storing one or more programs;

[0041] When one or more programs are executed by the one or more processors, the one or more processors implement a low-dose CT imaging method guided by semantic prior knowledge.

[0042] Based on the same inventive concept, the present invention also discloses a computer-readable medium on which a computer program is stored. When the program is executed by a processor, a low-dose CT imaging method guided by semantic prior knowledge is implemented.

[0043] The present invention is a little bit in that:

[0044] 1. Exploiting the significant differences in noise levels among different anatomical structures (such as bones, soft tissues, gas areas, etc.), semantic prior information is extracted through pre-trained visual models. These features can effectively distinguish different tissue types and capture specific noise patterns in different regions.

[0045] 2. A semantic perception module is proposed, which aims to align denoising features with semantic features. In the denoising process, the denoising network focuses on the detail information of the image to remove noise and preserve fine structures, while the semantic network focuses on capturing high-level anatomical semantic information. This module effectively integrates these two features to organically align detail features with semantic features, ensuring that the denoising network can better perceive the anatomical structure while denoising.

[0046] 3. A segmentation network is introduced after the denoising network, and a multi-task learning strategy is adopted to achieve joint optimization of denoising and semantic segmentation. This design not only improves the denoising network's perception of noise levels in different anatomical structures, but also better preserves key anatomical details, thereby providing clearer and more reliable image support for clinical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of the overall network structure of an embodiment of the present invention.

[0048] Figure 2 is a schematic diagram of a denoising network according to an embodiment of the present invention.

[0049] Figure 3 Schematic diagram of a segmentation network according to an embodiment of the present invention.

[0050] Figure 4 Schematic diagram of a semantic perception module according to an embodiment of the present invention.

[0051] Figure 5 It is an overall flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present invention will be described clearly and completely below in combination with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] Embodiment 1

[0054] This embodiment is mainly based on the difference characteristics of noise levels of different anatomical structures, introduces semantic prior information through a semantic segmentation model, and proposes a CT imaging method and device based on semantic prior guidance. The differences of different anatomical structures are fully considered, and the denoising process is supervised by a segmentation network. The results obtained by the present invention are more scientific and more accurate.

[0055] Step 1: Data preparation: Collect a large number of low-dose CT (LDCT) images and their corresponding normal-dose CT (NDCT) images to form a training data set. The implementation process is as follows:

[0056] Two public low-dose CT datasets released by the "NIHAAPM-Mayo Clinic Low-Dose CT Grand Challenge" in 2016 and 2020 were used. The Mayo 2016 dataset includes normal-dose abdominal CT images of 10 anonymous patients and corresponding simulated quarter-dose CT images. In this embodiment, images of 8 patients are selected as training sets, images of 1 patient are selected as validation sets, and images of 1 patient are selected as test sets. The Mayo 2020 dataset provides abdominal CT image data of 100 patients. In this embodiment, 20 patients are randomly selected for experiments, of which images of 16 patients are used as training sets, images of 2 patients are used as validation sets, and images of 2 patients are used as test sets. In order to facilitate subsequent processing, the data is preprocessed, the original intensity values ​​are truncated within the range of 2% to 98% of the initial Hounsfield unit (HU), and each original CT case is normalized to have a mean of zero and a variance of one. Those skilled in the art can construct a dataset according to actual conditions. Generally speaking, the more data samples, the more diseases and anatomical structures covered, and the stronger the generalization ability of the model.

[0057] Step 2: Semantic extraction: Use the pre-trained visual model to forward propagate the pre-processed NDCT image to extract the semantics of the anatomical structure. The implementation process is described as follows:

[0058] The visual branch in the pre-trained CLIP-Driven general model is used as the pre-trained visual model, which is abbreviated as "CLIP Vision" below. Specifically, "CLIP-Vision" consists of a visual encoder and a text-driven segmenter. The visual encoder uses Swin UNETR. Swin UNETR is a "U-shaped" network. The encoder part uses SwinTransformer, which directly uses 3D blocks and connects to the CNN-based decoder through skip connections of different resolutions. The pre-processed NDCT image first passes through the four stages of the encoder, each stage contains two transformer blocks, and the block merging layer is used between each stage to reduce the resolution of the feature by half. Then, after the four stages of the decoder, the representation extracted at each stage is input into the residual block, and then the processed features of each stage are upsampled through the deconvolution layer and connected with the features output by the residual block of the previous stage. Finally, the semantic label of NDCT is obtained through the text classifier. The text-driven segmenter consists of three consecutive convolutional layers, and finally outputs the probability of foreground and background of each category. The pre-trained visual model is denoted as g, and the semantic features extracted by inputting the NDCT image Y into g are denoted as Y s The formula is:

[0059] Y s =g(Y)

[0060] Those skilled in the art can select an appropriate pre-trained visual model based on actual results.

[0061] Step 3: Multi-task learning model construction: Design a joint network structure, including a denoising network and a segmentation network. The denoising network aims to reduce the noise in the image, while the semantic segmentation network measures the semantic distance between the predicted normal dose CT and the standard dose CT. Inspired by the pre-trained visual model, this scheme uses the standard dose CT as the input of the pre-trained visual model, and its segmentation result is used as the semantic label of the standard dose CT. The semantic segmentation network adopts an encoder-decoder structure. Both the encoder and the decoder use residual modules and use jump connections to retain the low-level features of the image. Downsampling is achieved by convolution with a stride of 2, and upsampling is achieved by deconvolution operations. The denoising network also adopts an encoder-decoder structure. Different from the encoder of the semantic segmentation network, the encoder of the denoising network uses a semantic perception module to fuse the extracted semantic priors with image features to distinguish different regions. In addition, a semantic loss function is introduced to measure the semantic distance between the standard dose CT and the predicted standard dose CT. This scheme can adaptively adjust the denoising strategy for different anatomical structures, better preserve the local details of the image, especially in low-noise areas, ensure that important anatomical structure information is not lost, and thus improve the diagnostic value of CT images.

[0062] The denoising network consists of a shallow feature extraction layer, a deep feature extraction layer, and an output layer. The shallow feature extraction layer and the output layer are both composed of a 3×3 convolution, which is used to expand or squeeze the channel to adapt to the input image size. The deep feature extraction layer consists of 3 semantic perception modules, which are composed of cross attention, allowing the denoising network to focus on semantic alignment.

[0063] Shallow feature extraction can represent

[0064] f l =conv(X)

[0065] Where conv represents the convolutional layer and X represents the LDCT image.

[0066] The semantic features are first used as query keys for the cross-attention module through the self-attention module, and the formula is expressed as:

[0067] Q=LN(SA(LN(GN(S))))

[0068] Among them, LN, GN, SA, and S represent layer normalization, group normalization, self-attention module, and semantic features. Then the LDCT features are convolved as K and V values. Finally, the semantic features transformed by cross attention and the original features f are concatenated as the final features.

[0069]

[0070] where f s and f represent the image features after fusion of semantic features and the output of the previous layer respectively. k and GELU represent the scaling factor and Gaussian error linear unit respectively. concat represents the concatenation operation.

[0071] The semantic features can be expressed as

[0072] S = conv(GELU(conv(Y s )))

[0073] where Y s Represents the semantic labels extracted by the pre-trained vision model on NDCT.

[0074] The output layer can be expressed as

[0075]

[0076] in and f d They represent the denoised LDCT and the output of the deep network feature extraction layer, respectively. Skip connections are used to accelerate the learning process of the network.

[0077] The segmentation network structure is mainly divided into encoder, bottleneck layer and decoder. The encoder consists of multiple convolutional layers and pooling layers. Each convolutional layer usually contains two 3×3 convolution operations and ReLU activation functions to extract image features and maintain spatial structure. The pooling layer reduces the spatial resolution of the feature map through a 2×2 maximum pooling operation, increases the receptive field, and captures a wider range of contextual information. Between the encoder and the decoder, the bottleneck layer further extracts deep features through a convolution block. The upsampling layer in the decoder uses deconvolution or bilinear interpolation to increase the spatial resolution of the feature map, and concatenates the corresponding encoder feature map with the decoder feature map through jump connections to retain detail information and enhance segmentation accuracy. Finally, the output of the decoder passes through a 1×1 convolution layer to map the feature map to the required number of output channels to generate the final segmentation map. The formula is as follows

[0078]

[0079] Where Seg represents the segmentation network.

[0080] The number of features in the convolutional layer is set to 64. The number of attention features in the semantic perception module is 64. Those skilled in the art can select these parameters according to actual effects.

[0081] Step 4: Loss function design: Design a comprehensive loss function, including denoising loss and segmentation loss, to ensure that image details and anatomical structure information are preserved as much as possible during the denoising process. The implementation process is described as follows:

[0082] The objective function of the denoising network uses mean square error, which is expressed as:

[0083]

[0084] The objective function of the segmentation network uses the Dice coefficient and cross entropy. For multi-category segmentation, the Dice coefficient formula is expressed as:

[0085]

[0086] The cross entropy objective function using focal loss can be expressed as:

[0087]

[0088] The overall objective function of the segmentation model is:

[0089] L s =L Dice +αL CE

[0090] Where C represents the number of categories, γ is an adjustable parameter, and α represents the weight factor for balancing different loss functions. In the experiment, C=32, γ=2, and α=0.5 are set.

[0091] Those skilled in the art can select these parameters according to actual effects.

[0092] Step 5: Alternating optimization strategy: Use alternating optimization to train the denoising network and the segmentation network. In each training cycle, first use the pre-trained visual model output as prior information and fuse it into the denoising network through the semantic perception module to guide the optimization of the parameters of the denoising network; then use the denoised image as input and the semantic prior knowledge as the prediction target to optimize the parameters of the segmentation network. In this way, the two networks can promote each other, thereby improving the overall performance. The implementation process is described as follows:

[0093] First, select paired LDCT and NDCT data from the dataset, then output NDCT to the pre-trained visual model to obtain semantic labels, output the semantic labels and LDCT together to the denoising network to obtain denoised LDCT, calculate the loss function according to the total loss of the segmentation model, and perform back propagation to optimize the parameters of the denoising network. Then output the denoised LDCT to the segmentation network to obtain the semantic labels after segmentation, calculate the loss function according to the total loss of the segmentation model, and perform back propagation to optimize the parameters of the denoising network and the segmentation network. Repeat this cycle until the loss function converges. The training hyperparameters are as follows: In this embodiment, the AdamW optimizer with a mini-batch size of 4 is used to train the network and β 1 , β 2 , ε are set to 0.9, 0.99, and 1e-8. The input image size is set to 512×512, and the network is trained for 100 epochs. The initial learning rate of the denoising network and the segmentation network is 1e-4, and gradually decays to 1e-6 using a cosine annealing strategy. This method is implemented based on PyTorch and trained and tested on an NVIDIA Tesla V100 GPU. Since the segmentation network only assists the denoising network in learning during training, the denoising network does not introduce additional computational costs during the testing phase.

[0094] Step 6: Model evaluation: After training, use an independent validation set to evaluate the model performance, including the denoising effect. By comparing with traditional denoising methods, the advantages of the present invention are verified. The implementation process is described as follows:

[0095] After the model is trained, the best model parameters are selected for testing according to the validation set divided in step 1 and the evaluation index. The performance of the present invention is observed on the test set.

[0096] The present invention uses three commonly used objective image quality evaluation indicators to quantitatively evaluate the denoising performance: Peak signal-to-noise ratio (PSNR), structural similarity index and root mean square error (RMSE). Among them, the higher the PSNR and SSIM, the lower the RMSE, the better the performance. The definitions of PSNR, SSIM and RMSE are as follows:

[0097]

[0098] Embodiment 2

[0099] Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; when one or more programs are executed by the one or more processors, the one or more processors implement the method described in Example 1.

[0100] Since the device introduced in the second embodiment of the present invention is an electronic device used to implement the low-dose CT imaging method based on semantic prior knowledge guidance in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, a person skilled in the art can understand the specific structure and deformation of the electronic device, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.

[0101] Embodiment 3

[0102] Based on the same inventive concept, the present invention further provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processor, the method described in the first embodiment is implemented.

[0103] Since the device introduced in the third embodiment of the present invention is a computer-readable medium used to implement the low-dose CT imaging method based on semantic prior knowledge guidance in the first embodiment of the present invention, a person skilled in the art can understand the specific structure and deformation of the electronic device based on the method introduced in the first embodiment of the present invention, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.

[0104] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A low-dose CT imaging method guided by semantic prior knowledge, characterized in that: The following steps are involved: The semantic features of low-dose CT images are obtained using large models as prior knowledge of anatomical structures; Designing a joint network structure, including a denoising network and a semantic segmentation network, wherein the denoising network is used to reduce noise in the CT image, and the semantic segmentation network uses the extracted semantic features to identify the anatomical structure; The mean square error is used as the loss in the denoising process, and the Dice coefficient and cross entropy are used as the loss in the segmentation process; Using the above loss, the denoising network and the segmentation network are trained using an alternating optimization method; The trained joint network structure is used for low-dose CT imaging.

2. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 1, characterized in that: The vision branch in the pre-trained CLIP-Driven general model is used to extract semantic features.

3. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 1, characterized in that: The denoising network includes a shallow feature extraction layer, a deep feature extraction layer, and an output layer in sequence; the shallow feature extraction layer and the output layer are both composed of convolutions; the deep feature extraction layer is composed of three semantic perception modules, and the semantic perception module is composed of cross attention.

4. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 3, characterized in that: The denoising process is as follows: Shallow feature extraction is expressed as: f l =conv(X) Where conv represents the convolutional layer, X represents the low-dose CT image, The semantic features are first used as query keys for the cross-attention module through the self-attention module, and the formula is expressed as: Q=LN(SA(LN(GN(S)))) Among them, LN, GN, SA, and S represent layer normalization, group normalization, self-attention module, and semantic features. Then the low-dose CT features are convolved as K and V values. Finally, the semantic features transformed by cross attention and the original features f are concatenated as the final features: where f s and f represent the image features after fusion of semantic features and the output of the previous layer, respectively. k and GELU represent the scale factor and Gaussian error linear unit respectively, and concat represents the concatenation operation; Semantic features can be expressed as: S=conv(GELU(conv(Y s ))) where Y s represents the semantic labels extracted by the pre-trained visual model on NDCT, The output layer can be expressed as: in and f d They represent the denoised low-dose CT and the output of the deep network feature extraction layer, respectively. Skip connections are used to accelerate the learning process of the network.

5. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 1, characterized in that: The segmentation network includes an encoder-decoder structure, both the encoder and the decoder adopt residual modules, and the encoder and the decoder use skip connections.

6. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 1, characterized in that: In each training cycle, the alternating optimization training method first uses the output of the pre-trained large model as prior knowledge, and integrates it into the denoising network through the semantic perception module to guide the optimization of the parameters of the denoising network; then uses the denoised image as input and the semantic prior knowledge as the prediction target to optimize the parameters of the segmentation network.

7. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 6, characterized in that: The training process of the alternating optimization is specifically as follows: First, pairs of low-dose CT images and normal-dose CT images are selected from the dataset. Then, the normal-dose CT images are output to the pre-trained visual model to obtain semantic labels. The semantic labels and low-dose CT images are output to the denoising network together to obtain denoised low-dose CT images. The loss function is calculated according to the following formula: Where Y represents the standard dose CT image, represents the denoised low-dose CT image; then back propagation is performed to optimize the parameters of the denoising network, and then the denoised normal-dose CT image is output to the segmentation network to obtain the semantic label after segmentation, and the loss function is calculated according to the following formula L s =L Dice +αL CE Where L CE Cross entropy objective function, L Dice represents the Dice coefficient, and α represents the weight factor for balancing different loss functions; then back propagation is performed to optimize the parameters of the denoising network and the segmentation network; this cycle is repeated until the loss function converges.

8. The low-dose CT imaging method based on semantic prior knowledge guidance according to claim 1, characterized in that: After the joint network structure training is completed, an independent validation dataset is used for performance evaluation, and image quality evaluation indicators are used to quantitatively evaluate the denoising performance, including peak signal-to-noise ratio, structural similarity index, and root mean square error.

9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Improved RegGAN low-dose CT image denoising method and related device

    CN115760651A

  • Noise removal method for low-dose X-ray CT image

    CN116630195A

  • Low-dose CT image denoising method based on self-supervised perception loss multi-scale convolutional neural network

    CN116645283A

  • Multimodal medical image segmentation method and system based on knowledge depolarization

    CN118072014A

  • Image segmentation method combined with self-supervised noise removal and model training method

    CN118429634A

Cited By

  • CT image noise reduction processing method and system based on deep neural network

    CN121304474A