Low-dose CT image denoising method and device based on self-supervised learning and medium
Through self-supervised learning methods, learning the denoising mode from a single low-dose CT image is solved, and the denoising effect without high-dose CT images is achieved, improving image quality and diagnostic accuracy.
Patent Information
- Application Number
- CN202510276314.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Due to noise interference, low-dose CT images have low image quality, which affects the accuracy of clinical diagnosis. The existing supervised learning methods rely on high-dose CT images as training data, making it difficult to effectively apply in clinical practice.
Using a self-supervised learning-based approach, denoising images are generated by learning the denoising mode from a single low-dose CT image, using pixel shuffling downsampling and denoising network training, without the need for real high-dose CT images as training data.
It realizes effective denoising of low-dose CT images without relying on high-dose CT images, improving image quality, and is suitable for scenes where data is scarce in medium and high-dose CT images in clinical practice, improving diagnosis accuracy.
Smart Images

Figure CN120219215A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a low-dose CT image denoising method, device and medium based on self-supervised learning. Background Art
[0002] Computed tomography (CT) is an advanced medical imaging technology that has been widely used in clinical diagnosis and disease monitoring. By the interaction of X-rays with different tissues and structures, CT can obtain tissue density information and generate accurate cross-sectional images, showing the detailed structures of various organs inside the human body. However, the use of high-dose X-rays may cause potential damage to the human body and even increase the risk of developing malignant tumors, and the increase in radiation dose is proportional to the patient's radiation exposure. Therefore, while improving image quality, reducing radiation risk has become an important goal in research and clinical applications. For this reason, low-dose computed tomography (LDCT) technology has emerged, aiming to reduce the radiation exposure of patients. However, due to the use of a lower X-ray dose in LDCT, the image quality is often affected by noise interference, thus affecting the accuracy of clinical diagnosis. Therefore, how to effectively reduce noise and improve image quality has become an important research topic in the current medical imaging field.
[0003] With the development of deep learning technology, more and more methods adopt deep learning architectures such as convolutional neural networks, generative adversarial networks, and autoencoders to learn the features of LDCT images and perform denoising. Through large-scale training data and deep network structures, deep learning methods can more effectively capture complex features in images, thereby achieving efficient denoising effects and restoring high-quality CT images, making them more suitable for medical diagnosis. The current image denoising field mainly trains on pairs of noisy input images and clean target images. Clinically, although high-dose and low-dose CT images of the same patient can be obtained, due to the actual clinical environment, there are still significant structural differences in the CT images of the same patient and the same part. Therefore, it is difficult to obtain the HDCT (high-dose CT) images corresponding to LDCT images, which greatly limits the application of supervised learning methods in clinical practice. Self-supervised learning provides a new method for low-dose CT image denoising and also makes it possible to perform denoising using only LDCT without HDCT as supervision information. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention provides a low-dose CT image denoising method, device and medium based on self-supervised learning, which uses a self-supervised learning framework to complete denoising only by processing a single low-dose CT image. By learning the denoising pattern from the noisy LDCT image itself, without the need for real groundtruth, the limitations of traditional methods that rely on real high-dose CT images as training data are avoided.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A low-dose computed tomography (CT) image denoising method based on self-supervised learning, comprising the following steps:
[0007] S1. Through pixel shuffle downsampling operation, the correlated pixel points in the LDCT image are assigned to multiple different sub-images with two different strides, generating two groups of sub-images;
[0008] S2. Generate a neighbor sub-image for each sub-image, forming a pair of neighbor sub-images, obtaining a first set of sub-image pairs and a second set of sub-image pairs;
[0009] S3. Use the first set of sub-image pairs and the second set of sub-image pairs to train the denoising network respectively. One of each pair of neighbor sub-images is used as the input of the denoising network, and the other is used as the prediction target; establish a loss function between the denoised image generated by the denoising network and the prediction target, and train the denoising network with the goal of minimizing the loss function to obtain two denoising networks with different parameters;
[0010] S4. Through pixel shuffle downsampling operation, the correlated pixel points in the LDCT image to be denoised are assigned to multiple different sub-images with two different strides, generating two groups of sub-images, and inputting the two groups of sub-images into the two trained denoising networks with different parameters to obtain two groups of denoised sub-images;
[0011] S5. Perform pixel shuffle upsampling operations on the two groups of denoised sub-images respectively, and recombine them into two denoised images with the same size as the original LDCT image;
[0012] S6. Perform linear weighted combination on the two denoised images to obtain the final reconstructed LDCT image.
[0013] To optimize the above technical solutions, the specific measures taken further include:
[0014] Further, S1 is specifically:
[0015] Divide the LDCT image into several first cells of r×r according to the first stride r. There are r×r pixel points in each first cell. Take one pixel point from each first cell and assign it to a sub-image until all the pixel points in the first cell are assigned, obtaining r 2 different first sub-images, defined as the first group of sub-images wherein, g 1i (y) represents the i-th first sub-image; y represents the LDCT image;
[0016] Divide the LDCT image into several second cells of s×s according to the second step length s. There are s×s pixel points in each second cell. Take one pixel point from each second cell and assign it to a sub-image until all pixel points in the second cell are assigned, obtaining s 2 different second sub-images, defined as the second group of sub-images where h 1j (y) represents the j-th second sub-image.
[0017] Furthermore, in S2, specifically generating a neighbor sub-image for each sub-image is as follows:
[0018] For the first group of sub-images, the pixel points in the i-th first sub-image come from the first cell. Randomly select pixel points from the pixel points in the first cell that have not been assigned to the i-th first sub-image. Randomly select one pixel point from each first cell to form a neighbor sub-image corresponding to the i-th first sub-image; the neighbor sub-images of all first sub-images are defined as the first group of neighbor sub-images In the formula, g 2i (y) represents the neighbor sub-image corresponding to the i-th first sub-image;
[0019] For the second group of sub-images, the pixel points in the j-th second sub-image come from the second cell. Randomly select pixel points from the pixel points in the second cell that have not been assigned to the j-th second sub-image. Randomly select one pixel point from each second cell to form a neighbor sub-image corresponding to the j-th second sub-image. The neighbor sub-images of all second sub-images are defined as the second group of neighbor sub-images In the formula, h 2j (y) represents the neighbor sub-image corresponding to the j-th second sub-image.
[0020] Furthermore, in S3, the specific process of establishing a loss function between the denoised image generated by the denoising network and the prediction target and training the denoising network with the goal of minimizing the loss function is as follows:
[0021] The first group of sub-images g1(y) is fed into the first denoising network f θ1 to generate the first denoised image f θ1 (g1(y)). The first group of neighbor sub-images g2(y) is used as the prediction target. The loss function of the first denoising network is as follows:
[0022]
[0023] In the formula, E 1 represents the optimization result of the first denoising network on multiple input samples, θ1 represents the parameters of the first denoising network; training the first denoising network with the goal of minimizing the loss function of the first denoising network; y represents the LDCT image;
[0024] The second group of sub - graphs h1(y) is fed into the second denoising network f θ2 Generate the second denoised image f θ2 (h1(y)), the second group of neighboring sub - graphs h2(y) is used as the prediction target, and the loss function of the second denoising network is as follows:
[0025]
[0026] In the formula, E 2 represents the optimization result of the second denoising network on multiple input samples, θ2 represents the parameters of the second denoising network; the second denoising network is trained with the goal of minimizing the loss function of the second denoising network.
[0027] Furthermore, in S6, the specific operation of linearly weighted combination of the two denoised images to obtain the final reconstructed LDCT image is as follows:
[0028] F output = a·F1+(1 - a)·F2
[0029] In the formula, F1 represents the denoised image output by the first denoising network, F2 represents the denoised image output by the second denoising network, a represents the weight, and F output represents the final reconstructed LDCT image.
[0030] The present invention also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for denoising low - dose computed tomography images based on self - supervised learning as described above is implemented.
[0031] The present invention also proposes a computer - readable storage medium storing a computer program, and the computer program enables a computer to execute the method for denoising low - dose computed tomography images based on self - supervised learning as described above.
[0032] The beneficial effects of the present invention are as follows:
[0033] The present invention does not rely on high - dose CT images as a reference and can complete denoising by only processing a single low - dose CT image. The present invention is particularly suitable for scenarios where high - dose CT image data is scarce in actual clinical practice, and at the same time improves the usability of low - dose CT images and the accuracy of clinical diagnosis. Brief Description of the Drawings
[0034] Figure 1 is a flow chart of the method for denoising low - dose computed tomography images based on self - supervised learning proposed by the present invention.
[0035] Figure 2 is a schematic diagram for generating sub - graphs and neighboring sub - graphs.
[0036] Figure 3 Schematic diagram of the denoising principle of the denoising network.
[0037] Figure 4 Schematic diagram of the LDCT image to be denoised in the embodiment.
[0038] Figure 5 Schematic diagram of the denoised LDCT image output in the embodiment. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0040] Embodiment 1
[0041] The present invention proposes a low-dose computed tomography image denoising method based on self-supervised learning. The flowchart of the method is as Figure 1 shown, and includes the following steps:
[0042] S1. Through the Pixel Shuffling Downsampling (PD) operation, the correlated pixel points in the LDCT (low-dose computed tomography) image are assigned to multiple different sub-images with two different step sizes, reducing the correlation of noise between pixels and generating two groups of sub-images; thereby reducing the dependence of noise in the image and facilitating subsequent denoising operations. S1 is specifically as follows:
[0043] Divide the LDCT image into several first cells of r×r according to the first step size r. There are r×r pixel points in each first cell. Take one pixel point from each first cell and assign it to a sub-image until all the pixel points in the first cell are assigned, obtaining r 2 different first sub-images, defined as the first group of sub-images wherein, g 1i (y) represents the i-th first sub-image; y represents the LDCT image; each LDCT image is converted to grayscale mode, resized to 512*512, and normalized. Ensure that the key features of the original image can be retained as much as possible during the denoising process, and prevent the loss of important image details during the noise removal process.
[0044] Divide the LDCT image into several second cells of s×s according to the second step length s. There are s×s pixel points in each second cell. Take one pixel point from each second cell and allocate it to a sub-image until all the pixel points in the second cell are allocated, obtaining s 2 different second sub-images, defined as the second group of sub-images where h 1j (y) represents the j-th second sub-image
[0045] S2. Generate a neighbor sub-image for each sub-image. The neighbor sub-image is similar but not identical to the sub-image, forming a pair of mutually neighboring sub-images, obtaining the first set of sub-image pairs and the second set of sub-image pairs; to ensure that each pixel can participate in the prediction of the denoising result. This method effectively avoids information loss that may be caused by processing a single sub-image, and at the same time helps to retain more image details. In S2, the specific method of generating a neighbor sub-image for each sub-image is as follows
[0046] For the first group of sub-images, the pixel points in the i-th first sub-image come from the first cell. Randomly select pixel points from the pixel points in the first cell that have not been allocated to the i-th first sub-image. Randomly select one pixel point from each first cell to form a neighbor sub-image corresponding to the i-th first sub-image; the neighbor sub-images of all first sub-images are defined as the first group of neighbor sub-images where g 2i (y) represents the neighbor sub-image corresponding to the i-th first sub-image
[0047] For the second group of sub-images, the pixel points in the j-th second sub-image come from the second cell. Randomly select pixel points from the pixel points in the second cell that have not been allocated to the j-th second sub-image. Randomly select one pixel point from each second cell to form a neighbor sub-image corresponding to the j-th second sub-image. The neighbor sub-images of all second sub-images are defined as the second group of neighbor sub-images where h 2j (y) represents the neighbor sub-image corresponding to the j-th second sub-image
[0048] To more clearly describe the principle of generating sub-images and neighbor sub-images, in this embodiment, the step length is taken as 2, as Figure 2 shown Figure 2 The left part shows the generation of 4 sub-images. The cell is 2×2. Four pixel points A, B, C, and D are outlined. Select pixel point A from all cells to form the first sub-image y 11 , select pixel point B from all cells to form the second sub-image y 12 , select pixel point C from all cells to form the third sub-image y 21 , select pixel point D from all cells to form the fourth sub-image y 22, to generate the first sub - figure y 11 of the neighboring sub - figure y' 11 as an example, as Figure 2 shown in the right - hand part, the light - blue pixel points represent the pixel points assigned to the first sub - figure y 11 From the remaining pixel points that have not been assigned to the first sub - figure, pixel points are randomly selected. The dark - blue pixel points represent the randomly selected pixel points. One pixel point is randomly selected from each 2×2 cell to form the first sub - figure y 11 of the neighboring sub - figure y' 11 , and the neighboring sub - figures of the other three sub - figures are generated in the same way.
[0049] S3. Use the first sub - figure pair set and the second sub - figure pair set to train the denoising network respectively. One of each pair of neighboring sub - figures is used as the input of the denoising network, and the other is used as the prediction target; establish a loss function between the denoised image generated by the denoising network and the prediction target to capture the error between each pair of neighboring sub - figures, ensuring that each pixel can be fully considered and optimized during the denoising process. The designed loss function can handle the complex correlations between pixels and avoid the performance limitations brought by only relying on the independent pixel assumption in traditional methods. Train the denoising network with the goal of minimizing the loss function to obtain two denoising networks with different parameters; in this embodiment, the denoising network adopts a U - Net structure, which is a classic network commonly used in image denoising and image segmentation. In addition to the U - Net structure, other network structures can also be used. The denoising principle of the denoising network is as Figure 3 . In S3, the specific process of establishing a loss function between the denoised image generated by the denoising network and the prediction target and training the denoising network with the goal of minimizing the loss function is as follows:
[0050] The first group of sub - figures g1(y) is fed into the first denoising network f θ1 to generate the first denoised image f θ1 (g1(y)). The first group of neighboring sub - figures g2(y) is used as the prediction target. The loss function of the first denoising network is as follows:
[0051]
[0052] In the formula, E 1 represents the optimization result of the first denoising network on multiple input samples, and θ1 represents the parameters of the first denoising network; train the first denoising network with the goal of minimizing the loss function of the first denoising network; y represents the LDCT image;
[0053] The second group of sub - figures h1(y) is fed into the second denoising network f θ2 to generate the second denoised image f θ2(h1(y)), the second group of neighbor sub - graphs h2(y) is used as the prediction target, and the loss function of the second denoising network is as follows:
[0054]
[0055] In the formula, E 2 represents the optimization result of the second denoising network on multiple input samples, and θ2 represents the parameters of the second denoising network; the second denoising network is trained with the goal of minimizing the loss function of the second denoising network.
[0056] S4. The LDCT image to be denoised is as Figure 4 shown. Through the pixel - shuffle down - sampling operation, the relevant pixel points in the LDCT image to be denoised are assigned to multiple different sub - graphs with two different strides, generating two groups of sub - graphs. The two groups of sub - graphs are input into two trained denoising networks with different parameters to obtain two groups of denoised sub - graphs; in this embodiment, the stride of one branch is 2, and the stride of the other branch is 4. The branch with a stride of 2 retains the structure and texture effect of the image, and the branch with a stride of 4 meets the requirement of the independent pixel noise hypothesis.
[0057] S5. Pixel - shuffle up - sampling operations are respectively performed on the two groups of denoised sub - graphs and recombined into two denoised images with the same size as the original LDCT image; ensuring that the generated images not only retain the key information of the original image but also effectively reduce the influence of noise.
[0058] S6. Linear weighted combination is performed on the two denoised images to obtain the final reconstructed LDCT image. The denoised LDCT image is as Figure 5 shown. The linear weighted combination is expressed by the formula:
[0059] F output = a·F1+(1 - a)·F3
[0060] In the formula, F1 represents the denoised image output by the first denoising network, F2 represents the denoised image output by the second denoising network, a represents the weight, and F output represents the final reconstructed LDCT image. By weighted - combining the denoising results of these two branches, the denoising effect can be enhanced while retaining the structure and texture of the image, thereby obtaining a denoising result closer to the real image.
[0061] By balancing the relationship between noise suppression and signal retention, the present invention avoids the problems of image structure loss or noise residue caused by improper stride setting in traditional methods, and thus achieves more ideal performance in the image denoising task.
[0062] Embodiment 2
[0063] The present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, a denoising method for low-dose computed tomography images based on self-supervised learning as described in Embodiment 1 is implemented.
[0064] Embodiment 3
[0065] The present invention provides a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the denoising method for low-dose computed tomography images based on self-supervised learning as described in Embodiment 1.
[0066] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0067] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed in the present application can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0068] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A low-dose CT image denoising method based on self-supervised learning, characterized in that: The following steps are involved: S1, through pixel shuffling and downsampling operation, the associated pixels in the LDCT image are distributed to multiple different sub-images with two different step lengths to generate two groups of sub-images; S2. Generate a neighbor subgraph for each subgraph to form a pair of subgraphs that are neighbors of each other, and obtain a first subgraph pair set and a second subgraph pair set; S3, using the first sub-image pair set and the second sub-image pair set to train the denoising network respectively, one of each pair of neighboring sub-images is used as the input of the denoising network, and the other is used as the prediction target; using the denoised image generated by the denoising network and the prediction target to establish a loss function, and training the denoising network with the goal of minimizing the loss function, to obtain two denoising networks with different parameters; S4, through pixel shuffling and downsampling operations, the associated pixels in the LDCT image to be denoised are distributed to a plurality of different sub-images with two different step lengths to generate two groups of sub-images, and the two groups of sub-images are input into two trained denoising networks with different parameters to obtain two groups of denoised sub-images; S5, performing pixel shuffling and upsampling operations on the two groups of denoised sub-images respectively, and recombining them into two denoised images with the same size as the original LDCT image; S6. Perform a linear weighted combination on the two denoised images to obtain a final reconstructed LDCT image.
2. The low-dose CT image denoising method based on self-supervised learning according to claim 1, characterized in that: S1 is specifically: The LDCT image is divided into a number of r×r first cells according to the first step length r. Each first cell has r×r pixels. One pixel is taken from each first cell and assigned to a sub-image until all the pixels in the first cell are assigned. 2 different first subgraphs, defined as the first group of subgraphs Among them, g 1i (y) represents the first sub-image of the i-th image; y represents the LDCT image; The LDCT image is divided into a number of s×s second cells according to the second step length s. Each second cell has s×s pixels. One pixel is taken from each second cell and assigned to a sub-image until all the pixels in the second cell are assigned. Then, s 2 different second subgraphs, defined as the second set of subgraphs Among them, h 1j (y) represents the j-th second sub-image.
3. The low-dose CT image denoising method based on self-supervised learning as claimed in claim 2, characterized in that: In S2, generating a neighbor subgraph for each subgraph is specifically as follows: For the first group of sub-graphs, the pixels in the i-th first sub-graph come from the first cell, and pixels are randomly selected from the pixels in the first cell that are not assigned to the i-th first sub-graph. Each first cell randomly selects a pixel to form a neighbor sub-graph corresponding to the i-th first sub-graph; the neighbor sub-graphs of all first sub-graphs are defined as the first group of neighbor sub-graphs. In the formula, g 2i (y) represents the neighbor subgraph corresponding to the i-th first subgraph; For the second group of sub-graphs, the pixels in the j-th second sub-graph come from the second cell, and pixels are randomly selected from the pixels in the second cell that are not assigned to the j-th second sub-graph. One pixel is randomly selected from each second cell to form a neighbor sub-graph corresponding to the j-th second sub-graph. The neighbor sub-graphs of all second sub-graphs are defined as the second group of neighbor sub-graphs. In the formula, h 2j (y) represents the neighbor subgraph corresponding to the j-th second subgraph.
4. The low-dose CT image denoising method based on self-supervised learning according to claim 1, characterized in that: In S3, the denoising image generated by the denoising network is used to establish a loss function with the predicted target, and the specific process of training the denoising network with the goal of minimizing the loss function is as follows: The first group of sub-images g1(y) is fed into the first denoising network f θ1 Generate the first denoised image f θ1 (g1(y)), the first group of neighbor subgraphs g2(y) is used as the prediction target, and the loss function of the first denoising network is as follows: In the formula, E 1 represents the optimization result of the first denoising network on multiple input samples, θ1 represents the parameters of the first denoising network; the first denoising network is trained with the goal of minimizing the loss function of the first denoising network; y represents the LDCT image; The second group of sub-images h1(y) is fed into the second denoising network f θ2 Generate the second denoised image f θ2 (h1(y)), the second group of neighbor subgraphs h2(y) is used as the prediction target, and the loss function of the second denoising network is as follows: In the formula, E 2 represents the optimization result of the second denoising network on multiple input samples, θ2 represents the parameters of the second denoising network; the second denoising network is trained with the goal of minimizing the loss function of the second denoising network.
5. The low-dose CT image denoising method based on self-supervised learning according to claim 1, characterized in that: In S6, the linear weighted combination of the two denoised images is performed to obtain the final reconstructed LDCT image: F output =a·F1+(1-a)·F2 In the formula, F1 represents the denoised image output by the first denoising network, F2 represents the denoised image output by the second denoising network, a represents the weight, and F output Represents the final reconstructed LDCT image.
6. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the low-dose CT image denoising method based on self-supervised learning as described in any one of claims 1 to 5 is implemented.
7. A computer-readable storage medium storing a computer program, characterized in that: The computer program enables a computer to execute the low-dose CT image denoising method based on self-supervised learning as described in any one of claims 1 to 5.
Citation Information
Cited By
Denoising enhancement method and device for low-dose CT image and storage medium
CN122289095A