Training method of image denoising model, image processing method and image processing system
By training the image denoising model, using multi-scale quantized autoencoder and noise mode information extraction module, combined with the generation adversarial network and anatomical perception discriminator, the generalization ability problem of low-dose CT images in different noise modes is solved, and better image denoising effect and anatomical structure retention are achieved.
Patent Information
- Application Number
- CN202510514007.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-11
AI Technical Summary
Existing low-dose CT image denoising models have challenges in generalization capabilities and are unable to effectively identify different noise patterns, resulting in poor performance when processing low-dose CT images of other noise patterns.
By acquiring sample standard images and sample noise-added images, using multi-scale quantized autoencoder and noise mode information extraction module, the image denoising model is trained, and combined with the generation of adversarial network and anatomical perception discriminator, the identification and denoising of noise patterns are achieved.
It improves the generalization ability of image denoising models, and can more robustly process low-dose CT images of different noise modes, improves image quality and maintains anatomical integrity and detail clarity.
Smart Images

Figure CN120298247A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for training an image denoising model, an image processing method, and an image processing system. Background Art
[0002] CT (Computed Tomography) is a medical imaging technology that performs tomographic scans of the human body using an X-ray beam and generates detailed images of the internal structure of the body with the aid of computer processing. In traditional CT scans, the current of the X-ray tube is relatively high to ensure the generation of high-quality images. However, long-term or frequent exposure to high-dose X-ray radiation increases the risk of cancer or other radiation-related health problems. Therefore, low-dose CT has emerged. Low-dose CT reduces the X-ray dose by strategies such as reducing the current of the X-ray tube and adjusting the exposure time, thereby reducing radiation exposure.
[0003] However, the images generated by low-dose CT have more noise and lower image quality compared to traditional CT images. Therefore, related technologies are needed to denoise low-dose CT images. Currently, significant progress has been made in learning-based low-dose CT denoising technologies. However, these methods still face challenges in terms of generalization ability. Current denoising methods usually perform single-mode modeling on simulated datasets with specific noise patterns. The denoising models obtained in this way perform poorly on data with other noise patterns. Therefore, how to train a low-dose CT denoising model that can recognize noise patterns and improve the generalization ability of the model to different noise patterns has become an urgent problem for technicians to solve. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for training an image denoising model, an image processing method, and an image processing system. One or more embodiments of this specification also relate to a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for training an image denoising model is provided, including:
[0006] Obtaining an image training sample, where the image training sample includes a sample standard image and a sample noisy image corresponding to the sample standard image;
[0007] Generate a sample reference image according to the sample standard image, and input the sample noise-added image and the sample reference image into an initial image denoising model to obtain an initial predicted image output by the initial image denoising model. Among them, the initial image denoising model determines noise pattern information according to the encoded feature information of the sample noise-added image and the sample reference image, and denoises the sample noise-added image according to the noise pattern information;
[0008] Adjust the model parameters of the initial image denoising model according to the initial predicted image and the sample standard image to obtain a reference image denoising model;
[0009] Input the sample noise-added image into the reference image denoising model to obtain a target predicted image output by the reference image denoising model. Among them, the reference image denoising model predicts noise pattern information according to the sample noise-added image, and denoises the sample noise-added image according to the noise pattern information;
[0010] Adjust the model parameters of the reference image denoising model according to the target predicted image and the sample standard image to obtain an image denoising model.
[0011] According to the second aspect of the embodiments of the present specification, there is provided an image processing method, including:
[0012] Receive an image processing task, where the image processing task carries a plurality of images to be processed corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
[0013] Input the plurality of images to be processed into the image denoising model to obtain a plurality of target images output by the image denoising model. Among them, the image denoising model is trained by the training method of the above image denoising model;
[0014] Determine whether there is an abnormal object in the target detection area according to the plurality of target images.
[0015] According to the third aspect of the embodiments of the present specification, there is provided a CT image processing method, including:
[0016] Receive a CT image processing task, where the CT image processing task carries a plurality of low-dose CT images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
[0017] Input the plurality of low-dose CT images into the CT image denoising model to obtain a plurality of denoised CT images output by the CT image denoising model. Among them, the CT image denoising model is trained by the training method of the above image denoising model;
[0018] Determine whether there is an abnormal object in the target detection area according to the multiple denoised CT images.
[0019] According to the fourth aspect of the embodiments of the present specification, an image processing system is provided, including a client and a server. Among them,
[0020] The client is configured to send a CT image processing task to the server, where the CT image processing task carries multiple low-dose CT images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
[0021] The server is configured to input the multiple low-dose CT images into a CT image denoising model to obtain multiple denoised CT images output by the CT image denoising model, where the CT image denoising model is trained by the above-mentioned image denoising model training method; determine whether there is an abnormal object in the target detection area according to the multiple denoised CT images.
[0022] According to the fifth aspect of the embodiments of the present specification, a computing device is provided, including:
[0023] A memory and a processor;
[0024] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.
[0025] According to the sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0026] According to the seventh aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0027] Through the method provided by the embodiments of the present specification, a noise pattern-aware image denoising model is provided. First, noise pattern information is provided according to a reference image, and then the denoising process is guided according to pre-trained prior knowledge, making the denoising of the noisy image more robust and having better generalization. When performing the guiding strategy of noise pattern information, a multi-scale quantization autoencoder is aligned with the image denoising model, thereby improving the modeling of the denoising model. Description of the Drawings
[0028] Figure 1 is a flowchart of a method for training an image denoising model provided by an embodiment of the present specification;
[0029] Figure 2 It is a schematic diagram of data processing of an initial image denoising model provided by an embodiment of this specification;
[0030] Figure 3 It is a schematic structural diagram of a noise pattern information extraction module provided by an embodiment of this specification;
[0031] Figure 4 It is a schematic structural diagram of a decoding unit provided by an embodiment of this specification;
[0032] Figure 5 It is a schematic diagram of the model calculating the loss value provided by an embodiment of this specification;
[0033] Figure 6 It is a schematic structural diagram of an attention-based feature fusion module provided by an embodiment of this specification;
[0034] Figure 7 It is a schematic structural diagram of a training device for an image denoising model provided by an embodiment of this specification;
[0035] Figure 8 It is an architecture diagram of a training system for an image denoising model provided by an embodiment of this specification;
[0036] Figure 9 It is a flowchart of an image processing method provided by an embodiment of this specification;
[0037] Figure 10 It is a flowchart of a CT image processing method provided by an embodiment of this specification;
[0038] Figure 11 It is a system schematic diagram of an image processing system provided by an embodiment of this specification;
[0039] Figure 12 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0040] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0041] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0042] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0043] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0044] First, the noun terms involved in one or more embodiments of this specification are explained.
[0045] CT (Computed Tomography): Computed Tomography, which uses a precisely collimated X-ray beam and a highly sensitive detector to perform successive cross-sectional scans around a certain part of the human body. It has the characteristics of fast scanning time and clear images and can be used for the examination of various diseases.
[0046] Low-dose CT (Low-dose Computed Tomography, LDCT): It is a CT imaging technology that uses a reduced X-ray radiation dose, mainly used to reduce the radiation exposure of patients. In traditional CT scans, the current of the X-ray tube is relatively high to ensure the generation of high-quality images. However, long-term or frequent exposure to high-dose X-ray radiation may increase the risk of cancer or other radiation-related health problems. To alleviate this problem, LDCT reduces the X-ray dose by strategies such as reducing the X-ray tube current and adjusting the exposure time, thereby reducing the radiation exposure.
[0047] Normal-dose Computed Tomography (NDCT): Conventional-dose CT refers to CT scans using standard radiation doses to provide high-quality images, which are typically used for detailed diagnostic evaluations.
[0048] Low-dose Computed Tomography Denoising refers to a series of image processing and restoration techniques aimed at removing noise and artifacts in LDCT images while maintaining the integrity of anatomical structures and the clarity of details as much as possible. The ultimate goal is to restore LDCT images to the quality level of normal-dose CT (NDCT).
[0049] Compared with traditional low-dose CT denoising techniques such as iterative reconstruction, learning-based low-dose CT denoising techniques have made significant progress. In specific scenarios, the image quality after denoising by relevant techniques has been significantly improved and can be comparable to that of normal-dose CT images. However, these methods still face challenges in terms of generalization ability. Specifically, there is no unified standard for the proportion of current low-dose CT, and different hospitals use different noise patterns. It is difficult to obtain paired NDCT and LDCT data for model training. Currently, only single-mode modeling can be performed on the obtained dataset under specific noise patterns, and the image denoising model obtained will perform poorly when processing LDCT with other noise patterns.
[0050] Based on this, in this specification, a training method for an image denoising model, an image processing method, and an image processing system are provided. This specification also relates to a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0051] See Figure 1 , Figure 1 shows a flowchart of a training method for an image denoising model provided according to an embodiment of this specification, which specifically includes the following steps.
[0052] Step 102: Obtain image training samples, where the image training samples include sample standard images and sample noisy images corresponding to the sample standard images.
[0053] Specifically, the training method of the image denoising model provided in this specification uses supervised training, which includes multiple image training samples. Among the image training samples, there are sample standard images and sample noisy images corresponding to the sample standard images. Herein, the sample standard image can be understood as the standard image, that is, the sample label in the training method of the image denoising model. The sample noisy image refers to the image obtained by adding noise to the sample standard image, and the sample noisy image is the sample data in the training method of the image denoising model.
[0054] In the method provided in the embodiments of this specification, the sample noisy image is input into the image denoising model for denoising to obtain a predicted image. The predicted image and the sample standard image are compared to obtain a loss value, and the model parameters of the image denoising model are adjusted based on the loss value.
[0055] In the auxiliary medical scenario, the sample standard image can be understood as the normal-dose CT (NDCT), and the sample noisy image can be understood as the low-dose CT (LDCT) corresponding to the normal-dose CT.
[0056] In practical applications, the acquisition cost of the training sample pair of the normal-dose CT and the low-dose CT is relatively high, and the number of currently available training sample pairs is small. To better obtain the training sample pairs, in the method provided in the embodiments of this specification, the image training samples are obtained through the following steps:
[0057] Obtain the sample standard image;
[0058] Add noise to the sample standard image based on different noise pattern information to obtain sample noisy images corresponding to different noise pattern information;
[0059] Construct image training samples according to the sample standard image and the sample noisy images.
[0060] In practical applications, the acquisition method of the sample standard image is more convenient. For example, in the auxiliary medical scenario, the acquisition method of the normal-dose CT is more convenient. Therefore, in the method provided in the embodiments of this specification, after obtaining the sample standard image, the sample noisy image can be generated by adding noise to the sample standard image. Thus, the sample noisy image and the sample standard image are composed into the image training samples for training the image denoising model.
[0061] The noise pattern information can be understood as the degree of noise when adding noise to the sample standard image. For example, adding 0.5 times noise, 1 time noise, 1.5 times noise, 2 times noise, etc. to the sample standard image. In the method provided in the embodiments of this specification, it is desired to train an image denoising model that can adapt to various noise patterns. Therefore, when constructing the image training samples, different noise pattern information can be used to add noise to the sample standard image, so as to obtain the sample noisy images corresponding to different noise patterns.
[0062] For example, still taking the auxiliary medical scenario as an example, the sample standard image is the normal-dose CT image. Different noise pattern information can be added to the normal-dose CT image, that is, adding different degrees of noise to the normal-dose CT to obtain the sample noisy images corresponding to different noise pattern information. Then, the image training samples can be constructed according to the sample standard image and the sample noisy images.
[0063] Through the method provided in the embodiments of this specification, starting from the sample standard image, adding noise to the sample standard image according to different noise pattern information, so as to obtain the sample noisy images corresponding to different noise pattern information, which solves the problem of difficult acquisition of image training samples and saves the acquisition cost of image training samples.
[0064] Step 104: Generate a sample reference image according to the sample standard image, and input the sample noisy image and the sample reference image into the initial image denoising model to obtain an initial prediction image output by the initial image denoising model. Wherein, the initial image denoising model determines the noise pattern information according to the encoded feature information of the sample noisy image and the sample reference image, and denoises the sample noisy image according to the noise pattern information.
[0065] In the method provided in the embodiments of this specification, when training the image denoising model to denoise the sample noisy image, during the training process, the reference image helps the image denoising model obtain the noise pattern information from the latent space difference between the sample noisy image and the reference image. Then, denoise the sample noisy image to obtain the corresponding prediction image.
[0066] In this process, if the sample standard image is directly used, it will cause the image denoising model to rely too much on the features of the sample standard image during the denoising process, and the effect will be poor in the subsequent application stage. Therefore, in the specific implementation provided in this specification, the sample standard image is not directly used to denoise the sample noisy image, but a sample reference image is generated through the sample standard image. The image clarity of the sample reference image is the same as that of the sample standard image, but there will be differences in the image content. In practical applications, the sample reference image can be obtained by performing data augmentation on the sample standard image.
[0067] In a specific embodiment provided in this specification, generating a sample reference image according to the sample standard image includes:
[0068] Determining an initial sample reference image according to the sample standard image;
[0069] Performing data augmentation on the initial sample reference image to obtain a sample reference image.
[0070] In the method provided in the embodiments of this specification, the corresponding initial sample reference image can be obtained according to the sample standard image first, and then the initial sample reference image can be adjusted through data augmentation to obtain a sample reference image.
[0071] For example, taking the auxiliary medical scenario as an example, the sample standard image can be a normal-dose CT image. In actual applications, there are usually multiple CT images for a certain region. When selecting one normal-dose CT as the sample standard image, other images adjacent to it can be selected as the initial sample reference image. Then, data augmentation is performed on the initial sample reference image through elastic transformation, rotation, translation, etc., so as to obtain a sample reference image.
[0072] After obtaining the sample reference image, the sample reference image and the sample noisy image can be input into the initial image denoising model for processing to obtain the initial prediction image output by the initial image denoising model. The initial image denoising model is the model in the first stage of the model training of the image denoising model. In the current stage, the model input of the initial image denoising model includes the sample noisy image and the sample reference image, so that the initial image denoising model can extract noise pattern information according to the sample noisy image and the sample reference image, and denoise the sample noisy image according to the noise pattern information to obtain the initial prediction image.
[0073] Furthermore, in a specific embodiment provided in this specification, the initial image denoising model includes an encoder, a noise pattern information extraction module, and a decoder;
[0074] See Figure 2 , Figure 2 shows the data processing schematic diagram of the initial image denoising model provided in an embodiment of this specification. As Figure 2 shown, the image denoising model includes an encoder, a noise pattern information extraction module, and a decoder. The encoder is used to extract image feature information, the noise pattern information extraction module is used to extract noise pattern information according to the image feature information extracted by the encoder, and the decoder is used to denoise the sample noisy image according to the noise pattern information. First, a sample reference image is generated according to the sample standard image, and the sample noisy image and the sample reference image are input into the initial image denoising model. After being processed by the encoder, the noise pattern information extraction module, and the decoder, an initial prediction image is generated.
[0075] Specifically, input the sample noisy image and the sample reference image into the initial image denoising model, and obtain the initial predicted image output by the initial image denoising model, including S1042 - S1046:
[0076] S1042. Input the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features that correspond one-to-one to the multi-scale noisy image features.
[0077] In this embodiment, input the sample noisy image and the sample reference image into the encoder for image feature extraction. The latent space features of the sample noisy image and the sample reference image are respectively extracted by the encoder. Specifically, the latent space features of the sample noisy image and the sample reference image are multi-scale noisy image features and multi-scale reference image features, and the multi-scale noisy image features and the multi-scale reference image features correspond one-to-one.
[0078] Specifically, in the method provided in the embodiments of this specification, the encoder is a multi-scale quantization autoencoder;
[0079] Inputting the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features that correspond one-to-one to the multi-scale noisy image features includes:
[0080] Obtain the scale number information corresponding to the multi-scale quantization autoencoder;
[0081] Input the sample noisy image and the sample reference image into the encoder to obtain the multi-scale noisy image features and multi-scale reference image features corresponding to the scale number information.
[0082] In this implementation method, the encoder is a multi-scale quantization autoencoder (Multi-scale Quantization Autoencoder), which is used to extract the latent space features of the sample noisy image and the sample reference image. Specifically, first obtain the scale number information of the multi-scale quantization autoencoder. When extracting the multi-scale image features of the sample noisy image and the sample reference image, generate the corresponding number of multi-scale noisy image features and multi-scale reference image features according to the scale number information. For example, if the scale number information of the multi-scale quantization autoencoder is k, then k different-scale noisy image features and k corresponding-scale reference image features will be obtained.
[0083] S1044. Input the multi-scale noisy image features and the multi-scale reference image features into the noise pattern information extraction module to obtain the noise pattern information corresponding to each scale of the noisy image features.
[0084] After obtaining the multi-scale noisy image features and the multi-scale reference image features, the multi-scale noisy image features and the multi-scale reference image features are input into the noise pattern information extraction module for processing, and the noise pattern information corresponding to the noisy image features at each scale is extracted through the noise pattern extraction module.
[0085] In practical applications, the multi-scale noisy image features and the multi-scale reference image features are in one-to-one correspondence. Inputting the multi-scale noisy image features and the multi-scale reference image features into the noise pattern information extraction module for processing means inputting the multi-scale noisy image features and the multi-scale reference image features into the noise pattern information extraction module in pairs to obtain the noise pattern information. The noise pattern information is used as guiding information for denoising operations in subsequent decoding processes.
[0086] In a specific embodiment provided in this specification, the multi-scale noisy image features and the multi-scale reference image features are correspondingly input into the noise pattern information extraction module to obtain the noise pattern information corresponding to the noisy image features at each scale, including:
[0087] Determine the target multi-scale noisy image features and the target multi-scale reference image features, where the target multi-scale noisy image features and the target multi-scale reference image features are in one-to-one correspondence;
[0088] Input the target multi-scale noisy image features and the target multi-scale reference image features into the noise pattern information extraction module to obtain the target noise pattern information corresponding to the target multi-scale noisy image features.
[0089] In this embodiment, the target multi-scale noisy image features are taken as an example for explanation. The target multi-scale noisy image features are any one of the multi-scale noisy image features. At the same time, the target multi-scale reference image features corresponding to the target multi-scale noisy image features are obtained.
[0090] Input the target multi-scale noisy image features and the target multi-scale reference image features into the noise pattern information extraction module for processing to obtain the target noise pattern information corresponding to the target multi-scale noisy image features.
[0091] See Figure 3 , Figure 3 shows a schematic structural diagram of the noise pattern information extraction module provided in an embodiment of this specification, as Figure 3As shown, the noise pattern information extraction module consists of a convolutional layer (Conv), an activation function (LRelu), a pooling layer (AvgPool), and a linear layer (Linear). The target multi-scale noisy image features and the target multi-scale reference image features are input into Conv1 for convolutional processing to extract image features. Then, after N three-layer convolutional processes, and finally, after Conv1 and average pooling operations, the output passes through the linear layer and then a difference process is performed. The difference result is the target noise pattern information.
[0092] The above Figure 3 is a schematic diagram of the processing of the noise pattern information extraction module. In practical applications, each pair of target multi-scale noisy image features and target multi-scale reference image features can be input into this noise pattern information extraction module for processing to obtain the noise pattern information corresponding to the noisy image features at each scale. In the method provided in the embodiments of this specification, the noise pattern information is captured from the difference in the encoded latent space of the multi-scale noisy image features and the multi-scale reference image features through the noise pattern information extraction module, which is used to provide guiding information in the subsequent denoising process.
[0093] S1046. Input the sample noisy image and each noise pattern information into the decoder to obtain the initial predicted image output by the decoder.
[0094] After obtaining each noise pattern information, denoise the sample noisy image through each noise pattern information to provide guiding information for the denoising process of the sample noisy image, and generate the initial predicted image output by the decoder.
[0095] In a specific implementation manner provided in this specification, the decoder includes an embedding unit, a plurality of decoding units, and an output unit. The number of decoding units is the same as the scale number information of the multi-scale quantization autoencoder;
[0096] Inputting the sample noisy image and each noise pattern information into the decoder to obtain the initial predicted image output by the decoder includes:
[0097] Input the sample noisy image into the embedding layer to obtain the image features to be decoded output by the embedding layer;
[0098] Input the image features to be decoded output by the previous processing unit and the target noise pattern information into the target decoding unit to obtain the image features to be decoded output by the target decoder, where the previous processing unit is the embedding unit or the decoding unit, and the target noise pattern information and the target decoding unit correspond one by one;
[0099] Input the image features to be decoded output by the last decoding unit into the output unit to obtain the initial predicted image.
[0100] In practical applications, the decoder includes an embedding unit, multiple decoding units, and an output unit. The embedding unit is used to perform an embedding process on the sample noisy image to obtain the image features to be decoded.
[0101] There are multiple decoding units, and the number of decoding units is the same as the scale number information of the multi-scale quantization sub-encoder. That is, the number of decoding units is the same as the number of each noise pattern information, and they correspond one-to-one with the multi-scale noisy image features.
[0102] In the method provided in the embodiments of this specification, the decoding unit is a multi-scale noise-guided decoding unit, which is used to control the denoising operation. The decoding unit is aligned with each noise pattern information to ensure that fine-grained noise features can guide the shallow layer of the denoising network, while coarser noise features are injected into the deep layer of the network to improve the accuracy and effect of denoising. For each decoding unit, its input is the image features to be decoded output by the previous processing unit and the target noise pattern information corresponding to the current decoding unit. The processing unit in the embodiments of this specification can be an embedding layer or the previous decoding unit. For example, for the first decoding unit, its input is the image features to be decoded output by the embedding layer and the target noise pattern information corresponding to the first decoding unit. For the second decoding unit and subsequent decoding units, its input is the image features to be decoded output by the previous decoding unit and the corresponding target noise pattern information.
[0103] See Figure 4 , Figure 4 shows a schematic structural diagram of the decoding unit provided in an embodiment of this specification. As Figure 4 shown, the decoding unit is composed of a standard layer (Norm), a linear layer (Linear), a convolutional layer (Conv), a deformable convolution (DConv), and an activation function (GELU). For the current decoding unit, its input is the image features to be decoded output by the previous processing unit and the target noise pattern information corresponding to the current decoding unit. After being processed by the decoding unit as Figure 4 shown, the image features to be decoded output by the current decoding unit are obtained.
[0104] In practical applications, each decoding unit also corresponds to a self-attention sub-layer, enabling each decoding unit to fuse noise pattern information of different scales, which is beneficial for denoising the image features to be decoded.
[0105] The image features to be decoded output by the last decoding unit are input to the output unit. After being processed by the output unit, the initial prediction image generated by the initial denoising model can be obtained and output.
[0106] Step 106: Adjust the model parameters of the initial image denoising model according to the initial prediction image and the sample standard image to obtain a reference image denoising model.
[0107] At this time, the initial image denoising model is still an untrained model, and it is necessary to further compare the predicted results with the labels, and then adjust the model parameters of the initial image denoising model. Specifically, calculate the loss value according to the initial predicted image and the sample standard image, and backpropagate according to the loss value to adjust the model parameters of the initial image denoising model.
[0108] Repeat the above model training steps of the initial image denoising model until the preset model training stop condition is reached. In the method provided in the embodiments of this specification, the model training stop condition includes performing model training for a preset number of rounds using image training samples, and / or the loss value is less than or equal to a preset loss value threshold. Thus, the reference image denoising model is obtained.
[0109] In the method provided in the embodiments of this specification, the input of the initial image denoising model includes the sample noisy image and the sample reference image. During the model training process, the initial image denoising model captures noise pattern information through the difference in the encoder space of the sample noisy image and the sample reference image, and inputs the noise pattern information into the decoder. In the decoder, the noise pattern information is used to provide guiding information for the denoising process of the sample noisy image, thereby improving the denoising effect and generalization ability of the image denoising model.
[0110] The reference image denoising model at this time is not the final image denoising model because the reference image denoising model at this time still needs to rely on the input sample reference image to extract noise pattern information and then guide the decoder to denoise. However, in actual applications, when denoising a noisy image, the corresponding reference image cannot be provided. Therefore, it is necessary to further train the reference image denoising model to generate noise pattern information.
[0111] Step 108: Input the sample noisy image into the reference image denoising model to obtain the target predicted image output by the reference image denoising model, where the reference image denoising model predicts noise pattern information according to the sample noisy image and denoises the sample noisy image according to the noise pattern information.
[0112] In the process of further denoising the reference image denoising model, it is only necessary to input the sample noisy image into the reference image denoising model. The reference image denoising model can predict noise pattern information according to the sample noisy image and guide the denoising of the sample noisy image according to the noise pattern information, so as to obtain the target predicted image.
[0113] In the method provided in the embodiments of this specification, the model structure of the reference image denoising model is the same as that of the initial image denoising model, and no further explanation is made for the model structure of the reference image denoising model here.
[0114] In a specific embodiment provided in this specification, the reference image denoising model includes an encoder, a noise pattern information extraction module, and a decoder. The encoder is used to generate simulated reference image features;
[0115] Inputting the sample noisy image into the reference image denoising model to obtain the target predicted image output by the reference image denoising model includes:
[0116] Inputting the sample noisy image into the encoder to obtain multi-scale noisy image features output by the encoder and multi-scale simulated reference image features corresponding one-to-one to the multi-scale noisy image features;
[0117] Inputting the multi-scale noisy image features and the multi-scale simulated reference image features into the noise pattern information extraction module correspondingly to obtain the noise pattern information corresponding to the noisy image features at each scale;
[0118] Inputting the sample noisy image and each noise pattern information into the decoder to obtain the target predicted image output by the decoder.
[0119] In the encoder of the reference image denoising model in the embodiments of this specification during the current model training, a visual autoregressive model adaptation method is used to adapt the visual autoregressive model in the previous stage of training, that is, the encoding ability of the encoder is used to simulate and generate the latent space information of the reference image.
[0120] Specifically, inputting the sample noisy image into the encoder to obtain multi-scale noisy image features output by the encoder and multi-scale simulated reference image features corresponding one-to-one to the multi-scale noisy image features. At this time, the prior information obtained by the encoder through the above steps is introduced into the multi-scale latent space of the encoder, so as to capture the distribution of the noise pattern. In practical applications, through the prior information of the previous multi-scale noisy image features and the target multi-scale noisy image features, after processing, the prior information corresponding to the target multi-scale noisy image features is obtained, and the prior information corresponding to the target multi-scale noisy image features is fused with the next multi-scale noisy image features to obtain the prior information of the next multi-scale noisy image features. The simulated reference image features are obtained through the prior information of each scale.
[0121] Inputting the multi-scale noisy image features and the multi-scale simulated reference image features into the noise pattern information extraction module to obtain the noise pattern information corresponding to the noisy image features at each scale, and then inputting the sample noisy image and each noise pattern information into the decoder to obtain the target predicted image output by the decoder.
[0122] In this embodiment, through adaptation by a pre-trained visual autoregressive model, a latent space prior for the denoising process is generated, thereby generating simulated reference image features. Then, the noise pattern information is determined based on the difference between the multi-scale noisy image features and the multi-scale simulated reference image features, thereby guiding the denoising of the sample noise model, and solving the problem of being unable to obtain a reference image.
[0123] Step 110: Adjust the model parameters of the reference image denoising model according to the target prediction image and the sample standard image to obtain an image denoising model.
[0124] After obtaining the target prediction image, the model loss value can be calculated based on the sample standard image and the target prediction image, and the model parameters of the reference denoising model can be continuously adjusted according to the model loss value until the model training stop condition is reached. Thus, the final image denoising model is obtained.
[0125] In a specific embodiment provided in this specification, adjusting the model parameters of the reference image denoising model according to the target prediction image and the sample standard image to obtain an image denoising model includes:
[0126] Calculate an image loss value, a contrast loss value, and a discriminative loss value based on the sample noisy image, the target prediction image, and the sample standard image;
[0127] Adjust the model parameters of the reference image denoising model according to the image loss value, the contrast loss value, and the discriminative loss value;
[0128] Continue to train the reference image denoising model until the model training stop condition is reached to obtain an image denoising model.
[0129] In practical applications, when calculating the model loss value based on the target prediction image and the sample standard image, the sample noisy image can also be specifically combined to calculate the image loss value, the contrast loss value, and the discriminative loss value. Among them, the image loss value is used to represent the difference information between the target prediction image and the sample standard image; the contrast loss value is used to represent the difference information of contrast learning between the target prediction image and the sample standard image and the sample noisy image; the discriminative loss value is used to combine the semantic information of the sample standard image to guide the discriminator to generate the difference information of the discrimination probability between the target prediction image and the sample standard image.
[0130] Adjust the model parameters of the reference image denoising model through the image loss value, the contrast loss value, and the discriminative loss value, and continue to train the reference image denoising model until the model training stop condition is reached, and then obtain an image denoising model.
[0131] Specifically, calculating an image loss value, a contrast loss value, and a discriminant loss value based on the sample noisy image, the target prediction image, and the sample standard image includes S1102 - S1106:
[0132] S1102. Calculate the image loss value according to the target prediction image and the sample standard image.
[0133] In the method provided in the embodiments of this specification, the image loss value L1 is calculated according to the target prediction image and the sample standard image. The image loss value is used to determine the difference information between the target prediction image and the sample standard image.
[0134] S1104. Construct positive and negative sample pairs at the pixel granularity according to the sample noisy image, the target prediction image, and the sample standard image, and calculate the contrast loss value according to the positive and negative sample pairs.
[0135] In the current denoising task of LDCT images, although the denoising effect has been improved to a certain extent, it usually relies on pixel-level constraints, which often limitedly reduces the global error at the cost of sacrificing the rationality of local anatomical structures, resulting in overly smooth textures, thereby masking small tissues and lesions.
[0136] The generative adversarial network learns the data distribution rather than pixel-level mapping. Applying the generative adversarial network to the denoising task of LDCT images provides an alternative solution for the LDCT denoising task. However, the traditional denoising methods of generative adversarial networks often fail to fully capture the relationship between noise features and anatomical semantics. This raises the need for a shift towards fine-grained anatomical-aware denoising.
[0137] Based on this, in the method provided in the embodiments of this specification, positive and negative sample pairs for semantic-guided contrast learning are constructed at the pixel granularity to achieve fine-grained contrast conversion. Specifically, positive and negative sample pairs at the pixel granularity are constructed according to the sample noisy image, the target prediction image, and the sample standard image, and the contrast loss value is calculated according to the positive and negative sample pairs.
[0138] In a specific implementation manner provided in this specification, constructing positive and negative sample pairs at the pixel granularity according to the sample noisy image, the target prediction image, and the sample standard image, and calculating the contrast loss value according to the positive and negative sample pairs includes:
[0139] Input the sample noisy image, the target prediction image, and the sample construction image into the feature extraction model to obtain the positive and negative sample pairs output by the feature extraction model. Among them, the feature extraction model uses the feature of the target pixel point in the target prediction image as the sample pixel feature, uses the feature of the target pixel point in the sample standard image as the sample pixel positive label, and uses the feature of any pixel point in the sample noisy image and the sample standard image as the sample pixel negative label;
[0140] Form positive and negative sample pairs according to the sample pixel feature, sample pixel positive label, and sample pixel negative label of the target pixel point, and calculate the contrast loss value according to the positive and negative sample pairs.
[0141] In this process, the feature extraction model that has been pre-trained can be used to extract features from the sample noisy image, the target prediction image, and the sample construction image. Based on the feature of the target pixel point in the target prediction image as the sample pixel feature, and using the feature of the target pixel point in the sample standard image as the sample pixel positive label; using the feature of the non-target pixel point in the sample standard image or the feature of the pixel point in the sample noisy image as the sample pixel negative label to construct positive and negative sample pairs. And based on the positive and negative sample pairs, contrast learning is performed to calculate the contrast loss value Lscl. The contrast learning based on semantics uses the features extracted by the pre-trained feature extraction model to construct contrast learning pairs to promote the segmentation consistency between the target prediction image and the sample standard image, thereby enhancing the denoising effect.
[0142] S1106. Obtain the standard image semantic information according to the sample standard image, and generate the prediction discrimination probability and the standard discrimination probability of the target prediction image and the sample standard image based on the standard image semantic information, and calculate the discrimination loss value through the prediction discrimination probability and the standard discrimination probability.
[0143] In order to implement the evaluation of the target detection area, in the method provided in the embodiments of this specification, an anatomical perception discriminator is also provided, which is used to determine the sample standard image and the target prediction image, and calculate the discrimination loss value according to the determination result.
[0144] Specifically, obtain the standard image semantic information corresponding to the sample standard image, and use the standard image semantic information to guide the anatomical perception discriminator to determine the sample standard image and the target prediction image. Obtain the prediction discrimination probability and the standard discrimination probability output by the anatomical perception discriminator. Calculate the discrimination loss value through the prediction discrimination probability and the standard discrimination probability. In the method provided in the embodiments of this specification, the standard image semantic information corresponding to the sample standard image can be extracted by using the pre-trained feature extraction model.
[0145] In a specific implementation provided in this specification, generating the prediction discrimination probability and the standard discrimination probability of the target prediction image and the sample standard image based on the standard image semantic information includes:
[0146] Input the target prediction image and the sample standard image into a discriminator respectively to obtain the prediction image features and the sample image features;
[0147] Input the prediction image features and the sample image features and the standard image semantic information into a feature fusion module respectively to obtain the prediction fusion image features and the sample fusion image features;
[0148] Input the prediction fusion image features and the sample fusion image features into the discriminator respectively to obtain the prediction discrimination probability and the standard discrimination probability output by the discriminator.
[0149] In this implementation, input the target prediction image and the sample standard image into the discriminator respectively to obtain the prediction image features and the sample image features. At the same time, in the method provided in the embodiments of this specification, an attention-based feature fusion module (AFF) is proposed. The attention-based feature fusion module is used to fuse the features in the discriminator and the standard image semantic information, and then output the fused features through the discriminator to obtain the prediction discrimination probability and the standard discrimination probability output by the discriminator. Thus, semantic-guided image discrimination information is realized.
[0150] See Figure 5 , Figure 5 shows a schematic diagram of calculating the loss value of the model provided in an embodiment of this specification. As Figure 5 shown, after the sample noisy image is processed by the reference image denoising model, a target prediction image is generated, and the target prediction image and the sample standard image calculate the image loss value.
[0151] In addition, input the sample noisy image, the target prediction image, and the sample standard image into a feature extraction network, construct positive and negative sample pairs at the pixel level in the feature extraction network, and calculate the contrast loss value according to the positive and negative sample pairs.
[0152] Input the sample standard image into the feature extraction network for feature extraction to obtain the standard image semantic information a of different scales l , a m , a h . Input the target prediction image and the sample standard image into the discriminator respectively to obtain the prediction image features and the sample image features output by the discriminator. If the input is the target prediction image, then obtain the prediction image features f l , f m , fn If the input is a sample standard image, the sample image feature f is obtained. l 、f m 、f n 。
[0153] The standard image semantic information and the predicted image feature, or the standard image semantic information and the sample image feature are input into the attention-based feature fusion module (AFF) to obtain the predicted fusion image feature and the sample fusion image feature. Then, the predicted fusion image feature and the sample fusion image feature are input into the discriminator to obtain the predicted discrimination probability and the standard discrimination probability output by the discriminator, and then the discrimination loss value is calculated according to the predicted discrimination probability and the standard discrimination probability.
[0154] See Figure 6 , Figure 6 which shows a schematic structural diagram of the attention-based feature fusion module provided by an embodiment of this specification. As Figure 6 shown, a * and f * are input into the AFF, and after being processed by the structure of the AFF as in Figure 6 , the fusion feature is obtained, where a * is the standard image semantic information, and f * is the predicted image feature or the sample image feature.
[0155] As Figure 6 shown, GroupNorm is a normalization technique used to stabilize the training process and improve the generalization ability of the model. GroupNorm divides the channels into several groups and normalizes the data within each group, thereby reducing the impact of batch size changes on model training.
[0156] ProjectionLayer is used to map feature information from high dimension to low dimension, or from low dimension to high dimension. Permute is used to reorder the dimensions of the tensor. LayerNorm normalizes the features by dimension. GeLu is an activation function.
[0157] Through the method provided by the embodiments of this specification, a noise pattern-aware image denoising model is provided. First, noise pattern information is provided according to the reference image, and then the denoising process is guided according to the pre-trained prior knowledge, making the denoising of the noisy image more robust and having better generalization. When performing the guiding strategy of the noise pattern information, the multi-scale quantization autoencoder is aligned with the image denoising model, thereby improving the modeling of the denoising model.
[0158] In the process of further training the image denoising model, a discriminator combined with semantic perception is used to achieve fine-grained semantic perception image denoising. At the same time, a generative adversarial learning network is combined to maintain consistency in image segmentation through positive and negative samples, and double negative samples are used to reduce noise and artifacts.
[0159] Corresponding to the above method embodiments, this specification also provides embodiments of a training device for an image denoising model. Figure 7 The structural schematic diagram of a training device for an image denoising model provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0160] An acquisition module 702, configured to acquire image training samples, where the image training samples include sample standard images and sample noisy images corresponding to the sample standard images;
[0161] A first training module 704, configured to generate a sample reference image according to the sample standard image, and input the sample noisy image and the sample reference image into an initial image denoising model to obtain an initial predicted image output by the initial image denoising model, where the initial image denoising model determines noise pattern information according to the encoded feature information of the sample noisy image and the sample reference image, and denoises the sample noisy image according to the noise pattern information;
[0162] A first parameter adjustment module 706, configured to adjust the model parameters of the initial image denoising model according to the initial predicted image and the sample standard image to obtain a reference image denoising model;
[0163] A second training module 708, configured to input the sample noisy image into the reference image denoising model to obtain a target predicted image output by the reference image denoising model, where the reference image denoising model predicts noise pattern information according to the sample noisy image, and denoises the sample noisy image according to the noise pattern information;
[0164] A second parameter adjustment module 710, configured to adjust the model parameters of the reference image denoising model according to the target predicted image and the sample standard image to obtain an image denoising model.
[0165] Optionally, the acquisition module 702 is further configured to:
[0166] Acquire sample standard images;
[0167] Add noise to the sample standard image based on different noise pattern information to obtain sample noisy images corresponding to different noise pattern information;
[0168] Construct an image training sample based on the sample standard image and the sample noisy image.
[0169] Optionally, the first training module 704 is further configured to:
[0170] Determine an initial sample reference image based on the sample standard image;
[0171] Perform data augmentation on the initial sample reference image to obtain a sample reference image.
[0172] Optionally, the initial image denoising model includes an encoder, a noise pattern information extraction module, and a decoder;
[0173] The first training module 704 is further configured to:
[0174] Input the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features that correspond one-to-one to the multi-scale noisy image features;
[0175] Input the multi-scale noisy image features and the multi-scale reference image features into the noise pattern information extraction module in a corresponding manner to obtain the noise pattern information corresponding to each scale of the noisy image features;
[0176] Input the sample noisy image and each piece of noise pattern information into the decoder to obtain an initial predicted image output by the decoder.
[0177] Optionally, the encoder is a multi-scale quantization autoencoder;
[0178] The first training module 704 is further configured to:
[0179] Obtain the scale number information corresponding to the multi-scale quantization autoencoder;
[0180] Input the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features corresponding to the scale number information.
[0181] Optionally, the first training module 704 is further configured to:
[0182] Determine target multi-scale noisy image features and target multi-scale reference image features, where the target multi-scale noisy image features and the target multi-scale reference image features correspond one-to-one;
[0183] Input the target multi-scale noisy image features and the target multi-scale reference image features into the noise pattern information extraction module to obtain the target noise pattern information corresponding to the target multi-scale noisy image features.
[0184] Optionally, the decoder includes an embedding unit, a plurality of decoding units, and an output unit, and the number of decoding units is the same as the scale number information of the multi-scale quantization auto-encoder;
[0185] The first training module 704 is further configured to:
[0186] Input the sample noisy image into the embedding layer to obtain the image features to be decoded output by the embedding layer;
[0187] Input the image features to be decoded output by the previous processing unit and the target noise pattern information into the target decoding unit to obtain the image features to be decoded output by the target decoder, where the previous processing unit is an embedding unit or a decoding unit, and the target noise pattern information and the target decoding unit are in one-to-one correspondence;
[0188] Input the image features to be decoded output by the last decoding unit into the output unit to obtain the initial predicted image.
[0189] Optionally, the reference image denoising model includes an encoder, a noise pattern information extraction module, and a decoder, and the encoder is used to generate simulated reference image features;
[0190] Inputting the sample noisy image into the reference image denoising model to obtain the target predicted image output by the reference image denoising model includes:
[0191] Input the sample noisy image into the encoder to obtain the multi-scale noisy image features output by the encoder and the multi-scale simulated reference image features corresponding to the multi-scale noisy image features;
[0192] Input the multi-scale noisy image features and the multi-scale simulated reference image features into the noise pattern information extraction module correspondingly to obtain the noise pattern information corresponding to each scale of noisy image features;
[0193] Input the sample noisy image and each noise pattern information into the decoder to obtain the target predicted image output by the decoder.
[0194] Optionally, the second training module 708 is further configured to:
[0195] Calculate an image loss value, a contrast loss value, and a discriminant loss value according to the sample noisy image, the target predicted image, and the sample standard image;
[0196] Adjust the model parameters of the reference image denoising model according to the image loss value, the contrast loss value, and the discriminant loss value;
[0197] Continue to train the reference image denoising model until the model training stop condition is reached, and obtain an image denoising model.
[0198] Optionally, the second training module 708 is further configured to:
[0199] Calculate an image loss value according to the target prediction image and the sample standard image;
[0200] Construct positive and negative sample pairs at the pixel granularity according to the sample noisy image, the target prediction image, and the sample standard image, and calculate a contrast loss value according to the positive and negative sample pairs;
[0201] Obtain the semantic information of the standard image according to the sample standard image, and generate the prediction discrimination probability and the standard discrimination probability of the target prediction image and the sample standard image based on the semantic information of the standard image, and calculate the discrimination loss value through the prediction discrimination probability and the standard discrimination probability.
[0202] Optionally, the second training module 708 is further configured to:
[0203] Input the sample noisy image, the target prediction image, and the sample construction image into the feature extraction model to obtain the positive and negative sample pairs output by the feature extraction model. Among them, the feature extraction model uses the feature of the target pixel point in the target prediction image as the sample pixel feature, uses the feature of the target pixel point in the sample standard image as the sample pixel positive label, and uses the feature of any pixel point in the sample noisy image and the sample standard image as the sample pixel negative label;
[0204] Form positive and negative sample pairs according to the sample pixel feature, the sample pixel positive label, and the sample pixel negative label of the target pixel point, and calculate a contrast loss value according to the positive and negative sample pairs.
[0205] Optionally, the second training module 708 is further configured to:
[0206] Input the target prediction image and the sample standard image into the discriminator respectively to obtain the prediction image feature and the sample image feature;
[0207] Input the prediction image feature and the sample image feature and the semantic information of the standard image into the feature fusion module respectively to obtain the prediction fusion image feature and the sample fusion image feature;
[0208] Input the prediction fusion image feature and the sample fusion image feature into the discriminator respectively to obtain the prediction discrimination probability and the standard discrimination probability output by the discriminator.
[0209] The above is a schematic solution of a training device for an image denoising model according to this embodiment. It should be noted that the technical solution of the training device for the image denoising model belongs to the same concept as the technical solution of the above-mentioned image denoising model training method. For the details not described in the technical solution of the training device for the image denoising model, reference can be made to the description of the technical solution of the above-mentioned image denoising model training method.
[0210] See Figure 8 , Figure 8 FIG. shows an architecture diagram of a training system for an image denoising model provided by an embodiment of this specification. The training system for the image denoising model may include a client 100 and a server 200;
[0211] The client 100 is used to send an image training sample to the server 200, where the image training sample includes a sample standard image and a sample noisy image corresponding to the sample standard image;
[0212] The server 200 is used to generate a sample reference image according to the sample standard image, and input the sample noisy image and the sample reference image into an initial image denoising model to obtain an initial predicted image output by the initial image denoising model. The initial image denoising model determines noise pattern information according to the encoded feature information of the sample noisy image and the sample reference image, and denoises the sample noisy image according to the noise pattern information; adjusts the model parameters of the initial image denoising model according to the initial predicted image and the sample standard image to obtain a reference image denoising model; inputs the sample noisy image into the reference image denoising model to obtain a target predicted image output by the reference image denoising model. The reference image denoising model predicts noise pattern information according to the sample noisy image, and denoises the sample noisy image according to the noise pattern information; adjusts the model parameters of the reference image denoising model according to the target predicted image and the sample standard image to obtain an image denoising model; sends the model parameters of the image denoising model to the client 100;
[0213] The client 100 is further used to receive the model parameters sent by the server 200 and construct an image denoising model according to the model parameters.
[0214] The training system for the image denoising model may include multiple clients 100 and a server 200. Among them, the client 100 may be referred to as an edge device, and the server 200 may be referred to as a cloud device. Communication connections can be established between multiple clients 100 through the server 200. In the training scenario of the image denoising model, the server 200 is used to provide training services for the image denoising model between multiple clients 100. Multiple clients 100 can be used as senders or receivers respectively to achieve communication through the server 200.
[0215] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the training scenario of the image denoising model, it can be that the user publishes a data stream to the server 200 through the client 100, and the server 200 generates the model parameters of the image denoising model according to the data stream, and pushes the model parameters of the image denoising model to other clients that have established communication.
[0216] Among them, a connection is established between the client 100 and the server 200 through a network. The network provides the medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the server 200.
[0217] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini program, a lightweight application program), or a cloud application, etc. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in a computing device and needs to rely on the device or certain APPs in the device to run, etc. The computing device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the computing device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0218] The server 200 may include servers that provide various services, such as a server that provides communication services for multiple clients, or a server for background training that supports models used on the client, or a server that processes data sent by the client, and so on. It should be noted that the server 200 may be implemented as a distributed server cluster composed of multiple servers, or may be implemented as a single server. The server may also be a server of a distributed system, or a server combined with a blockchain. The server may also be a cloud server such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN, Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0219] It is worth noting that the training method of the image denoising model provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have a similar function to the server, so as to execute the training method of the image denoising model provided in the embodiments of this specification. In other embodiments, the training method of the image denoising model provided in the embodiments of this specification may also be jointly executed by the client and the server.
[0220] See Figure 9 , Figure 9 shows a flowchart of an image processing method provided in an embodiment of this specification, which specifically includes the following steps:
[0221] Step 902: Receive an image processing task, where the image processing task carries multiple images to be processed corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.
[0222] Step 904: Input the multiple images to be processed into the image denoising model to obtain multiple target images output by the image denoising model, where the image denoising model is trained by the above-mentioned training method of the image denoising model.
[0223] Step 906: Determine whether there is an abnormal object in the target detection area according to the multiple target images.
[0224] See Figure 10 , Figure 10 shows a flowchart of a CT image processing method provided in an embodiment of this specification, which specifically includes the following steps:
[0225] Step 1002: Receive a CT image processing task, where the CT image processing task carries multiple low-dose CT images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area.
[0226] Step 1004: Input the multiple low-dose CT images into a CT image denoising model to obtain multiple denoised CT images output by the CT image denoising model, where the CT image denoising model is trained by the above-mentioned image denoising model training method.
[0227] Step 1006: Determine whether there is an abnormal object in the target detection area according to the multiple denoised CT images.
[0228] See Figure 11 , Figure 11 shows a system schematic diagram of an image processing system provided in an embodiment of this specification. The image processing system includes a client 1102 and a server 1104, where
[0229] The client 1102 is used to send a CT image processing task to the server, where the CT image processing task carries multiple low-dose CT images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area;
[0230] The server 1104 is used to input the multiple low-dose CT images into a CT image denoising model to obtain multiple denoised CT images output by the CT image denoising model, where the CT image denoising model is trained by the above-mentioned image denoising model training method; determine whether there is an abnormal object in the target detection area according to the multiple denoised CT images.
[0231] Figure 12 shows a structural block diagram of a computing device 1200 provided in an embodiment of the present application. The components of the computing device 1200 include but are not limited to a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 through a bus 1230, and a database 1250 is used to store data.
[0232] The computing device 1200 further includes an access device 1240, which enables the computing device 1200 to communicate via one or more networks 1260. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1240 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0233] In one embodiment of the present application, the above components of the computing device 1200 and Figure 12 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 12 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0234] The computing device 1200 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1200 can also be a mobile or stationary server.
[0235] Wherein, the processor 1220 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above image denoising model training method, image processing method, and CT image processing method are implemented.
[0236] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above-mentioned methods for training an image denoising model, image processing method, and CT image processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above-mentioned methods for training an image denoising model, image processing method, and CT image processing method.
[0237] An embodiment of this specification also provides a computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the above-mentioned methods for training an image denoising model, image processing method, and CT image processing method.
[0238] The embodiments in this specification are all described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computer-readable storage medium, since it is basically similar to the embodiments of the methods for training an image denoising model, image processing method, and CT image processing method, the description is relatively simple. For the relevant parts, reference can be made to the partial descriptions of the embodiments of the methods for training an image denoising model, image processing method, and CT image processing method.
[0239] An embodiment of this specification also provides a computer program product including computer programs / instructions, which, when executed by a processor, implement the steps of the above-mentioned methods for training an image denoising model, image processing method, and CT image processing method.
[0240] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above-mentioned methods for training an image denoising model, image processing method, and CT image processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above-mentioned methods for training an image denoising model, image processing method, and CT image processing method.
[0241] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0242] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0243] It should be noted that the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0244] In the above embodiments, the descriptions of the various embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0245] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A training method for an image denoising model, comprising: Obtaining image training samples, wherein the image training samples include sample standard images and sample noisy images corresponding to the sample standard images; Generating a sample reference image according to the sample standard image, and inputting the sample noisy image and the sample reference image into an initial image denoising model to obtain an initial prediction image output by the initial image denoising model, wherein the initial image denoising model determines noise pattern information according to the encoded feature information of the sample noisy image and the sample reference image, and denoises the sample noisy image according to the noise pattern information; Adjusting the model parameters of the initial image denoising model according to the initial prediction image and the sample standard image to obtain a reference image denoising model; Inputting the sample noisy image into the reference image denoising model to obtain a target prediction image output by the reference image denoising model, wherein the reference image denoising model predicts noise pattern information according to the sample noisy image and denoises the sample noisy image according to the noise pattern information; Adjusting the model parameters of the reference image denoising model according to the target prediction image and the sample standard image to obtain an image denoising model.
2. The method according to claim 1, wherein the image training samples are obtained through the following steps: Obtaining sample standard images; Adding noise to the sample standard image based on different noise pattern information to obtain sample noisy images corresponding to different noise pattern information; Constructing image training samples according to the sample standard images and the sample noisy images.
3. The method according to claim 1, wherein generating a sample reference image according to the sample standard image comprises: Determining an initial sample reference image according to the sample standard image; Performing data augmentation on the initial sample reference image to obtain a sample reference image.
4. The method according to claim 1, wherein the initial image denoising model comprises an encoder, a noise pattern information extraction module, and a decoder; Inputting the sample noisy image and the sample reference image into the initial image denoising model to obtain an initial prediction image output by the initial image denoising model, comprising: Inputting the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features corresponding one-to-one to the multi-scale noisy image features; Inputting the multi-scale noisy image features and the multi-scale reference image features correspondingly into the noise pattern information extraction module to obtain noise pattern information corresponding to each scale of noisy image features; Inputting the sample noisy image and each noise pattern information into the decoder to obtain an initial prediction image output by the decoder.
5. The method according to claim 4, wherein the encoder is a multi-scale quantization autoencoder; Inputting the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features corresponding one-to-one to the multi-scale noisy image features, comprising: Obtaining scale number information corresponding to the multi-scale quantization autoencoder; Input the sample noisy image and the sample reference image into the encoder to obtain multi-scale noisy image features and multi-scale reference image features corresponding to the scale number information.
6. The method according to claim 4, wherein inputting the multi-scale noisy image features and the multi-scale reference image features into the noise pattern information extraction module to obtain the noise pattern information corresponding to the noisy image features at each scale, includes: Determine the target multi-scale noisy image features and the target multi-scale reference image features, where the target multi-scale noisy image features and the target multi-scale reference image features correspond one by one; Input the target multi-scale noisy image features and the target multi-scale reference image features into the noise pattern information extraction module to obtain the target noise pattern information corresponding to the target multi-scale noisy image features.
7. The method according to claim 6, wherein the decoder includes an embedding unit, a plurality of decoding units, and an output unit, and the number of decoding units is the same as the scale number information of the multi-scale quantization autoencoder; Input the sample noisy image and each noise pattern information into the decoder to obtain the initial predicted image output by the decoder, including: Input the sample noisy image into the embedding layer to obtain the image features to be decoded output by the embedding layer; Input the image features to be decoded output by the previous processing unit and the target noise pattern information into the target decoding unit to obtain the image features to be decoded output by the target decoder, where the previous processing unit is the embedding unit or the decoding unit, and the target noise pattern information and the target decoding unit correspond one by one; Input the image features to be decoded output by the last decoding unit into the output unit to obtain the initial predicted image.
8. The method according to claim 1, wherein the reference image denoising model includes an encoder, a noise pattern information extraction module, and a decoder, and the encoder is used to generate simulated reference image features; Input the sample noisy image into the reference image denoising model to obtain the target predicted image output by the reference image denoising model, including: Input the sample noisy image into the encoder to obtain the multi-scale noisy image features output by the encoder and the multi-scale simulated reference image features corresponding one by one to the multi-scale noisy image features; Input the multi-scale noisy image features and the multi-scale simulated reference image features into the noise pattern information extraction module to obtain the noise pattern information corresponding to the noisy image features at each scale; Input the sample noisy image and each noise pattern information into the decoder to obtain the target predicted image output by the decoder.
9. The method according to claim 1, adjusting the model parameters of the reference image denoising model according to the target predicted image and the sample standard image to obtain an image denoising model, including: Calculate an image loss value, a contrast loss value, and a discriminant loss value according to the sample noisy image, the target predicted image, and the sample standard image; Adjust the model parameters of the reference image denoising model according to the image loss value, the contrast loss value, and the discriminant loss value; Continue to train the reference image denoising model until the model training stop condition is reached, and obtain an image denoising model.
10. The method according to claim 9, calculating an image loss value, a contrast loss value, and a discrimination loss value according to the sample noisy image, the target prediction image, and the sample standard image, including: Calculating an image loss value according to the target prediction image and the sample standard image; Constructing positive and negative sample pairs at the pixel granularity according to the sample noisy image, the target prediction image, and the sample standard image, and calculating a contrast loss value according to the positive and negative sample pairs; Obtaining standard image semantic information according to the sample standard image, generating prediction discrimination probabilities and standard discrimination probabilities of the target prediction image and the sample standard image based on the standard image semantic information, and calculating a discrimination loss value through the prediction discrimination probability and the standard discrimination probability.
11. The method according to claim 10, constructing positive and negative sample pairs at the pixel granularity according to the sample noisy image, the target prediction image, and the sample standard image, and calculating a contrast loss value according to the positive and negative sample pairs, including: Inputting the sample noisy image, the target prediction image, and the sample construction image into a feature extraction model to obtain positive and negative sample pairs output by the feature extraction model, where the feature extraction model uses the feature of the target pixel point in the target prediction image as the sample pixel feature, uses the feature of the target pixel point in the sample standard image as the sample pixel positive label, and uses the feature of any pixel point in the sample noisy image and the sample standard image as the sample pixel negative label; Forming positive and negative sample pairs according to the sample pixel feature, the sample pixel positive label, and the sample pixel negative label of the target pixel point, and calculating a contrast loss value according to the positive and negative sample pairs.
12. The method according to claim 10, generating prediction discrimination probabilities and standard discrimination probabilities of the target prediction image and the sample standard image based on the standard image semantic information, including: Inputting the target prediction image and the sample standard image into a discriminator respectively to obtain a prediction image feature and a sample image feature; Inputting the prediction image feature and the sample image feature and the standard image semantic information into a feature fusion module respectively to obtain a prediction fusion image feature and a sample fusion image feature; Inputting the prediction fusion image feature and the sample fusion image feature into a discriminator respectively to obtain a prediction discrimination probability and a standard discrimination probability output by the discriminator.
13. An image processing method, including: Receiving an image processing task, where the image processing task carries multiple images to be processed corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormal object in the target detection area; Inputting the multiple images to be processed into an image denoising model to obtain multiple target images output by the image denoising model, where the image denoising model is trained by the training method according to any one of claims 1-12; Determining whether there is an abnormal object in the target detection area according to the multiple target images.
14. A CT image processing method, comprising: Receiving a CT image processing task, wherein the CT image processing task carries a plurality of low-dose CT images corresponding to a target detection region, and the image processing task is used to detect whether there is an abnormal object in the target detection region; Inputting the plurality of low-dose CT images into a CT image denoising model to obtain a plurality of denoised CT images output by the CT image denoising model, wherein the CT image denoising model is trained by the training method according to any one of claims 1-12; Determining whether there is an abnormal object in the target detection region according to the plurality of denoised CT images.
15. An image processing system, comprising a client and a server, wherein, The client is configured to send a CT image processing task to the server, wherein the CT image processing task carries a plurality of low-dose CT images corresponding to a target detection region, and the image processing task is used to detect whether there is an abnormal object in the target detection region; The server is configured to input the plurality of low-dose CT images into a CT image denoising model to obtain a plurality of denoised CT images output by the CT image denoising model, wherein the CT image denoising model is trained by the training method according to any one of claims 1-12; and determining whether there is an abnormal object in the target detection region according to the plurality of denoised CT images.
16. A computer-readable storage medium storing computer programs / instructions, which when executed by a processor, implement the steps of the method according to any one of claims 1 to 14.
17. A computer program product comprising computer programs / instructions, which when executed by a processor, implement the steps of the method according to any one of claims 1 to 14.
Citation Information
Cited By
Personalized household appliance product design method and device based on visual autoregression model
CN120495593A
Personalized home appliance product design method and device based on visual autoregressive model
CN120495593B