Image depth neighbor downsampling method, unsupervised learning denoising network model acquisition method and denoising method

Through the image depth nearest neighbor downsampling method and the DDPM generation model, combined with Noise2noise loss and regular terms, an unsupervised denoising deep neural network model was established, which solved the problem of inadequate calculation efficiency and performance in image denoising, and improved training efficiency and denoising performance.

CN120236162APending Publication Date: 2025-07-01SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235932.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art has problems that both computational efficiency and high performance cannot be taken into account in the image denoising process, and there is an imbalance in training efficiency and denoising performance of the unsupervised denoising model.

Method used

The image depth nearest neighbor downsampling method is used to obtain the image downsampling tag pair through window slices, sub-window selection, block pixel averaging and random nearest neighbor sampling. At the same time, a large number of training images are generated using the DDPM generation model, and a deep neural network model framework with unsupervised denoising is established through the combination of Noise2noise loss and regular terms.

Benefits of technology

The training efficiency and denoising performance of the unsupervised denoising model are improved, and the problem of unbalanced training efficiency and denoising performance is alleviated. It is suitable for image denoising tasks under hardware equipment limitations and computing resource constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236162A_ABST
    Figure CN120236162A_ABST
Patent Text Reader

Abstract

The invention relates to an image depth neighbor down-sampling method, an unsupervised learning denoising network model acquisition method and a denoising method. The down-sampling method comprises the following steps: determining a slicing strategy, acquiring a block window, obtaining a sub-window, dividing the sub-window into a plurality of blocks, averaging the pixel value of each block, and obtaining the average pixel value of the plurality of blocks; recombining the average pixel value of the plurality of blocks according to the relative position information of the plurality of blocks in the sub-window to obtain a pixel set; in the pixel set, adopting a random neighbor sampling strategy to randomly select two neighbor pixels, and traversing the whole image to obtain a plurality of groups of two neighbor pixels; two pixels of a plurality of groups of neighbors are converted into a down-sampling tag pair of an approximate original image. According to the method, the DDPM is adopted to generate the model, the problem that the number of training data sets is insufficient is solved, and the problem that the training efficiency and the denoising performance of an unsupervised denoising model are unbalanced is solved through an image depth neighbor downsampling method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and machine learning, and in particular to an image deep nearest neighbor downsampling method, a method for obtaining an unsupervised learning denoising network model, and a denoising method. Background Art

[0002] In today's digital age, images, as an important information carrier, play a crucial role in numerous fields. However, during the processes of image generation, acquisition, transmission, and storage, images are inevitably disturbed by noise, resulting in a decline in image quality.

[0003] Currently, traditional image denoising methods such as designing filters or establishing noise models have problems such as blurred outputs, the need for manual debugging and parameter settings, or the inability to balance computational efficiency and high performance, and it is difficult to improve denoising performance by tuning parameters. Compared with traditional methods, deep learning-based methods have gradually shown their advantages in the field of image denoising, as they do not require complex noise modeling and cumbersome manual parameter tuning. However, in real scenarios, collecting a large number of paired noisy-clean training data pairs is extremely challenging and expensive, which limits the practical application of supervised denoising schemes.

[0004] Unsupervised methods have gained the favor of researchers because they do not require any clean image pairs to complete model training. However, due to hardware device limitations and computing power resource constraints, there is an imbalance problem in the training efficiency and denoising performance of unsupervised denoising models. Summary of the Invention

[0005] To solve the above defects, the present invention proposes an image deep nearest neighbor downsampling method, a method for obtaining an unsupervised learning denoising network model, and a denoising method.

[0006] The technical solution adopted by the present invention is an image deep nearest neighbor downsampling method, and the method includes:

[0007] (a) Determining a slicing strategy according to the feature information of the original image;

[0008] (b) Performing window slicing on the original image according to the slicing strategy to obtain block windows, and traversing the entire image;

[0009] (c) Selecting the central region of the block window to obtain a sub-window;

[0010] (d) Dividing the sub-window into several blocks, averaging the pixel values of each block to obtain the average pixel values of several blocks;

[0011] (e) Recombining the average pixel values of several blocks according to the relative position information of the several blocks in the sub-window to obtain a pixel set;

[0012] (f), in the pixel set, adopt a random nearest neighbor sampling strategy, randomly select two nearest neighbor pixels, and traverse the entire image to obtain several groups of two nearest neighbor pixels;

[0013] (g), several groups of two nearest neighbor pixels are transformed into downsampled label pairs of an approximate original image.

[0014] Furthermore, the slicing strategy described in (a) specifically includes:

[0015] Crop the original image into N regions according to the difference in information complexity, where N > 0;

[0016] Determine the size of the block window according to the size information and / or information complexity of each region.

[0017] Furthermore, it is characterized in that

[0018] The size of the block window increases as the size of each region increases;

[0019] The size of the block window increases as the information complexity of each region decreases.

[0020] The present invention also discloses a method for obtaining an unsupervised learning denoising network model, and the method includes:

[0021] S100, obtain a large number of training images;

[0022] S200, establish a framework of a deep neural network model for unsupervised denoising according to the Neighbor2Neighbor work;

[0023] S300, downsample the training images by using the above image deep nearest neighbor downsampling method, and obtain label pairs;

[0024] S400, input the label pairs obtained in S300 into the network model framework in S200 for training to obtain a trained denoising network model.

[0025] Furthermore, S100 specifically includes: obtaining a large number of training images through the generative model DDPM.

[0026] Furthermore, the obtaining of a large number of training images through the generative model DDPM specifically includes:

[0027] S110, train the DDPM generative model by inputting the original training images;

[0028] S120, use the DDPM generative model trained in S110 to generate a large number of training images.

[0029] Further, after S120, the following steps are also included:

[0030] S130. Perform data augmentation on the obtained training images to obtain more training images; the data augmentation can adopt one or several means of rotation, cropping and recombination, scaling, brightness adjustment, blurring or contrast adjustment.

[0031] Further, in S200, establishing a deep neural network model framework based on unsupervised denoising includes: determining the loss term as the Noise2noise loss and the regularization term.

[0032] Further, based on the determined loss term Noise2noise loss and the regularization term, the loss function is:

[0033]

[0034] where g1(y) and g2(y) are obtained by performing two downsamplings on the image to be denoised. This part is the Noise2Noise loss, that is, L rec , and the second part L reg is constrained by the regularization term, and γ is the loss coefficient.

[0035] The present invention also discloses an unsupervised learning denoising method, which inputs the image to be denoised into the above-obtained trained unsupervised learning denoising network model to obtain the denoised image.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. In the image depth nearest neighbor downsampling method of the present invention, from the original image through window slicing, sub-window selection to block pixel averaging, it is a process of gradually refining information. Through the depth averaging operation, the information of multiple pixels is condensed into the average pixel value of a block, reducing the data dimension. At the same time, random nearest neighbor sampling selects a part of the nearest neighbor pixels from a large number of pixel sets, further reducing the data volume. This is very beneficial for the training of the unsupervised denoising model under the hardware device limitations and computing power resource constraints, which can accelerate the training speed and improve the training efficiency. Through the depth averaging operation, more representative eigenvalue is extracted from the pixel information of the original image, retaining the important features of the image while reducing the data volume; random nearest neighbor sampling utilizes the spatial correlation of pixels to capture the local structural features of the image. The retention of these features helps the unsupervised denoising model to better learn the inherent pattern of the image, so as to more accurately restore the original information of the image during the denoising process, improving the denoising performance and alleviating the problem of imbalance between training efficiency and denoising performance.

[0038] 2. The image depth nearest neighbor downsampling method in the present invention not only meets the requirements of the unsupervised denoising model for adjacent pixels and very similar appearances of paired images, but also the depth nearest neighbor downsampling discards some redundant information, improves the algorithm efficiency, and is expected to play a more significant role in high-resolution biological microscopy imaging.

[0039] 3. The DDPM generation model can generate a large number of training images, effectively expanding the data volume. Especially in some fields where it is difficult to obtain enough real data, such as medical images of rare diseases, industrial inspection images of special scenarios, etc., DDPM can provide additional training data and improve the training effect of the model. The images generated by DDPM are diverse. It can simulate the distribution of real data to a certain extent and generate images containing various different features. When training an image denoising model, diverse training images can enable the model to learn a wider combination of noise and image features, thereby improving the generalization ability of the model and enabling it to have better denoising performance when facing different types of noise and various complex images. When collecting real data, there may be problems with data bias. For example, some types of images are collected more, while other types of images are collected less. Using DDPM to generate training images can compensate for this bias to a certain extent. By generating various types of images, the training data becomes more balanced, which helps the model learn more comprehensive image features and improve the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention will be described in detail below in conjunction with the embodiments and the drawings, where:

[0041] Figure 1 is a flowchart of an image depth nearest neighbor downsampling method;

[0042] Figure 2 is a flowchart of a method for obtaining an unsupervised learning denoising network model;

[0043] Figure 3 is a flowchart of obtaining a large number of training images through the generation model DDPM;

[0044] Figure 4 is an example schematic diagram of the image depth nearest neighbor downsampling method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below in conjunction with the drawings. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar components or components with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0046] In one embodiment, an image depth nearest neighbor downsampling method, see Figure 1 , the method includes:

[0047] (a) Determine a slicing strategy according to the feature information of the original image. Specifically, it refers to formulating the way of slicing the image based on the overall size of the original image (such as the number of pixels in the length and width of the image, etc.), the local feature distribution (such as the distribution of textures, colors, etc. in different regions of the image), etc., to provide a basis for subsequent processing.

[0048] (b) Perform window slicing on the original image according to the slicing strategy to obtain block windows and traverse the entire image. Specifically, it means dividing the original image into individual block windows according to the determined slicing strategy and performing such operations on the entire image to comprehensively process each part of the image.

[0049] (c) Select the central region of the block window to obtain a sub-window. Specifically, it refers to selecting the central region from each block window to obtain a sub-window. The central region can often better represent the characteristics of the window, so that subsequent processing can focus on the key parts.

[0050] (d) Divide the sub-window into several blocks, average the pixel values of each block, and obtain the average pixel values of several blocks. Specifically, it means further dividing each sub-window into several small blocks, calculating the average value of all pixel values within each small block, and obtaining the average pixel values of each block. This step is to initially integrate pixel information.

[0051] (e) Recombine the average pixel values of several blocks according to the relative position information of the several blocks in the sub-window to obtain a pixel set. Specifically, it refers to recombining the average pixel values of these blocks according to their relative position information in the sub-window to form a pixel set, maintaining the local spatial structure information of the image.

[0052] (f) In the pixel set, adopt a random nearest neighbor sampling strategy, randomly select two neighboring pixels, and traverse the entire image to obtain several groups of two neighboring pixels. Specifically, it means using the random nearest neighbor sampling strategy in the obtained pixel set, randomly selecting two adjacent pixels, and performing such operations on the entire image to obtain multiple groups of two neighboring pixels.

[0053] (g) Convert several groups of two neighboring pixels into approximate downsampling label pairs of the original image. Specifically, it refers to converting the obtained several groups of two neighboring pixels into approximate downsampling label pairs of the original image for subsequent operations such as model training.

[0054] In the image depth nearest neighbor downsampling method of this embodiment, from the original image through window slicing, sub-window selection to block pixel averaging, it is a process of gradually refining information. Through depth averaging operation, the information of multiple pixels is condensed into the average pixel value of a block, reducing the data dimension. At the same time, random nearest neighbor sampling selects a part of the nearest neighbor pixel pairs from a large number of pixel sets, further reducing the data volume. This is very beneficial for the training of unsupervised denoising models under hardware device limitations and computing power resource constraints, which can accelerate the training speed and improve the training efficiency.

[0055] Through depth averaging operation, more representative eigenvalue is extracted from the pixel information of the original image, retaining the important features of the image while reducing the data volume; random nearest neighbor sampling utilizes the spatial correlation of pixels to capture the local structural features of the image. The retention of these features helps the unsupervised denoising model better learn the inherent pattern of the image, so as to more accurately restore the original information of the image during the denoising process, improve the denoising performance, and alleviate the problem of imbalance between training efficiency and denoising performance.

[0056] The image depth nearest neighbor downsampling method in the present invention not only meets the requirements of the unsupervised denoising model for adjacent pixels and very similar appearance of paired images, but also the depth nearest neighbor downsampling discards some redundant information, improves the algorithm efficiency, and is expected to play a more significant role in high-resolution biological microscopy imaging.

[0057] Determining the slicing strategy according to the original image feature information enables subsequent processing to better adapt to the characteristics of the image itself, and can process different feature regions more pertinently, which helps to improve the denoising effect. During the processing, whether it is the selection of sub-windows or the recombination of block average pixel values, attention is paid to retaining the spatial structure information of the image, so that the details of the image can be better restored in applications such as denoising, avoiding the blurred output problem that may occur in traditional methods.

[0058] In one embodiment, the slicing strategy described in (a) specifically includes:

[0059] First, the original image is cropped into N regions according to the difference in information complexity, N>0. Specifically, it means that the original image is segmented into N different regions (N is greater than 0) according to the difference in its information complexity. The information complexity can be measured from aspects such as the texture, color change, and detail richness of the image. For example, the part of the image with a large amount of details and complex textures has a higher information complexity; while the relatively smooth and single-color region has a lower information complexity. In this way, different feature parts in the image are distinguished. If the information complexity of each region of the original image is similar, the original image does not need to be cropped, N = 0. If the information complexity of each region of the original image is different, it can be cropped first into regions with single information and regions with relatively complex information.

[0060] Then, determine the size of the partitioning window according to the size information of each region. Specifically, for each partitioned region, based on its own size information such as length and width and other dimensional information, determine the size of the partitioning window. It can be directly and flexibly adapted to different region sizes, without complex analysis and with a small amount of calculation, and can efficiently retain the features and details of each region. For example, for a region with high information complexity and a large size, a partitioning window with a small size can be set to more finely capture the detailed features; while for a region with low information complexity and relatively smooth, a partitioning window with a large size can be adopted to reduce the processing workload and effectively represent the features of the region. This method can better adapt to the size characteristics of the region, ensure that each region can be properly processed, and avoid information loss or overprocessing problems caused by inappropriate window sizes.

[0061] In other embodiments, the size of the partitioning window can also be determined according to the information complexity of each region. Specifically, for each partitioned region, based on its own information complexity, determine the size of the partitioning window. The information complexity of the region directly reflects the complexity of the image content. By determining the size of the partitioning window in this way, it can better fit the characteristics of the image content itself. This method can better adapt to the information complexity characteristics of the region. Whether it is a simple region or a complex region, it can be properly processed, avoiding information loss or overprocessing problems, and greatly improving the quality and efficiency of image analysis and processing.

[0062] In one embodiment, the size of the partitioning window increases as the size of each region increases; the size of the partitioning window increases as the information complexity of each region decreases. A large partitioning window is used in a region with a large size and low information complexity, reducing the number of partitions. For example, in a large-area solid-color background region, the large partitioning window significantly reduces the number of processed blocks, thereby reducing the calculation amount, shortening the processing time, and improving the efficiency of the entire image downsampling. A small partitioning window is used in a region with a small size or high information complexity to more comprehensively and carefully capture the detailed information of these regions. For example, in the processing of lesion sites in medical images, a small partitioning window can accurately extract the minute features of the lesion region, avoiding missing important information due to an overly large window, which is helpful for subsequent accurate analysis and processing of the image. It is applicable to denoising of ultra-large-size biological microscopic images and can have an adaptive optimization space with remarkable effects.

[0063] Dynamically adjusting the size of the partitioning window according to the different characteristics of the region realizes the reasonable allocation of computing resources. For regions that do not require fine processing, less computing resources are allocated; while for regions that need to be focused on, more computing resources are invested, which generally improves the resource utilization efficiency and enables the system to better complete the image downsampling task under limited resource conditions.

[0064] In one embodiment, referring to Figure 4 , taking an image with a size of 512 * 512 pixels as an example, the window size is selected as 6×6. Since 512 is not a multiple of 6, it needs to be processed into a size that is close to and divisible by 6, that is, 510 * 510 pixels. The processing method is to discard the edge part (2 pixels are discarded on each side). In this way, the size of the processed image can meet the conditions of the downsampling operation, ensuring that the subsequent downsampling process in the neural network can proceed smoothly, and avoiding errors or affecting the processing effect caused by size mismatch. At this time, the sub-window size is 4×4, and the sub-window is divided into 4 blocks of 2×2, and the effect is to solve the problem of the imbalance between the training efficiency and the denoising performance of the unsupervised denoising model.

[0065] In other embodiments, the size of the divided window can also be determined according to feature information such as the resolution of the image, the noise level of the image, and the semantic content of the image. For example, if the image resolution is high and contains rich detail information, a smaller divided window can be considered in the same area to process the image more carefully and retain more details. When the noise level of the image is high, in order to better suppress noise and retain the effective information of the image, a smaller divided window can be used in the area with larger noise.

[0066] The above method was simulated and verified in synthetic experiments and real image experiments with different noise distributions, and the experimental results confirmed that the sampling method proposed in this application effectively overcomes the balance problem between training efficiency and denoising performance.

[0067] In one embodiment, a method for obtaining an unsupervised learning denoising network model, referring to Figure 2 , the method includes:

[0068] S100. Obtain a large number of training images. The source of the images can be obtained according to the field of its application. For example, if the denoising network model is used to denoise biomedical images, medical images can be downloaded from a public image database, or images can be collected using professional equipment. For example, in the medical field, CT and MRI equipment are used to obtain human images. The collected images serve as basic data in the subsequent model training and provide an information source for the model to learn.

[0069] S200. Establish a deep neural network model framework for unsupervised denoising based on the Neighbor2Neighbor work. Referring to the working principle of Neighbor2Neighbor, construct a deep neural network model framework for unsupervised denoising. This framework usually includes multiple neural network layers, such as convolutional layers, pooling layers, deconvolutional layers, etc. The convolutional layer extracts features by sliding a convolutional kernel over the image. The pooling layer is used to reduce the size of the feature map and reduce the computational amount. The deconvolutional layer restores the size of the image during the image reconstruction stage. Each layer cooperates with each other and gradually forms the ability to denoise the image by continuously learning the features in the training images.

[0070] S300. Downsample the training image using the image depth neighbor downsampling method in the above embodiment and obtain label pairs. These label pairs are used as part of the training data to guide the training of the model.

[0071] S400. Input the label pairs obtained in S300 into the network model framework in S200 for training to obtain a trained denoising network model. During the training process, the model continuously adjusts the parameters in the network, such as the weights of the convolutional kernel, etc., according to the information contained in the label pairs to minimize the difference between the model prediction result and the label pairs. After multiple rounds of training, the model gradually learns how to remove the noise in the image, thus obtaining a trained denoising network model.

[0072] The method for obtaining an unsupervised learning denoising network model in this embodiment obtains label pairs through a unique image depth neighbor downsampling method, reducing the amount of data. Under the limitations of hardware devices and computing power resources, the training efficiency is improved. In the downsampling method, the slicing strategy and the size of the partitioning window are determined according to various features of the image, which can better adapt to the characteristics of the image itself. Different regions with different features can be processed more specifically, helping the model learn more accurate image features, thereby improving the denoising effect.

[0073] In addition, the unsupervised denoising model framework based on a deep neural network combines the powerful feature learning ability of deep learning and can automatically learn the patterns of noise and the inherent features of the image from the training images. Compared with traditional denoising methods, it does not require complex noise modeling and cumbersome manual parameter tuning and has an advantage in denoising performance. A large number of diverse training images enable the model to learn rich image features and noise patterns, making the trained denoising network model have good generalization ability and be able to handle different types of noise and image denoising tasks in various scenarios.

[0074] In a real - world scenario, the number of images available for training in public datasets is limited, especially in the biomedical field. Biomedical image acquisition is difficult and costly, so the amount of material available for deep - learning training is small, which poses high requirements for medical image denoising. Therefore, in one embodiment, S100 specifically includes: obtaining a large number of training images through the generative model DDPM. DDPM (Denoi sing Diffus ion Probabi list ic Models) is a generative model based on the diffusion process. Its core idea is to simulate a forward diffusion process that gradually adds noise to the data, and then generate new data by learning the reverse denoising process. For example, in the field of medical images, there may be a lack of sufficient real medical images of various diseases. Using the DDPM model, by inputting random noise and performing the reverse denoising operation of the model, medical images with different disease characteristics can be generated as training images. These generated images can cover various lesion conditions and image features, enriching the diversity of training data.

[0075] Through the DDPM generative model, a large number of training images can be generated based on a small number of original biomedical images, effectively expanding the amount of data and making up for the shortage of training materials in unsupervised denoising. Especially in some fields where it is difficult to obtain enough real data, such as medical images of rare diseases, industrial inspection images of special scenarios, etc., DDPM can provide additional training data and improve the training effect of the model. The images generated by DDPM are diverse. It can simulate the distribution of real data to a certain extent and generate images containing various different features. When training an image denoising model, diverse training images can enable the model to learn a wider combination of noise and image features, thereby improving the generalization ability of the model and enabling it to have better denoising performance when facing different types of noise and various complex images. When collecting real data, there may be problems with data bias. For example, some types of images are collected more, while other types are collected less. Using DDPM to generate training images can make up for this bias to a certain extent. By generating various types of images, the training data becomes more balanced, which helps the model learn more comprehensive image features and improve the performance of the model.

[0076] Further, obtaining a large number of training images through the generative model DDPM, see Figure 3 , specifically includes:

[0077] S110. Train the DDPM generation model by inputting the original training images. Collect a series of original training images that should cover as many features and variations as possible that may appear in subsequent application scenarios. Construct the basic architecture of the DDPM generation model, which includes neural network components for learning the parameters of the forward diffusion process and the reverse denoising process. Input the original training images into the initialized DDPM model and start training. During the training process, the model continuously adjusts its own parameters to minimize the difference between the generated noisy images and the real noisy images (in the forward process), and to minimize the difference between the denoised images and the original input images (in the reverse process). Through multiple rounds of iterative training, the model gradually learns the feature distribution and noise patterns of the original training images.

[0078] S120. Use the DDPM generation model trained in S110 to generate a large number of training images. Prepare a large amount of random noise that conforms to the standard normal distribution, and these noises serve as the starting points for generating new images. Input the prepared random noise into the trained DDPM generation model. The model, according to the reverse denoising process learned during the training phase, starts from pure noise and gradually removes the noise. After multiple time-step iterations, a large number of new training images are generated. These newly generated images are similar to the original training images in terms of features and distribution, but also have a certain degree of diversity because different input random noises result in different generated images.

[0079] Train the DDPM model with the original training images so that the generated images can closely meet the requirements of the actual application scenario. Input different random noises into the trained DDPM model to generate a large number of images, which greatly enriches the diversity of the training data. Generating a large number of training images can effectively solve the problem of insufficient data volume.

[0080] In one embodiment, after S120, it further includes: S130. Perform data augmentation on the obtained training images to obtain more training images; data augmentation can adopt rotation, cropping and recombination, scaling, brightness adjustment, blurring, or contrast adjustment, etc. Specifically, rotate the image at a certain angle, crop different-sized and -positioned regions from the original image, and then recombine these cropped regions into new images. Perform magnification or reduction operations on the image. Change the brightness value of the image to make it brighter or darker. Blur the image using methods such as Gaussian blur or mean blur. Increase or decrease the contrast of the image to make the bright parts brighter, the dark parts darker, or make the colors of the image more dull, etc.

[0081] Through various data augmentation means, the training data can be made more diverse, covering more image variation situations. Abundant training data can reduce the model's dependence on specific training samples and lower the risk of model overfitting. The rotation, illumination change, blurring, etc. simulated by data augmentation are similar to the situations that may be encountered during image acquisition in the real world. By performing these processes on the images, the model can be made to adapt in advance to the variations in various real scenarios, improving the robustness and accuracy of the model in practical applications.

[0082] In one embodiment, in S200, a deep neural network model framework based on unsupervised denoising is established, including: determining the loss term as Noise2noise loss and the regularization term. The core assumption of Noise2noise loss is that the noise is not completely random and there is an inherent correlation between noise images, and this correlation can be utilized to achieve denoising. Overfitting will cause the model to perform well on the training set, but its performance will drop significantly on the test set or in practical applications. The regularization term constrains the parameters of the model, enabling the model to learn more general and generalized features instead of just memorizing specific patterns in the training data, preventing the model from overfitting during the training process.

[0083] By considering both Noise2noise loss and the regularization term simultaneously, the model can converge more stably during the training process. Noise2noise loss guides the model to learn in the direction of denoising, while the regularization term restricts the complexity of the model, avoiding oscillations or instability during the training process, ensuring the smooth progress of the training process, and reducing the waste of training time and resources.

[0084] In one embodiment, based on the determined loss term Noise2noise loss and the regularization term, the loss function is:

[0085]

[0086] where g1(y) and g2(y) are obtained by performing two downsamplings on the image to be denoised. This part is the Noise2Noise loss, i.e., L rec . The second part L reg is constrained by the regularization term. γ is the loss coefficient, and the loss coefficient is used to adjust the weight of the regularization term in the entire loss function. According to different datasets and task requirements, the value of γ can be flexibly adjusted to balance the fitting ability and generalization ability of the model.

[0087] In one embodiment, in S200, a deep neural network model framework based on unsupervised denoising is established, and U-net is used as the denoising backbone network. U-net is a convolutional neural network structure with a unique symmetric U-shaped structure, mainly composed of an encoder (downsampling path) and a decoder (upsampling path). The feature maps of different stages of the encoder are concatenated with the feature maps of the corresponding stages of the decoder through skip connections in the middle. Using U-net as the backbone network of the unsupervised denoising deep neural network, its encoder and decoder fuse multi-scale features through skip connections, which can not only retain image details but also grasp the overall structure, making the denoised image more realistic and natural; the end-to-end structure adapts to the unsupervised loss function, without the need for a large number of clean samples, and has strong adaptability to images of different types and noise levels; at the same time, its computational efficiency is relatively high, and the structure can be flexibly adjusted, and the depth and number of channels can be changed according to task requirements to adapt to denoising tasks of different complexities.

[0088] In one embodiment, an unsupervised learning denoising method inputs the image to be denoised into the trained unsupervised learning denoising network model obtained in the above embodiment to obtain the denoised image. Aiming at the problem that the number of training image training sets for the unsupervised learning denoising model is limited, high-quality images generated by DDPM are used to expand the training materials, solving the problem of insufficient number of paired training data sets. At the same time, a more efficient downsampling method is adopted in the subsequent process of obtaining label pairs, solving the problem that it is difficult to balance the network training efficiency and denoising performance, and improving the training efficiency of the denoising model. This downsampling method not only meets the requirements that the image pairs are pixel-adjacent and have similar appearances, but also the depth-nearest neighbor downsampling discards some redundant information, avoiding a strong dependence on the noise distribution assumption, which effectively overcomes the balance problem between training efficiency and denoising performance.

[0089] In the description of this specification, if terms such as "Embodiment 1", "this embodiment", "in one embodiment", etc. appear, it means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention or the invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example; moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.

[0090] In the description of this specification, terms such as "connection", "installation", "fixation", "setting", "having", etc. are understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0091] In the description of this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0092] The above description of the embodiments is to enable those of ordinary skill in the art to understand and apply the technology of this case. Obviously, those who are familiar with the technology in this field can easily make various modifications to these examples and apply the general principles described here to other embodiments without creative labor. Therefore, this case is not limited to the above embodiments. For the following several types of modifications, they should all be within the protection scope of this case: ① A new technical solution implemented based on the technical solution of the present invention and combined with the existing common knowledge, and the technical effect produced by this new technical solution does not exceed the technical effect of the present invention; ② An equivalent replacement of some features of the technical solution of the present invention using well-known technologies, and the technical effect produced is the same as the technical effect of the present invention; ③ Expansion based on the technical solution of the present invention, and the substantial content of the expanded technical solution does not exceed the technical solution of the present invention; ④ An equivalent transformation made using the content of the specification and drawings of the present invention, directly or indirectly applied to other related technical fields.

Claims

1. A method for image depth neighbor downsampling, characterized in that: The method comprises: (a) Determine the slicing strategy based on the feature information of the original image; (b) performing window slicing on the original image according to the slicing strategy to obtain a block window, and traversing the entire image; (c) selecting the central area of ​​the block window to obtain a sub-window; (d) dividing the sub-window into a plurality of blocks, averaging the pixel values ​​of each block, and obtaining an average pixel value of the plurality of blocks; (e) reorganizing the average pixel values ​​of the plurality of blocks according to the relative position information of the plurality of blocks in the sub-window to obtain a pixel set; (f) In the pixel set, a random neighbor sampling strategy is adopted to randomly select two neighboring pixels, and the entire image is traversed to obtain several groups of two neighboring pixels; (g) Several groups of neighboring two pixels are transformed into downsampled label pairs that approximate the original image.

2. The downsampling method according to claim 1, characterized in that: The slicing strategy described in (a) specifically includes: The original image is cropped into N regions according to the difference in information complexity, N>0; The size of the partitioning window is determined according to the size information and / or the information complexity of each region.

3. The downsampling method according to claim 2, characterized in that: The size of the block window increases as the size of each region increases; The size of the binning window increases as the information complexity of each region decreases.

4. A method for obtaining an unsupervised learning denoising network model, characterized in that: The method comprises: S100, obtaining a large number of training images; S200, establish a deep neural network model framework for unsupervised denoising based on Neighbor2Neighbor work; S300, downsampling the training image using the image depth nearest neighbor downsampling method according to any one of claims 1 to 3, and obtaining a label pair; S400, training the network model framework of input S200 with the labels obtained in S300 to obtain a trained denoising network model.

5. The method for obtaining a model according to claim 4, characterized in that: The S100 specifically includes: obtaining a large number of training images by generating a model DDPM.

6. The method for obtaining a model according to claim 5, characterized in that: The method of obtaining a large number of training images by generating the model DDPM specifically includes: S110, training the DDPM generation model by inputting the original training image; S120 , using the DDPM generation model trained in S110 to generate a large number of training images.

7. The method for obtaining a model according to claim 6, characterized in that: The S120 further includes: S130, performing data enhancement on the obtained training images to obtain more training images; the data enhancement may be performed by one or more means of rotation, cropping and reorganization, scaling, brightness adjustment, blurring or contrast adjustment.

8. The method for obtaining a model according to any one of claims 4 to 7, characterized in that: In S200, a deep neural network model framework based on unsupervised denoising is established, including: determining the loss term as Noise2noise loss and a regularization term.

9. The method for obtaining a model according to claim 8, characterized in that: Based on the determined loss term Noise2noise loss and regularization term, the loss function is: Among them, g1(y) and g2(y) are obtained by performing two downsampling on the denoised image. This part is the Noise2Noise loss, that is, L rec , the second part L reg The regularization term is used to constrain it, and γ is the loss coefficient.

10. An unsupervised learning denoising method, characterized in that: The image to be denoised is input into the trained unsupervised learning denoising network model obtained in any one of claims 4 to 9 to obtain a denoised image.