Image noise processing network training method, image noise processing method and device

Through the image noise processing network training method, a uniform noise reference image is generated using cumulative identification and weights, which solves the problem of uneven image noise distribution and achieves the improvement of image quality.

CN119991489AActive Publication Date: 2025-05-13INTELLINDUST INFORMATION TECH (SHENZHEN) CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510109566.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

In the prior art, image noise is unevenly distributed between the moving area and the resting area, resulting in more obvious noise after the time domain enhancement, making it difficult to effectively denoise through preset parameters, affecting image quality.

Method used

An image noise processing network training method is adopted, by obtaining sample images containing motion areas, fusing images and calculating cumulative marks to determine weights, generating a reference image with uniform noise, and adjusting parameters of the image noise processing network to achieve noise uniformization processing.

Benefits of technology

This improves the uniformity of the noise distribution in the image and enhances the image quality, making the noise levels in the moving and resting areas more consistent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991489A_ABST
    Figure CN119991489A_ABST
Patent Text Reader

Abstract

The invention provides an image noise processing network training method, and an image noise processing method and device, and relates to the technical field of computer vision. The method comprises the following steps: acquiring a plurality of sample images which comprise a motion area and have noise; the motion area in each sample image is a content change area relative to the previous sample image; fusing the sample image with the previous sample image to obtain a fused image; calculating the ratio of the accumulated identifier to the total number of the sample images to obtain a weight; the cumulative identifier is in positive correlation with the number of times that the pixel position continuously belongs to the non-motion area in the adjacent sample images; calculating the weighted sum of the fused image and the true value image according to the weight, and taking the weighted sum as a reference image; the true value image and the sample image are consistent in image content and have no noise; inputting the fused image into an image noise processing network to obtain a result image; and adjusting the network parameters based on the difference between the result image and the reference image. According to the invention, the uniformity of noise distribution in the image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an image noise processing network training method, an image noise processing method and a device. Background Art

[0002] For a target image frame in a video, the camera can use a reference image frame in the video that is adjacent to the target image frame to perform temporal enhancement on the target image frame. Specifically, the reference image frame is fused with the target image frame to obtain an enhanced target image frame. The area where the content of the target image frame changes relative to the reference image frame is called a motion area, and the area where the content does not change is called a static area.

[0003] In the prior art, the camera can perform denoising on the image frame by using preset parameters. However, after time domain enhancement, the noise in the moving area will be more obvious than that in the static area, and the noise difference between the moving area and the static area is large, and the noise distribution in the enhanced target image frame is uneven. If the preset parameters are used to denoise the enhanced target image frame, that is, the same denoising intensity is used for the same image frame, it is difficult to improve the image quality. For example, if the denoising intensity for the moving area is too small, more noise will remain, or if the denoising intensity for the static area is too large, it will cause content loss.

[0004] It can be seen that how to improve the uniformity of noise distribution in images is an urgent problem to be solved. Summary of the invention

[0005] The purpose of the embodiments of the present invention is to provide an image noise processing network training method, an image noise processing method and a device to improve the uniformity of noise distribution in an image. The specific technical solution is as follows:

[0006] In a first aspect, an embodiment of the present invention provides an image noise processing network training method, the method comprising:

[0007] Acquire multiple sample images containing motion regions and having noise to obtain a sample image sequence; wherein the motion region in each sample image is: a region where the content of the sample image changes relative to the previous sample image;

[0008] For each sample image except the first sample image in the sample image sequence, the sample image is fused with a sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image;

[0009] Calculate the ratio of the cumulative identification of each pixel position in the sample image to the total number of sample images to obtain the weight of the pixel position; wherein the cumulative identification of a pixel position is positively correlated with the number of times that the pixel position continuously belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image;

[0010] According to the obtained weights, a weighted sum of the fused image and the true value image corresponding to the sample image is calculated as a reference image with uniform noise corresponding to the sample image; wherein the true value image corresponding to the sample image has the same image content as the sample image and is noise-free;

[0011] The fused image corresponding to the sample image is input into the image noise processing network of the initial structure for noise uniformization processing to obtain a result image corresponding to the sample image;

[0012] Based on the difference between the result image corresponding to the sample image and the reference image, the network parameters of the image noise processing network of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image noise processing network.

[0013] Optionally, the acquiring of a plurality of sample images containing a motion region and having noise includes:

[0014] Acquire an original image sequence including a plurality of original images of uniform size and without noise, and a motion image having a size smaller than each original image;

[0015] Using the motion image, replacing an image region of each original image that is the same size as the motion image, to obtain an image to be used corresponding to each original image; wherein, in the original image sequence, positions of the replaced image regions in any two adjacent original images are different;

[0016] Adding noise to each of the obtained images to be used respectively to obtain a plurality of sample images containing motion areas and having noise;

[0017] The true value image corresponding to each sample image is the image to be used when obtaining the sample image.

[0018] Optionally, before adding noise to each of the obtained images to be used to obtain a plurality of sample images containing motion areas and having noise, the method further includes:

[0019] Acquiring noise parameters of a sensor that collects the original image;

[0020] The step of adding noise to each of the obtained images to be used to obtain a plurality of sample images containing motion areas and having noise includes:

[0021] Based on the noise parameters, noise is added to each of the obtained images to be used, so as to obtain a plurality of sample images containing motion areas and with noise.

[0022] Optionally, using the moving image to replace an image region in each original image that is the same size as the moving image to obtain an image to be used corresponding to each original image includes:

[0023] Using the motion image, replacing an image region in each original image that has the same size as the motion image, to obtain an intermediate image corresponding to each original image;

[0024] The brightness of the intermediate image corresponding to each original image is reduced to obtain the image to be used corresponding to each original image.

[0025] Optionally, the cumulative identification of each pixel position in the first sample image in the sample image sequence is 1;

[0026] For each sample image other than the first sample image in the sample image sequence, the cumulative identifier of each pixel position in the sample image is determined by the following steps:

[0027] Determine a motion region and a non-motion region contained in a reference image corresponding to the sample image in a reference image sequence, and use them as a reference motion region and a reference non-motion region, respectively; wherein the reference image sequence is the sample image sequence or a true value image sequence composed of true value images corresponding to each of the plurality of sample images; and the position sequence of the reference image corresponding to the sample image in the reference image sequence is consistent with the position sequence of the sample image in the sample image sequence;

[0028] For each pixel position in the sample image, if the pixel position belongs to the reference motion area in the corresponding reference image, determining that the cumulative flag of the pixel position in the sample image is 1;

[0029] If the pixel position belongs to the reference non-motion area in the corresponding reference image, then 1 is added to the cumulative identification of the pixel position in the sample image before the sample image to obtain the cumulative identification of the pixel position in the sample image.

[0030] Optionally, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork;

[0031] The step of inputting the fused image corresponding to the sample image into the image noise processing network of the initial structure for noise uniformization processing to obtain the result image corresponding to the sample image comprises:

[0032] Inputting the fused image corresponding to the sample image into the encoding subnetwork to obtain a first feature vector;

[0033] Based on the first feature vector and the decoding sub-network, obtain a second feature vector;

[0034] Based on the second feature vector and the output sub-network, a result image corresponding to the sample image is obtained.

[0035] Optionally, the encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork;

[0036] The step of inputting the fused image corresponding to the sample image into the encoding subnetwork to obtain a first feature vector comprises:

[0037] Inputting the fused image corresponding to the sample image into the first convolutional sub-network to obtain a first sub-feature vector;

[0038] Inputting the first sub-feature vector into the second convolutional sub-network to obtain a second sub-feature vector;

[0039] Inputting the second sub-feature vector into the third convolutional sub-network to obtain a first feature vector;

[0040] Based on the first feature vector and the decoding sub-network, a second feature vector is obtained, including:

[0041] Inputting the first feature vector into the fourth convolutional sub-network to obtain a third sub-feature vector;

[0042] Inputting the third sub-feature vector and the second sub-feature vector into the fifth convolutional sub-network to obtain a second feature vector;

[0043] Based on the second feature vector and the output sub-network, a result image corresponding to the sample image is obtained, including:

[0044] The second feature vector and the first sub-feature vector are input into the output sub-network to obtain a result image corresponding to the sample image.

[0045] In a second aspect, an embodiment of the present invention provides an image noise processing method, the method comprising:

[0046] Get the image to be processed;

[0047] The image to be processed is input into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed; wherein the image noise processing network is trained based on the above-mentioned image noise processing network training method.

[0048] Optionally, the method further includes:

[0049] The resulting image is subjected to denoising using the preset denoising parameters to obtain an enhanced image.

[0050] Optionally, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork;

[0051] The image to be processed is input into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed, including:

[0052] Inputting the image to be processed into the encoding subnetwork to obtain a third feature vector;

[0053] Based on the third feature vector and the decoding sub-network, obtaining a fourth feature vector;

[0054] Based on the fourth eigenvector and the output subnetwork, a result image corresponding to the image to be processed is obtained.

[0055] Optionally, the encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork;

[0056] Inputting the image to be processed into the encoding subnetwork to obtain a third feature vector includes:

[0057] Inputting the image to be processed into the first convolutional sub-network to obtain a fourth sub-feature vector;

[0058] Inputting the fourth sub-feature vector into the second convolutional sub-network to obtain a fifth sub-feature vector;

[0059] Inputting the fifth sub-feature vector into the third convolutional sub-network to obtain a third feature vector;

[0060] Based on the third feature vector and the decoding sub-network, a fourth feature vector is obtained, including:

[0061] Inputting the third feature vector into the fourth convolutional sub-network to obtain a sixth sub-feature vector;

[0062] Inputting the sixth sub-feature vector and the fifth sub-feature vector into the fifth convolutional sub-network to obtain a fourth eigenvector;

[0063] Based on the fourth eigenvector and the output subnetwork, a result image corresponding to the image to be processed is obtained, including:

[0064] The fourth eigenvector and the fourth sub-eigenvector are input into the output subnetwork to obtain a result image corresponding to the image to be processed.

[0065] In a third aspect, an embodiment of the present invention provides an image noise processing network training device, the device comprising:

[0066] The first acquisition module is used to acquire a plurality of sample images containing motion regions and having noise, and obtain a sample image sequence; wherein the motion region in each sample image is: a region where the content of the sample image changes relative to the previous sample image;

[0067] a fusion module, configured to fuse, for each sample image except the first sample image in the sample image sequence, the sample image with a sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image;

[0068] A first calculation module is used to calculate the ratio of the cumulative identification of each pixel position in the sample image to the total number of sample images to obtain the weight of the pixel position; wherein the cumulative identification of a pixel position is positively correlated with the number of times that the pixel position continuously belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image;

[0069] A second calculation module is used to calculate the weighted sum of the fused image and the true value image corresponding to the sample image according to the obtained weights as a reference image with uniform noise corresponding to the sample image; wherein the true value image corresponding to the sample image has the same image content as the sample image and is noise-free;

[0070] The first input module is used to input the fused image corresponding to the sample image into the image noise processing network of the initial structure for noise uniformization processing to obtain a result image corresponding to the sample image;

[0071] The adjustment module is used to adjust the network parameters of the image noise processing network of the initial structure based on the difference between the result image corresponding to the sample image and the reference image until the preset convergence condition is reached to obtain a trained image noise processing network.

[0072] In a fourth aspect, an embodiment of the present invention provides an image noise processing device, the device comprising:

[0073] A second acquisition module, used for acquiring an image to be processed;

[0074] The second input module is used to input the image to be processed into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed; wherein the image noise processing network is trained based on the above-mentioned image noise processing network training device.

[0075] An embodiment of the present invention further provides an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0076] Memory, used to store computer programs;

[0077] The processor is used to implement the above-mentioned image noise processing network training method and image noise processing method when executing the program stored in the memory.

[0078] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the image noise processing network training method and the image noise processing method are implemented.

[0079] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the above-mentioned image noise processing network training method and image noise processing method when executed by a processor.

[0080] Beneficial effects of the embodiments of the present invention:

[0081] In an embodiment of the present invention, the sample image is noisy, and the sample image is fused with the previous sample image, and the noise in the motion area in the obtained fused image is more, and the noise in the non-motion area is less. The cumulative mark of a pixel position is positively correlated with the number of times the pixel position belongs to the non-motion area in the corresponding adjacent sample image. That is, the cumulative mark of the pixel position in the motion area in the sample image is the smallest, and the cumulative mark of the pixel position in the non-motion area will be accumulated on the basis of the cumulative mark of the previous sample image, that is, the cumulative mark of the pixel position in the non-motion area is always kept large. Since the weight is the ratio of the cumulative mark to the total number of sample images, the weight of the motion area is small, and the weight of the non-motion area is large. Therefore, by fusing the fused image with uneven noise distribution and the true value image without noise according to the weight, the more noise in the motion area can be weakened, and the less noise in the non-motion area can be maintained, so as to obtain a reference image with uniform noise, and the reference image with uniform noise is used as the target of the image noise processing network. Based on the image noise processing network, the fused image with uneven noise is processed to uniform noise to obtain the result image. The network parameters are adjusted based on the gap between the result image and the target, so that the output of the image noise processing network tends to an image with uniform noise, which can improve the uniformity of noise distribution in the image.

[0082] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0084] Figure 1 A schematic diagram of a flow chart of a first image noise processing network training method provided by an embodiment of the present invention;

[0085] Figure 2 A schematic diagram of a flow chart of a second image noise processing network training method provided by an embodiment of the present invention;

[0086] Figure 3 A schematic diagram of a flow chart of a cumulative identification determination process provided by an embodiment of the present invention;

[0087] Figure 4 A schematic diagram of the principle of the cumulative identification determination process provided by an embodiment of the present invention;

[0088] Figure 5A schematic diagram of a flow chart of a third image noise processing network training method provided by an embodiment of the present invention;

[0089] Figure 6 A schematic flow chart of a first image noise processing method provided by an embodiment of the present invention;

[0090] Figure 7 A schematic flow chart of a second image noise processing method provided by an embodiment of the present invention;

[0091] Figure 8 A schematic diagram of the structure of an image noise processing network provided by an embodiment of the present invention;

[0092] FIG9 (a) and FIG9 (b) are schematic diagrams showing the effect of an image noise processing method provided by an embodiment of the present invention;

[0093] Fig.10 A schematic diagram of the structure of an image noise processing network training device provided by an embodiment of the present invention;

[0094] Fig.11 A schematic diagram of the structure of an image noise processing device provided by an embodiment of the present invention;

[0095] Fig.12 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0096] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field based on the present invention belong to the scope of protection of the present invention.

[0097] In order to more clearly understand the embodiments of the present invention, the prior art will be described in detail below.

[0098] For images or videos captured under dim light conditions (for example, light intensity less than 100 lux) (i.e., images or videos with low illumination and strong noise), temporal enhancement is an effective and necessary technical means. Specifically, the image can be fused with a reference image with the same content to reduce the noise level, reduce the occurrence of image jitter, and improve the signal-to-noise ratio and image quality.

[0099] However, the time domain enhancement effects of the moving area and the static area in the image are very different. If the static area is fused with multiple reference images (that is, the accumulation of long sequence images), the noise can be reduced; if the moving area is superimposed with high intensity, it will cause obvious motion blur. That is to say, after time domain enhancement, the noise in the moving area will be more obvious than that in the static area (for example, the noise granularity in the moving area is much stronger than that in the static area, and the noise in the static area is lighter). The noise difference between the moving area and the static area is large, and the noise levels in different areas are obviously uneven.

[0100] In the prior art, the Image Signal Processor (ISP) in the camera can use the spatial noise reduction method to improve the image quality. Specifically, it can perform denoising on the image frame through preset parameters. If the preset parameters are used to denoise the image with uneven noise, that is, the same denoising strength is used for each area in the same image, it is difficult to improve the image quality. For example, if the denoising strength for the moving area is too small, more noise will remain, or if the denoising strength for the static area is too large, it will cause content loss. In other words, the processing difficulty is relatively large in the ISP processing link.

[0101] It can be seen that how to improve the uniformity of noise distribution in images is an urgent problem to be solved.

[0102] In order to improve the uniformity of noise distribution in an image, an embodiment of the present invention provides an image noise processing network training method, an image noise processing method and an image noise processing device.

[0103] The following first introduces an image noise processing network training method provided by an embodiment of the present invention.

[0104] The image noise processing network training method embodied in the embodiment of the present invention can be applied to an electronic device. For example, the electronic device can be a remote server or a local computer.

[0105] An image noise processing network training method provided by an embodiment of the present invention may include the following steps:

[0106] Acquire multiple sample images containing motion regions and having noise to obtain a sample image sequence; wherein the motion region in each sample image is: a region where the content of the sample image changes relative to the previous sample image;

[0107] For each sample image except the first sample image in the sample image sequence, the sample image is fused with the sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image;

[0108] The ratio of the accumulated mark of each pixel position in the sample image to the total number of sample images is calculated to obtain the weight of the pixel position; wherein the accumulated mark of a pixel position is positively correlated with the number of times the pixel position continuously belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image;

[0109] According to the obtained weights, a weighted sum of the fused image and the true value image corresponding to the sample image is calculated as a reference image with uniform noise corresponding to the sample image; wherein the true value image corresponding to the sample image has the same image content as the sample image and is noise-free;

[0110] The fused image corresponding to the sample image is input into the image noise processing network of the initial structure for noise uniformization processing to obtain a result image corresponding to the sample image;

[0111] Based on the difference between the result image corresponding to the sample image and the reference image, the network parameters of the image noise processing network of the initial structure are adjusted until the preset convergence conditions are reached, thereby obtaining a trained image noise processing network.

[0112] In an embodiment of the present invention, the sample image is noisy, and the sample image is fused with the previous sample image, and the noise in the motion area in the obtained fused image is more, and the noise in the non-motion area is less. The cumulative mark of a pixel position is positively correlated with the number of times the pixel position belongs to the non-motion area in the corresponding adjacent sample image. That is, the cumulative mark of the pixel position in the motion area in the sample image is the smallest, and the cumulative mark of the pixel position in the non-motion area will be accumulated on the basis of the cumulative mark of the previous sample image, that is, the cumulative mark of the pixel position in the non-motion area is always kept large. Since the weight is the ratio of the cumulative mark to the total number of sample images, the weight of the motion area is small, and the weight of the non-motion area is large. Therefore, by fusing the fused image with uneven noise distribution and the true value image without noise according to the weight, the more noise in the motion area can be weakened, and the less noise in the non-motion area can be maintained, so as to obtain a reference image with uniform noise, and the reference image with uniform noise is used as the target of the image noise processing network. Based on the image noise processing network, the fused image with uneven noise is processed to uniform noise to obtain the result image. The network parameters are adjusted based on the gap between the result image and the target, so that the output of the image noise processing network tends to an image with uniform noise, which can improve the uniformity of noise distribution in the image.

[0113] Combine the following Figure 1 An image noise processing network training method provided by an embodiment of the present invention is introduced. Figure 1As shown, the method may include steps S101-S106.

[0114] S101, acquiring a plurality of sample images containing motion areas and having noise, to obtain a sample image sequence.

[0115] The motion region in each sample image is a region where the content of the sample image changes relative to the previous sample image.

[0116] It is understandable that before the captured video image frames are processed using ISP, the video image frames may be subjected to noise uniformity processing, and an embodiment of the present invention provides an image noise processing network training method to train a network capable of performing noise uniformity processing on video image frames. In order to simulate video image frames with uneven noise in motion areas and non-motion areas for network training, the electronic device may obtain multiple sample images containing motion areas and with noise to obtain a sample image sequence, and the motion area in each sample image is: an area where the content of the sample image changes relative to the previous sample image.

[0117] If the signal-to-noise ratio of an image is less than a first preset threshold, the image is considered to be a noisy image, and if the signal-to-noise ratio of an image is greater than a second preset threshold, the image is considered to be a noise-free image; wherein the first preset threshold is less than the second preset threshold, for example, the first preset threshold is 20 decibels, and the second preset threshold is 55 decibels. Each sample image is a noisy image, and the signal-to-noise ratio of each sample image may be less than the first preset threshold. The sizes of each sample image may be consistent for subsequent fusion processing. There may be no motion area in the first sample image in the sample image sequence, but there is a motion area in the second sample image relative to the first sample image.

[0118] Optionally, in one implementation, the image acquisition device can acquire multiple images containing motion areas and without noise under good acquisition conditions, and the electronic device can add noise to the acquired multiple images to obtain a sample image sequence. Exemplarily, the staff can adjust the analog gain of the image sensor of the image acquisition device to the lowest analog gain in advance, and the image acquisition device acquires sample videos under normal lighting conditions (for example, normal lighting conditions are 300-500 lux of light intensity) according to a moderate exposure time (for example, a moderate exposure time is 1 / 500th of a second to 1 / 120th of a second). When the analog gain of the image sensor is the lowest analog gain, the amplification factor of the electrical signal acquired by the image sensor is the smallest, and the introduced noise is also relatively small; the noise of the image acquired under normal lighting conditions is less than the noise of the image acquired under dark light conditions (for example, dark light conditions are 1-100 lux of light intensity); when acquiring sample videos, acquiring according to a moderate exposure time can avoid the occurrence of unclear image content caused by overexposure or underexposure of image brightness. Among them, normal lighting conditions can be regarded as conditions of moderate brightness. Moderate brightness and exposure can maintain the linearity of RAW format data as much as possible (that is, the numerical value of the pixel value in the image is approximately proportional to the brightness), so as to divide the pixel value of each pixel position in the image by a specified factor to reduce the image brightness.

[0119] Optionally, in another implementation, the image acquisition device can capture a noise-free image under good acquisition conditions. The electronic device can copy the captured image, add a motion image, and then add noise to obtain a sample image sequence. This implementation will be described in detail in subsequent embodiments.

[0120] Optionally, in another implementation, the image acquisition device may directly acquire multiple sample images containing motion areas and with noise under poor acquisition conditions, and the electronic device may use the acquired multiple sample images as a sample image sequence. The poor acquisition conditions may be the opposite of the good acquisition conditions, for example, under poor acquisition conditions, the analog gain of the image sensor is not the lowest analog gain, the image is acquired under dark light conditions, and / or the image is not acquired according to the exposure time of the clock.

[0121] The sample image may be in RAW format, which contains the original information collected without processing, such as the original color information and detail information used for image enhancement. Image processing of RAW images can better process image frames collected under dark light conditions and improve image quality. In addition, in the process of adding synthetic noise to RAW images, it may not be affected by operations such as ISP and image compression. Adding noise according to parameters modeled by physical characteristics can simulate noisy images captured under dark light conditions as much as possible.

[0122] The total number of sample images in the sample image sequence may be referred to as a sequence length (Length of Sequence, LEN_SEQ). The total number of sample images in the sample image sequence may be a preset number, for example, the preset number may be 32 or 64.

[0123] S102 , for each sample image except the first sample image in the sample image sequence, fuse the sample image with a sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image.

[0124] It is understandable that the electronic device can fuse each sample image except the first sample image in the sample image sequence with the sample image that is located before the sample image in the sample image sequence. Since there are motion areas in the sample images, there will be more noise in the motion areas in the fused image, and less noise in the non-motion areas, that is, the noise granularity in the motion areas is significantly higher than that in the non-motion areas, and the noise distribution in the fused image is uneven. Specifically, the sample image can be weightedly fused with the pixel values ​​at the same pixel position in the sample image that is located before the sample image in the sample image sequence. The weight of the weighted fusion can be obtained by processing the sample image with a convolutional neural network that learns the fusion weights of different areas; optionally, the weight of the weighted fusion can also be set based on experience.

[0125] Optionally, during the fusion process, each sample image except the first and second sample images can be fused with the fusion result of the previous sample image. For example, the second sample image is fused with the first sample image to obtain a fused image of the second sample image, and the third sample image is fused with the fused image of the second sample image to obtain a fused image of the third sample image. In an embodiment of the present invention, the sample image can be fused with the fused image of the previous sample image. Since the fused image of the previous sample image has already been fused with all the previous sample images, it is possible to avoid fusing multiple images and improve the fusion efficiency. Moreover, the fused image fused with all the previous sample images can more comprehensively describe the motion trajectory and simulate the motion area caused by real motion.

[0126] S103, calculating the ratio of the accumulated identification of each pixel position in the sample image to the total number of sample images, and obtaining the weight of the pixel position.

[0127] Among them, the cumulative identification of a pixel position is positively correlated with the number of times the pixel position belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image.

[0128] It is understandable that each pixel position is the same in each sample image, but the cumulative mark of each pixel position in each sample image may be different. For example, the cumulative mark of pixel position a of sample image 1 may be different from the cumulative mark of pixel position a of sample image 2. In other words, each sample image has its own cumulative mark. For each sample image except the first sample image, the electronic device may accumulate the number of times that the sample image belongs to the non-motion area in the adjacent sample images before the sample image to obtain the cumulative mark. In other words, the cumulative mark of the pixel position in the non-motion area in the sample image is larger, and the cumulative mark of the pixel position in the motion area is smaller.

[0129] Exemplarily, the cumulative identification of pixel positions that continuously belong to the non-motion area in the previous adjacent sample images will increase, and the cumulative identification of pixel positions that belong to the motion area will not increase. For example, if a pixel position belongs to the non-motion area in the 5th sample image, and if the 4th sample image belongs to the motion area, then the cumulative identification of the pixel position in the 5th sample image is 1; if the 4th sample image belongs to the non-motion area and the 3rd sample image belongs to the motion area, then the cumulative identification of the pixel position in the 5th sample image is 2; if the 4th sample image belongs to the non-motion area, the 3rd sample image belongs to the non-motion area, and the 2nd sample image belongs to the motion area, then the cumulative identification of the pixel position in the 5th sample image is 3.

[0130] The electronic device may use the ratio of the accumulated identification of each pixel position in the sample image to the total number of the sample image as the weight of the pixel position to obtain the weight of each pixel position in the sample image. Exemplarily, the weight may be calculated according to the weight formula, which is: ;in, represents the weight of each pixel position, Indicates the cumulative mark of the pixel position, Indicates the total number of sample images.

[0131] Since the cumulative mark of the pixel position in the non-moving area is larger, and the cumulative mark of the pixel position in the moving area is smaller, the weight of the pixel position in the non-moving area is larger, which can retain the low-noise state of the non-moving area, and the weight of the pixel position in the moving area is smaller, which can weaken the high-noise state of the moving area.

[0132] It should be noted that if a pixel position belongs to the motion area in a sample image, even if the pixel position in the adjacent sample images before the sample image belongs to the non-motion area, the accumulated mark of the pixel position in the sample image will be re-accumulated. In other words, the accumulated mark is not limited by the continuous non-motion state in the past. As long as the current pixel position shows motion characteristics (that is, it belongs to the motion area), it will be re-accumulated starting from the current sample image.

[0133] The cumulative mark of a pixel position reflects the motion state of the sample image relative to the pixel position in the previous sample image. What is concerned is the motion change of the pixel position in the previous image sequence. In this way, the motion state change of the pixel position can be captured timely and accurately, and then the amount of noise in the fused image can be determined. Specifically, the larger the cumulative mark, the longer the non-motion state is maintained, the fewer motion areas are fused, and the less noise in the fused image; the smaller the cumulative mark, the shorter the non-motion state is maintained, the more motion areas are fused, and the more noise in the fused image.

[0134] Since the cumulative mark of each pixel position in each sample image reflects the motion state of the pixel position in the sample image before the sample image, in order to more accurately compare the difference between the reference image and the fused image obtained based on the cumulative mark, in step S102, a sample image is fused with the previous sample image. Specifically, a sample image can be fused with the fusion result of the previous moment (that is, the fused image of the previous sample image) by adopting a recursive mode to obtain the fusion result of the current moment (that is, the fused image of the sample image), instead of fusing the sample image with all the sample images in the entire sample image sequence or other sample images.

[0135] S104, calculating the weighted sum of the fused image and the true value image corresponding to the sample image according to the obtained weights, as a reference image with uniform noise corresponding to the sample image.

[0136] The true value image corresponding to the sample image is consistent with the image content of the sample image and has no noise.

[0137] It is understandable that the electronic device can multiply the pixel value of each pixel position in the fused image by the obtained weight (which can be called the first weight), multiply the pixel value of each pixel position in the true value image by another part of the weight (which can be called the second weight), and add the pixel values ​​of the same pixel position in the product of the two parts to obtain the reference image. Among them, the sum of the first weight and the second weight is 1, that is, the weight of the motion area in the true value image is larger, which can enhance the low noise state of the motion area, and the weight of the non-motion area is smaller, which can retain the low noise state of the non-motion area.

[0138] In the process of calculating the weighted sum, the pixel value of each pixel position in the fused image is multiplied by the obtained weight, so that the moving area in the fused image can be multiplied by a smaller weight and the non-moving area can be multiplied by a larger weight; the pixel value of each pixel position in the true value image is multiplied by another part of the weight, so that the moving area in the true value image can be multiplied by a larger weight and the non-moving area can be multiplied by a smaller weight. Since the true value image and the sample image have the same image content and are noise-free, the above-mentioned process of calculating the weighted sum can weaken the high-noise situation in the moving area in the fused image as much as possible, retain the low-noise situation in the non-moving area in the fused image, and obtain a reference image with uniform noise.

[0139] S105, inputting the fused image corresponding to the sample image into the image noise processing network of the initial structure to perform noise uniformization processing, and obtaining a result image corresponding to the sample image.

[0140] It can be understood that the electronic device can input the fused image corresponding to the sample image into the image noise processing network of the initial structure, perform noise uniformization processing on the fused image corresponding to the sample image according to the network parameters of the image noise processing network of the initial structure, and obtain the result image corresponding to the sample image.

[0141] S106, based on the difference between the result image corresponding to the sample image and the reference image, adjusting the network parameters of the image noise processing network of the initial structure until a preset convergence condition is reached, thereby obtaining a trained image noise processing network.

[0142] It can be understood that the reference image has uniform noise. The reference image can be used as the training target of the image noise processing network, and the network parameters can be continuously adjusted to make the result image tend to the reference image, that is, the output of the image noise processing network tends to an image with uniform noise.

[0143] The electronic device may compare the difference between the result image corresponding to the sample image and the reference image. Exemplarily, a loss value between the result image corresponding to the sample image and the reference image may be calculated according to a loss function as the difference between the result image corresponding to the sample image and the reference image.

[0144] Optionally, in one implementation, the fused image of each sample image in the sample image sequence can be input into the image noise processing network of the initial structure to obtain the result image corresponding to the sample image; then, based on the use of a preset loss function, the loss value between the result image corresponding to each sample image and the reference image is calculated, and the sum of the loss values ​​is taken as the total loss value, and the total loss value of the sample image sequence is used to adjust the network parameters. Exemplarily, the preset loss function can be: ;in, represents the loss value, Represents the total number of sample images in the sample image sequence, Indicates the sample image sequence sample images, Indicates The result image corresponding to the sample images is Indicates The reference image corresponding to the sample image.

[0145] Optionally, in another implementation, network parameters may be adjusted based on a loss value between a result image corresponding to a sample image in the sample image sequence and a reference image, that is, network parameters may be adjusted using a loss value.

[0146] It can be understood that the network parameters of the image noise processing network of the initial structure can be adjusted based on the obtained loss value to determine whether the preset convergence conditions are met. If the preset convergence conditions are met, the training is terminated to obtain the trained image noise processing network. If the preset convergence conditions are not met, the process returns to step S101 to continue training the image noise processing network.

[0147] Optionally, the preset convergence condition can be pre-set as the loss value is less than the preset loss threshold. If the loss value is less than the preset loss threshold, it can be determined that the preset convergence condition is met; if the loss value is greater than the preset loss threshold, it can be determined that the preset convergence condition is met. Optionally, in another implementation, the preset convergence condition can be pre-set as the number of training times reaches the preset number of iterations. If the number of training times reaches the preset number of iterations, the training is terminated. In the solution of the present invention, the image noise processing network can be trained according to the preset convergence condition to ensure the effectiveness of the trained image noise processing network.

[0148] In an embodiment of the present invention, the sample image is noisy, and the sample image is fused with the previous sample image, and the noise in the motion area in the obtained fused image is more, and the noise in the non-motion area is less. The cumulative mark of a pixel position is positively correlated with the number of times the pixel position belongs to the non-motion area in the corresponding adjacent sample image. That is, the cumulative mark of the pixel position in the motion area in the sample image is the smallest, and the cumulative mark of the pixel position in the non-motion area will be accumulated on the basis of the cumulative mark of the previous sample image, that is, the cumulative mark of the pixel position in the non-motion area is always kept large. Since the weight is the ratio of the cumulative mark to the total number of sample images, the weight of the motion area is small, and the weight of the non-motion area is large. Therefore, by fusing the fused image with uneven noise distribution and the true value image without noise according to the weight, the more noise in the motion area can be weakened, and the less noise in the non-motion area can be maintained, so as to obtain a reference image with uniform noise, and the reference image with uniform noise is used as the target of the image noise processing network. Based on the image noise processing network, the fused image with uneven noise is processed to uniform noise to obtain the result image. The network parameters are adjusted based on the gap between the result image and the target, so that the output of the image noise processing network tends to an image with uniform noise, which can improve the uniformity of noise distribution in the image.

[0149] Optionally, in one embodiment, Figure 1 Based on the image noise processing network training method shown in Figure 2 As shown, step S101 includes steps S1011-S1013.

[0150] S1011, acquiring an original image sequence including a plurality of original images with consistent sizes and without noise, and a motion image with a size smaller than each original image.

[0151] It is understandable that the image acquisition device can acquire a sequence of original images containing multiple original images of the same size and without noise under good acquisition conditions, and the electronic device can obtain the image acquired by the image acquisition device. Optionally, the image acquisition device can acquire an original image without noise under good acquisition conditions, and then the electronic device can copy the original image to obtain a sequence of original images containing multiple original images of the same size and without noise. The acquired original image without noise can be called label data (ground-truth, gt) under bright light.

[0152] The image acquisition device can acquire a moving image smaller than the original image; or acquire an image of the same size as the original image, and the electronic device can cut out a moving image smaller than the original image from the image. Among them, the moving image can be called a moving foreground image, and the original image can be called a background image. In order to simulate the effect of the moving image moving in the original image, the size of the moving image is smaller than the size of the original image, and the original image will not be blocked. There is a movable space for the moving image in the original image. For example, the size of the moving image may not exceed 20% of the size of the original image. The moving image and the original image can be images acquired by the same image acquisition device under the same acquisition conditions (for example, the same analog gain of the sensor, the same light intensity, and the same exposure time); the moving image and the original image can have different contents. For example, the moving image and the original image can be acquired in different scenes. For example, the moving image can be acquired on the road, and the original image can be acquired in the park.

[0153] The total number of original images in the original image sequence is the same as the total number of sample images in the sample image sequence.

[0154] S1012, using the motion image, replacing the image region in each original image that is the same size as the motion image, to obtain an image to be used corresponding to each original image.

[0155] Among them, in the original image sequence, the positions of the replaced image regions in any two adjacent original images are different.

[0156] It can be understood that the electronic device can determine an image area in each original image that is consistent with the size of the motion image, and the positions of the image areas in any two adjacent original images are different, and then replace the determined image area with the motion image to obtain the image to be used corresponding to each original image.

[0157] Optionally, the electronic device may determine and replace different image areas for every two arbitrary original images in the original image sequence.

[0158] Optionally, in another implementation, the electronic device may determine and replace the image area in each original image in sequence according to the order of the original images in the original image sequence. Exemplarily, the electronic device may randomly determine and replace the image area in the first original image, and then move the image area determined in the first original image in the second original image to determine and replace the image area in the second original image, and so on, until all the original images in the original image sequence are replaced. The movement may be performed according to the rules of simulated translational motion. For example, the staff may pre-set the movement angle and movement amount to move the image area. Optionally, if the image area exceeds the edge of the original image, the image area exceeding the edge may be set in the original image opposite to the edge; for example, if half of the image area exceeds the left edge of the original image, the exceeding area may be set in the area of ​​the right edge of the original image.

[0159] In the embodiment of the present invention, the original images in the original image sequence may be replaced in sequence to ensure that the positions of the replaced image regions in any two adjacent original images are different, so as to ensure that there is a motion region in the sample image subsequently obtained based on the image to be used.

[0160] Optionally, in one implementation, step S1012 includes steps S1012a-S1012b.

[0161] S1012a, using the motion image, replacing the image area in each original image that is the same size as the motion image, to obtain an intermediate image corresponding to each original image.

[0162] S1012b, reducing the brightness of the intermediate image corresponding to each original image to obtain an image to be used corresponding to each original image.

[0163] It is understandable that the electronic device can use the moving image to replace the image area in the original image to obtain the intermediate image corresponding to each original image. The electronic device can then reduce the brightness of the intermediate image to simulate the low-brightness image collected under dark light conditions, so that a network can be trained to uniformly distribute noise to the image with uneven noise under dark light conditions.

[0164] For example, since the mean value of all pixels in an image can represent the brightness of the image, dividing the pixel value of the image by a specified value can reduce the brightness of the image in a geometric ratio with the specified value as a ratio, so that the pixel value of each pixel position of the intermediate image can be divided by the same specified value to obtain the image to be used, that is, the GT image under dark light conditions. The specified value can be preset, such as 10, 20, 50, or 100.

[0165] It is understandable that the order of replacing the moving image and reducing the brightness can be adjusted, for example, first reducing the brightness of the original image and the moving image, and then replacing the original image after the brightness is reduced. In this regard, the embodiment of the present invention is only used as an example for illustration and is not specifically limited.

[0166] S1013, adding noise to each of the obtained images to be used, to obtain a plurality of sample images containing motion areas and with noise.

[0167] It can be understood that the added noise is used to simulate the noise in the image captured under dark light conditions. If the noise is added first and then the brightness is reduced, the brightness reduction process will disturb the noise added first, causing the noise characteristics to be destroyed and the noise of the image captured under dark light conditions cannot be simulated. In the embodiment of the present invention, the brightness is reduced first and then the noise is added, so that the added noise can be prevented from being disturbed, the noise of the image captured under dark light conditions can be simulated as much as possible, and the effect of the trained image noise processing network on the noise uniformity of the image captured under dark light conditions is improved.

[0168] It is understandable that the images in an image sequence can be considered as a set. For example, the set: ; Indicates A true value image, Indicates Sample images.

[0169] Optionally, in one implementation, before step S1013 in the image noise processing network training method, the method further includes step A1, and step S1013 includes step A2.

[0170] A1, obtain the noise parameters of the sensor that collects the original image.

[0171] A2, based on the noise parameter, adding noise to each of the obtained images to be used, to obtain a plurality of sample images containing motion areas and with noise.

[0172] It is understood that the noise parameters can be pre-calibrated. For example, the manufacturer of the sensor can calibrate the noise parameters of the sensor under laboratory conditions. In other words, the noise parameters of a sensor are fixed and will not change with changes in the illumination conditions during acquisition.

[0173] The electronic device can obtain the noise parameters of the sensor that captures the original image, and then add noise based on the obtained noise parameters. The added noise simulates the noise generated by the sensor itself under dark light conditions, so that the sample image can simulate the image captured by the sensor under dark light conditions, thereby improving the effect of the trained image noise processing network in making the noise uniform for images captured under dark light conditions.

[0174] Optionally, in one implementation, step A2 may include steps a1-a3.

[0175] a1, when the noise parameter includes the mean value of Poisson noise, noise is added to each of the obtained images to be used according to the noise parameter to obtain a plurality of sample images containing motion areas and with noise.

[0176] It can be understood that when the noise parameters include the mean of Poisson noise, the pixel value of each pixel in the image to be used can be randomly changed according to the Poisson distribution, and the mean of the pixel values ​​of each pixel after the change is the mean of Poisson noise, simulating the shot noise generated under dark light conditions.

[0177] a2. When the noise parameters include the mean and standard deviation of Gaussian noise, noise is added to each of the obtained images to be used according to the noise parameters to obtain a plurality of sample images containing motion areas and with noise.

[0178] It can be understood that when the noise parameters include the mean and standard deviation of Gaussian noise, random pixel values ​​can be generated according to the mean and standard deviation of Gaussian noise, and for each pixel point in the image to be used, the generated random pixel value can be superimposed with the pixel value of the pixel point; for each pixel point in the image to be used, the generated random pixel value can be superimposed with the pixel value of the pixel value to simulate the Gaussian noise generated under dark light conditions.

[0179] a3. When the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, noise is added to each image to be used according to the noise parameters to obtain multiple sample images containing motion areas and with noise.

[0180] It can be understood that, when the noise parameters include the mean of Poisson noise and the mean and standard deviation of Gaussian noise, Poisson noise and Gaussian noise can be added at the same time to simulate the shot noise and Gaussian noise generated under dark light conditions, ensuring that the added noise is the noise generated by the sensor itself under dark light conditions, so that the sample image can simulate the image captured by the sensor under dark light conditions, thereby improving the effect of the trained image noise processing network in making the noise uniform for images captured under dark light conditions.

[0181] Optionally, in this embodiment, the true value image corresponding to each sample image is the image to be used when obtaining the sample image.

[0182] It can be understood that since the true value image corresponding to each sample image is consistent with the image content of the sample image and is noise-free, and the image to be used when obtaining each sample image is also consistent with the image content of the sample image and no noise has been added, the true value image corresponding to each sample image can be used as the image to be used when obtaining the sample image when determining the true value image used in the reference image.

[0183] In the prior art, a noisy sample image is directly obtained, and then a denoising network is used to denoise the sample image to obtain a true value image. However, the true value image obtained using the existing denoising means still contains noise. If the true value image with noise is used to determine the reference image, redundant noise will be introduced into the reference image, resulting in a low uniformity of noise distribution in the reference image. In an embodiment of the present invention, the original image without noise is replaced and denoised to obtain a sample image. The image without noise to be used generated in the above process can be directly used as the true value image, avoiding directly obtaining a noisy sample image and then denoising the sample image to obtain a true value image, which can improve the uniformity of noise in the reference image.

[0184] Optionally, in one embodiment, the cumulative flag of each pixel position in the first sample image in the sample image sequence is 1, such as Figure 3 As shown, for each sample image except the first sample image in the sample image sequence, the process of determining the cumulative identification of each pixel position in the sample image includes steps S301-S303.

[0185] S301 , determining a motion region and a non-motion region included in a reference image corresponding to the sample image in a reference image sequence, and using them as a reference motion region and a reference non-motion region, respectively.

[0186] Among them, the reference image sequence is a sample image sequence or a true value image sequence composed of true value images corresponding to multiple sample images; the position order of the reference image corresponding to the sample image in the reference image sequence is consistent with the position order of the sample image in the sample image sequence.

[0187] It can be understood that if the position order of a sample image in a sample image sequence is the same as the position order of a true value image in a true value image sequence, then the motion area of ​​the sample image and the true value image are the same, the non-motion area is also the same, and the cumulative identification of each pixel position in the sample image is also the same as the cumulative identification of each pixel position in the true value image.

[0188] The electronic device may determine the cumulative identification based on the motion area in the sample image sequence, or may determine the cumulative identification based on a true value image sequence consisting of true value images corresponding to each of the plurality of sample images. The true value image sequence consisting of true value images corresponding to each of the plurality of sample images may be an image sequence consisting of a plurality of images to be used in the above-mentioned embodiment.

[0189] It can be understood that if the motion region in the sample image is generated based on the motion image in the aforementioned embodiment, the motion region in a sample image may include the region of the motion image in the previous sample image and the region of the motion image in the sample image.

[0190] S302 : for each pixel position in the sample image, if the pixel position belongs to a reference motion region in the corresponding reference image, determine that the cumulative flag of the pixel position in the sample image is 1.

[0191] S303, if the pixel position belongs to the reference non-motion area in the corresponding reference image, then add 1 to the cumulative mark of the pixel position in the sample image before the sample image to obtain the cumulative mark of the pixel position in the sample image.

[0192] It can be understood that if the reference image sequence is a sample image sequence, the reference image is the corresponding sample image; if the reference image sequence is a true value image sequence, the reference image is the corresponding true value image. The following description is taken as an example that the reference image sequence is a sample image sequence.

[0193] For each pixel position in the sample image, if the pixel position belongs to the reference motion region in the corresponding reference image, that is, the pixel position belongs to the motion region in the sample image, the cumulative identification of the pixel position in the sample image can be determined as 1. If the pixel position belongs to the reference non-motion region in the corresponding reference image, that is, the pixel position belongs to the non-motion region in the sample image, 1 can be added to the cumulative identification of the pixel position in the sample image before the sample image to obtain the cumulative identification of the pixel position in the sample image. For example, for a pixel position, if the pixel position belongs to the motion region in the 4th sample image, the cumulative identification of the pixel position in the 4th sample image is determined as 1; if the pixel position belongs to the non-motion region in the 4th sample image, and the cumulative identification of the pixel position in the 3rd sample image is 2, the cumulative identification of the pixel position in the 4th sample image can be determined as 3. Wherein, there is no motion region in the first sample image, and the cumulative identification of each pixel position in the first sample image can be set to 1.

[0194] like Figure 4As shown, a continuous video is simulated by a sample image sequence. The first sample image simulates the image at the initial moment (t=0), the tth sample image simulates the image at the t+1 moment, and the t+1th sample image simulates the image at the t+2 moment. The large rectangle in the figure represents the t+1th sample image, the solid small rectangle represents the motion image in the t+1th sample image, and the dotted small rectangle represents the motion image in the tth sample image. The area occupied by the solid small rectangle and the dotted small rectangle is the motion area, and the cumulative mark of each pixel position in the motion area is set to 1; for the cumulative mark of the pixel position in the tth sample image that is the same as the non-motion area in the t+1th sample image, add 1 to obtain the cumulative mark of each pixel position in the non-motion area in the t+1th sample image. Among them, the cumulative mark can also be called mask information or the accumulated number (n_counts). For example, the cumulative identification of the pixel positions in all areas of the first sample image is 1, the cumulative identification of the pixel positions in the moving area of ​​the second sample image is 1, the cumulative identification of the pixel positions in the non-moving area of ​​the second sample image is 2, the cumulative identification of the pixel positions in the moving area of ​​the third sample image is 1, and the cumulative identification of each pixel position in the non-moving area of ​​the third sample image is 1 plus the cumulative identification in the second sample image. If the cumulative identification of a pixel position in the non-moving area of ​​the third sample image in the second sample image is 1, then the cumulative identification of the pixel position in the third sample image is 2; if the cumulative identification of a pixel position in the non-moving area of ​​the third sample image in the second sample image is 2, then the cumulative identification of the pixel position in the third sample image is 3.

[0195] Optionally, in one embodiment, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork. Figure 1 Based on the image noise processing network training method shown in Figure 5 As shown, step S105 includes steps S1051-S1053.

[0196] S1051, input the fused image corresponding to the sample image into the encoding sub-network to obtain a first eigenvector.

[0197] S1052: Obtain a second feature vector based on the first feature vector and the decoding sub-network.

[0198] S1053: Based on the second eigenvector and the output subnetwork, obtain a result image corresponding to the sample image.

[0199] It can be understood that the image noise processing network includes: an encoding subnetwork (encoder), a decoding subnetwork (decoder), and an output subnetwork. The encoding subnetwork includes a structure of multiple convolutional layers stacked and a downsampling structure, which can extract the features of the fused image and compress the fused image into a low-dimensional feature space to obtain a first feature vector. The decoding subnetwork includes a stack of multiple convolutional layers and an upsampling structure. Based on the first feature vector and the decoding subnetwork, the first feature vector output by the encoding subnetwork can be reconstructed back to the space with the fused image to obtain a second feature vector. The feature vector representing the features of the fused image can be obtained through the encoder-decoder structure. The output subnetwork can include a convolutional layer without an activation (relu) layer. Based on the second feature vector and the output subnetwork, a result image consistent with the dimension of the fused image can be obtained.

[0200] Optionally, in one implementation, the encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork.

[0201] Step S1051 includes S1051a-S1051c.

[0202] S1051a, input the fused image corresponding to the sample image into the first convolutional sub-network to obtain a first sub-feature vector.

[0203] S1051b, input the first sub-feature vector into the second convolutional sub-network to obtain a second sub-feature vector.

[0204] S1051c, input the second sub-feature vector into the third convolutional sub-network to obtain the first feature vector.

[0205] It can be understood that the fused image can be input into the first convolution subnetwork in the encoding subnetwork for feature extraction to obtain the first sub-feature vector. Then, the first sub-feature vector is input into the second convolution subnetwork for feature extraction to obtain the second sub-feature vector. Then, the second sub-feature vector is input into the third convolution subnetwork for feature extraction to obtain the first feature vector. Among them, the first convolution subnetwork includes a convolution layer with a stride of 1 and an activation layer. The convolution layer with a stride of 1 can convert the image into a feature map, retaining more image features for more accurate feature extraction in the future; the nonlinear activation layer can enable the network to learn the complex features of nonlinear information. The second convolution subnetwork and the third convolution subnetwork both include a convolution layer with a stride of 2, a convolution layer with a stride of 1, and an activation layer. The convolution layer with a stride of 2 has the effect of halving the resolution.

[0206] Step S1052 includes S1052a-S1052b.

[0207] S1052a, input the first feature vector into the fourth convolutional sub-network to obtain a third sub-feature vector.

[0208] S1052b, input the third sub-feature vector and the second sub-feature vector into the fifth convolutional sub-network to obtain a second feature vector.

[0209] Step S1053 includes S1053a, inputting the second feature vector and the first sub-feature vector into the output sub-network to obtain a result image corresponding to the sample image.

[0210] It can be understood that both the second convolutional subnetwork and the fifth convolutional subnetwork include a convolutional layer with a stride of 2, a convolutional layer with a stride of 1, and an activation layer. The third sub-feature vector and the second sub-feature vector can be input into the fifth convolutional subnetwork through a skip connection to obtain a second feature vector. The third sub-feature vector and the second sub-feature vector can be input into the fifth convolutional subnetwork through a skip connection structure to obtain a second feature vector. The features of the earlier layers contain more details and local information, while the features of the deeper layers are more abstract and global. Through the skip connection, the decoder can simultaneously utilize feature information of different scales during the reconstruction process, thereby better evenly fusing the noise of the image.

[0211] Combine the following Figure 6 An image noise processing method provided by an embodiment of the present invention is introduced. Figure 6 As shown, the method may include steps S601-S602.

[0212] S601, obtaining an image to be processed.

[0213] It is understandable that the image noise processing method can be applied to an electronic device that can obtain a video captured by a camera and can enhance the image frames in the captured video. Exemplarily, the electronic device can be a camera or a server. Optionally, when the server can have the functions of training the network and enhancing the image at the same time, the image noise processing method and the image noise processing network training method can be the same execution subject; optionally, when the server can only have the functions of training the network or processing image noise, the image noise processing method and the image noise processing network training method can be different execution subjects.

[0214] Optionally, in one implementation, the image sensor of the image acquisition device can capture image frames of the video frame by frame and send the captured image frames to the image processing chip in the image acquisition device frame by frame. The image processing chip can perform time domain enhancement on the image frames received frame by frame and use the time domain enhanced images as images to be processed.

[0215] The captured video image frames may be images with a low signal-to-noise ratio captured under dark light conditions.

[0216] S602: Input the image to be processed into a pre-trained image noise processing network for noise equalization processing to obtain a result image corresponding to the image to be processed.

[0217] The image noise processing network is trained based on the image noise processing network training method provided in the above-mentioned embodiment.

[0218] It is understandable that the trained image noise processing network can improve the uniformity of noise in the image to be processed. After obtaining a result image with high noise uniformity, the ISP can be used to perform subsequent processing on the result image.

[0219] Optionally, in one implementation, the method further includes: performing noise reduction processing on the obtained result image using preset noise reduction parameters to obtain an enhanced image.

[0220] It is understandable that the spatial noise reduction method can be used in electronic devices through ISP. Specifically, the preset noise reduction parameters can be used to perform noise reduction processing on the resulting image. Since the resulting image is an image with improved noise uniformity, the resulting image with higher noise uniformity can be processed with the same denoising intensity, thereby reducing the noise in the image as much as possible, retaining the details in the image, and improving the image quality.

[0221] Exemplarily, the preset noise reduction parameter can be a preset window size for smoothing the noise in each area of ​​the image. For each pixel in the image, the nearby pixels of the pixel are determined based on the preset window size, and the average pixel value of the nearby pixels is used to replace the pixel value of the pixel to reduce the noise in the image frame.

[0222] In the embodiment of the present invention, since the output of the image noise processing network trained using the aforementioned embodiment tends to be an image with uniform noise, the uniformity of noise distribution in the image to be processed can be improved.

[0223] Optionally, in one embodiment, Figure 6 Based on the image noise processing method shown in FIG, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork; Figure 7 As shown, step S602 includes steps S6021-S6023.

[0224] S6021, input the image to be processed into the encoding sub-network to obtain a third eigenvector.

[0225] S6022: Obtain a fourth eigenvector based on the third eigenvector and the decoding subnetwork.

[0226] S6023: Based on the fourth eigenvector and the output sub-network, obtain a result image corresponding to the image to be processed.

[0227] Optionally, in one implementation, the encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork;

[0228] Step S6021 includes steps S6021a-S6021c.

[0229] S6021a, inputting the image to be processed into the first convolutional sub-network to obtain a fourth sub-feature vector;

[0230] S6021b, input the fourth sub-feature vector into the second convolutional sub-network to obtain a fifth sub-feature vector;

[0231] S6021c, input the fifth sub-feature vector into the third convolutional sub-network to obtain a third feature vector;

[0232] Step S6022 includes steps S6022a-S6022b.

[0233] S6022a, input the third eigenvector into the fourth convolutional sub-network to obtain a sixth sub-eigenvector;

[0234] S6022b, input the sixth sub-eigenvector and the fifth sub-eigenvector into the fifth convolutional sub-network to obtain a fourth eigenvector.

[0235] Step S6023 includes S6023a, inputting the fourth eigenvector and the fourth sub-eigenvector into the output subnetwork to obtain a result image corresponding to the image to be processed.

[0236] It can be understood that during the two processes of training and applying the image noise processing network, the network structure of the image noise processing network will not change. The specific implementation method of enabling the network structure in the image noise processing network to process the input content can refer to the aforementioned embodiment and will not be elaborated here.

[0237] Figure 8 This is a schematic diagram of the structure of the image noise processing network. Figure 8As shown, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork. The encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork. The first convolutional subnetwork includes multiple convolutional layers and activation layer combinations with a step size of 1, and each convolutional layer and activation layer combination with a step size of 1 includes a convolutional layer with a step size of 1 and an activation layer (a). The second convolutional subnetwork includes a convolutional layer with a step size of 2 and an activation layer (b), as well as multiple convolutional layers and activation layer combinations with a step size of 1. The third convolutional subnetwork is the same as the second convolutional subnetwork. The fourth convolutional subnetwork includes a convolutional layer with a step size of 2 (c) and multiple convolutional layers and activation layer combinations with a step size of 1. The fourth convolutional subnetwork is the same as the fifth convolutional subnetwork. Among them, the size of the convolutional layers in the image noise processing network is 3×3. The output subnetwork contains multiple stride-1 convolutional layers combined with activation layers and a stride-1 convolutional layer (d).

[0238] The electronic device can use the fused image as input data, input it into the first convolution subnetwork of the encoding subnetwork, and obtain the first sub-feature vector; then, input the first sub-feature vector into the second convolution subnetwork, and obtain the second sub-feature vector; then, input the second sub-feature vector into the third convolution subnetwork, and obtain the first feature vector. Then, input the first feature vector into the fourth convolution subnetwork of the decoding subnetwork, and obtain the third sub-feature vector; then, input the third sub-feature vector and the second sub-feature vector into the fifth convolution subnetwork, and obtain the second feature vector; then, input the second feature vector and the first sub-feature vector into the output subnetwork, and superimpose the output content of the output subnetwork with the input data to obtain the result image corresponding to the sample image as the output data.

[0239] FIG9 (a) and FIG9 (b) are visualization effect diagrams of RAW format images after simple ISP processing based on an embodiment of the present invention. As shown in FIG9 (a), the image is processed according to the time domain enhancement method provided by the prior art. Since the head part moves, the area of ​​the head part is the moving area, and the other parts of the picture are static areas. The black dots represent noise, and the noise in the moving area is larger than the noise in the static area, that is, the granularity of the noise in the moving area is stronger than that in the static area, and the noise distribution is uneven. As shown in FIG9 (b), the image is processed according to the image noise processing method provided by the embodiment of the present invention in the aforementioned embodiment, and then processed using the time domain enhancement method. Although the area of ​​the head part moves, the area of ​​the head part is the moving area, and the other parts are static areas, the noise level of the moving area is consistent with that of the static area, the noise uniformity is high, the visual perception is better, and it is more conducive to subsequent ISP processing, which can reflect the application value of the present invention in improving image quality.

[0240] The embodiment of the present invention also provides an image noise processing network training device, such as Fig.10 As shown, the device comprises:

[0241] The first acquisition module 1010 is used to acquire a plurality of sample images containing motion regions and having noise, and obtain a sample image sequence; wherein the motion region in each sample image is: a region where the content of the sample image changes relative to the previous sample image;

[0242] A fusion module 1020 is used for fusing, for each sample image except the first sample image in the sample image sequence, the sample image with a sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image;

[0243] The first calculation module 1030 is used to calculate the ratio of the cumulative identification of each pixel position in the sample image to the total number of sample images to obtain the weight of the pixel position; wherein the cumulative identification of a pixel position is positively correlated with the number of times that the pixel position continuously belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image;

[0244] The second calculation module 1040 is used to calculate the weighted sum of the fused image and the true value image corresponding to the sample image according to the obtained weights as the reference image with uniform noise corresponding to the sample image; wherein the true value image corresponding to the sample image has the same image content as the sample image and has no noise;

[0245] The first input module 1050 is used to input the fused image corresponding to the sample image into the image noise processing network of the initial structure to perform noise uniformization processing to obtain a result image corresponding to the sample image;

[0246] The adjustment module 1060 is used to adjust the network parameters of the image noise processing network of the initial structure based on the difference between the result image corresponding to the sample image and the reference image until a preset convergence condition is reached to obtain a trained image noise processing network.

[0247] Optionally, the first acquisition module 1010 includes:

[0248] A first acquisition unit is used to acquire an original image sequence including a plurality of original images of uniform size and without noise, and a motion image of a size smaller than each original image;

[0249] a replacement unit, configured to use the motion image to replace an image region of each original image having a size consistent with that of the motion image, so as to obtain an image to be used corresponding to each original image; wherein, in the original image sequence, positions of the replaced image regions in any two adjacent original images are different;

[0250] An adding unit, used for adding noise to each of the obtained images to be used, to obtain a plurality of sample images containing motion areas and with noise;

[0251] The true value image corresponding to each sample image is the image to be used when obtaining the sample image.

[0252] The device also includes:

[0253] A second acquisition module, used to acquire noise parameters of a sensor that acquires the original image;

[0254] The adding unit is specifically used to add noise to each of the obtained images to be used based on the noise parameters, so as to obtain a plurality of sample images containing motion areas and with noise.

[0255] a replacement unit, specifically configured to use the motion image to replace an image region in each original image that has the same size as the motion image, so as to obtain an intermediate image corresponding to each original image;

[0256] The brightness of the intermediate image corresponding to each original image is reduced to obtain the image to be used corresponding to each original image.

[0257] Optionally, the cumulative identification of each pixel position in the first sample image in the sample image sequence is 1;

[0258] For each sample image other than the first sample image in the sample image sequence, the cumulative identifier of each pixel position in the sample image is determined by the following steps:

[0259] Determine a motion region and a non-motion region contained in a reference image corresponding to the sample image in a reference image sequence, and use them as a reference motion region and a reference non-motion region, respectively; wherein the reference image sequence is the sample image sequence or a true value image sequence composed of true value images corresponding to each of the plurality of sample images; and the position sequence of the reference image corresponding to the sample image in the reference image sequence is consistent with the position sequence of the sample image in the sample image sequence;

[0260] For each pixel position in the sample image, if the pixel position belongs to the reference motion area in the corresponding reference image, determining that the cumulative flag of the pixel position in the sample image is 1;

[0261] If the pixel position belongs to the reference non-motion area in the corresponding reference image, then 1 is added to the cumulative identification of the pixel position in the sample image before the sample image to obtain the cumulative identification of the pixel position in the sample image.

[0262] Optionally, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork;

[0263] The first input module 1050 includes:

[0264] A first input unit, used for inputting a fused image corresponding to the sample image into the encoding subnetwork to obtain a first feature vector;

[0265] A first obtaining unit, configured to obtain a second feature vector based on the first feature vector and the decoding subnetwork;

[0266] The second obtaining unit is used to obtain a result image corresponding to the sample image based on the second feature vector and the output sub-network.

[0267] Optionally, the encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork;

[0268] A first input unit, specifically used to input the fused image corresponding to the sample image into the first convolutional sub-network to obtain a first sub-feature vector;

[0269] Inputting the first sub-feature vector into the second convolutional sub-network to obtain a second sub-feature vector;

[0270] Inputting the second sub-feature vector into the third convolutional sub-network to obtain a first feature vector;

[0271] A first obtaining unit, specifically configured to input the first feature vector into the fourth convolutional sub-network to obtain a third sub-feature vector;

[0272] Inputting the third sub-feature vector and the second sub-feature vector into the fifth convolutional sub-network to obtain a second feature vector;

[0273] The second obtaining unit is specifically configured to input the second feature vector and the first sub-feature vector into the output sub-network to obtain a result image corresponding to the sample image.

[0274] In an embodiment of the present invention, the sample image is noisy, and the sample image is fused with the previous sample image, and the noise in the motion area in the obtained fused image is more, and the noise in the non-motion area is less. The cumulative mark of a pixel position is positively correlated with the number of times the pixel position belongs to the non-motion area in the corresponding adjacent sample image. That is, the cumulative mark of the pixel position in the motion area in the sample image is the smallest, and the cumulative mark of the pixel position in the non-motion area will be accumulated on the basis of the cumulative mark of the previous sample image, that is, the cumulative mark of the pixel position in the non-motion area is always kept large. Since the weight is the ratio of the cumulative mark to the total number of sample images, the weight of the motion area is small, and the weight of the non-motion area is large. Therefore, by fusing the fused image with uneven noise distribution and the true value image without noise according to the weight, the more noise in the motion area can be weakened, and the less noise in the non-motion area can be maintained, so as to obtain a reference image with uniform noise, and the reference image with uniform noise is used as the target of the image noise processing network. Based on the image noise processing network, the fused image with uneven noise is processed to uniform noise to obtain the result image. The network parameters are adjusted based on the gap between the result image and the target, so that the output of the image noise processing network tends to an image with uniform noise, which can improve the uniformity of noise distribution in the image.

[0275] The embodiment of the present invention also provides an image noise processing device, such as Fig.11 As shown, the device comprises:

[0276] The second acquisition module 1110 is used to acquire an image to be processed;

[0277] The second input module 1120 is used to input the image to be processed into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed; wherein the image noise processing network is trained based on the image noise processing network training device provided in the above embodiment.

[0278] Optionally, the device further comprises:

[0279] The denoising module is used to perform denoising processing on the result image using preset denoising parameters to obtain an enhanced image.

[0280] Optionally, the image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork;

[0281] The second input module 1120 includes:

[0282] A second input unit, used for inputting the image to be processed into the encoding sub-network to obtain a third feature vector;

[0283] A third obtaining unit, configured to obtain a fourth feature vector based on the third feature vector and the decoding sub-network;

[0284] The fourth obtaining unit is used to obtain a result image corresponding to the image to be processed based on the fourth eigenvector and the output subnetwork.

[0285] Optionally, the encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork;

[0286] A second input unit, specifically configured to input the image to be processed into the first convolutional sub-network to obtain a fourth sub-feature vector;

[0287] Inputting the fourth sub-feature vector into the second convolutional sub-network to obtain a fifth sub-feature vector;

[0288] Inputting the fifth sub-feature vector into the third convolutional sub-network to obtain a third feature vector;

[0289] A third obtaining unit is specifically used to input the third feature vector into the fourth convolutional sub-network to obtain a sixth sub-feature vector;

[0290] Inputting the sixth sub-feature vector and the fifth sub-feature vector into the fifth convolutional sub-network to obtain a fourth eigenvector;

[0291] The fourth obtaining unit is specifically configured to input the fourth eigenvector and the fourth sub-eigenvector into the output subnetwork to obtain a result image corresponding to the image to be processed.

[0292] In the embodiment of the present invention, since the output of the image noise processing network trained using the aforementioned embodiment tends to be an image with uniform noise, the uniformity of noise distribution in the image to be processed can be improved.

[0293] The embodiment of the present invention further provides an electronic device, such as Fig.12 As shown, it includes a processor 1201 , a communication interface 1202 , a memory 1203 and a communication bus 1204 , wherein the processor 1201 , the communication interface 1202 , and the memory 1203 communicate with each other via the communication bus 1204 .

[0294] Memory 1203, used for storing computer programs;

[0295] The processor 1201 is used to implement the image noise processing network training method and the image noise processing method in the above-mentioned embodiment when executing the program stored in the memory 1203.

[0296] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0297] The communication interface is used for communication between the above electronic device and other devices.

[0298] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0299] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0300] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned image noise processing network training methods and image noise processing methods are implemented.

[0301] In another embodiment provided by the present invention, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any one of the image noise processing network training methods and image noise processing methods in the above embodiments.

[0302] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0303] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0304] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0305] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A method for training an image noise processing network, characterized in that: The method comprises: Acquire multiple sample images containing motion regions and having noise to obtain a sample image sequence; wherein the motion region in each sample image is: a region where the content of the sample image changes relative to the previous sample image; For each sample image except the first sample image in the sample image sequence, the sample image is fused with a sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image; Calculate the ratio of the cumulative identification of each pixel position in the sample image to the total number of sample images to obtain the weight of the pixel position; wherein the cumulative identification of a pixel position is positively correlated with the number of times that the pixel position continuously belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image; According to the obtained weights, a weighted sum of the fused image and the true value image corresponding to the sample image is calculated as a reference image with uniform noise corresponding to the sample image; wherein the true value image corresponding to the sample image has the same image content as the sample image and is noise-free; The fused image corresponding to the sample image is input into the image noise processing network of the initial structure for noise uniformization processing to obtain a result image corresponding to the sample image; Based on the difference between the result image corresponding to the sample image and the reference image, the network parameters of the image noise processing network of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image noise processing network.

2. The method according to claim 1, characterized in that: The step of acquiring a plurality of sample images containing a motion region and having noise comprises: Acquire an original image sequence including a plurality of original images of uniform size and without noise, and a motion image having a size smaller than each original image; Using the motion image, replacing an image region of each original image that is the same size as the motion image, to obtain an image to be used corresponding to each original image; wherein, in the original image sequence, positions of the replaced image regions in any two adjacent original images are different; Adding noise to each of the obtained images to be used respectively to obtain a plurality of sample images containing motion areas and having noise; The true value image corresponding to each sample image is the image to be used when obtaining the sample image.

3. The method according to claim 2, characterized in that Before adding noise to each of the obtained images to be used to obtain a plurality of sample images containing motion areas and having noise, the method further includes: Acquiring noise parameters of a sensor that collects the original image; The step of adding noise to each of the obtained images to be used to obtain a plurality of sample images containing motion areas and having noise includes: Based on the noise parameters, noise is added to each of the obtained images to be used, so as to obtain a plurality of sample images containing motion areas and with noise.

4. The method according to claim 2, characterized in that: The using of the motion image to replace an image region in each original image that is the same size as the motion image to obtain an image to be used corresponding to each original image includes: Using the motion image, replacing an image region in each original image that has the same size as the motion image, to obtain an intermediate image corresponding to each original image; The brightness of the intermediate image corresponding to each original image is reduced to obtain the image to be used corresponding to each original image.

5. The method according to claim 1, characterized in that The cumulative mark of each pixel position in the first sample image in the sample image sequence is 1; For each sample image other than the first sample image in the sample image sequence, the cumulative identifier of each pixel position in the sample image is determined by the following steps: Determine a motion region and a non-motion region contained in a reference image corresponding to the sample image in a reference image sequence, and use them as a reference motion region and a reference non-motion region, respectively; wherein the reference image sequence is the sample image sequence or a true value image sequence composed of true value images corresponding to each of the plurality of sample images; and the position sequence of the reference image corresponding to the sample image in the reference image sequence is consistent with the position sequence of the sample image in the sample image sequence; For each pixel position in the sample image, if the pixel position belongs to the reference motion area in the corresponding reference image, determining that the cumulative flag of the pixel position in the sample image is 1; If the pixel position belongs to the reference non-motion area in the corresponding reference image, then 1 is added to the cumulative identification of the pixel position in the sample image before the sample image to obtain the cumulative identification of the pixel position in the sample image.

6. The method according to any one of claims 1 to 5, characterized in that: The image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork; The step of inputting the fused image corresponding to the sample image into the image noise processing network of the initial structure for noise uniformization processing to obtain the result image corresponding to the sample image comprises: Inputting the fused image corresponding to the sample image into the encoding subnetwork to obtain a first feature vector; Based on the first feature vector and the decoding sub-network, obtain a second feature vector; Based on the second feature vector and the output sub-network, a result image corresponding to the sample image is obtained.

7. The method according to claim 6, characterized in that The encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork; The step of inputting the fused image corresponding to the sample image into the encoding subnetwork to obtain a first feature vector includes: Inputting the fused image corresponding to the sample image into the first convolutional sub-network to obtain a first sub-feature vector; Inputting the first sub-feature vector into the second convolutional sub-network to obtain a second sub-feature vector; Inputting the second sub-feature vector into the third convolutional sub-network to obtain a first feature vector; Based on the first feature vector and the decoding sub-network, a second feature vector is obtained, including: Inputting the first feature vector into the fourth convolutional sub-network to obtain a third sub-feature vector; Inputting the third sub-feature vector and the second sub-feature vector into the fifth convolutional sub-network to obtain a second feature vector; Based on the second feature vector and the output sub-network, a result image corresponding to the sample image is obtained, including: The second feature vector and the first sub-feature vector are input into the output sub-network to obtain a result image corresponding to the sample image.

8. A method for processing image noise, characterized in that: The method comprises: Get the image to be processed; The image to be processed is input into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed; wherein the image noise processing network is trained based on any one of the methods of claims 1-7 above.

9. The method according to claim 8, characterized in that The method further comprises: The resulting image is subjected to denoising using the preset denoising parameters to obtain an enhanced image.

10. The method according to claim 8, characterized in that The image noise processing network includes: an encoding subnetwork, a decoding subnetwork and an output subnetwork; The image to be processed is input into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed, including: Inputting the image to be processed into the encoding subnetwork to obtain a third feature vector; Based on the third feature vector and the decoding sub-network, obtaining a fourth feature vector; Based on the fourth eigenvector and the output subnetwork, a result image corresponding to the image to be processed is obtained.

11. The method according to claim 10, characterized in that The encoding subnetwork includes a first convolutional subnetwork, a second convolutional subnetwork, and a third convolutional subnetwork, and the decoding subnetwork includes a fourth convolutional subnetwork and a fifth convolutional subnetwork; Inputting the image to be processed into the encoding subnetwork to obtain a third feature vector includes: Inputting the image to be processed into the first convolutional sub-network to obtain a fourth sub-feature vector; Inputting the fourth sub-feature vector into the second convolutional sub-network to obtain a fifth sub-feature vector; Inputting the fifth sub-feature vector into the third convolutional sub-network to obtain a third feature vector; Based on the third feature vector and the decoding sub-network, a fourth feature vector is obtained, including: Inputting the third feature vector into the fourth convolutional sub-network to obtain a sixth sub-feature vector; Inputting the sixth sub-feature vector and the fifth sub-feature vector into the fifth convolutional sub-network to obtain a fourth eigenvector; Based on the fourth eigenvector and the output subnetwork, a result image corresponding to the image to be processed is obtained, including: The fourth eigenvector and the fourth sub-eigenvector are input into the output subnetwork to obtain a result image corresponding to the image to be processed.

12. An image noise processing network training device, characterized in that: The device comprises: The first acquisition module is used to acquire a plurality of sample images containing motion regions and having noise, and obtain a sample image sequence; wherein the motion region in each sample image is: a region where the content of the sample image changes relative to the previous sample image; a fusion module, configured to fuse, for each sample image except the first sample image in the sample image sequence, the sample image with a sample image located before the sample image in the sample image sequence to obtain a fused image corresponding to the sample image; A first calculation module is used to calculate the ratio of the cumulative identification of each pixel position in the sample image to the total number of sample images to obtain the weight of the pixel position; wherein the cumulative identification of a pixel position is positively correlated with the number of times that the pixel position continuously belongs to the non-motion area in the corresponding adjacent sample image; in the sample image sequence, the adjacent sample image corresponding to the pixel position is located before the sample image and is adjacent to the sample image; A second calculation module is used to calculate the weighted sum of the fused image and the true value image corresponding to the sample image according to the obtained weights as a reference image with uniform noise corresponding to the sample image; wherein the true value image corresponding to the sample image has the same image content as the sample image and is noise-free; The first input module is used to input the fused image corresponding to the sample image into the image noise processing network of the initial structure for noise uniformization processing to obtain a result image corresponding to the sample image; The adjustment module is used to adjust the network parameters of the image noise processing network of the initial structure based on the difference between the result image corresponding to the sample image and the reference image until the preset convergence condition is reached to obtain a trained image noise processing network.

13. An image noise processing device, characterized in that: The device comprises: A second acquisition module, used for acquiring an image to be processed; The second input module is used to input the image to be processed into a pre-trained image noise processing network for noise uniformization processing to obtain a result image corresponding to the image to be processed; wherein the image noise processing network is trained based on the image noise processing network training device in claim 12 above.

14. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-11 when executing a program stored in a memory.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

16. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Equipment and method used for multi-frame fusion of strong noise images

    CN103985106A

  • Method and Apparatus for Noise Reduction

    CN110120015A

  • Multi-viewpoint linked tracking system using particle filter in object area detected through background modeling and method for same

    KR101422084B1

  • Image processing apparatus, image processing method, and program

    US20180338156A1

  • Preserving detail in denoised images for content generation systems and applications

    US20240096050A1