Information processing device, information processing method, and program
The described method for SPAD sensors addresses the loss of positive peak signals by using a two-stage process of interpolation and noise reduction, ensuring effective image restoration in low-light conditions.
Patent Information
- Application Number
- JP2024049992
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-26
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2044-03-26
AI Technical Summary
Existing noise reduction methods for SPAD sensors in low-light environments result in the loss of positive peak signals, which are essential for high-gain imaging, as they are mistakenly identified as noise and removed during conventional noise reduction processes.
An information processing device and method that includes a discrimination unit to identify positive peak signals, an interpolation unit to broaden these signals, and a noise reduction unit using CNN to maintain signal integrity while reducing noise, employing a two-stage process of interpolation followed by noise reduction.
The method effectively restores image quality by minimizing signal loss and noise, particularly in low-light conditions, maintaining the integrity of positive peak signals in images captured with SPAD sensors.
Smart Images

Figure 0007681147000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In recent years, a type of image sensor called a SPAD (Single Photon Avalanche Diode) sensor has been developed and installed in cameras.
[0003] CMOS sensors are well known as camera sensors, but when they read out light as an electrical signal, they also mix in noise that reduces image quality. On the other hand, although SPAD sensors do not generate readout noise, they do generate shot noise and dark current noise, so when the signal is amplified, the noise is also amplified, and the amount of noise is more noticeable the higher the sensitivity (gain) at which images are taken.
[0004] As a method for reducing noise, for example, Patent Document 1 discloses a technique for smoothing the signal level of a local region after removing impulse noise. Patent Document 2 discloses a technique for performing noise reduction processing in the spatial direction for each frame, calculating the amount of motion between frames, and performing noise reduction processing in the time direction based on the amount of motion.
[0005] Furthermore, as a method for reducing noise in a scanning electron microscope using machine learning, a technology has been disclosed in which artificial noise is added to a training image and an anisotropic filter is applied in the scanning direction to stretch the noise in a specific direction, as in Patent Document 3. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] JP 2010-092461 A [Patent Document 2] JP 2010-171808 A [Patent Document 3] JP 2023-170078 A Summary of the Invention [Problem to be solved by the invention]
[0007] The difference between CMOS sensors and SPAD sensors becomes evident when shooting at high gain. When shooting at high gain with a camera equipped with a SPAD sensor in a low-light environment, an image is obtained in which there are sparse pixels that were able to count photons, and positive peak signals are sparsely present due to avalanche amplification. This localized positive peak signal is a signal that can be obtained when there are pixels that were just able to count a few photons, and is useful as subject information. This signal is unique to imaging devices that have photon-counting image sensors such as SPAD sensors.
[0008] However, when the noise reduction processes of Patent Documents 1 and 2 are performed on such an image, the positive peak signals are judged as singular points, that is, noise, and effective signals that should be retained tend to be reduced, that is, lost.
[0009] Furthermore, even if, as in Patent Document 3, noise specific to the imaging device is added to the teacher image, and then the resulting student and teacher images are processed using an anisotropic filter, they are paired, and machine learning is used to learn to make the student image closer to the teacher image, signal loss will still occur.
[0010] This is because there are two conflicting tasks: whether to remove the signal (= noise reduction processing) or to generate a new signal in a place where no signal exists and interpolate it (= interpolation processing), making it difficult to derive an optimal solution. Also, if the signal is made directional by an anisotropic filter, the signal may remain after noise reduction, resulting in unnatural processing results.
[0011] Therefore, the present invention restores degradation while reducing the loss of positive peak signals in an image. [Means for solving the problem]
[0012] In order to solve this problem, for example, an information processing device of the present invention has the following arrangement. A discrimination means for acquiring an image including a target pixel which is a pixel having a local positive peak signal; an interpolation means for interpolating an interpolation signal into surrounding pixels which are pixels surrounding the target pixel of the image; a reduction means for reducing noise in an image into which the interpolated signal is interpolated; has. Effect of the Invention
[0013] According to the present invention, degradation restoration can be performed while reducing the loss of positive peak signals in an image. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing system according to an embodiment. [Diagram 2] FIG. 1 is a block diagram showing the functional configuration of an information processing system according to a first embodiment. [Diagram 3] FIG. 4 is a diagram for explaining a positive peak signal. [Figure 4] A diagram explaining the structure of CNN and the flow of learning and inference. [Diagram 5] 3A to 3C are views for explaining signal interpolation and noise reduction according to the first embodiment. [Figure 6] FIG. 4 is a flowchart showing a process flow according to the first embodiment. [Figure 7] FIG. 11 is a block diagram showing the functional configuration of an information processing system according to a second embodiment. [Figure 8] 10A to 10C are views for explaining signal interpolation and noise reduction according to a second embodiment. [Figure 9] FIG. 11 is a flowchart showing a process flow according to the second embodiment. [Figure 10] FIG. 11 is a block diagram showing the functional configuration of an information processing system according to a third embodiment. [Figure 11] FIG. 13 is a diagram for explaining the flow of inference and learning according to the third embodiment. [Figure 12] FIG. 13 is a diagram for explaining a degradation imparting method. [Figure 13] FIG. 11 is a flowchart showing the flow of information processing according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0016] <CNNについて> First, a description will be given of a Convolutional Neural Network (CNN) used in information processing technology generally using deep learning, which is used in the following embodiment. CNN is a technology that repeats convolution of a filter generated by training with image data, followed by nonlinear calculation. The filter is also called a local receptive field. Image data obtained by convolving the filter with image data and then performing nonlinear calculation is called a feature map. Furthermore, learning is performed using training data (training images or data sets) consisting of a pair of input image data and output image data. Simply put, learning is the generation of filter values that can convert input image data into corresponding output image data with high accuracy from training data. This will be described in detail later.
[0017] When image data has RGB color channels or when a feature map is composed of multiple pieces of image data, the filter used for convolution also has multiple channels accordingly. In other words, a convolution filter is expressed as a four-dimensional array that includes the number of channels in addition to the vertical and horizontal sizes and number of sheets. The process of nonlinear calculation after convolving a filter with image data (or a feature map) is expressed in units of layers, such as the nth layer feature map or the nth layer filter. In addition, for example, a CNN that repeats filter convolution and nonlinear calculation three times has a three-layer network structure. Such nonlinear calculation processing can be formulated as the following equation (1).
[0018]
number
[0019]
number
[0020] As networks using CNN, ResNet in the field of image recognition and its application SRCNN (Super Resolution CNN) in the field of super-resolution are well-known. In both cases, by stacking multiple layers of CNN and performing convolution of filters multiple times, high-precision processing is achieved. For example, ResNet features a network structure with a shortcut path for convolutional layers, enabling the realization of a 152-layer deep network and achieving high-precision recognition approaching the human recognition rate. The reason why processing becomes more accurate with multi-layer CNNs is simply that by repeating non-linear operations multiple times, a non-linear relationship between input and output can be expressed.
[0021] <Learning of CNN> Next, the learning of CNN will be explained. The learning of CNN is generally performed by minimizing the objective function represented by the following equation (3) for the training data consisting of a pair of input training image (student image) data and corresponding output training image (teacher image) data.
[0022]
Equation
[0023] In the above equation (3), L is a loss function that measures the error between the correct answer and its estimation. Also, Y i is the i-th output training image data, and X i is the i-th input training image data. Also, F is a function that collectively represents the operations (Equation 1) performed in each layer of the CNN. Also, θ is the network parameter (filter and bias). Also, ||Z|| 2is the L2 norm, or more simply, the square root of the sum of the squares of the elements of vector Z. Additionally, n is the total number of training data used for learning. Since the total number of training data is generally large, in the Stochastic Gradient Descent method (SGD), a portion of the training image data is randomly selected and used for learning. This reduces the computational load in learning using a large amount of training data. Additionally, various methods are known for minimizing (optimizing) objective functions, including the momentum method, the AdaGrad method, the AdaDelta method, and the Adam method. The Adam method is given by the following equation (4).
[0024]
number
[0025] In the above formula (4), θ i t is the i-th network parameter at the t-th iteration, and g is θ i t is the gradient of the loss function L with respect to . Additionally, m and v are moment vectors, α is the base learning rate, β1 and β2 are hyperparameters, and ε is a small constant. Note that since there are no guidelines for selecting an optimization method in learning, essentially any method can be used, but it is known that differences in convergence between methods will result in differences in learning time.
[0026] In this embodiment, information processing (image processing) is performed to reduce degradation of still images using the above-mentioned CNN. The degradation factor of an image is noise. The degradation restoration process in this embodiment is a process of generating or restoring an image without degradation (or with very little degradation) from an image with degradation, and will be referred to as degradation restoration process in the following description.
[0027] (First embodiment: approach from the inference side (still image)) In the first embodiment, a method for performing degradation restoration with reduced signal loss will be described by performing an interpolation process to broaden the positive peak signal of the input image data in the first stage, and performing a process to reduce noise in the input image data in the second stage.
[0028] <Example of information processing system configuration> Fig. 1 is a diagram showing an example of a hardware configuration of an information processing system according to a first embodiment. The information processing system shown in Fig. 1 includes an edge device 100 that performs degradation restoration (hereinafter, degradation restoration inference) and a cloud server 200 that performs learning to generate learning data and restore image quality degradation (hereinafter, degradation restoration learning). The edge device 100 and the cloud server 200 are examples of information devices, and are connected to each other via the Internet so that they can transmit and receive data to each other.
[0029] <Edge device hardware configuration> The edge device 100 of this embodiment acquires RAW image data of a Bayer array input from the imaging device 10 as an input image to be subjected to degradation restoration processing. In the following description, the term "image" may include image and image data. Then, noise in the RAW image data is reduced by executing an information processing application program installed in advance. The edge device 100 is an information processing device. The edge device 100 has a CPU 101, a RAM 102, a ROM 103, a large-capacity storage device 104, a general-purpose I / F 105, a network I / F 106, and a system bus 107. I / F is an abbreviation for interface. The components of the edge device 100 are connected to each other via the system bus 107 so as to be able to transmit and receive data to and from each other. The edge device 100 is connected to the imaging device 10, the input device 20, the external storage device 30, and the display device 40 via the general-purpose I / F 105 so as to be able to transmit and receive data to and from each other.
[0030] The CPU 101 is an abbreviation of Central Processing Unit, and is an arithmetic processing device. Instead of or in addition to the CPU 101, the edge device 100 may have other processors such as an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), and a QPU (Quantum Processing Unit). The CPU 101 executes programs stored in the ROM 103 and the large-capacity storage device 104, etc., using the RAM 102 as a work memory. In this way, the CPU 101 realizes various functions and centrally controls each component of the edge device 100 via the system bus 107. The edge device 100 may realize various functions by the CPU 101 and other processors. A part or all of each function of the edge device 100 may be realized by one or more circuits, such as an ASIC (Application Specific Integrated Circuit) and a PLD (Programmable Logic Device) including an FPGA (Field Programmable Gate Array).
[0031] The RAM 102 is an abbreviation for Random Access Memory, and is a memory that can be read and written at high speed. The RAM 102 temporarily stores the programs executed by the CPU 101, parameters required for executing the programs, and the like.
[0032] ROM 103 is an abbreviation for Read Only Memory, and is a non-volatile storage device that can retain data even when power is not being supplied.
[0033] The mass storage device 104 is a non-volatile secondary storage device such as a hard disk drive (HDD) and a solid state drive (SSD), and stores various data handled by the edge device 100. The mass storage device 104 stores data sent via the system bus 107 based on instructions from the CPU 101, and reads out the stored data and transfers it to the CPU 101, etc.
[0034] The general-purpose I / F 105 is a serial bus interface such as USB, IEEE1394, or HDMI (registered trademark). The edge device 100 acquires data from the external storage device 30 via the general-purpose I / F 105. The external storage device 30 is, for example, various storage media such as a memory card, a CF card, an SD card, or a USB memory. The general-purpose I / F 105 accepts a user instruction from an input device 20 such as a mouse or a keyboard, and outputs the instruction to the CPU 101 or the like. The general-purpose I / F 105 outputs image data or the like processed by the CPU 101 to the display device 40. The display device 40 is, for example, an image display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. The general-purpose I / F 105 acquires data of a captured image such as a RAW image to be subjected to degradation restoration processing from the imaging device 10, and outputs the data to the CPU 101.
[0035] The network I / F 106 is an interface for connecting to a network such as the Internet. The network I / F 106 transmits data to an external device such as a cloud server 200 via the network based on an instruction from the CPU 101, for example.
[0036] <Cloud server hardware configuration> The cloud server 200 of this embodiment is an information processing device that provides cloud services on the Internet. More specifically, the cloud server 200 generates learning data and performs degradation restoration learning, and trains a model that stores the network parameters and network structure of the learning result to generate a trained model. Then, the cloud server 200 provides the trained model in response to a request from the edge device 100. The cloud server 200 has a CPU 201, a ROM 202, a RAM 203, a mass storage device 204, and a network I / F 205, and the respective components are connected to each other by a system bus 206.
[0037] The CPU 201 reads out control programs stored in the ROM 202, the mass storage device 204, etc., to realize various functions and execute various processes, thereby controlling the overall operation of the cloud server 200. Instead of or in addition to the CPU 201, the cloud server 200 may have other processors such as an MPU, a GPU, and a QPU.
[0038] The ROM 202 stores the programs executed by the CPU 201 and the like.
[0039] The RAM 203 is used as a main memory for the CPU 201 and as a temporary storage area such as a work area.
[0040] The large-capacity storage device 204 is a large-capacity secondary storage device such as an HDD or SSD that stores image data, various programs, and parameters required for executing the programs.
[0041] The network I / F 205 is an interface for connecting to the Internet, and provides the above-mentioned network parameters in response to a request from the web browser of the edge device 100.
[0042] <Configuration of imaging device> The imaging device 10 has, for example, a SPAD (Single Photon Avalanche Diode) sensor as an imaging element. The SPAD sensor is an element that amplifies charges generated by photoelectric conversion by avalanche amplification and outputs them as an electrical signal. Avalanche amplification is a phenomenon in which electrons accelerated by an electric field in an impurity diffusion region of a PN junction collide with lattice atoms and break their bonds, and new electrons thus generated collide with other lattice atoms and break their bonds, repeating this process, thereby multiplying the current. The SPAD sensor is an imaging element that uses a photon counting method. The SPAD sensor discretely counts the number of photons and eliminates the influence of electrical noise (read noise), thereby converting the detected slight light into a signal and amplifying it. This allows the SPAD sensor to capture an object even in the dark.
[0043] The edge device 100 and the cloud server 200 have other components in addition to those described above, but their description will be omitted here. In this embodiment, it is assumed that the cloud server 200 downloads a trained model, which is a result of performing training data generation and degradation restoration learning, to the edge device 100, and the edge device 100 performs degradation restoration inference on the input image data 114 to be processed. The above-mentioned system configuration is an example, and is not limited to this. For example, the functions of the cloud server 200 may be subdivided, and the training data generation and degradation restoration learning may be performed by separate devices. Alternatively, the imaging device 10, which has both the functions of the edge device 100 and the cloud server 200, may perform all of the training data generation and degradation restoration learning.
[0044] <System Functional Blocks> Next, the functional configuration of the entire information processing system in this embodiment will be described with reference to Fig. 2. Fig. 2 is a functional block diagram showing the functional configuration of the information processing system. The cloud server 200 holds a trained model 210 that has been trained to restore degradation that occurs in the imaging device 10 by degradation restoration learning. The cloud server 200 transmits the trained model 210 to the edge device 100 in response to a request from the edge device 100 or the like. Since the learning method is not the main focus of the present embodiment, a detailed description thereof will be omitted.
[0045] The configuration shown in FIG. 2 can be modified or changed as appropriate. For example, one functional unit may be divided into multiple functional units, or two or more functional units may be integrated into one functional unit. The configuration shown in FIG. 2 may be realized by two or more devices. In this case, the devices are connected via a circuit or a wired or wireless network, and perform data communication with each other to perform cooperative operations, thereby realizing each process according to this embodiment.
[0046] The following describes in detail each of the functional units of the edge device 100. As shown in FIG.
[0047] The signal discrimination unit 111 acquires input image data 114 including pixels of local positive peak signals (an example of target pixels). For example, the signal discrimination unit 111 may acquire input image data 114 captured by a photon counting type image sensor in a low illuminance environment and free of black floating. The signal discrimination unit 111 determines whether or not pixels of local positive peak signals are sparsely present in each local region of the input image data 114. In this embodiment, a RAW image captured by a Bayer array color filter is used. Here, the positive peak signal will first be described with reference to FIG. 3.
[0048] Positive peak signals in low gain shooting and high gain shooting will be described with reference to Fig. 3. Fig. 3 is a diagram for explaining positive peak signals.
[0049] Figure 3(A) shows an image taken with low gain shooting in a sufficiently bright environment. In such an environment, photons are easy to detect, so low gain shooting can produce sufficiently clear results. However, in a low-light environment that is darker than starlight, it is necessary to increase the gain (sensitivity) to shoot.
[0050] FIG. 3B is a diagram showing an image captured with an increased gain in a dark, low-illumination environment. Note that FIG. 3B captures almost the same location as FIG. 3A. The positive peak signal may be a signal that spans multiple pixels. Note that the positive peak signal may be a signal of a single pixel. The pixels surrounding the target pixel having the local positive peak signal have a value close to the output level (black level) of the imaging device when there is no incident light. The signal close to the black level may be, for example, a signal whose signal intensity (also called magnitude) is equal to or less than a predetermined level threshold. The positive peak signal is, for example, a signal that exists within the rectangle in FIG. 3B. This is because there are pixels that can detect even one photon and pixels that cannot detect one, and the former undergo avalanche amplification and gain, resulting in a sparse local positive peak signal in the image. In addition, the SPAD sensor has the characteristic that there is no read noise that cannot be ignored during high-gain shooting, and the gain is mainly applied only to the signal, so that the black level does not change even during high-gain shooting. On the other hand, when an image capture device with a CMOS sensor captures an image at high gain, a phenomenon known as black floating occurs, in which pixels that should be at a black level become bright due to the effects of readout noise, and the peak signal becomes buried in the noise.
[0051] The signal discrimination unit 111 refers to settings such as a gain value and an exposure value set in the imaging device 10, and discriminates whether or not the acquired image has a positive peak signal based on the settings.
[0052] The lower the illuminance, the higher the gain setting must be set in order to see the subject, and the higher the gain, the more noticeable the positive peak signal becomes.
[0053] When a positive peak signal is included in the input image data 114, the signal interpolation unit 112 interpolates a signal (an example of an interpolated signal) to surrounding pixels that are pixels surrounding a target pixel of the input image. For example, the signal interpolation unit 112 spreads the local positive peak signal to the periphery and interpolates the signal to surrounding pixels at the black level and surrounding pixels close to the black level.
[0054] The signal interpolation unit 112 performs smoothing by weighted averaging of pixels included in the periphery of a target pixel of a local positive peak signal using spatial correlation information between the target pixel and multiple surrounding pixels. The periphery of the target pixel may be within a predetermined set area. The signal interpolation unit 112 performs smoothing on all pixels of the input image data 114, thereby interpolating the positive peak signal of the entire image to the surrounding pixels.
[0055] The aim of this interpolation is to reduce the difficulty of the task of the interpolation process and to specialize it for the task of noise reduction process performed in the subsequent noise reduction unit 113. Therefore, in smoothing, the signal interpolation unit 112 maintains the positive peak signal as much as possible so that signal information is transmitted to pixels close to the black level around the positive peak signal. The signal interpolation unit 112 applies, for example, edge-preserving filtering as smoothing that maintains the peak signal.
[0056] It is desirable that the surrounding pixels to be interpolated have no anisotropy, assuming that the peak signal is isotropic. This is to prevent the peak signal from having directionality and to prevent unnatural signals from remaining in the results of the noise reduction process at the subsequent stage.
[0057] The noise reduction unit 113 performs processing to reduce noise (hereinafter also referred to as noise reduction processing) on input image data 114 in which an interpolated signal generated from a positive peak signal or the like is interpolated by the signal interpolation unit 112. The noise reduction unit 113 uses CNN to repeat convolution calculations and nonlinear calculations using filters expressed by equations (1) and (2) multiple times, and outputs output image data 115 using a model that has learned degradation restoration processing.
[0058] Fig. 4(A) is a diagram for explaining the structure of CNN and the flow of learning and inference. As shown in Fig. 4(A), the CNN used in this embodiment has a plurality of filters 401. The noise reduction unit 113 sequentially applies the filters 401 to this input data and calculates a feature map (not shown). Then, the noise reduction unit 113 outputs output image data 115 having the same number of channels as the input image data 114 from the final filter.
[0059] The flow up to this point will be described with reference to Fig. 5. Fig. 5 is a diagram for explaining signal interpolation and noise reduction according to the first embodiment. In Fig. 5(A) to Fig. 5(C), the horizontal axis represents pixel position, and the vertical axis represents pixel value of each pixel. Note that the vertical axis may represent luminance.
[0060] Fig. 5(A) shows signal values (pixel values) of pixels in one horizontal line in a local region of input image data. The black circles in the figure represent original signals that exist from the beginning. In a region 501 surrounded by a dotted line, for example, pixel 502 has a constant pixel value and the surrounding pixels have pixel values close to the black level, and is a target pixel of a sparse local positive peak signal. Pixel 502 may be a single pixel or multiple pixels.
[0061] Fig. 5(B) shows the result when an NxN (e.g., N=3) smoothing filter is applied as the interpolation process performed by the signal interpolation unit 112 under the conditions of Fig. 5(A). The black circles in the figure represent the original signals that existed from the beginning, and the white circles represent the interpolated signals that have been created by spreading the original signals due to the smoothing process. In this way, the positive peak signal spreads to the surroundings, and as a result, information of the original signal is added to pixels close to the black level.
[0062] Fig. 5(C) shows the result when the noise reduction unit 113 performs noise reduction processing using a trained model under the conditions of Fig. 5(B). The black triangles in the figure represent the noise-reduced output signal. The noise reduction processing reduces the variation of each pixel, improving the visibility of edges (dashed lines in the figure) that were difficult to see in the input image.
[0063] <Processing flow of the entire system> Next, various processes performed in the information processing system of this embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the flow of processes in the information processing system of this embodiment. Each process shown in Fig. 6 is executed by each functional unit in Fig. 2 which is realized by CPU 101 executing a computer program for image processing according to this embodiment. However, all or part of the functional units shown in Fig. 2 may be implemented in hardware. Below, an explanation will be given along with the flowchart in Fig. 6. In the following explanation, the symbol "S" means a processing step.
[0064] In S601, the signal discrimination unit 111 acquires input image data 114 to be processed. The signal discrimination unit 111 may directly acquire input image data captured by the imaging device 10, for example, or may read input image data captured in advance and stored in the large-capacity storage device 104.
[0065] In S602, the signal discrimination unit 111 discriminates whether or not a local positive peak signal is included in the input image data 114. If the signal discrimination unit 111 determines that a local positive peak signal is included, the process proceeds to S603, and if it determines that a local positive peak signal is not included, the process proceeds to S604.
[0066] In S603, the signal interpolation unit 112 spatially interpolates the signal of the input image when a target pixel of a local positive peak signal is included in the input image data 114. For example, the signal interpolation unit 112 interpolates and smoothes the signal in the input image data 114 based on a weighted average smoothing process so that the signal spreads to pixels close to the black level that are pixels surrounding the pixel of the local positive peak signal.
[0067] In S604, the noise reduction unit 113 reduces noise in the input image data 114 whose signals have been interpolated by the signal interpolation unit 112 through smoothing processing. Then, the noise reduction unit 113 outputs the image data after noise reduction as output image data 115.
[0068] The above is the overall flow of the processing carried out in the information processing system of this embodiment.
[0069] In the first embodiment, a method for restoring degraded input images while reducing signal loss has been described by performing a two-stage process in which a positive peak signal of the input image data is interpolated around the signal and noise in the input image data is reduced.
[0070] The characteristic of having a positive peak signal that occurs during high-gain shooting is unique to imaging devices equipped with SPAD sensors, and until now, when shooting in low-light environments, the signal in the output image has often been lost.
[0071] On the other hand, as in the present embodiment, the degradation restoration task consisting of the interpolation process and the noise reduction process is executed in the order of the interpolation process and the noise reduction process, thereby reducing the loss of local positive peak signals in the image, while maintaining the signal, and reducing the noise, thereby achieving degradation restoration. In addition, a use case of the present embodiment is assumed to be a case where handheld shooting is performed in a low-illumination environment.
[0072] In this embodiment, the presence or absence of a local positive peak signal is determined based on the setting value of the imaging device 10, and therefore the determination process can be performed at high speed.
[0073] In this embodiment, the input image data is described as data of a RAW image captured with a Bayer color filter array, but image data of other color filter arrays may be used. The input image data may be any one of RGB image data obtained by developing a RAW image, YUV image data converted from the RGB image data, and JPEG image data compressed.
[0074] In this embodiment, the presence or absence of a target pixel having a local positive peak signal is determined based on the setting value of the imaging device 10, but the determination method is not limited to this. For example, the signal determination unit 111 may determine the presence or absence of a local positive peak signal based on either a pixel value or brightness. Specifically, the signal determination unit 111 may compare a pixel value difference between a target pixel having a positive peak signal of the input image data 114 to be determined and a plurality of pixels located in a predetermined area around the target pixel with a pixel threshold value set in advance, and if there are many pixels that are equal to or greater than the pixel threshold value, determine the target pixel as a target pixel of a local positive peak signal, and perform this determination for all pixels. Also, the signal determination unit 111 may compare a brightness difference between a target pixel of the input image data 114 to be determined and a plurality of pixels in a predetermined area around the target pixel with a brightness threshold value set in advance, and if there are many pixels that are equal to or greater than the brightness threshold value, determine the target pixel as a target pixel of a local positive peak signal, and perform this determination for all pixels. As a result, this embodiment can determine the presence or absence of a local positive peak signal with high accuracy for each input image.
[0075] Here, the signal discrimination unit 111 may determine whether or not there is a majority by majority vote, or may determine whether or not there is a majority based on whether or not the ratio of pixels that are equal to or greater than the threshold in a predetermined region exceeds a ratio threshold (e.g., 70%). In other words, the noise reduction unit 113 determines that the greater the ratio of positive peak signals in a predetermined region around the pixel of interest, the more localized the positive peak signal is, and leaves it as a target pixel.
[0076] In the present embodiment, the signal interpolation unit 112 performs smoothing on all pixels by taking a weighted average of the surrounding pixels as the interpolation process on the surrounding pixels of the positive peak signal. However, the signal interpolation unit 112 may perform similar smoothing only on the target pixel determined to be a local positive peak signal. Alternatively, the signal interpolation unit 112 may copy the target pixel determined to be a local positive peak signal to a nearby pixel included in a predetermined surrounding setting area, and interpolate the interpolated signal to the surrounding pixel. At this time, the strength of the interpolated signal may be determined by setting a weight according to the distance between the pixel having the positive peak signal and the pixel to be copied. For example, the weight W when the distance is D is set as W=C×1 / D. Here, C is a predetermined correction coefficient. The interpolation process may be performed by a Gaussian filter.
[0077] In the present embodiment, the process of reducing noise using CNN has been described, but the present invention is not limited to this. For example, edge-preserving, patch-based, or other rule-based noise reduction processes may be applied.
[0078] (Second embodiment: Approach from the inference side (video)) In the first embodiment, the input image data is one frame, but with one frame, it is difficult to distinguish between a locally generated positive peak signal and impulse noise, and there is a possibility that impulse noise will remain. Impulse noise is a sudden change in pixel value due to scratches on the sensor, and is displayed as random pixels close to black or white, and is also called salt and pepper noise.
[0079] In the second embodiment, an example of degradation restoration processing that improves the accuracy of distinguishing between the two by using multiple frames of input image data will be described. Note that, among the basic configurations of the information processing system, the description of the configurations common to the configurations described in the first embodiment will be omitted, and the following description will focus on the differences.
[0080] Fig. 7 is a block diagram showing the functional configuration of the information processing system according to the second embodiment. As shown in Fig. 7, an edge device 700 of the information processing system according to the second embodiment has a signal discrimination unit 701, a spatial direction signal interpolation unit 702, a time direction signal interpolation unit 703, and a noise reduction unit 704.
[0081] The signal discrimination unit 701 acquires an input image group 705 including images of a plurality of frames adjacent (or consecutive) in time series (or time). The signal discrimination unit 111 determines whether or not a local positive peak signal exists in at least one of time and space for each acquired frame based on the camera setting value. However, at this point, it is difficult for the signal discrimination unit 701 to distinguish whether the signal is a positive peak signal to be left or an impulse noise to be removed. Therefore, the signal discrimination unit 701 may refer to positive peak signals at the same pixel position in a plurality of frames adjacent to the frame of interest, and determine whether or not the positive peak signal is a signal to be left according to the frequency of occurrence of the positive peak signal. For example, the signal discrimination unit 701 may determine a positive peak signal present in many frames as a positive peak signal to be left, and a signal present in only a few frames as a positive peak signal not to be left, i.e., an impulse noise.
[0082] The spatial direction signal interpolation unit 702 performs the same interpolation process as in the first embodiment on frames including positive peak signals in the input image group 705, and interpolates signals around the positive peak signals.
[0083] The time direction signal interpolation unit 703 performs the interpolation process in the time direction, which has been performed in the spatial direction, on frames including positive peak signals in the input image group 705. Specifically, the time direction signal interpolation unit 703 performs the interpolation process by taking a weighted average of signals at the same pixel position in M (e.g., M=3) frames that are temporally adjacent (or consecutive) to the target frame.
[0084] The noise reduction unit 704 reduces noise in the images in which the positive peak signals have been interpolated by the spatial direction signal interpolation unit 702 and the time direction signal interpolation unit 703 , and outputs an output image group 706 .
[0085] The flow up to this point will be described in detail with reference to Fig. 8. Fig. 8 is a diagram for explaining signal interpolation and noise reduction according to the second embodiment.
[0086] Fig. 8(A) shows the pixel values of one horizontal line of pixels in a frame at each time T. With reference to the same pixel position in three frames, signals that exist in many frames are true signals including positive peak signals (black diamonds in Fig. 8(A)), while signals that exist only in a few frames are impulse noise (white diamonds in Fig. 8(A)).
[0087] Fig. 8(B) shows the result of performing the interpolation process in the spatial direction by the spatial direction signal interpolation unit 702 and the interpolation process in the time direction by the time direction signal interpolation unit 703 under the conditions of Fig. 8(A). This shows that the peak signal is spread around in the spatial and time directions and is interpolated to pixels with close black levels.
[0088] Fig. 8(C) shows the result when the noise reduction unit 704 performs noise reduction processing under the conditions of Fig. 8(B). By performing noise reduction processing by the noise reduction unit 704, the variation in each pixel is reduced. This improves the visibility of edges (dashed lines in the figure) that were difficult to see at the time of input image.
[0089] <Processing flow of the entire system> Next, various processes performed in the information processing system of the second embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of processes in the information processing system of the second embodiment. Each process shown in Fig. 9 is executed by each functional unit realized by CPU 101 executing a computer program corresponding to each process.
[0090] In S901, the signal discrimination unit 701 acquires the input image group 705 to be processed.
[0091] In S902, the signal discrimination unit 701 discriminates whether or not a local positive peak signal is included in either space or time for each frame of the input image group 705. If the signal discrimination unit 701 determines that each frame includes a positive peak signal, the process proceeds to S903, and if it determines that each frame does not include a positive peak signal, the process proceeds to S905.
[0092] In S903, the spatial direction signal interpolation unit 702 performs a weighted average smoothing process on the input images of frames of the input image group 705 that contain positive peak signals, thereby interpolating an interpolated signal generated from the peak signals in the spatial direction.
[0093] In S904, the time direction signal interpolation unit 703 performs a weighted average smoothing process on an input image of a frame containing a positive peak signal among each frame of the input image group 705, by referring to signals at the same pixel position in other chronologically adjacent frames, thereby interpolating an interpolated signal generated from the peak signal or the like in the time direction.
[0094] In S905, the noise reduction unit 704 reduces noise in each input image of the input image group 705 in which the interpolation signals have been interpolated by the spatial direction signal interpolation unit 702 and the time direction signal interpolation unit 703. Then, the noise reduction unit 704 outputs all frames after noise reduction as an output image group 706.
[0095] The above is the overall flow of the processing carried out in the information processing system of this embodiment.
[0096] In the second embodiment, when any frame of an input image group contains a local positive peak signal, an interpolation process is performed to expand the positive peak signal in the spatial direction to interpolate an interpolated signal, and an interpolation process is performed to expand the positive peak signal in the time direction to interpolate an interpolated signal. In this embodiment, after performing the two interpolation processes, noise in the input image group is reduced, thereby realizing degradation restoration with reduced signal loss.
[0097] In this way, in this embodiment, the accuracy of distinguishing between a true signal and impulse noise, which is difficult to do with a single frame, is improved by using multiple frames adjacent in time. As a result, this embodiment makes it possible to maintain the signal and reduce noise while suppressing the influence of impulse noise. In addition, possible use cases of this embodiment include generating one high-quality frame from multiple frames captured in chronological order at high gain in a low-illumination environment, and making surveillance video in a low-illumination environment clearer.
[0098] In this embodiment, the weighted average is used as a process for interpolating signals in the time direction, but the present invention is not limited to this, and a process for interpolating signals that exist exclusively may be used. Although this increases the risk of leaving impulse noise, it works in the direction of leaving peak signals, and therefore has the effect of minimizing information loss.
[0099] In this embodiment, an example in which the positive peak signal is spread to the periphery by explicitly dividing the signal into the spatial direction and the time direction has been described, but the method may be performed in the spatial-temporal direction. In this case, the embodiment may perform the interpolation process only on the target pixel having the positive peak signal and its surrounding pixels.
[0100] (Third embodiment: Approach from the learning side) In the first and second embodiments, an example was described in which a positive peak signal is expanded and then noise reduction using a neural network is performed. This utilizes a model that has learned noise reduction using a pair of a teacher image and an image to which noise has been added. As mentioned at the beginning, in an image that has a local positive peak signal, it is difficult to distinguish whether the task is a noise reduction task or an interpolation task. In the first and second embodiments, preprocessing is performed so that the signal is not lost during inference, but the accuracy of degradation restoration can be improved by performing preprocessing during learning as well.
[0101] In the third embodiment, noise is added to the teacher image to reproduce a local positive peak signal, and a student image is generated by spreading the peak signal in at least one of the spatial and temporal directions. In the third embodiment, an example is described in which a teacher image and a student image are paired and learning is performed to reduce noise, and a model is generated. Note that a description of the contents common to the configurations given in the first and second embodiments, such as the basic configuration of the information processing system, will be omitted, and the following description will focus on the differences.
[0102] <System Functional Blocks> The functional configuration of the entire information processing system in this embodiment will be described with reference to Fig. 10. Fig. 10 is a block diagram showing the functional configuration of the information processing system according to the third embodiment. The information processing system of this embodiment has an edge device 1000 and a cloud server 1010.
[0103] A detailed description will be given of each functional unit of the edge device 1000. The edge device 1000 has a signal discrimination unit 1001, an inference-use signal interpolation unit 1002, and an inference unit 1003. The inference unit 1003 includes an inference-use noise amount estimating unit 1004 and an inference-use noise reducing unit 1005.
[0104] The signal discriminator 1001 and the inference-use signal interpolator 1002 have the same functions as the signal discriminator 111 and the signal interpolator 112 of the first embodiment, respectively.
[0105] The inference unit 1003 estimates the amount of noise in the input image data 1006 using the trained model 1022 received from the cloud server, and performs degradation restoration inference based on the estimation result. The trained model 1022 is a neural network that estimates the amount of noise and performs degradation restoration. The degradation restoration inference is performed by an inference noise amount estimation unit 1004 and an inference noise reduction unit 1005.
[0106] The inference unit 1003 will be described in detail with reference to Fig. 11(A). Fig. 11(A) is a diagram showing the flow of processing in the inference unit 1003. The inference noise amount estimation unit 1004 acquires input image data 1006 and estimates the amount of noise in the input image data 1006 using the trained model 1022. A neural network is used to estimate the amount of noise. As shown in Fig. 11(A), the inference noise amount estimation unit 1004 inputs the input image data 1006 to a first CNN 1101, and repeats a convolution operation and a nonlinear operation using filters represented by formulas (1) and (2) multiple times to output a noise amount estimation result 1102 for the input image data.
[0107] Next, the processing of the first CNN 1101 will be described with reference to Fig. 11(A) and Fig. 4(A). The first CNN 1101 is composed of a plurality of filters 401 that perform the calculation of the above-mentioned formula (1). The noise amount estimation unit for inference 1004 inputs the input image data 1006 to this CNN. Next, the noise amount estimation unit for inference 1004 sequentially applies the filters 401 to the input image data 1006 to calculate a feature map. Thereafter, the noise amount estimation unit for inference 1004 outputs the result of applying the last filter 401 to the noise reduction unit for inference 1005 as a noise amount estimation result 1102. The noise amount estimation result 1102 has the same channel as the input image data 1006.
[0108] The noise reduction unit for inference 1005 receives the noise amount estimation result 1102, and performs restoration processing for degradation of the input image data 1006 based on the estimation result. Specifically, the noise reduction unit for inference 1005 inputs the input image data 1006 and the noise amount estimation result 1102 to a second CNN 1103. Then, the noise reduction unit for inference 1005 repeats convolution calculation and nonlinear calculation using filters expressed by equations (1) and (2) multiple times, and outputs output image data 1007 after restoration processing.
[0109] Next, the processing of the second CNN 1103 will be described with reference to FIG. 11(A) and FIG. 4(B). As shown in FIG. 4(B), the second CNN 1103 is composed of a plurality of filters 401 and a concatenation layer 402. First, the noise reduction unit for inference 1005 inputs to this CNN a concatenation or addition of the input image data 1006 and the noise amount estimation result 1102 in the channel direction. Next, the noise reduction unit for inference 1005 sequentially applies the filters 401 to this input data to calculate a feature map. Next, the noise reduction unit for inference 1005 concatenates the feature map and the input data in the channel direction by the concatenation layer 402. Furthermore, the noise reduction unit for inference 1005 sequentially applies the filters 401 to the concatenation result, and outputs the output image data 1007 having the same number of channels as the input image data 1006 from the final filter.
[0110] Next, each functional unit of the cloud server 1010 will be described. The cloud server 1010 trains a model for restoring an image to generate a trained model 1022, and passes it to the edge device 1000. As shown in Fig. 10, the cloud server 1010 has a training data generation unit 1011 and a training unit 1012. The training data generation unit 1011 includes a degradation adding unit 1013 and a training signal interpolation unit 1014. The training unit 1012 includes a training noise amount estimation unit 1015, a training noise reduction unit 1016, an error calculation unit 1017, and a model update unit 1018.
[0111] The degradation adding unit 1013 adds noise according to a binomial distribution to teacher image data extracted from a group of undegraded teacher images to generate student image data. At least a part of the noise signal may be a reproduction of a local positive peak signal. In the present embodiment, the degradation adding unit 1013 analyzes the physical characteristics of the imaging device based on the settings of the imaging device and generates student image data by adding noise corresponding to a wider range of degradation amounts than the amount of degradation that can occur in the imaging device as a degradation element to the teacher image data. The reason for adding a wider range of degradation amounts than the analysis result is to provide a margin to increase robustness since the range of degradation amounts varies depending on individual differences in the imaging device.
[0112] FIG. 12 is a diagram for explaining the degradation adding method. That is, as shown in FIG. 12, the degradation adding unit 1013 adds noise 1105 based on the physical property analysis result 1020 of the imaging device as a degradation element to teacher image data 1107 extracted from the teacher image group 1019, thereby generating student image' data 1202. Note that a part of the noise 1105 may be a signal that reproduces a local positive peak signal. The physical property analysis result 1020 of the imaging device may be data indicating the relationship between the variance of the noise and the brightness, for example, as shown in FIG. 12. Then, the learning signal interpolation unit 1014 performs an interpolation process similar to that of the signal interpolation unit 112 on the student image' data 1202, thereby generating student image data 1106. Finally, the degradation adding unit 1013 and the learning signal interpolation unit 1014 use the pair of teacher image data 1107 and student image data 1106 as learning data. The learning data generation unit 1011 adds a degradation element to each teacher image data of the teacher image group 1019 and performs an interpolation process to generate a student image group made up of multiple student image data, thereby generating learning data 1104.
[0113] The teacher image group 1019 stores various types of image data, such as nature photos including landscapes or animals, portraits such as portraits or sports photos, and photos of man-made objects such as architecture and products. The physical property analysis result 1020 of the imaging device includes the amount of noise for each sensitivity generated by the imaging sensor built into the camera (imaging device). By using these, it is possible to estimate the degree of image quality degradation that occurs for each shooting condition. In other words, the degradation adding unit 1013 can generate an image equivalent to the image obtained at the time of shooting by adding the degradation estimated under a certain shooting condition to the teacher image data.
[0114] The learning unit 1012 acquires network parameters 1021 to be applied to the CNN for degradation restoration learning, initializes the weights of the CNN using the network parameters, and then performs degradation restoration learning using the learning data generated by the degradation adding unit 1013. The network parameters 1021 include initial values of the parameters of the neural network, and hyperparameters indicating the structure and optimization method of the neural network. The degradation restoration learning in the learning unit 1012 is performed by a learning noise amount estimating unit 1015, a learning noise reducing unit 1016, an error calculating unit 1017, and a model updating unit 1018.
[0115] FIG. 11B is a diagram showing the flow of processing in the learning unit 1012.
[0116] The learning noise amount estimation unit 1015 receives learning data 1104 from the learning data generation unit 1011, and estimates the amount of noise from noise 1105 added to student image data 1106 included in the student image group of the learning data 1104. Specifically, the learning noise amount estimation unit 1015 first inputs the student image data 1106 to a first CNN 1101, and repeats convolution calculation and nonlinear calculation using filters expressed by equations (1) and (2) multiple times, and outputs a noise amount estimation result 1108.
[0117] The learning noise reduction unit 1016 receives student image data 1106 and a noise amount estimation result 1108 estimated by the learning noise amount estimation unit 1015, and performs noise reduction processing on the student image data 1106. Specifically, the learning noise reduction unit 1016 first inputs the student image data 1106 and the noise amount estimation result 1108 to a second CNN 1103, and repeats convolution calculations and nonlinear calculations using filters expressed by equations (1) and (2) multiple times to output a restoration result 1111.
[0118] The error calculation unit 1017 inputs the added noise 1105 and the noise amount estimation result 1108 to a first loss process 1109 which is a loss function calculation, and calculates the error therebetween.
[0119] Next, the model update unit 1018 inputs the error calculated by the error calculation unit 1017 to a first update process 1110, and updates the network parameters related to the first CNN 1101 so that the error becomes smaller (minimum).
[0120] Moreover, the error calculation unit 1017 inputs the teacher image data 1107 and the noise amount estimation result 1108 to a second loss processing 1112 to calculate the error therebetween. Then, the model update unit 1018 inputs the calculated error to a second update processing 1113 to update the network parameters related to the second CNN 1103 so as to reduce the error.
[0121] Here, the added noise 1105, the student image data 1106, and the noise amount estimation result 1108 all have the same number of pixels.
[0122] The learning noise amount estimation unit 1015 and the learning noise reduction unit 1016 calculate errors at different times, but update network parameters at the same time. The first CNN and the second CNN used in the learning unit 1012 may be the same neural networks as the first CNN and the second CNN used in the inference unit 1003, respectively.
[0123] <Processing flow of the entire system> Next, various processes performed in the information processing system of the third embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing the flow of processes in the information processing system of the third embodiment. Fig. 13(A) is a flowchart showing the flow of processes performed in the cloud server 1010. Fig. 13(B) is a flowchart showing the flow of processes performed in the edge device 1000. Each process in Fig. 13(A) and Fig. 13(B) is performed by each functional unit shown in Fig. 10 which is realized by CPU 101 or CPU 201 executing a computer program.
[0124] 13A, a description will be given of the flow of processing for learning a model performed by the cloud server 1010. It is assumed that the cloud server 1010 has previously stored a group of teacher images 1019 prepared by a user or the like, and a physical property analysis result 1020 of an imaging device, such as the characteristics of an imaging sensor and sensitivity at the time of shooting.
[0125] In S1301, the degradation adding unit 1013 acquires data of the teacher image group 1019 stored in the cloud server 1010 and the physical property analysis result 1020 of the imaging device.
[0126] In S1302, the degradation adding unit 1013 adds noise based on the physical property analysis result 1020 of the imaging device to the teacher image data of the teacher image group 1019 input in S1301 to generate student image data. Note that the degradation adding unit 1013 may add noise in an amount measured in advance based on the physical property analysis result 1020 of the imaging device in a preset order or in a random order.
[0127] In S1303, the learning signal interpolation unit 1014 interpolates signals in the student image data to which noise has been added. The signal interpolation here may be performed by the same interpolation process as in S603.
[0128] In S1304, the training noise amount estimation unit 1015 and the training noise reduction unit 1016 acquire network parameters 1021 to be applied to the CNN for degradation restoration training. As described above, the network parameters here include the initial values of the neural network parameters, and hyperparameters indicating the structure and optimization method of the neural network.
[0129] In S1305, the learning noise amount estimation unit 1015 initializes the weights of the CNN using the received network parameters 1021, and then estimates the amount of noise in the student image data generated in S1302. Then, the learning noise reduction unit 1016 reduces noise in the student image data based on the noise amount estimation result.
[0130] In S1306, the error calculation unit 1017 calculates the error between the noise amount estimation result and the teacher image data in accordance with the loss function shown in equation (3).
[0131] In S1307, the model update unit 1018 updates the network parameters so that the error obtained in S1306 becomes smaller (minimum).
[0132] In S1308, the learning unit 1012 determines whether or not to end the learning. If the learning unit 1012 determines not to end the learning, the process of the cloud server 1010 returns to S1305, and learning is performed using other student image data and teacher data by the process from S1305 onwards. On the other hand, if the learning unit 1012 determines to end the learning, this process ends.
[0133] Next, the flow of processing performed by the edge device 1000 according to the third embodiment will be described with reference to the flowchart in Fig. 13(B). It is assumed that the cloud server 1010 has transmitted the trained model 1022 and the input image data 1006 to the edge device 1000 in advance.
[0134] In S1309, the signal discrimination unit 1001 acquires the trained model 1022 trained in the cloud server 1010 and the input image data 1006.
[0135] In S1310, the signal discrimination unit 1001 discriminates whether or not there are sparse local positive peak signals in the input image data 1006. If the signal discrimination unit 1001 determines that there are local positive peak signals, the process proceeds to S1311. If the signal discrimination unit 1001 determines that there are no local positive peak signals, the process proceeds to S1312.
[0136] In S1311, the inference signal interpolation unit 1002 interpolates an interpolation signal to pixels around a positive peak signal of the input image.
[0137] In S1312, the noise amount for inference estimation unit 1004 estimates the amount of noise in the input image using the received trained model 1022. Then, the noise reduction unit for inference 1005 reduces the noise in the input image based on the estimation result.
[0138] The above is the overall flow of the processing performed in the information processing system of the third embodiment.
[0139] In the third embodiment, a method for improving the accuracy of degradation restoration was described by combining not only preprocessing on the inference side but also an approach from the learning side in which positive peak signals are reproduced and learned using an interpolated image expanded in the spatiotemporal direction.
[0140] This embodiment is premised on the use of a neural network, but can perform more effective degradation restoration than the first and second embodiments. Although it takes time because processing is performed when creating learning data, it does not affect inference, so it is possible to maintain the same processing speed as when using a neural network in the first and second embodiments.
[0141] In the third embodiment, the amount of noise in the student image data immediately after degradation is applied is estimated, but the amount of noise in the student image data whose signals have been interpolated may be estimated.
[0142] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for implementing one or more of the functions.
[0143] The above-mentioned embodiments are merely examples of the implementation of the present invention, and the technical scope of the present invention should not be interpreted as being limited by these. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.
[0144] (Other Examples) The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0145] The disclosure of this specification includes the following information processing device, information processing method, and program. (Item 1) A discrimination means for acquiring an image including a target pixel which is a pixel having a local positive peak signal; an interpolation means for interpolating an interpolation signal into surrounding pixels which are pixels surrounding the target pixel of the image; a reduction means for reducing noise in an image into which the interpolated signal is interpolated; 13. An information processing device comprising: (Item 2) The determining means determines whether the target pixel exists. 2. The information processing device according to item 1, (Item 3) The determining means determines whether the target pixel exists based on a setting value of an imaging device that captured the image. 3. The information processing device according to item 1 or 2. (Item 4) The discrimination means is A pixel value difference between a pixel value of a pixel of interest that is a pixel to be discriminated and pixel values of pixels surrounding the pixel of interest is compared with a predetermined pixel threshold value; and, A luminance difference between the luminance of the pixel of interest and the luminance of pixels surrounding the pixel of interest is compared with a predetermined luminance threshold value; The presence or absence of the target pixel is determined based on at least one of the above. 4. The information processing device according to any one of items 1 to 3, (Item 5) The image obtained by the determination means is free of black floating. 5. The information processing device according to any one of items 1 to 4, (Item 6) The image acquired by the discrimination means is captured by a photon counting type image sensor. 6. The information processing device according to any one of items 1 to 5, (Item 7) The image acquired by the discrimination means is taken in a low-illumination environment. 7. The information processing device according to any one of items 1 to 6, (Item 8) The local positive peak signal is over a number of pixels, The signals of the surrounding pixels are black level signals below a predetermined level threshold to obtain an image. 8. The information processing device according to any one of items 1 to 7, (Item 9) The interpolation means interpolates the interpolation signal to the surrounding pixels based on a positive peak signal included in a predetermined set region around the target pixel. 10. The information processing device according to any one of items 1 to 9, (Item 10) The interpolation means interpolates the interpolated signal to surrounding pixels in at least one of a spatial direction, a temporal direction, and a time-space direction. 10. The information processing device according to item 9, (Item 11) The interpolation means interpolates the interpolated signal based on a weighted average of signals in the set area around the target pixel. 11. The information processing device according to item 10, (Item 12) The interpolation means generates the interpolated signal by duplicating the signal of the target pixel to surrounding pixels in a predetermined set area. 12. The information processing device according to any one of items 1 to 11, (Item 13) The interpolation means determines the intensity of the interpolated signal based on the distance between the target pixel and the surrounding pixels. 13. The information processing device according to item 12, (Item 14) The reduction unit reduces the amount of the positive peak signal remaining as the proportion of the positive peak signal within a predetermined region increases. 14. The information processing device according to any one of items 1 to 13, (Item 15) When images of a plurality of frames are used, the reduction means determines whether or not to retain the positive peak signal depending on the frequency of occurrence of the positive peak signal at the same pixel position in each frame of the image. 15. The information processing device according to item 14, (Item 16) The reduction means uses a neural network trained using training data based on images with noise that reproduces positive peak signals. 16. The information processing device according to any one of items 1 to 15, (Item 17) The learning data is an image before noise is added, and then the interpolated signal is added around the noise after adding noise according to a binomial distribution. 17. The information processing device according to item 16, (Item 18) The discrimination means acquires an image captured by a SPAD (Single Photon Avalanche Diode) sensor. 18. The information processing device according to any one of items 1 to 17, (Item 19) The discrimination means discriminates a pixel of interest, which is a pixel to be discriminated, as the target pixel when signals of pixels spatially located around the pixel of interest are at a predetermined black level. 19. The information processing device according to any one of items 1 to 18, (Item 20) The discrimination means is Acquire multiple consecutive images in time series, When a signal of a pixel adjacent in time to a pixel of interest that is at the same pixel position as the pixel of interest that is the pixel to be discriminated is at a predetermined black level, the pixel of interest is discriminated as the pixel of interest. 20. The information processing device according to any one of items 1 to 19, (Item 21) An information processing device that trains a model, a degradation imparting means for imparting noise including a local positive peak signal to an image; a signal interpolation means for generating an image, as learning data, by interpolating an interpolated signal around a pixel of the local positive peak signal of the image to which noise has been added; A learning means for learning the model using the learning data; 13. An information processing device comprising: (Item 22) Obtaining an image including a target pixel that is a pixel of a local positive peak signal; Interpolating an interpolation signal to surrounding pixels that are pixels surrounding the target pixel of the image; the interpolated signal reduces noise in the interpolated image; 23. An information processing method comprising: (Item 23) A program for causing a computer to function as each of the means of the information processing device according to any one of items 1 to 20.
[0146] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0147] 100, 700, 1000... edge device, 200, 1010... cloud server, 10... imaging device, 114, 1006... input image data, 210, 1022... learned model, 111, 701, 1001... signal discrimination unit, 112... signal interpolation unit, 113, 704... noise reduction unit, 502... pixel, 702... spatial direction signal interpolation unit, 703... time direction signal interpolation unit, 1002... inference signal interpolation unit, 1004... inference noise amount estimation unit, 1005... inference noise reduction unit, 1011... learning data generation unit, 1012... learning unit, 1013... degradation imparting unit, 1014···Learning signal interpolation unit.
Claims
1. A discrimination means for acquiring an image including a target pixel which is a pixel having a local positive peak signal; an interpolation means for interpolating an interpolation signal into surrounding pixels which are pixels surrounding the target pixel of the image; a reduction means for reducing noise in an image into which the interpolated signal is interpolated; 13. An information processing device comprising:
2. The determining means determines whether the target pixel exists.
2. The information processing apparatus according to claim 1,
3. The determining means determines whether the target pixel exists based on a setting value of an imaging device that captured the image.
2. The information processing apparatus according to claim 1,
4. The discrimination means is A pixel value difference between a pixel value of a pixel of interest that is a pixel to be discriminated and pixel values of pixels surrounding the pixel of interest is compared with a predetermined pixel threshold value; and, A luminance difference between the luminance of the pixel of interest and the luminance of pixels surrounding the pixel of interest is compared with a predetermined luminance threshold value; The presence or absence of the target pixel is determined based on at least one of the above.
2. The information processing apparatus according to claim 1,
5. The image obtained by the determination means is free of black floating.
2. The information processing apparatus according to claim 1,
6. The image acquired by the discrimination means is captured by a photon counting type image sensor.
2. The information processing apparatus according to claim 1,
7. The image acquired by the discrimination means is taken in a low-illumination environment.
2. The information processing apparatus according to claim 1,
8. The local positive peak signal is over a number of pixels, The signals of the surrounding pixels are black level signals below a predetermined level threshold to obtain an image.
2. The information processing apparatus according to claim 1,
9. The interpolation means interpolates the interpolation signal to the surrounding pixels based on a positive peak signal included in a predetermined set region around the target pixel.
2. The information processing apparatus according to claim 1,
10. The interpolation means interpolates the interpolated signal to surrounding pixels existing in at least one of a spatial direction, a temporal direction, and a time-space direction.
10. The information processing apparatus according to claim 9,
11. The interpolation means interpolates the interpolated signal based on a weighted average of signals in the set area around the target pixel.
11. The information processing apparatus according to claim 10,
12. The interpolation means generates the interpolated signal by duplicating the signal of the target pixel to surrounding pixels in a predetermined set area.
2. The information processing apparatus according to claim 1,
13. The interpolation means determines the intensity of the interpolated signal based on the distance between the target pixel and the surrounding pixels.
13. The information processing apparatus according to claim 12.
14. The reduction unit reduces the amount of the positive peak signal remaining as the proportion of the positive peak signal within a predetermined region increases.
2. The information processing apparatus according to claim 1,
15. When images of a plurality of frames are used, the reduction means determines whether or not to retain the positive peak signal depending on the frequency of occurrence of the positive peak signal at the same pixel position in each frame of the image.
15. The information processing apparatus according to claim 14,
16. The reduction means uses a neural network trained using training data based on images with noise that reproduces positive peak signals.
2. The information processing apparatus according to claim 1,
17. The learning data is an image before noise is added, noise according to a binomial distribution is added, and then the interpolated signal is interpolated around the noise.
17. The information processing apparatus according to claim 16,
18. The discrimination means acquires an image captured by a SPAD (Single Photon Avalanche Diode) sensor.
2. The information processing apparatus according to claim 1,
19. The discrimination means discriminates a pixel of interest, which is a pixel to be discriminated, as the target pixel when signals of pixels spatially located around the pixel of interest are at a predetermined black level.
2. The information processing apparatus according to claim 1,
20. The discrimination means is Acquire multiple consecutive images in time series, When a signal of a pixel that is at the same pixel position as a pixel of interest that is a pixel to be discriminated and adjacent in time is at a predetermined black level, the pixel of interest is discriminated as the pixel of interest.
2. The information processing apparatus according to claim 1,
21. An information processing device that trains a model, a degradation imparting means for imparting noise including a local positive peak signal to an image; a signal interpolation means for generating an image, as learning data, by interpolating an interpolated signal around a pixel of the local positive peak signal of the image to which noise has been added; A learning means for learning the model using the learning data; 13. An information processing device comprising:
22. Obtaining an image including a target pixel that is a pixel of a local positive peak signal; Interpolating an interpolation signal to surrounding pixels that are pixels surrounding the target pixel of the image; the interpolated signal reduces noise in the interpolated image; 23. An information processing method comprising:
23. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 20.
Citation Information
Patent Citations
Photon counting laser 3D detection imaging method based on sparse representation
CN109613556A
LDCT image denoising and classifying method based on self-supervised and supervised combined training
CN113538260A
Image processing apparatus and method of processing image
JP2010092461A
Processor for reducing moving image noise, and image processing program
JP2010171808A
Image noise reduction method
JP2023170078A