Information processing apparatus, learning apparatus, and program
The information processing apparatus addresses the issue of fluctuating artifacts in moving images by using two neural networks to process image data, where the second network incorporates previous outputs to reduce fluctuations, resulting in improved image stability and quality.
Patent Information
- Application Number
- JP2023139002
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2043-08-29
AI Technical Summary
Existing image processing technologies for moving images using neural networks often result in artifacts that fluctuate in series of images, leading to undesirable visual effects.
An information processing apparatus that sequentially applies image processing to a series of image data using two neural networks. The second neural network uses previous image data outputs from both itself and the first neural network as inputs to generate image data that reduces fluctuations in the processed images.
Effectively suppresses the manifestation of fluctuations in a series of images, improving the stability and quality of processed moving images.
Smart Images

Figure 0007700186000008 
Figure 0007700186000009 
Figure 0007700186000010
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, a learning apparatus, and a program.
Background Art
[0002] In recent years, in image processing technologies that perform various processes on images such as still images and moving images (for example, image processing technologies for improving image quality), development of methods using neural networks has been actively carried out. As an example of such a method, there is one that uses a neural network to realize high-image-quality image processing such as noise removal, blur removal, and super-resolution. In addition, in time-series processing for moving images, voices, etc., a method applying a recurrent neural network has been proposed. A recurrent neural network has a recursive structure that uses the feature amount generated in the processing at the previous time as the input for the processing at the next time in time-series processing, and it is possible to obtain feature amounts effective for time-series processing. Non-Patent Document 1 discloses a technique for improving image quality by using a network having a recursive structure in noise removal processing for moving images.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the technology of performing image processing on a moving image using a neural network, after sequentially performing image processing on each of a series of frame images (still images) that make up the moving image, the frame images after image processing are connected along the time series to generate a moving image after image processing. When image processing is performed on a series of images having continuity along the time series such as a moving image, there may be a phenomenon in which artifacts generated in the structure of an object or flat portions of the background appear to fluctuate in the series of images. Although Patent Document 1 and Non-Patent Document 2 disclose techniques for suppressing the fluctuations that appear in such moving images, there is a demand for realizing a technique that can more preferably suppress the appearance of fluctuations.
[0006] In view of the above problems, an object of the present invention is to suppress the manifestation of fluctuations in a series of images accompanying the application of image processing to a series of images having continuity along the time series in a more preferable manner.
Means for Solving the Problems
[0007] The information processing apparatus according to the present invention sequentially performs predetermined image processing on each of a series of image data having continuity along the time series with respect to the input image data First neural network Among the first and second neural networks, the first neural networkBy inputting thereto, first image processing means for generating the first image data after the predetermined image processing, and for the first image data sequentially output from the first image processing means along the time series, the Second second image processing means for generating the second image data after the predetermined image processing based on a neural network, wherein the second image processing means includes the first image data output from the first image processing means at a first time or a second time before the first time, and the second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, as input to the neural network at the first time, and generates the second image data corresponding to the first time. Second It is characterized by generating the second image data corresponding to the first time as input to the neural network at the first time.
Advantages of the Invention
[0008] According to the present invention, it is possible to suppress the manifestation of fluctuations in a series of images associated with the application of image processing to a series of images having continuity along the time series in a more suitable manner.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0010] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted. Further, the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the illustrated configurations.
[0011] <First Embodiment> Hereinafter, a first embodiment of the present disclosure will be described. First, with reference to FIG. 1, an example of the hardware configuration of the information processing apparatus 100 according to the present embodiment will be described. The information processing apparatus 100 includes a CPU (Central Processing Unit) 101, a memory 102, an input unit 103, a storage unit 104, an output unit 105, and a communication unit 106. These components of the information processing apparatus 100 are connected via a bus so as to be able to transmit and receive information to and from each other.
[0012] The CPU 101 controls the operation of the entire information processing apparatus 100 and realizes the functions provided by the information processing apparatus 100 by expanding and executing various programs stored in the storage unit 104 and the like in the memory 102. The memory 102 is also used as a storage area for temporarily holding various data, such as a work area of the CPU 101. The memory 102 can be realized by, for example, a RAM (Random Access Memory) or the like. The storage unit 104 is a storage area for storing various programs and various data. The storage unit 104 can be realized by, for example, an auxiliary storage device represented by a ROM (Read Only Memory) or an HDD (Hard Disk Drive).
[0013] The input unit 103 serves as an input interface for receiving instructions from the user. The input unit 103 can be realized by, for example, an input device such as a mouse, a keyboard, or a touch panel. The output unit 105 serves as an output interface for presenting various information to the user. Note that the configuration of the output unit 105 can be appropriately changed according to the information output method. For example, the output unit 105 can be realized by a display device such as a display, and the information to be presented can be displayed as an image in a predetermined display area. As another example, the output unit 105 can be realized by an acoustic output device such as a speaker, and the information to be presented can be output as an acoustic signal such as voice or electronic sound. The communication unit 106 serves as a communication interface for connecting the information processing apparatus 100 to a network such as the Internet or a LAN (Local Area Network). Note that the configuration of the communication unit 106 can be appropriately changed according to the type of the connected network and the applied communication method.
[0014] Note that since the hardware configuration of the learning apparatus 120 according to the present embodiment is substantially the same as that of the above-described information processing apparatus 100, a detailed description thereof is omitted. Also, when the program stored in the storage unit 104 is expanded in the memory 102 and executed by the CPU 101, the functional configurations described later with reference to FIGS. 2 and 7 and the processes described later with reference to FIG. 3 and the like are realized.
[0015] Referring to FIG. 2, an example of the functional configurations of the information processing apparatus 100 and the learning apparatus 120 according to the present embodiment will be described. The information processing apparatus 100 takes image data of an image such as a moving image as input, performs various image processes on the image, and outputs the image data of the image after the image process. Examples of the image processes performed on the input image by the information processing apparatus 100 may include image processes for improving the image quality of the target image such as noise removal, blur removal, and super-resolution. Further, as another example of the image processes performed on the input image by the information processing apparatus 100, image processes such as converting the painting style of the input image to another painting style may be included. Thus, the type of the image process performed on the input image by the information processing apparatus 100 is not particularly limited as long as it is an image process for performing some kind of processing on the input image. In the present embodiment, for the sake of convenience, various explanations will be made by focusing on the case where the image process for improving the image quality as exemplified above is detailed for the input image. The information processing apparatus 100 includes an image acquisition unit 110, a first network unit 111, and a second network unit 112. Further, the information processing apparatus 100 is connected to the imaging apparatus 200 via a predetermined transmission path (for example, a cable or a network) so as to be able to transmit and receive data to and from each other. The imaging apparatus 200 is an imaging apparatus having an imaging optical system and an imaging element, and outputs the image data of an image (for example, a moving image or a still image) corresponding to the imaging result of the subject to the image acquisition unit 110 of the information processing apparatus 100. In the present embodiment, for the sake of convenience, the imaging apparatus 200 sequentially outputs the image data of the moving image corresponding to the imaging result of the subject to the image acquisition unit 110.
[0016] The image acquisition unit 110 corresponds to an input interface for inputting image data of an image to be subjected to image processing by the information processing apparatus 100 into the information processing apparatus 100 as an input image. In the present embodiment, the image acquisition unit 110 is configured to acquire, from the imaging apparatus 200, image data of a moving image corresponding to the imaging result of the imaging apparatus 200. The first network unit 111 includes a neural network that, by inputting image data, performs predetermined image processing related to processing of the image represented by the image data and generates image data representing the image after the image processing. Specifically, the first network unit 111 causes the neural network to generate image data after image processing by inputting the image data acquired by the image acquisition unit 110 into the neural network. The second network unit 112 has a neural network that performs the same image processing as the neural network included in the first network unit 111 on the input image data. Under such a configuration, the second network unit 112 has a recursive structure in which the output from the first network unit 111 and the output of the neural network that the second network unit 112 itself has are used as inputs to the neural network. As described above, the second network unit 112 causes the neural network to generate image data after image processing by inputting the image data output from the first network unit 111 into the neural network that the second network unit 112 itself has. Note that, in the present embodiment, the image data after image processing output from the first network unit 111 corresponds to an example of "first image data", and the image data after image processing output from the second network unit 112 corresponds to an example of "second image data". The learning apparatus 120 performs learning of the neural networks included in the first network unit 111 and the second network unit 112, respectively. Note that details of an example of the configuration of the learning apparatus 120 will be described separately later.
[0017] Next, with reference to FIGS. 2 and 3(a), an example of the processing of the information processing apparatus 100 according to the present embodiment will be described, particularly focusing on the processing when performing predetermined image processing related to image processing on a target image using a neural network. In the present embodiment, as the above image processing, an example in which noise removal processing for generating an image with reduced noise influence (ideally an image without noise) by removing noise from an image containing noise (hereinafter also referred to as a noise-added image) is applied will be described. Note that the image processing performed by the information processing apparatus 100 on the target image is not limited to noise removal, and other high-quality image processing such as super-resolution and blur removal may be applied, and other image processing other than the high-quality image processing may be applied as long as it is image processing related to image processing.
[0018] In S301, the image acquisition unit 110 acquires, as input image data, the image data of a noise-added image in which noise is manifested in the image from the imaging device 200, and inputs the image data of the noise-added image to the first network unit 111. The image acquisition unit 110 sequentially acquires the image data of the images sequentially acquired by the imaging device 200 along the time series, and uses the image indicated by the image data as the input image to the first network unit 111. In the present embodiment, for convenience, it is assumed that the image captured by the imaging device 200 contains noise. Examples of such noise include noise generated in the imaging element that is amplified and manifested when the sensor sensitivity is increased.
[0019] In S302, the first network unit 111 generates a high-quality image by performing image processing on the input noise-added image. The high-quality image here corresponds to an image with reduced influence of the noise (hereinafter also referred to as a noise-removed image) by removing the noise from the noise-added image. As an example of a neural network for realizing high-quality image processing, Non-Patent Document 3 proposed a Convolutional Neural Network (CNN) for realizing noise removal. The CNN is composed of a number of convolutional layers and activation functions. In particular, as a neural network for realizing high-quality image processing such as noise removal and super-resolution, a network called U-Net with a U-shaped structure is used. Even in the network disclosed in Non-Patent Document 3, noise removal is performed using U-Net. Also in this embodiment, similarly to the neural network disclosed in Non-Patent Document 3, a structure based on U-Net is used.
[0020] Here, with reference to FIG. 4, an example of the configuration of the neural network used in this embodiment will be described. The series of neural networks shown in FIG. 4 includes two neural networks, which are applied to a first network unit 111 that performs noise removal on an input image and a second network unit 112 that similarly performs noise removal on the input image, respectively. Also, in the series of neural networks shown in FIG. 4, temporally consecutive images (for example, a series of frame images constituting a moving image) are sequentially input, and predetermined image processing (for example, image processing related to noise removal) is performed on each of the input images. Here, taking the time when the image processing is executed as s, an example will be described in the case where the image processing is sequentially executed for each of the times s = 0, 1,... along the time series. Also, hereinafter, apart from the variable s, a variable t will be used as a variable representing time.
[0021] The image processing when s = 0 corresponds to the first image processing in the image processing that is executed every moment. First, in the image processing when s = 0, the noise-added image 401 at t = 0 is acquired. When the image data of the acquired noise-added image 401 is input to the first network unit 111, the first network unit 111 executes image processing for removing the noise manifested in the image on the noise-added image 401. Then, the first network unit 111 outputs the noise-removed image generated as a result of the image processing as the image data of the high-quality image 402 at t = 0. Next, the second network unit 112 combines the high-quality image 402 (noise-removed image) indicated by the image data output from the first network unit 111 and the high-quality image indicated by the image data output by itself at the previous time in the channel direction. Moreover, the second network unit 112 performs image processing on the combined image using the combined image as an input. In the image processing when s = 0, there is no high-quality image corresponding to the output from the second network unit 112 at the previous time. Therefore, in such a case, the second network unit 112 generates a dummy input (dummy high-quality image 403), combines it with the high-quality image 402 output from the first network unit 111, and uses the combined image as an input to itself. The second network unit 112 generates a noise-removed image by performing image processing on the aforementioned input (combined image), and outputs the image data of the high-quality image with the noise-removed image as the high-quality image at t = 0.
[0022] Next, when s = 1, the image data of the noisy image 405 with t = 1 is acquired. The first network unit 111 generates a high-quality image 406 (denoised image) obtained by removing noise from the image indicated by the input image data, and outputs the image data of the high-quality image 406. The second network unit 112 combines, in the channel direction, the high-quality image 406 indicated by the image data output from the first network unit 111 and the high-quality image 404 indicated by the image data that is its own output in the previous-time image processing, and uses the combined image as an input to itself. In the present disclosure, the channel corresponds to a color channel typified by RGB or the like. Also, the high-quality image 404 that is the output of the second network unit 112 in the previous-time image processing corresponds to the denoised image that the second network unit 112 outputs as the high-quality image with t = 0 as a result of the image processing at time s = 0. The second network unit 112 generates a denoised image by performing image processing on the aforementioned input, and outputs the image data of the denoised image as the high-quality image with t = 1.
[0023] Referring to FIG. 5, an example of the configuration of the first network unit 111 will be described in more detail, particularly focusing on the configuration of the neural network applied to the first network unit 111. The first network unit 111 uses a U-Net composed of an encoder that generates feature amounts while compressing an image and a decoder that restores an image from the compressed feature amounts as a neural network that performs image processing on the input image. The neural network first generates feature quantities (feature quantities extracted from the image) with different resolutions and numbers of channels from the input image 501 in the encoder. The neural network applies the process 511 of applying the convolution operation and the relu function to the input image 501 multiple times to generate the feature quantity 512. By performing the pooling process 513 on the generated feature quantity 512, the resolution is reduced, and by repeating the convolution operation and the relu function again, a feature quantity with an increased number of channels is obtained. Also, the feature quantity 512 generated at this time is used during the later image restoration process and is skip-connected 514 with another feature quantity generated while being upsampled. By repeating a series of processes, the compressed feature quantity is restored to an image while reducing the number of channels and increasing the resolution by the transposed convolution operation 515. At this time, the feature quantity upsampled by the transposed convolution process is skip-connected with the feature quantity generated by the encoder, and a plurality of convolution operations, the application of the relu function, and the transposed convolution process are repeated. Finally, the image data of the high-quality image 502 having the desired resolution and number of channels is output. The high-quality image 502 that is the output of the first network unit 111 corresponds to an image from which noise has been removed from the input image 501. In this embodiment, it is assumed that the neural network configured as described above is used. However, as long as it is possible to perform image processing related to image processing such as high-quality conversion on the input image, the structure of the neural network to be applied is not particularly limited.
[0024] Here, referring to FIG. 3 again. In S303, the second network unit 112 determines whether there is image data of the high-quality image that is its own output at the previous time. If the second network unit 112 determines in S303 that there is image data of the high-quality image that is its own output at the previous time, the process proceeds to S304. In this case, the second network unit 112 acquires the image data of the high-quality image that is its own output at the previous time in S304. On the other hand, when the second network unit 112 determines in S303 that there is no image data of the high-quality image that is its own output at the previous time, the process proceeds to S305. In this case, the second network unit 112 generates image data of a dummy high-quality image in S305.
[0025] The second network unit 112 uses the image data of the high-quality image output by itself in the image processing executed at the previous time as the input in its own image processing at the current time. The information processing apparatus according to the present embodiment sequentially acquires the image data of the input noise-added image and executes image processing on the noise-added image at each moment. However, in the image processing at the start time among the image processing executed at each moment, there is no image data of the high-quality image output by the second network unit 112 as the result of the image processing at the previous time. Therefore, after generating the image data of the dummy high-quality image, the second network unit 112 uses the image data of the dummy high-quality image as the input instead of the image data of the high-quality image output as the result of the image processing at the previous time. Here, the processing of the second network unit 112 will be described with reference to FIG. 4. In the image processing with s = 0, the second network unit 112 attempts to acquire the image data of the high-quality image (noise removal image) with t = -1 output as the result of its own image processing at s = -1 as its own input. However, there is no image data of the high-quality image (noise removal image) with t = -1 corresponding to the result of the image processing at s = -1. Therefore, the second network unit 112 generates the image data of a dummy high-quality image 403 to be used as the input instead of the high-quality image with t = -1. As the dummy high-quality image to be generated, for example, an image having the same width, height, and number of channels as the high-quality image output by the second network unit 112 and having a value of 0 for each pixel can be applied. Note that for the dummy high-quality image, as long as the width, height, and number of channels are the same as those of the high-quality image used as the input, the value of each pixel is not necessarily limited to 0, and for example, a value determined during learning may be applied.
[0026] Here, referring to FIG. 3 again. In S306, the second network unit 112 combines, in the channel direction, the high-quality image indicated by the image data output by itself in the image processing at the previous time and the high-quality image at the current time indicated by the image data output by the first network unit 111. The second network unit 112 uses the combined image described above as an input, executes image processing to generate a high-quality image (for example, a noise-removed image), and outputs the image data of the high-quality image. The image data of the high-quality image (noise-removed image) output at this time is used as an input to the second network unit 112 in the image processing executed at the next time.
[0027] Referring to FIG. 6, an example of the configuration of the second network unit 112 will be described in more detail, particularly focusing on the configuration of the neural network applied to the second network unit 112. Similar to the first network unit 111, the second network unit 112 uses a U-Net as a neural network that performs image processing on the input image. However, since the second network unit 112 uses the high-quality image, which is its own output at the previous time, as an input, it is configured to perform a combining process 611 for using its own output as a recursive input. At this time, before the combining process 611, the high-quality image 602 is subjected to multiple applications of convolution + relu functions to generate feature amounts from the high-quality image that will be the input for the image processing at the next time, and these feature amounts may be used as inputs. Also, in the image processing at the current time, the high-quality image 502 output by the first network unit 111 and the high-quality image 601 output by the second network unit 112 in the image processing at the previous time are combined in the channel direction. The image generated by this combination is applied as the input to the second network unit 112 at the current time. The second network unit 112 generates a high-quality image 602 by performing compression of feature amounts and restoration processing to the image on the above input in the same manner as the first network unit 111, and outputs the image data of the high-quality image 602. The high-quality image 602 indicated by the image data corresponding to the output of the second network unit 112 is an image with higher image quality than the high-quality image 502 indicated by the image data corresponding to the output of the first network unit 111.
[0028] Note that in this embodiment, the case where the neural network configured as described above is used has been described, but the structure of the neural network used in the first network unit 111 and the second network unit 112 is not necessarily limited. That is, the structure of the neural network is not limited as long as it is possible to use the image data of the image output as a result of the image processing at the previous time as the input to the second network unit 112 and perform image processing related to image enhancement such as high-quality conversion on the input image.
[0029] As described above, the information processing apparatus of the present invention acquires an image after image processing (for example, a high-quality image) by performing image processing on time-series images sequentially acquired at each moment.
[0030] Next, with reference to FIG. 2, an example of the functional configuration of the learning apparatus according to the present embodiment will be described. The learning apparatus 120 performs learning of a neural network applied to a first network unit that performs image processing such as the above-described high-quality conversion and a second network unit that performs the same image processing as the image processing. The learning apparatus 120 includes a database unit 121, an image acquisition unit 110, a degraded image generation unit 122, a first network unit 111, a second network unit 112, a difference calculation unit 126, an error calculation unit 124, and a learning unit 125.
[0031] The database unit 121 stores a large amount of image data of a moving image after a change manifested in the image such as noise is restored by the target image processing. As a specific example, when noise removal processing is applied as the target image processing, the database unit 121 stores image data of a clean moving image without noise. The series of moving image data stored in the database unit 121 corresponds to learning data (teacher data) used for learning a neural network applied to the first network unit and the second network unit. The image acquisition unit 110 extracts image data of a moving image used for learning from among the series of moving image data stored in the database unit 121. Moreover, the image acquisition unit 110 sequentially acquires, along the time series, the image data of a series of frame images (that is, a series of still images continuously arranged in a case series) constituting the moving image indicated by the extracted image data as the image data of the true value image. The degraded image generation unit 122 generates a degraded image with degraded image quality from the true value image based on the image data of the true value image acquired by the image acquisition unit 110. For example, when noise removal processing is applied as the target image processing, the degraded image generation unit 122 generates a noise-added image with noise added to the true value image as the degraded image. Note that the image data of the degraded image corresponds to an example of "processed image data".
[0032] The difference calculation unit 126 generates a difference image between two temporally consecutive true value images extracted from the moving image used for input. As a specific example, the difference calculation unit 126 may generate a difference image between these true value images based on the true value images corresponding to the current time and the previous time, respectively. Further, the difference calculation unit 126 generates a difference image between two temporally consecutive high-quality images generated by subjecting two degraded images generated from each of the two temporally consecutive true value images to image processing by the second network unit 112. As a specific example, the difference calculation unit 126 may generate a difference image between these high-quality images based on the high-quality images corresponding to the current time and the previous time, respectively. Note that the difference image between the two true value images corresponds to an example of the "first difference", and the difference between the two high-quality images (in other words, the difference between the two images after image processing) corresponds to an example of the "second difference". The error calculation unit 124 calculates an error value using the difference image between two temporally consecutive true value images as the true value and the difference image between two high-quality images corresponding to the two temporally consecutive true value images. The learning unit 125 performs neural network learning for each of the first network unit 111 and the second network unit 112 using the error value generated by the error calculation unit 124. Specifically, the learning unit 125 updates the parameters of the neural network applied to each of the first network unit 111 and the second network unit 112 based on the error value generated by the error calculation unit 124.
[0033] Next, with reference to FIG. 3(b) together with FIG. 2, an example of the processing of the learning device 120 according to the present embodiment will be described, particularly focusing on the processing related to the learning of the neural network applied to the information processing device according to the present embodiment.
[0034] In FIG. 3(b), the series of processes indicated by the loop end S311 shows the processes related to the learning of the neural networks applied to the first network unit 111 and the second network unit 112, respectively. By this series of processes, the update of parameters such as the weights and biases of the neural network is repeatedly executed for each piece of image data of the moving image corresponding to the learning data. Specifically, at the start of learning, initial parameter values are given to the target neural network, and thereafter, by repeatedly performing learning, the parameters of the neural network are updated. Also, the series of processes related to the learning of the neural network indicated by the loop end S311 is repeatedly executed until the termination condition is satisfied. As a specific example, when the number of learning times of the neural network reaches a predetermined number, it may be determined that the termination condition is satisfied, and the execution of the series of processes related to the learning of the neural network indicated by the loop end S311 may be terminated.
[0035] Next, the content of the series of processes indicated by the loop end S311 will be described below. In S312, the degraded image generation unit 122 acquires one of the pieces of image data of the series of moving images stored in the database unit 121. As described above, the database unit 121 stores a large amount of image data of clean moving images without noise. That is, the database unit 121 stores the image data of the time-series images (for example, a series of images continuous in time series like a series of frame images constituting a moving image) that are the true values when learning the neural network. In S313, the degradation image generation unit 122 generates a degradation image corresponding to the time-series images obtained from the image data of one moving image acquired in S312. In this embodiment, it is assumed that a noise-added image in which noise is added to a clean image is generated as the degradation image. As a specific example, the degradation image generation unit 122 may generate a noise-added image by analyzing and modeling the noise characteristics of the sensor and adding the modeled noise to the target image. As the noise, for example, noise generated in the process of converting photons detected by the sensor into digital signals can be assumed. Examples of noise sources include photon shot noise, readout noise, dark current noise, quantization error, etc., and it is assumed that they are modeled. The degradation image generation unit 122 associates the generated degradation image with the original clean image, and then assigns time information to form one set, and stores the set of the data of these images in the database unit 121.
[0036] The series of processes indicated by the loop end S314 shows a series of processes related to the training of a neural network, which is executed for the image data of one moving image acquired in S312. That is, a series of pairs of clean images and degradation images generated from the image data of the target moving image in S313 are sequentially acquired along the time series, and for each pair, a series of processes related to the training of the neural network indicated by the loop end S314 are executed. The content of the series of processes indicated by the loop end S314 will be specifically described below.
[0037] In S315, the image acquisition unit 110 acquires two sets of the clean image and the degraded image generated in S313 in a temporally consecutive manner. At this time, among the two sets to be acquired, the image acquisition unit 110 acquires, as the set of an earlier time, the same set as the latest set among a series of sets acquired when image processing was executed at the previous time. However, if there is no set corresponding to the time to be acquired, the image acquisition unit 110 excludes the set from the acquisition target. In the above manner, the image acquisition unit 110 acquires sets for as many acquirable times as possible. Next, the image acquisition unit 110 uses the clean image included in the acquired set as the true value image, and outputs the true value image and the degraded image included in the set to each component of the learning device 120. Specifically, the image acquisition unit 110 acquires one degraded image from the set with a newer time and outputs the degraded image to the first network unit 111. Next, the image acquisition unit 110 sequentially acquires two true value images from the two sets and outputs them to the difference calculation unit 123 and the error calculation unit 124. Note that when only one set can be acquired, for example, when image processing is executed at the first time, the image acquisition unit 110 treats the images at the times that could not be acquired as non-existent. As a specific example, consider an example in the case where image processing is sequentially performed on each of the time-series images of t = 0, 1,.... At this time, let s be the time when image processing is executed. In the image processing at the start time s = 0, an attempt is made to acquire the degraded image at t = 0 and the true value images at t = -1, 0 from the sets at times t = -1, 0. However, in the case of s = 0, since there is no true value image at t = -1, only the true value image at t = 0 is acquired. Also, in the image processing at the next time s = 1, the degraded image at t = 1 and the true value images at t = 0, 1 are acquired from the sets at times t = 0, 1.
[0038] Next, through a series of processes from S302 to S306, image processing is performed on the degraded image input to the first network unit 111 at S315, and a noise-removed image is obtained as a high-quality image. Since the series of processes from S302 to S306 are substantially the same as the series of processes from S302 to S306 shown in Fig. 3(a), detailed description thereof is omitted. The second network unit 112 outputs the high-quality image output as a result of the image processing at the current time and the high-quality image at the previous time used for the input to the difference calculation unit 123 and the error calculation unit 124. Also, the second network unit 112 holds the high-quality image corresponding to the current time, which is the input to itself in the image processing at the next time. Also, as described above, when there is no high-quality image at the previous time, the second network unit 112 uses a dummy image instead of the high-quality image to be input to itself. In this case, the second network unit 112 uses, as the dummy image, an image having the same width, height, and number of channels as the corresponding high-quality image and with the value of each pixel being 0. However, for the dummy image, it is not necessarily the case that the value of each pixel is 0 as long as the width, height, and number of channels are the same as those of the high-quality image used as the input. Also, when the information processing apparatus 100 executes image processing, for example, when a dummy image is used for the input to the second network unit 112, a dummy high-quality image determined at the time of learning may be used as the dummy image.
[0039] In S316, the difference calculation unit 123 generates a ground truth difference image using the ground truth image at the previous time acquired from the image acquisition unit 110 and the ground truth image at the current time. Also, the difference calculation unit 123 generates a difference image of the high-quality image using the high-quality image at the previous time and the high-quality image at the current time. Here, the process related to the generation of each of the above difference images by the difference calculation unit 123 will be described in more detail. The difference calculation unit 123 generates a difference image using two images that are temporally continuous. The ground truth image I at the previous time t and the ground truth image I at the current time t-1 The ground truth difference image I calculated usingdiff is expressed by the relational expression shown below as equation (1).
[0040]
number
[0041] In addition, the high-resolution image I^ t-1 and the high-resolution image I^ at the current time. t The difference image I^ of the improved image is calculated using diff is expressed by the following relational expression shown as Equation (2). Note that in this disclosure, the notation "I^" refers to the letter "I" with a hat attached.
[0042]
number
[0043] In S317, the error calculation unit 124 calculates the error used to update the parameters of the neural network. Here, the process of calculating the error by the error calculation unit 124 will be described in more detail. The error calculation unit 124 calculates an error used for learning to enable the neural network to execute high quality image processing. Specifically, the error calculation unit 124 first calculates an error (hereinafter also referred to as an image reconstruction error) between a high quality image at the current time and a true value image at the current time, which is a true value corresponding to the high quality image. Next, the error calculation section 124 calculates a difference error between frames using the difference image of the true image generated in S316 and the difference image of the image-enhanced image. These errors are added together to obtain the final error.
[0044] Image reconstruction error L i is the high-resolution image I^ at the current time. t and the true image I t Using these, the L1 error is calculated based on the relational expression shown below as equation (3).
[0045]
Number
[0046] Inter-frame difference error L f is calculated as the L1 error represented by the relational expression shown as Equation (4) below based on the difference image I^ of the high-quality image diff and the difference image I of the true value diff Note that M shown in Equation (4) is a mask represented by the relational expression shown as Equation (5) below and created according to the difference value of the difference image of the true value. The mask M is created so that its value becomes 0 or 1 according to a predetermined threshold value th. As shown as Equation (4) below, the inter-frame difference error L f is calculated by taking the product of the L1 error calculated based on the above difference image and the elements of the mask M. For example, when a small value such as th = 0.01 is set and the mask M is created, the error value is calculated for the portions with little change, that is, the stationary regions with little movement, between images that are continuous in time series. In other words, the mask M weights the error for portions with less change between the true value images to have a greater weight Also, from the characteristics of the mask M described above, it is also possible to control the calculation target of the error value by adjusting the values (in other words, weights) of the elements according to each condition in the mask M. As a specific example, by inverting 0 and 1 shown as values according to the threshold value th in the following Equation (5), it is also possible to set the portions with greater change (that is, the portions with large movement) between the true value images as the calculation target of the error value
[0047]
Number
[0048] The image reconstruction error L described above i and the inter-frame difference error L fBased on this, the error L that is finally used to update the parameters of the neural network is calculated. For example, the error L is represented by the relational expression shown as Equation (6) below. Note that the coefficient α shown in Equation (6) is a term for adjusting the value of the inter-frame error L f is a term for adjusting the value.
[0049]
Equation
[0050] In S318, the learning unit 125 uses the error calculated in S317 to update the parameters of the neural networks of the first network unit 111 and the second network unit 112 respectively by the error backpropagation method.
[0051] As described above, by repeatedly executing the series of processes of S311 to S318, learning of a neural network that performs image processing such as high-quality conversion on the image indicated by the target image data is performed.
[0052] As described above, according to the present embodiment, when learning a neural network that performs image processing on each of a series of images that are temporally continuous, using each of the temporally continuous images as a sequential input, the correspondence relationship between the temporally continuous images, that is, the spatio-temporal information, is easily learned. This is due to inputting the image after image processing at the previous time (for example, a high-quality image) to the second stage of the two-stage neural network that executes the same image processing task. Here, the reason for achieving the above-described effects will be described by focusing on the case where noise removal processing is performed on the target image. In the present embodiment, first, by the image processing using the first-stage neural network, a noise-removed image with some noise removed from the input image is output. Further, the second-stage neural network receives as input the noise-removed image with some noise removed in the first stage and the noise-removed image that is the output of the second-stage neural network at the previous time. By applying such a configuration, the second-stage neural network is trained with an image having less temporally continuous noise as input, so that it becomes easier to learn the correspondence relationship between the two images. In addition, in the present embodiment, learning of a neural network that performs image processing such as high image quality improvement on an input image is performed using an inter-frame difference error. With such a configuration, even in a situation where image processing is performed on a series of images having continuity along a time series such as a moving image, it is possible to expect an effect of further suppressing the manifestation of a phenomenon in which artifacts generated in the image appear to fluctuate.
[0053] <Second Embodiment> Hereinafter, a second embodiment of the present disclosure will be described. In the above-described first embodiment, an example of a recurrent neural network in which an image (for example, a high image quality-improved image) after image processing at the previous time is input to the second stage of a two-stage neural network that performs the same image processing on the input image was described. In the present embodiment, an example of a neural network that receives a plurality of images as input and outputs a plurality of images on which image processing such as high image quality improvement is performed will be described.
[0054] With reference to FIG. 7(a), an example of the functional configuration of the information processing apparatus 700 according to the present embodiment will be described. The information processing apparatus 700 according to the present embodiment is different from the information processing apparatus 100 shown in FIG. 2(a) in that it has a storage unit 104. The memory unit 104 is a storage area for holding image data of an image after image processing (e.g., a high-quality image) generated by the image processing constantly executed by the information processing apparatus 700 according to the present embodiment. Note that, since other components are substantially the same as those of the information processing apparatus 100 shown in FIG. 2(a), detailed description thereof will be omitted. Also, hereinafter, for convenience, as in the first embodiment described above, various descriptions will be made focusing on the case where noise removal processing for removing noise manifested in a target image is applied as the image processing by the information processing apparatus 700.
[0055] FIG. 8 shows an example of the configuration of the neural network used in the present embodiment. FIG. 8 shows the states of input and output of a series of neural networks at the processing time s of the image processing constantly executed by the information processing apparatus 700 according to the present embodiment. In the present embodiment, for convenience, the start time of the image processing executed by the information processing apparatus 700 is set to s = 0, and the time when the next image processing is executed is set to s = 1.
[0056] Here, with reference to FIGS. 3(a), 7(a), and 8, an example of the processing of the information processing apparatus 700 according to the present embodiment will be described, focusing particularly on the processing when a predetermined image processing related to image processing is performed on a target image using a neural network.
[0057] In S301, the image acquisition unit 110 acquires, as input image data, the image data of a noise-added image in which noise is manifested in the image from the imaging device 200, and inputs the image data of the noise-added image to the first network unit 111. At this time, the image acquisition unit 110 acquires the image data of each of three images that are temporally continuous as the input image, and sets the three images indicated by the series of image data as the input images to the first network unit 111. The information processing apparatus 700 according to this embodiment is designed such that, in three images that are temporally continuous and acquired by the image acquisition unit 110, the same image as the image at the last time in the image processing at the previous time is used as the image at the first time. Specifically, the three noise-added images indicated by a series of image data acquired by the image acquisition unit 110 in the image processing with s = 0 are the image 800 at time t = 0, the image 801 at time t = 1, and the image 802 at time t = 2. Also, the three noise-added images indicated by a series of image data acquired by the image acquisition unit 110 in the image processing with s = 1 are the image 802 at time t = 2, the image 811 at time t = 3, and the image 814 at time t = 4. Thus, in the example shown in FIG. 8, in each image processing with s = 0 and 1, the image acquisition unit 110 acquires the image data of each of the three noise-added images to be input to each image processing so that the image 802 (noise-added image) at time t = 2 becomes the input to the neural network.
[0058] In S302, the first network unit 111 executes image processing with the noise-added image as the input, and outputs the image data of the noise-removed image corresponding to the result of the image processing to the second network unit 112. Also, at this time, the first network unit 111 also outputs the image data of the noise-removed image corresponding to the result of the image processing to the storage unit 104. Thereby, the image data of the noise-removed image output from the first network unit 111 is stored in the storage unit 104. Note that, as the neural network applied to the first network unit 111 according to this embodiment, a neural network having the same structure as the example shown in FIG. 5 is applied. However, in this embodiment, three temporally continuous noise-added images are applied as the input to the neural network, and the image data of three noise-removed images (image with enhanced image quality) corresponding to the input image is output from the neural network. Specifically, in the image processing with s = 0, as the input noisy image, the image 800 at time t = 0, the image 801 at time t = 1, and the image 802 at time t = 2 are used. At this time, the first network unit 111 outputs the noise removal images 803 at time t = 0, 804 at time t = 1, and 805 at time t = 2 as the high-quality images corresponding to the input images respectively. The same applies to the image processing with s = 1. However, as the noise removal images output by the first network unit 111 for the input of the overlapping noisy image at t = 2 in the image processing at each time, different noise removal images 805 and 815 are obtained. In addition, in the image processing with s = 1, the image data of the noise removal image 804 at time t = 1 and the noise removal image 816 at time t = 2, which were output by the first network unit 111 at the previous time s = 0, are stored in the storage unit 104 for input to the second network unit 112.
[0059] In S303, the second network unit 112 determines whether there is image data of the noise removal image (high-quality image) that is its own output at the previous time. If the second network unit 112 determines in S303 that there is image data of the noise removal image that is its own output at the previous time, the process proceeds to S304. In this case, the second network unit 112 acquires the image data of the noise removal image that is its own output at the previous time in S304. On the other hand, if the second network unit 112 determines in S303 that there is no image data of the noise removal image (high-quality image) that is its own output at the previous time, the process proceeds to S305. In this case, the second network unit 112 generates the image data of a dummy high-quality image in S305. The second network unit 112 determines whether there is image data corresponding to the output of the first network unit 111 at the previous time and the output of the second network unit 112 at the previous time as a high-quality image (noise-removed image) corresponding to the previous time. Regarding the image data of these high-quality images corresponding to the previous time, if they exist, they are stored in the storage unit 104. Specifically, in the image processing with s = 0, the second network unit 112 attempts to acquire the image data of the noise-removed image (high-quality image) at t = -1 output from the first network unit 111 at s = -1 and the high-quality image at t = 0 as its own input. Also, the second network unit 112 attempts to acquire the image data of the noise-removed image (high-quality image) at t = -2 output from itself at s = -1. However, there is no image data of the noise-removed image corresponding to the result of the image processing with s = -1. Therefore, the second network unit 112 generates dummy image data of the high-quality images corresponding to the outputs of the first network unit 111 and itself as the image data of a dummy high-quality image to be input instead of the image data of the high-quality image corresponding to the result of the image processing with s = -1. Also, in the image processing with s = 1, the second network unit 112 attempts to acquire the image data of the high-quality image (noise-removed image 804) at t = 1 output from the first network unit 111 in the image processing with s = 0 and the high-quality image (noise-removed image 805) at t = 2. Also, the second network unit 112 attempts to acquire the image data of the high-quality image (noise-removed image 811) at t = 0 output from itself in the image processing with s = 0.
[0060] In S304, the second network unit 112 acquires, from the storage unit 104, the image data of the noise-removed image (high-quality image) at the previous time that serves as its input. Specifically, in the image processing with s = 1, the second network unit 112 acquires, as its input, the image data of the noise-removed image (high-quality image) with t = 1 and the high-quality image with t = 2 output from the first network unit 111 in the image processing with s = 0. The second network unit 112 also acquires the image data of the noise-removed image with t = 0 output by itself in the image processing with s = 0.
[0061] In S305, the second network unit 112 generates the image data of a dummy high-quality image that serves as its input. Specifically, when there is no image data of the high-quality image (noise-removed image) at the previous time serving as the input of the second network unit 112, the second network unit 112 generates a dummy input. The dummy high-quality image generated at this time has the same width, height, and number of channels as each of the high-quality image output by the first network unit 111 and the high-quality image output by the second network unit 112, and is an image with the value of each pixel being 0. However, for the dummy high-quality image, as long as its width, height, and number of channels are the same as those of the high-quality image used as the input, it is not necessarily required that the value of each pixel be 0.
[0062] In S306, the second network unit 112 generates a feature amount indicating the features of the image by using the high-quality image or the dummy high-quality image at the previous time and the high-quality image output by the first network unit 111. Note that the details of the feature amount generated at this time will be described separately later in conjunction with the structure of the neural network applied to the second network unit 112. The second network unit 112 executes image processing using the generated feature amount as the input, generates a noise-removed image (high-quality image) at the time corresponding to the input feature amount, and outputs the image data of the noise-removed image. The image data of the noise-removed image output at this time is stored in the storage unit 104 because it is used as the input to the second network unit 112 in the image processing that the information processing apparatus 100 executes at each moment.
[0063] Next, an example of the structure of the neural network applied to the second network unit 112 will be described. For the neural network applied to the second network unit 112, a structure similar to the configuration shown in FIG. 6 is used. The second network unit 112 takes as input the feature amounts generated using four temporally consecutive high-quality images. Specifically, when image processing is executed at a certain time, the second network unit 112 acquires two images in the order of older times out of the three high-quality images output as a series of image data from the first network unit 111. Also, the second network unit 112 acquires two images in the order of newer times out of the three high-quality images output as a series of image data from the first network unit 111 in the image processing at the previous time. Further, the second network unit 112 acquires the image at the central time out of the three high-quality images that it output as a series of image data in the image processing at the previous time. Then, the second network unit 112 synthesizes the high-quality image at the current time and the high-quality image at the previous time, and when there are images at the same time, it synthesizes those high-quality images. However, when there is no high-quality image at the previous time, the second network unit 112 uses a dummy high-quality image at the previous time. Then, the second network unit 112 combines the results of synthesizing a plurality of target high-quality images in the channel direction to obtain a feature amount as an input to itself, and by performing image processing on the feature amount, it generates a high-quality image as its output. At this time, the second network unit 112 generates image data of a noise removal image for the time when the three-time high-quality images, which are the output of the first network unit 111 used as the input, are further enhanced in quality as an output according to the result of the image processing. The image data of the series of noise removal images output from the second network unit 112 at this time is stored in the storage unit 104 because it is used in the image processing at the next time. Note that, as described above, the feature amount in the present embodiment can also be said to correspond to a composite image in which a plurality of images (a plurality of images that may include dummy images in part) are combined. Therefore, when image data is described in the present disclosure, unless otherwise particularly limited, it shall be construed to include both data indicating the image itself and data indicating the feature amount (in other words, the composite image obtained by combining a plurality of images) generated based on a plurality of images as described above.
[0064] With reference to FIG. 8, an example of the processing of the second network unit 112 will be described in more detail with a specific example. In the image processing with s = 0, the second network unit 112 attempts to acquire the image data of the noise reduction image (high-quality image) with t = -1, the noise reduction image with t = 0, and the noise reduction image with t = 1, which are the outputs of the first network unit 111. Specifically, the second network unit 112 first acquires the image data of the high-quality image (noise reduction image 803) with t = 0 and the high-quality image (noise reduction image 804) with t = 1, which are the outputs of the first network unit 111 in the image processing with s = 0. Next, the second network unit 112 attempts to acquire the image data of the high-quality image with t = -1 and the high-quality image with t = 0, which are the outputs of the first network unit 111 in the image processing at the previous time, i.e., s = -1. However, the image processing with s = -1 has not been executed and the image data of the target high-quality image does not exist. Therefore, the second network unit 112 acquires the image data of dummies 807 and 808 instead of the image data of the target high-quality image. In addition, the second network unit 112 attempts to acquire the image data of the high-quality image with t = -2, which is its own output in the image processing with s = -1, but similarly, the image processing has not been executed and the image data of the target high-quality image does not exist. Therefore, the second network unit 112 acquires the image data of dummy 806 instead of the image data of the target high-quality image.
[0065] The second network unit 112 synthesizes the images corresponding to the same time among the images obtained as its own input as described above. Specifically, the second network unit 112 synthesizes the dummy 808 at t = 0 in the image processing with s = -1 and the noise-removed image 803 (high-quality image) at t = 0 which is the output of the first network unit in the image processing with s = 0. Regarding the synthesis of the dummy 808 and the noise-removed image 803, for example, an image is generated by adding the pixel values of each image for each channel and then using the average value of the values for each channel as the pixel value. For the above image synthesis method, for example, the synthesis method determined during the learning of the neural network is applied. Also, the above image synthesis method is merely an example, and the method is not particularly limited as long as the result obtained by the synthesis can maintain the same width, height, and number of channels as the high-quality image of the synthesis source.
[0066] The second network unit 112 generates a feature amount by combining the dummy 806 at t = -2, the dummy 807 at t = -1, the synthesis result 809 at t = 0, and the noise-removed image 804 at t = 1 obtained as described above in the channel direction, and uses the feature amount as the input to itself. At this time, regarding the time corresponding to the output image of the first network unit used as the input to the second network unit 112, the image output by the second network unit 112 at this time is treated as the output of the second network unit 112 at this time. Specifically, the second network unit 112 performs image processing on the feature amount obtained as its own input described above, and outputs the image data of each of the noise-removed image 810 (high-quality image) at t = -1, the noise-removed image 811 at t = 0, and the noise-removed image 812 at t = 1. However, regarding the noise-removed image 810 (high-quality image) at t = -1, since the input image is a dummy, the output result is also a dummy image.
[0067] In the image processing with s = 1, the second network unit 112 acquires the image data of each of the noise-removed image at t = 1 which is the output of the first network unit 111, the noise-removed image at t = 2, and the noise-removed image at t = 3. Specifically, the second network unit 112 first acquires the image data of the noise removal image 815 (image with enhanced quality) at t = 2, which is the output of the first network unit 111 at s = 1, and the noise removal image 816 at t = 3. Next, the second network unit 112 acquires the image data of the noise removal images 804 and 805 (images with enhanced quality) at t = 1 and t = 2, which are the outputs of the first network unit 111 in the image processing at the previous time, i.e., s = 0. Also, the second network unit 112 acquires the image data of the noise removal image 811 (image with enhanced quality) at t = 0, which is its own output in the image processing at s = 0. Then, the second network unit 112 synthesizes the images at the same time in the series of input images indicated by each of the acquired series of image data. Specifically, the second network unit 112 synthesizes the noise removal image 805 (image with enhanced quality) at t = 2 in the image processing at s = 0 and the noise removal image 815 (image with enhanced quality) at t = 2, which is the output of the first network unit 111 in the image processing at s = 1. Then, the second network unit 112 combines the acquired noise removal image 811 at t = 0, the noise removal image at t = 1, the synthesis result 818 at t = 2, and the noise removal image 816 at t = 3 in the channel direction to generate a feature amount, and uses the feature amount as an input to itself. Then, the second network unit 112 performs image processing on the above-mentioned feature amount that serves as an input to itself, and outputs the image data of the noise removal image 819 (image with enhanced quality) at t = 1, the noise removal image 820 at t = 2, and the noise removal image 821 at t = 3.
[0068] As described above, the information processing apparatus according to the present embodiment sequentially performs image processing on the images (for example, frame images constituting a moving image) acquired at each moment, thereby acquiring images after image processing (for example, images with enhanced quality such as noise removal images).
[0069] Next, a learning apparatus that performs learning of the neural networks applied to each of the first network unit 111 and the second network unit 112 in the present embodiment will be described.
[0070] With reference to FIG. 7(b), an example of the functional configuration of the learning device 701 according to the present embodiment will be described. The learning device 701 differs from the learning device 120 shown in FIG. 2(b) in that it has a storage unit 104. The storage unit 104 is a storage area for holding image data of images after image processing (for example, high-quality images) generated by the image processing constantly executed by the learning device 701 according to the present embodiment. Note that since the other components are substantially the same as those of the learning device 120 shown in FIG. 2(b), detailed description thereof will be omitted.
[0071] Next, with reference to FIGS. 3 and 7(b), an example of the processing of the learning device 701 according to the present embodiment will be described, focusing particularly on the processing related to the learning of the neural network applied to the information processing device 700 according to the present embodiment. Note that in the present embodiment, detailed description of the processing substantially the same as that of the first embodiment described above will be omitted.
[0072] Similar to the first embodiment described above, in FIG. 3(b), the series of processes indicated by the loop end S311 show the processes related to the learning of the neural network applied to each of the first network unit 111 and the second network unit 112. By this series of processes, the update of parameters such as the weights and biases of the neural network is repeatedly executed for each piece of image data of the moving image corresponding to the learning data.
[0073] In S312, the degraded image generation unit 122 acquires one of the series of moving image data stored in the database unit 121. In S313, the degraded image generation unit 122 generates a degraded image corresponding to the time-series image obtained from the one moving image data acquired in S312. In the present embodiment, it is assumed that a noise-added image in which noise is added to a clean image is generated as the degraded image. The degraded image generation unit 122 associates the generated degraded image with the original clean image, and then assigns time information to form one set, and stores the set of these image data in the database unit 121.
[0074] The series of processes indicated by the loop end S314 shows a series of processes related to the learning of a neural network, which is executed for the image data of one moving image acquired at S312. That is, a series of pairs of clean images and degraded images generated from the image data of the target moving image at S313 are sequentially acquired along the time series, and for each pair, a series of processes related to the learning of the neural network indicated by the loop end S314 is executed. Hereinafter, the content of the series of processes indicated by the loop end S314 will be specifically described.
[0075] In S315, the image acquisition unit 110 acquires five sets of the clean image and the degraded image generated at S313 so as to be continuous in time series. At this time, among the five sets to be acquired, the image acquisition unit 110 acquires, as the two sets in order from the oldest time, the same ones as the two sets at the latest times among the series of sets acquired when image processing was executed at the previous time. However, if there is no set corresponding to the time to be acquired, the image acquisition unit 110 excludes the set from the acquisition target. In the above manner, the image acquisition unit 110 acquires sets for as many times as can be acquired. Next, the image acquisition unit 110 uses the clean image included in the acquired set as the true value image, and outputs the true value image and the degraded image included in the set to each component of the learning device 120. Specifically, the image acquisition unit 110 acquires three degraded images in order from the set with the latest time, and outputs the three degraded images to the first network unit 111. Next, the image acquisition unit 110 acquires four true value images in order from the set with the second latest time, and outputs them to the difference calculation unit 123 and the error calculation unit 124. Note that when the image acquisition unit 110 can acquire only three sets, for example, when image processing at the first time is executed, the images at the times that could not be acquired are treated as non-existent.
[0076] Here, consider an example where image processing is sequentially performed on each of the time-series images with t = 0, 1, 2, 3, 4, 5,.... At this time, when the time at which the image processing is executed is s, in the image processing at the start time s = 0, an attempt is made to acquire the image data of the degraded images at t = 0, 1, 2 and the true-value images at t = -2, -1, 0, 1, 2 from the set of times t = -2, -1, 0, 1, 2. However, in the case of s = 0, since there is no image data of the true-value images at t = -2, -1, only the image data of the true-value images at t = 0, 1 will be acquired. Also, in the image processing at the next time s = 1, an attempt will be made to acquire the image data of the degraded images at t = 2, 3, 4 and the true-value images at t = -2, -1, 0, 1 from the set of times t = 1, 2, 3, 4, 5.
[0077] Next, through a series of processes from S302 to S306, image processing is performed on the degraded image input to the first network unit 111 at S315, and a noise-removed image is obtained as a high-quality image. Note that since the series of processes from S302 to S306 are substantially the same as the series of processes from S302 to S306 shown in Fig. 3(a), detailed description thereof is omitted. Also, when a dummy image is used as the high-quality image at the previous time at this time, for example, a dummy image having the same width, height, and number of channels as the corresponding high-quality image and with the value of each pixel being 0 is used. However, for the dummy image, it is not necessarily required that the value of each pixel be 0 as long as the width, height, and number of channels are the same as those of the high-quality image used as the input. Also, when there are a plurality of images corresponding to the same time for the image input to the second network unit 112, the combined image of these images is used as the input to the second network unit 112. As a method for combining a plurality of images in this case, for example, a method of generating an image in which the pixel values of each image are added together for each channel and then the average value of the values for each channel is used as the pixel value can be mentioned. Of course, the combined method of the said image is merely an example, and the method is not particularly limited as long as the result obtained by the combination can maintain the same width, height, and number of channels as the high-quality image before combination.
[0078] In S316, the difference calculation unit 123 generates a difference image of true values using the time-series continuous true value images acquired from the image acquisition unit 110. At this time, the difference calculation unit 123 generates a difference image by calculating the difference between two adjacent time points in the same manner as in the first embodiment described above. The difference calculation unit 123 according to this embodiment acquires four time-series continuous high-quality images using three time-series continuous high-quality images output in the image processing at a certain time and one of the three high-quality images output in the image processing at the previous time. The difference calculation unit 123 generates three difference images by calculating the difference between the images in the same manner as in the first embodiment using these four high-quality images. As a specific example, in the image processing at time s = 1, the image data of three noise removal images (high-quality images) corresponding to t = 2, 3, and 4 are output. Also, in the image processing at the previous time s = 0, the image data of three noise removal images (high-quality images) corresponding to t = -1, 0, and 1 are output. In this case, the difference calculation unit 123 generates a difference image of three high-quality images by calculating the difference between two adjacent time points using the noise removal images (high-quality images) of t = 1, 2, 3, and 4. Further, the difference calculation unit 123 acquires the image data of the true value image at the time corresponding to the difference image of the high-quality image. Then, the difference calculation unit 123 generates three difference images of true values in the same manner as when generating the difference image of the high-quality image described above.
[0079] In S317, the error calculation unit 124 calculates an error used for updating the parameters of the neural network. The image reconstruction error L i is the high-quality image I^ t which is the output of the second network unit 112, and the true value image I tUsing these, it is calculated as the L1 error based on the relational expression shown as Equation (7) below. At this time, the error calculation unit 124 uses the three high-quality images output in one image process and the true value image corresponding to that time, and sets the sum of the L1 errors calculated between the images at each time as the final error value.
[0080]
Equation
[0081] The inter-frame difference error Lf is, in the same manner as in the first embodiment, the difference image I^ of the high-quality image diff and the difference image I of the true value diff Based on these, it is calculated as the L1 error represented by the relational expression shown as Equation (8) below. Note that M shown in Equation (8) is the mask created according to the difference value of the difference image of the true value, which was shown as Equation (5) in the first embodiment. However, since there are three difference images, the error calculation unit 124 sets the sum of the L1 errors calculated between the difference image of the high-quality image corresponding to each time interval and the difference image of the true value as the final error value.
[0082]
Equation
[0083] Based on the image reconstruction error L i and the inter-frame difference error L f described above, the error L finally used for updating the parameters of the neural network is calculated based on the relational expression shown as Equation (6) in the same manner as in the first embodiment described above.
[0084] In S318, the learning unit 125 updates the parameters of the neural networks of the first network unit 111 and the second network unit 112 respectively by the error backpropagation method using the error calculated in S317.
[0085] As described above, by repeatedly executing the series of processes of S311 to S318, learning of a neural network that performs image processing such as image quality improvement on the image indicated by the target image data is performed.
[0086] As described above, according to the present embodiment, when learning a neural network that performs image processing on each of a series of images that are temporally continuous, sequentially using the images as input, the correspondence relationship between the temporally continuous images, that is, the spatio-temporal information, is easily learned. Therefore, even in a situation where image processing is performed on a series of images having continuity along a time series such as a moving image, an effect of further suppressing the manifestation of a phenomenon in which artifacts generated in the images appear to fluctuate can be expected. Further, according to the present embodiment, since a plurality of images are collectively targeted for image processing, an effect of improving the processing speed related to image processing can be expected as compared with a case where a series of images (for example, a moving image) after image processing are generated while sequentially executing image processing.
[0087] <Modification Example 1> Modification Example 1 of each embodiment of the present disclosure will be described below. In this modification example, as a derivative form of the second embodiment described above, an example of a configuration will be described in which, in image processing at consecutive times, the high-quality image that is the output of the second network unit 112 does not include the high-quality image at the same time.
[0088] FIG. 9 is a diagram showing an example of the configuration of a neural network applied to the first network unit 111 and the second network unit 112 according to this modified example. In the configuration of the neural network according to the second embodiment described above, in each image processing at consecutive times s = 0, 1, a noise-removed image (high-quality image) at t = 1, which is the output of the second network unit 112, is output. On the other hand, in the configuration shown in FIG. 9, as the high-quality images that are the output of the second network unit 112, in the image processing at s = 0, a dummy noise-removed image 908 at t = -1, a noise-removed image 909 at t = 0, and a noise-removed image 910 at t = 1 are obtained. Also, in the image processing at s = 1, as the high-quality images that are the output of the second network unit 112, a noise-removed image 917 at t = 2, a noise-removed image 918 at t = 3, and a noise-removed image 919 at t = 4 are obtained.
[0089] With such a configuration, it is possible to eliminate the duplication of the output high-quality images between the image processes at consecutive times. With such a configuration, when a series of images (for example, a moving image) is generated while sequentially executing image processing, compared to the configuration according to the second embodiment, it is possible to increase the number of images output in one image process. Thereby, an improvement in the processing speed related to the image processing can be expected.
[0090] <Modified Example 2> A second modified example of each embodiment of the present disclosure will be described below. In this modified example, an example of a derivative form of the neural network applied to the second network unit 112 described above in the first and second embodiments will be described. For example, FIG. 10 is a diagram showing an example of the configuration of a neural network applied to the second network unit 112 according to this modified example. In the neural network shown in FIG. 10, the high-quality image 1000 of the second network unit 112 at the previous time used for the input is combined in the channel direction with the high-quality image that is the image processing result of the second network unit 112 at the current time by a skip connection 1002. Thereafter, the second network unit 112 applies convolution and a relu function to the combined high-quality image, and outputs the result as a new high-quality image 1003. By applying such a configuration, the high-quality image 1000, which is the image processing result at the previous time, is directly synthesized with the high-quality image that is the image processing result at the current time, and as a result, an improvement in the effect of suppressing fluctuations can be expected.
[0091] <Modification Example 3> Modification Example 3 of each embodiment of the present disclosure will be described below. In this modification example, as an example of a form of derivation of the neural network described above in the first and second embodiments, a configuration in which a noise estimation network unit for estimating noise manifested in the input image is added will be described. Since the noise estimation network unit disclosed in Non-Patent Document 3 can be applied, a detailed description thereof will be omitted. For example, FIG. 11 is a diagram showing an example of the configuration of the neural network according to this modification example. The neural network shown in FIG. 11 is different from the neural network described in Modification Example 1 in that a noise estimation network unit 1120 is added.
[0092] The image processing at time s = 0 will be described. First, the noise estimation network unit 1120 performs noise estimation by taking as input the noisy images at t = 0, the noisy image 1101 at t = 1, and the noisy image 1102 at t = 2, which are time-series images serving as the input. As a result of the noise estimation, the noise estimation network unit 1120 outputs a noise map for each of the input time-series images. In the example shown in FIG. 10, the noise map 1103 at t = 0, the noise map 1104 at t = 1, and the noise map 1105 at t = 2 are output. Note that the noise map in this modification example is a map in which the noise intensity of the input noisy image is associated as a pixel value with the position in the noisy image.
[0093] After that, the first network unit 111 executes image processing by taking as input the image data of the feature amount obtained by combining the noisy image and the noise map in the channel direction for each time along the time series. Specifically, for t = 0, the noisy image 1100 and the noise map 1103 are combined in the channel direction. Similarly, for t = 1, the noisy image 1101 and the noise map 1104 are combined, and for t = 2, the noisy image 1102 and the noise map 1105 are combined. As described above, the image data showing the feature amounts 1106, 1107, and 1108 obtained by further combining in the channel direction the feature amounts in which the noisy image and the noise map are combined for each time along the time series becomes the input for the image processing by the first network unit 111. Note that the processing until the second network unit 112 outputs the image data of the image after image processing (for example, the high-quality image) is substantially the same as that in each of the above-described embodiments and other modification examples.
[0094] By applying such a configuration, it becomes possible for the neural network to learn the characteristics of the image shown in the noise map. Therefore, for example, it becomes possible to expect an improvement in the image quality of the image output as a result of image processing. Also, by adjusting the intensity of the noise map, it becomes possible to adjust the intensity of noise removal for the input image by the neural network.
[0095] <Modification Example 4> A fourth modification of each embodiment of the present disclosure will be described below. In the first and second embodiments described above, a plurality of still images (for example, frame images) are acquired from a moving image in time series and sequentially used as targets for image processing by a neural network, and an example of a learning method for updating the parameters of the neural network was described. In this modification, an example of another learning method different from the learning method will be described. Specifically, in the learning method according to this modification, first, image processing and error calculation processing that are sequentially executed for each of a series of time-series images acquired from a moving image are completed. Then, the sum of the calculated series of errors is taken, and the calculated error as a result may be used for learning of the neural network, that is, updating of the parameters of the neural network. By applying such a configuration, for example, it becomes possible to make the neural network learn information acquired over a longer period of time, so an effect of further improving the image quality of the image output as a result of image processing can be expected.
[0096] <Modification Example 5> Modification example 5 of each embodiment of the present disclosure will be described below. In the above-described first and second embodiments, an example was described in the case where high-quality image processing such as noise removal processing for removing noise manifested in the input image is applied as the image processing performed on the input image by the neural network. On the other hand, the type of image processing performed on the input image by the neural network is not necessarily limited to noise removal processing only, and any image processing that performs some kind of processing on the target image can be applied. As a specific example, high-quality image processing for deterioration such as aberration, compression, low resolution, and defect, and contrast reduction due to weather effects such as fog, haze, snow, and rain during imaging may be applied.
[0097] Here, an example will be described in the case where blur removal processing for removing the blur manifested in the target image is applied as the image processing performed by the neural network on the input image. In this case, the image data of the time-series image that is input to the neural network is the image data of an image in which so-called blur in which the subject or the like is blurred is manifested. Also, in this case, the neural network performs image processing for removing (or reducing) the blur manifested in the input image, and outputs the image data of the image (high-quality image) on which the image processing has been performed. In addition, other configurations and processes can be realized by applying the same configurations and processes as those of the above-described first and second embodiments. Further, the degradation image generation unit 122 of the learning device may be configured to generate a blurred image as the degradation image. Specifically, the degradation image generation unit 122 may be configured to apply blur modeled from blur caused by the movement of the subject, blur occurring outside the depth of field of view, etc. to the target image. Then, by applying the image with blur applied by the degradation image generation unit 122 as the degradation image in the above-described first and second embodiments, the neural network may be trained. By performing the learning of the neural network as described above, for example, it becomes possible to obtain a high-quality image in which blurring is suppressed while realizing blur removal as high-quality image processing.
[0098] As described above, according to the technology according to each embodiment of the present disclosure, not limited to high-quality image processing such as noise removal processing, even in a situation where various image processing is performed, it is possible to make the neural network learn the correspondence relationship between temporally consecutive images. Further, due to such characteristics, according to the technology according to each embodiment of the present disclosure, it is also possible to suppress the manifestation of shaking even in a situation where various image processing is performed on the neural network for temporally consecutive images.
[0099] <Other Embodiments> The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0100] Further, the disclosure of the present embodiment includes the following configurations and programs. (Configuration 1) By sequentially using each of a series of image data having continuity along a time series as an input to a neural network that performs predetermined image processing on the input image data, a first image processing means for generating first image data after the predetermined image processing; and second image processing means for generating second image data after the predetermined image processing based on the neural network, targeting the first image data sequentially output from the first image processing means along the time series. The second image processing means uses the first image data output from the first image processing means at a first time or a second time before the first time, and the second image data generated by the second image processing means targeting the first image data output from the first image processing means at the second time, as an input to the neural network at the first time, to generate the second image data corresponding to the first time. An information processing apparatus characterized by this. (Configuration 2) The neural network is based on an error between a first difference between true value image data corresponding to a first time and a second time before the first time among a series of true value image data having continuity along a time series, and a second difference between the second image data generated by performing the predetermined image processing on the processed image data generated by subjecting the true value image data corresponding to the first time and the second time to processing restored by the predetermined image processing. The information processing apparatus according to Configuration 1, characterized in that it is learned. (Configuration 3) The information processing apparatus according to Configuration 2, characterized in that the second image data corresponding to the output of the second image processing means at the second time is used for calculating the error. (Configuration 4) The first image processing means generates, as the first image data, the image or one or more pieces of image data indicating a feature amount generated from the image, which has been subjected to the predetermined image processing, by using the image or one or more pieces of image data indicating a feature amount generated from the image as image data to be input to the neural network. The information processing apparatus according to any one of Configurations 1 to 3, characterized in that. (Configuration 5) The second image processing means uses, as input to the neural network at the first time, one or more pieces of image data indicating an image or a feature amount generated from the image, which are output as the first image data from the first image processing means at the first time, and the second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, and generates the second image data corresponding to the first time. The information processing apparatus according to claim 4, characterized in that. (Configuration 6) The first image processing means generates, as the first image data, the plurality of images or a series of image data indicating a plurality of feature amounts generated from the plurality of images, which have been subjected to the predetermined image processing, by using a series of image data indicating a plurality of images or a plurality of feature amounts generated from the plurality of images that are temporally continuous as image data to be input to the neural network. The information processing apparatus according to any one of Configurations 1 to 3, characterized in that. (Configuration 7) The second image processing means uses, as input to the neural network at the first time, a series of image data indicating a plurality of images or a plurality of feature amounts generated from the plurality of images that are temporally continuous, which are output as the first image data from the first image processing means at the first time, and the second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, and generates the second image data corresponding to the first time. The information processing apparatus according to Configuration 6, characterized in that. (Configuration 8) The second image processing means applies, as the second image data to be input to the neural network at the first time, the second image data generated by the second image processing means to the first image data output from the first image processing means at the second time. The information processing apparatus according to any one of Configurations 1 to 7, characterized in that. (Configuration 9) The second image processing means applies, as the first image data to be input to the neural network at the first time, the first image data output from the first image processing means at the first time. The information processing apparatus according to any one of Configurations 1 to 8, characterized in that. (Configuration 10) The second image processing means applies, as the first image data to be input to the neural network at the first time, the first image data output from the first image processing means at the second time. The information processing apparatus according to any one of Configurations 1 to 8, characterized in that. (Configuration 11) The second image processing means combines the second image data generated by the second image processing means at the second time with the output of the neural network at the second time and performs convolution as the first image data to be input to the neural network at the first time, thereby generating the second image data corresponding to the first time. The information processing apparatus according to any one of Configurations 1 to 10, characterized in that. (Configuration 12) The predetermined image processing is a process of removing noise manifested in the target image. The information processing apparatus according to any one of Configurations 1 to 11, characterized in that. (Configuration 13) It has an estimation means for estimating the noise manifested in the image represented by each of the series of image data, and the first image processing means generates the first image data by applying the noise estimation result by the estimation means as an input to the neural network. The information processing apparatus according to any one of Configurations 1 to 12, characterized in that. (Configuration 14) Image data generation means for generating a series of processed image data by performing first image processing on each of a series of true value image data having continuity along a time series; and for each of the series of processed image data, sequentially using the input data as an input to a neural network that performs second image processing, thereby generating first image data after the second image processing; first image processing means; second image processing means for generating second image data after the second image processing based on the neural network, targeting the first image data sequentially output from the first image processing means along a time series; difference calculation means for calculating a first difference between the true value image data corresponding to each of a first time and a second time before the first time, and a second difference between the second image data generated by performing the second image processing on the processed image data generated from the true value image data corresponding to each of the first time and the second time; error calculation means for calculating an error between the first difference and the second difference; and learning means for performing learning of the neural network for each of the first image processing means and the second image processing means based on the error. The second image processing means generates the second image data corresponding to the first time by using, as an input to the neural network at the first time, the first image data output from the first image processing means at the first time or the second time, and the second image data generated by the second image processing means targeting the first image data output from the first image processing means at the second time. Learning device characterized by the above. (Configuration 15) The learning device according to Configuration 14, wherein the error calculation means calculates an L1 error between the first difference and the second difference as the error. (Configuration 16) The learning device according to Configuration 14 or 15, wherein the learning means performs learning of the neural network based on the error backpropagation method using the error. (Configuration 17) The error calculation unit performs weighting such that the weight for the error at a location with less change between the true value image data corresponding to the first time and the second time is larger. The learning apparatus according to any one of Configurations 14 to 16, characterized in that. (Configuration 18) The error calculation unit performs weighting such that the weight for the error at a location with less change between the true value image data corresponding to the first time and the second time is larger. The learning apparatus according to any one of Configurations 14 to 16, characterized in that. (Program 1) A program for causing a computer to function as an information processing apparatus having: a first image processing means for generating first image data after the predetermined image processing by sequentially inputting each of a series of image data having continuity along a time series as an input to a neural network that performs the predetermined image processing on the input image data; and a second image processing means for generating second image data after the predetermined image processing based on the neural network, targeting the first image data sequentially output from the first image processing means along the time series, wherein the second image processing means generates the second image data corresponding to the first time by using, as an input to the neural network at the first time, the first image data output from the first image processing means at the first time or at a second time before the first time and the second image data generated by the second image processing means targeting the first image data output from the first image processing means at the second time. (Program 2) A processed image data generation means for generating a series of processed image data by subjecting a computer to a first image process on each of a series of true value image data having continuity along a time series; a first image process means for generating first image data after the second image process by using, as an input to a neural network that sequentially subjects each of the series of processed image data to a second image process on the input data; a second image process means for generating second image data after the second image process based on the neural network for each of the first image data sequentially output from the first image process means along a time series; a difference calculation means for calculating a first difference between the true value image data corresponding to each of a first time and a second time before the first time, and a second difference between the second image data generated by subjecting the processed image data generated from the true value image data corresponding to each of the first time and the second time to the second image process; an error calculation means for calculating an error between the first difference and the second difference; and a learning means for performing learning of the neural network for each of the first image process means and the second image process means based on the error, wherein the second image process means generates the second image data corresponding to the first time by using, as an input to the neural network at the first time, the first image data output from the first image process means at the first time or the second time and the second image data generated by the second image process means for the first image data output from the first image process means at the second time. A program for functioning as a learning device.
Explanation of Signs
[0101] 100 Information processing apparatus 111 First network section 112 Second network section 120 Learning device
Claims
1. First image processing means for generating first image data after the predetermined image processing by inputting each of a series of image data having continuity along a time series to the first neural network among the first neural network and the second neural network that sequentially perform the predetermined image processing on the input image data; Second image processing means for generating second image data after the predetermined image processing based on the second neural network for the first image data sequentially output from the first image processing means along the time series; characterized by comprising; The second image processing means: The first image data output from the first image processing means at a first time or a second time before the first time, The second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, are used as inputs to the second neural network at the first time to generate the second image data corresponding to the first time. An information processing apparatus characterized by the above.
2. The first neural network and the second neural network are: The first difference between the true value image data corresponding to the first time and the second time before the first time among a series of true value image data having continuity along the time series, Based on the error between the first difference and the second difference between the second image data generated by subjecting the processed image data generated by subjecting the true value image data corresponding to the first time and the second time to the processing restored by the predetermined image processing to the predetermined image processing. Learned The information processing apparatus according to claim 1, characterized by the above.
3. When calculating the second difference, a plurality of second image data that are temporally continuous and include the second image data corresponding to the output of the second image processing means at the first time and the second image data corresponding to the output of the second image processing means at the second time are used. The information processing apparatus according to claim 2, characterized by the above.
4. The first image processing means generates, as the first image data, the image or one or more pieces of image data indicating a feature amount generated from the image, on which the predetermined image processing has been performed, by using the image or one or more pieces of image data indicating a feature amount generated from the image as image data input to the first neural network. The information processing apparatus according to claim 1, characterized in that.
5. The second image processing means one or more pieces of image data indicating an image or a feature amount generated from the image, which are output as the first image data from the first image processing means at a first time, the second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, as input to the second neural network at the first time, generates the second image data corresponding to the first time The information processing apparatus according to claim 4, characterized in that.
6. The first image processing means generates, as the first image data, a series of image data indicating a plurality of images that are temporally continuous or a plurality of feature amounts generated from each of the plurality of images, on which the predetermined image processing has been performed, by using the series of image data as image data input to the first neural network. The information processing apparatus according to claim 1, characterized in that.
7. The second image processing means a series of image data indicating a plurality of images that are temporally continuous or a plurality of feature amounts generated from the plurality of images, which are output as the first image data from the first image processing means at a first time, the second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, as input to the second neural network at the first time, generates the second image data corresponding to the first time The information processing apparatus according to claim 6, characterized in that.
8. The second image processing means as the second image data to be input to the second neural network at the first time, Applying the second image data generated by the second image processing means to the first image data output from the first image processing means at the second time The information processing apparatus according to claim 1, characterized in that.
9. The second image processing means As the first image data to be input to the second neural network at the first time Applying the first image data output from the first image processing means at the first time The information processing apparatus according to claim 1, characterized in that.
10. The second image processing means As the first image data to be input to the second neural network at the first time Applying the first image data output from the first image processing means at the second time The information processing apparatus according to claim 1, characterized in that.
11. The second image processing means The second image data generated by the second image processing means at the second time, which is to be input to the second neural network at the first time By combining and performing convolution with the output of the second neural network at the second time Generating the second image data corresponding to the first time The information processing apparatus according to claim 1, characterized in that.
12. The information processing apparatus according to claim 1, characterized in that the predetermined image processing is a process of removing noise manifested in the target image.
13. Having an estimation means for estimating noise manifested in the images respectively indicated by each of the series of image data The first image processing means generates the first image data by applying the noise estimation result by the estimation means as an input to the first neural network The information processing apparatus according to claim 1, characterized in that.
14. Image data generation means for generating a series of processed image data by performing first image processing on each of a series of true value image data having continuity along a time series For each of the series of processed image data, by using the input data as an input to the first neural network among the first neural network and the second neural network that sequentially perform a second image process on the input data, first image processing means for generating first image data after the second image process; Second image processing means for generating second image data after the second image process based on the second neural network, targeting the first image data sequentially output from the first image processing means along a time series; Difference calculation means for calculating a first difference between the true value image data corresponding to each of a first time and a second time prior to the first time, and a second difference between the second image data generated by performing the second image process on the processed image data generated from the true value image data corresponding to each of the first time and the second time; Error calculation means for calculating an error between the first difference and the second difference; Learning means for performing learning of the first neural network and the second neural network, targeting the first image processing means and the second image processing means respectively based on the error; comprising; The second image processing means uses, as an input to the second neural network at the first time, the first image data output from the first image processing means at the first time or the second time and the second image data generated by the second image processing means, targeting the first image data output from the first image processing means at the second time to generate the second image data corresponding to the first time. A learning device characterized by the above.
15. The learning device according to claim 14, wherein the error calculation means calculates an L1 error between the first difference and the second difference as the error.
16. The learning device according to claim 14, wherein the learning means performs learning of the first neural network and the second neural network based on an error backpropagation method using the error.
17. The error calculation means performs weighting such that the weight for the error at a location with less change between the true value image data corresponding to each of the first time and the second time becomes larger. The learning device according to claim 14, characterized in that.
18. The error calculation means performs weighting such that the weight for the error at a location with less change between the true value image data corresponding to each of the first time and the second time becomes larger. The learning device according to claim 14, characterized in that.
19. A computer, First image processing means for generating first image data after the predetermined image processing by inputting each of a series of image data having continuity along a time series to the first neural network out of the first neural network and the second neural network that sequentially perform predetermined image processing on the input image data; Second image processing means for generating second image data after the predetermined image processing based on the second neural network for the first image data sequentially output from the first image processing means along a time series; having, The second image processing means, The first image data output from the first image processing means at a first time or a second time before the first time, The second image data generated by the second image processing means for the first image data output from the first image processing means at the second time, as an input to the second neural network at the first time, to generate the second image data corresponding to the first time A program for causing a computer to function as an information processing apparatus characterized by the above.
20. A computer, Processing image data generation means for generating a series of processed image data by performing first image processing on each of a series of true value image data having continuity along a time series; First image processing means for generating first image data after the second image processing by inputting each of the series of processed image data to the first neural network out of the first neural network and the second neural network that sequentially perform second image processing on the input data; Second image processing means for generating second image data after the second image processing based on the second neural network, targeting the first image data sequentially output from the first image processing means along the time series; Difference calculation means for calculating a first difference between the true value image data corresponding to each of a first time and a second time prior to the first time, and a second difference between the second image data generated by subjecting the processed image data generated from the true value image data corresponding to each of the first time and the second time to the second image processing; Error calculation means for calculating an error between the first difference and the second difference; Learning means for performing learning of the first neural network and the second neural network for each of the first image processing means and the second image processing means based on the error; comprising; The second image processing means: the first image data output from the first image processing means at the first time or the second time; and the second image data generated by the second image processing means, targeting the first image data output from the first image processing means at the second time, as an input to the second neural network at the first time, generate the second image data corresponding to the first time A program for causing a learning device to function, characterized by the above.
Citation Information
Patent Citations
Video upsampling using one or more neural networks
JP2022547517A
Neural network training based on consistency loss
JP2023035928A
Image blending using one or more neural network
JP2023091703A
Frame-Recurrent Video Super-Resolution
US20190206026A1
Image restoration method and apparatus, and electronic device
US20230177652A1