Image processing apparatus and learning method for image processing apparatus

The image processing device and learning method address the challenge of naturally reproducing image movement by using motion information as learning data and error calculation in the loss function, resulting in effective high-resolution image processing.

JP2025071579APending Publication Date: 2025-05-08CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023181863
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing super-resolution image processing methods using machine learning fail to naturally reproduce the movement of original moving images, especially when reconstructing time series images, and require estimation and compensation of motion between frames.

Method used

An image processing device and learning method that incorporates motion information between frames as learning data, calculates errors from estimated images and motion information, and uses these errors as a loss function to learn image processing parameters, ensuring natural reproduction of image movement.

Benefits of technology

The method effectively guarantees the natural movement of original moving images during reconstruction, achieving high-resolution image processing with a simple structure and without the need for additional motion estimation and compensation steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071579000001_ABST
    Figure 2025071579000001_ABST
Patent Text Reader

Abstract

To provide a learning method for an image processing apparatus that naturally reproduces the movement of original moving images.SOLUTION: An image processing parameter learning method comprises: image processing means; means that acquires, from time-series images, a plurality of original images, target images corresponding to the plurality of original images, and information on the movement between two images of the target images, as learning data; means that calculates a first error from an estimated image obtained by inputting at least one image of the plurality of acquired original images to the image processing means and the target image corresponding to an input image; means that estimates information on the movement between first and second estimated images obtained by acquiring first and second original images from learning data acquisition means and inputting them to the image processing means; means that calculates a second error from the estimated movement information and information on the movement between target images corresponding to the first and second original images acquired by the learning data acquisition means; and learning means that learns the parameter of the image processing means, by using the error determined by first error calculation means and the error determined by second error calculation means, as a loss function.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device that uses machine learning to increase the resolution of an image, and a learning method thereof. [Background technology]

[0002] Conventionally, methods of enlarging an image using bilinear interpolation or the like have been commonly used to improve image resolution. However, when such a simple numerical data interpolation method is applied to an image, there is a problem that the enlarged image becomes blurred and the sharpness decreases. In order to solve such problems, a method of performing super-resolution processing using machine learning has been proposed in recent years. In particular, super-resolution processing using a neural network using deep learning has significantly improved image quality (see Non-Patent Document 1). In addition, a method has been proposed for further improving image quality in moving images by using images from adjacent frames (Non-Patent Document 2). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Dong et al. Image Super-Resolution Using Deep Convolutional Networks. Proceedings of the European Conference on Computer Vision, 2014 [Non-Patent Document 2] Caballero et al. Realtime video super-resolution with spatio-temporal networks and motion compensation. IEEE Conference on Computer Vision and Pattern Recognition, 2017 [Non-Patent Document 3] Shi et al. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. Proceedings of the European Conference on Computer Vision, 2016 [Non-Patent Document 4] Xie et al. Image Denoising and Inpainting with Deep Neural Network. Conference on Neural Information Processing Systems, 2012 Summary of the Invention [Problem to be solved by the invention]

[0004] In the method proposed in Non-Patent Document 2, an image is output that is obtained by performing super-resolution processing on a single target frame in each frame of a video. Therefore, when a time-series image is reconstructed from the processing result, the movement of the original video image is not guaranteed, and the movement may be reproduced unnaturally. In addition, there is also a problem that when using images of adjacent frames, it is necessary to estimate and compensate for the movement between frames.

[0005] SUMMARY OF THE PRESENT EMBODIMENT In order to solve the above problems, an object of the present invention is to provide an image processing device which increases the resolution of moving images and reproduces the movement of the original moving images in a natural manner, and a learning method thereof. [Means for solving the problem]

[0006] In order to achieve the above object, a learning method for an image processing device of the present invention has the following arrangement.

[0007] That is, the learning method of the image processing device of the present invention comprises an image processing means, a means for acquiring a plurality of original images and target images corresponding to each of the plurality of original images, and motion information between two of the target images as learning data from a time-series image, a means for calculating a first error from an estimated image obtained by inputting at least one image of the plurality of original images acquired by the learning data acquisition means into the image processing means and a target image corresponding to the input image, a means for estimating motion information between first and second estimated images obtained by acquiring a first and second original image from the learning data acquisition means and inputting each of them into the image processing means, a means for calculating a second error from the estimated motion information and motion information between target images corresponding to the first and second original images acquired by the learning data acquisition means, and a learning means for learning parameters of the image processing means using the error calculated by the first error calculation means and the error calculated by the second error calculation means as a loss function. Effect of the Invention

[0008] According to the present invention, in learning image processing parameters in an image processing device, not only the error in the estimated image of the image processing means but also the estimated error of the motion between multiple frames in a moving image are learned as a loss function, so that when a time-series image is reconstructed from the results of image processing, the motion of the original moving image can be guaranteed. [Brief description of the drawings]

[0009] [Figure 1] 1 is a diagram showing the functional configuration of an entire system for performing processing and learning of an image processing device according to an embodiment of the present invention. [Diagram 2] 1 is a diagram illustrating a functional configuration of an image processing apparatus according to an embodiment of the present invention. [Diagram 3] FIG. 2 is a diagram illustrating a functional configuration of a learning device according to an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating an example of the configuration of an image processing unit that performs high-resolution processing. [Diagram 5] FIG. 11 is a diagram showing a flow of learning processing of the image processing device. [Figure 6] FIG. 2 is a diagram showing a processing flow of the image processing device. [Figure 7] 1 is a diagram showing a hardware configuration of an entire system for performing processing and learning of an image processing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0011] 1 is a diagram showing the functional configuration of an entire system for performing image processing and learning in an image processing device according to an embodiment of the present invention. The image processing device 100 acquires image data to be processed and outputs the image processing results. The learning device 200 learns parameters for the image processing device 100 to perform image processing.

[0012] 7 is a diagram showing the hardware configuration of the entire system for performing processing and learning of the image processing device according to this embodiment. The system for performing processing and learning of the image processing device includes an arithmetic processing device 1, a storage device 2, an input device 3, and an output device 4. Each device is configured to be able to communicate with each other, and is connected by a bus or the like.

[0013] The arithmetic processing device 1 controls the operations of the image processing device 100 and the learning device 200, executes programs stored in the storage device 2, and is composed of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The storage device 2 is a storage device such as a magnetic storage device or a semiconductor memory, and stores programs loaded based on the operation of the arithmetic processing device 1, data that must be stored for a long time, and the like. In this embodiment, the arithmetic processing device 1 performs processing according to the procedures of the programs stored in the storage device 2, thereby realizing the functions of the image processing device 100 and the learning device 200 and the processing related to the flowcharts described later. The storage device 2 also stores images to be processed by the image processing device 100 according to the embodiment of the present invention, processing results, and learning data to be processed by the learning device 200.

[0014] The input device 3 is a mouse, a keyboard, a touch panel device, a button, etc., and inputs various instructions. The input device 3 also includes an imaging device such as a camera. The output device 4 is a liquid crystal panel, an external monitor, etc., and outputs various information.

[0015] The hardware configuration of the entire system is not limited to the above configuration. For example, the image processing device 100 may include an I / O device for communicating between various devices. For example, the I / O device is an input / output unit such as a memory card or a USB cable, or a wired or wireless transmission / reception unit.

[0016] 2 is a diagram showing the functional configuration of the image processing device 100. As shown in the figure, the processing and functions of the image processing device of the present invention are realized by an image acquisition means 110, an image processing means 120, and a moving image reconstruction means .

[0017] The image acquisition means 110 acquires image data from a time-series video image captured by the camera of the input device 3 .

[0018] The image processing means 120 processes the image data acquired by the image acquisition means 110 and outputs a high-resolution image.

[0019] The video reconstructing means 130 obtains a plurality of frames of high-resolution images obtained by the image processing means 120, and reconstructs a time-series video image. The reconstructed video image is output to the output device 4, such as a monitor.

[0020] 3 is a diagram showing the functional configuration of a learning device 200. As shown in the figure, the processes and functions of the learning device of the present invention are realized by learning data storage means 210, learning data acquisition means 220, first error calculation means 230, motion estimation means 240, second error calculation means 250, and parameter learning means 260.

[0021] The learning data storage means 210 stores and holds learning data for the learning device 200 to perform learning. The learning data consists of a set of a low-resolution original image and a high-resolution target image corresponding to each original image for each frame of a time-series video. For example, this set of original image and target image is reduced (e.g., reduced by half) by bilinear interpolation to obtain the original image, with a single image constituting the video as the target image. The learning data also holds motion information between the target images of two adjacent frames. The motion information is previously associated with position coordinates as a motion vector for each pixel of the target image at the previous and next times.

[0022] Here, the video images serving as learning data can be obtained by capturing various scenes with a camera. The motion information between frames can be obtained by, for example, using the Lucas-Kanade method to obtain the motion vector between two images for each pixel. Alternatively, the motion information may be obtained by a person associating characteristic points such as corners in the images between the two images. In addition, the video images serving as learning data may be pseudo-created by performing geometric transformations such as translation, enlargement, reduction, and rotation on a single still image, and the transformation information used when the video images were created may be used to obtain the motion vectors and use them as the motion information. Even when a sufficient amount of actually captured video images cannot be obtained, the amount of learning data can be sufficiently expanded by utilizing geometric transformations.

[0023] The learning data acquiring means 220 acquires the learning data stored in the learning data storage means 210 .

[0024] The first error calculation means 230 calculates a first error from a high-resolution estimated image obtained by inputting the low-resolution original image acquired by the learning data acquisition means 220 into the image processing means 120 of the image processing device 100 and a high-resolution target image corresponding to the original image.

[0025] The motion estimation means 240 estimates motion information between high-resolution estimated images obtained by inputting the low-resolution original images of two adjacent frames acquired by the learning data acquisition means 220 into the image processing means 120 of the image processing device 100.

[0026] The second error calculation means 250 calculates a second error from the motion information estimated by the motion estimation means 240 and the motion information between the target images corresponding to the original images of two adjacent frames acquired by the learning data acquisition means 220.

[0027] The parameter learning means 260 learns the parameters of the image processing means 120 of the image processing device 100 using the first error calculated by the first error calculation means 230 and the second error calculated by the second error calculation means 250 as a loss function.

[0028] The learning and image processing operations of the image processing device according to the embodiment of the present invention will be described below. Here, the image processing means 120 increases the resolution of the input image, for example using the method proposed in Non-Patent Document 1. Figure 4 shows the configuration. The image enlargement process 1210 takes an RGB image with a resolution of H x W as input and enlarges the image to a resolution of 2H x 2W by bicubic interpolation. The convolutional neural network 1220 is a neural network consisting of three convolutional layers, and performs processing so that the image enlarged by the image enlargement process 1210 becomes the target high-resolution image.

[0029] FIG. 5 shows the flow of the learning process of the image processing device.

[0030] The learning data acquisition means 220 acquires the learning data stored in the learning data storage means 210 (S201). As described above, the learning data consists of a set of low-resolution source images and high-resolution target images corresponding to each of the source images. In the present invention, in order to guarantee the motion of the original video when a time-series image is reconstructed from the result of image processing, N pairs of source and target images of two adjacent frames are acquired. In addition, motion information between the target images of two adjacent frames is also acquired.

[0031] The image processing means 120 of the image processing device 100 inputs the original image of the learning data acquired in S201 and obtains a high-resolution estimated image (S202). In this embodiment, as described above, the image processing means 120 estimates a high-resolution image by the image enlargement process and the convolutional neural network shown in Fig. 4. A high-resolution image is obtained for each of the N pairs of original images of the learning data.

[0032] The first error calculation means 230 calculates a first error from the high-resolution estimated image acquired in S202 and the target image of the learning data acquired in S201 (S203). The estimated image and the target image as the correct answer of one pair of the n-th (1≦n≦N) learning data are respectively represented as I n,k , I' n,k Then the first error E1 n,k is calculated by (Equation 1). Here, k represents one of the two learning data that make up a pair and takes the value 1 or 2. |I n,k -I' n,k | is Image I n,k and I' n,k It represents the sum of the absolute values ​​of the differences in pixel values ​​of corresponding pixels that make up E1 n,k = |I n,k -I' n,k | (Formula 1)

[0033] The motion estimation means 240 estimates motion information between the high-resolution estimated images of the two adjacent frames acquired in S202 (S204). For example, the motion information can be obtained by estimating a motion vector between the two images for each pixel using the Lucas-Kanade method. Alternatively, the motion vector can be estimated by inputting the two images using a neural network.

[0034] The second error calculation means 250 calculates a second error from the motion information estimated in S204 and the motion information in the learning data acquired in S201 (S205). The motion vector estimated from the n-th pair of learning data and the correct motion vector are respectively f n , f' n Then the second error E2 n can be calculated using Equation 2. Here, |f n -f' n | is the motion vector f n and f' n represents the sum of the absolute values ​​of the differences between the corresponding horizontal and vertical components that make up E2 n = |f n -f' n | (Formula 2)

[0035] In this embodiment, the first error and the second error are calculated as the absolute value of the difference between the estimated value and the target value as shown in (Equation 1) and (Equation 2). However, other methods, such as squaring the difference, may be used.

[0036] The parameter learning means 260 calculates a total loss value of the first error calculated in S203 and the second error calculated in S205 (S206). The loss function L is as shown in (Equation 3). Here, L1 and L2 are the sums of the first error and the second error for the learning data, respectively. Also, λ is a weighting coefficient representing the weighting of the first error and the second error in the loss function. Σ represents the sum for the learning data. L = L1+λ L2 (Eq. 3) L1 = ΣE1 n,k , L2 = ΣE2 n

[0037] In addition, in (Equation 3), the first error sum L1 may be found in both cases where k=1 and k=2, that is, for two adjacent frames, or may be found as the sum for only one of the images.

[0038] In order to suppress overlearning of the neural network, a regularization term such as the sum of absolute values ​​of the parameters to be obtained may be added to (Equation 3).

[0039] The parameter learning means 260 updates the parameters of the neural network, which is the image processing means, based on the loss value calculated in S206 (S207). The parameters to be updated are the weight coefficients of the convolution layer of the neural network. The parameter updating is performed by the backpropagation method or the like.

[0040] Through steps S201 to S207, learning for the learning data acquired in S201 is completed.

[0041] The parameter learning means 260 judges whether or not learning has been completed according to a preset termination condition (S208). As the termination condition, a method is used in which learning data for accuracy verification is prepared separately from learning data for parameter update, and a loss value is calculated by performing the above-mentioned processes from S201 to S206, and whether or not the loss value has become equal to or less than a predetermined value is judged. Alternatively, the judgment may be made based on the number of times steps from S201 to S207 are repeated. If it is judged that learning has not been completed, the process returns to S201.

[0042] The learning of the image processing device has been described above. The following describes the processing of the image processing device using the parameters learned by the above-mentioned learning method. Figure 6 shows the flow of processing of the image processing device.

[0043] The image acquisition means 110 acquires a single image data for each frame from a time-series video image to be subjected to high-resolution processing (S301).

[0044] The image processing means 120 inputs the image data acquired in S301 and obtains a high-resolution estimated image (S302). The estimated image is obtained by the image enlargement process and the convolutional neural network shown in Fig. 4 as described above. Here, the parameters of the convolutional neural network are obtained by the learning method described above.

[0045] As shown in FIG. 6, the process of S301 to S302 is repeated for each frame of the moving image acquired by the image acquisition means 110.

[0046] The video reconstructing means 303 reconstructs a time-series video from the multiple high-resolution images obtained in S302 (S303). The reconstructed video is output to the output device 4, such as a monitor.

[0047] As described above, in the embodiment of the present invention, not only the estimation error of the image processing means but also the estimation error of the motion between multiple frames in the video is learned as a loss function. Then, using the learned neural network parameters, the image processing device estimates a high-resolution image for each frame of the video, and the video is reconstructed. By performing such a learning process, it is possible to obtain a high-resolution video that guarantees the motion of the original video and reproduces natural motion. Furthermore, in the image processing of the above-mentioned embodiment, it is not necessary to perform a process of estimating and compensating for the motion between frames as in the method of Non-Patent Document 2, and it is possible to realize high-resolution video with a simple configuration.

[0048] In this embodiment, the second error is calculated by estimating the motion vector as the motion information, but other methods may be used. For example, geometric transformation parameters such as translation, enlargement / reduction, and rotation that represent the motion of the entire image may be used as the motion information. In this case, the geometric transformation parameters of two adjacent frames are stored in advance as learning data, the geometric transformation parameters between the estimated images are estimated in S204, and the difference between the geometric transformation parameters is calculated as the second error in S205. In addition, not only the speed but also the acceleration may be used as the motion information. In this case, the motion vector and the acceleration vector are calculated from three consecutive frames. The motion vectors and the acceleration vectors of the three consecutive frames are stored in advance as learning data. Then, the motion vector and the acceleration vector between the estimated images are estimated in S204, and the difference between the motion vector and the acceleration vector is calculated as the second error in S205.

[0049] Moreover, the learning method of the Generative Adversarial Networks can also be applied to the learning method of the present invention. In the learning of the Generative Adversarial Network, in the learning of the Generator that generates images, first, a Discriminator that identifies images generated by the Generator is learned. Then, the Generator is trained to deceive the trained Discriminator, so that it generates images that are indistinguishable from the real thing. When applying to the above-mentioned embodiment, the loss function L during learning of the convolutional neural network 1220, which is the Generator, can be calculated by the following (Equation 4). L = L1+λ·L2+λ′·L3 (Equation 4) L3 = ΣE3 n,k

[0050] However, L3 is an adversarial loss to deceive the discriminator, and E3 n,kis the loss for each estimated image. In (Equation 4), when a high-resolution estimated image estimated by the image processing means 120 is input to the Discriminator, if it is identified as a true high-resolution image, E3 n,k = 0, if it is identified as being estimated by the Generator, E3 n,k = 1. λ′ is a weighting coefficient that represents the weighting of the adversarial loss in the loss function.

[0051] In the above explanation, the method proposed in Non-Patent Document 1 was used as the image processing means for increasing the resolution of the input image, but other methods may be used. For example, the method proposed in Non-Patent Document 3 uses Subpixel Convolution without using image enlargement processing to obtain a high-resolution image from the input image using only a neural network.

[0052] Another feature of the present invention is that a second error based on motion information is used as a loss function when learning the parameters of the image processing means, and it is also possible to combine this with the method proposed in Non-Patent Document 2.

[0053] Although the present invention has been described above as an example of increasing the resolution of an image, the present invention can also be applied to other image processing. For example, in Non-Patent Document 4, noise reduction processing and image restoration of an input image are performed by a neural network.

[0054] Even in such a neural network, when image processing is applied to each frame of a video and a time series image is reconstructed from the processing results, there is a possibility that the motion will be reproduced unnaturally. In learning the parameters of the neural network, by learning not only the image estimation error but also the motion estimation error between multiple frames as a loss function, it is possible to guarantee the motion of the original video, as in the embodiment of the present invention. [Industrial Applicability]

[0055] Needless to say, the object of the present invention can also be achieved by supplying a recording medium (or storage medium) on which the program code of the software for realizing the functions of the above-mentioned embodiments is recorded to a system or device, and having the computer (or CPU or MPU) of the system or device read and execute the program code stored in the recording medium. In this case, the program code itself read from the recording medium will realize the functions of the above-mentioned embodiments, and the recording medium on which the program code is recorded constitutes the present invention.

[0056] Furthermore, it goes without saying that the functions of the above-mentioned embodiments are not only realized by the computer executing the program code it has read, but also that the functions of the above-mentioned embodiments are realized by an operating system (OS) running on the computer performing some or all of the actual processing based on the instructions of the program code.

[0057] Furthermore, it goes without saying that this also includes cases where the program code read from the recording medium is written into a memory provided on a function expansion card inserted into a computer or a function expansion unit connected to a computer, and then a CPU provided on the function expansion card or function expansion unit performs some or all of the actual processing based on the instructions of the program code, thereby realizing the functions of the above-mentioned embodiments.

[0058] When the present invention is applied to the above-mentioned recording medium, the recording medium stores program code corresponding to the processing flow described above. [Explanation of symbols]

[0059] 1 Processing Unit 2 Storage device 3 Input Devices 4 Output Device 100 Image processing device 110 Image Acquisition Means 120 Image processing means 130 Video image reconstruction means 200 Learning Device

Claims

1. a means for acquiring, from a time-series image, a plurality of original images, target images corresponding to the plurality of original images, and motion information between two of the target images as learning data; a means for calculating a first error from an estimated image obtained by inputting at least one image of the plurality of original images acquired by the learning data acquisition means into the image processing means and a target image corresponding to the input image; a means for estimating motion information between first and second estimated images obtained by acquiring a first and second original image from the learning data acquisition means and inputting each of the first and second original images into the image processing means; a means for calculating a second error from the estimated motion information and motion information between target images corresponding to the first and second original images acquired by the learning data acquisition means; and a learning means for learning parameters of the image processing means using an error calculated by the first error calculation means and an error calculated by the second error calculation means as a loss function.

2. 2. The learning method according to claim 1, wherein said image processing means performs any one of a process for increasing the resolution of the input image, a process for reducing noise, and a process for restoring an image.

3. 2. The method according to claim 1, wherein the motion information is a motion vector.

4. 2. The method according to claim 1, wherein said image processing means comprises a convolutional neural network, and said learning means learns parameters of the convolutional neural network.

5. 5. An image processing device comprising image processing means trained by the training method according to any one of claims 1 to 4.

6. 6. The image processing device according to claim 5, further comprising: a means for acquiring an image from a time series of images; and a moving image reconstructing means for reconstructing a time series of images from an estimated image obtained by inputting each of a plurality of images acquired by the image acquiring means into the image processing means.