Image processing device, learning device, image processing method, and program

The image processing apparatus addresses the challenge of improving recognition module accuracy by using reinforcement learning to optimize image improvement processes, resulting in enhanced image quality tailored for specific recognition tasks.

JP2025080559APending Publication Date: 2025-05-26CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023193797
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-26

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to improve the accuracy of external recognition modules, particularly when dealing with noisy or degraded images, as they require faithful reproduction of high-quality images which may not always be optimal for recognition tasks.

Method used

An image processing apparatus that includes image data acquisition, image improvement processing, and policy generation using reinforcement learning to optimize the image improvement process specifically for enhancing the accuracy of external recognition modules.

Benefits of technology

The proposed solution enables the generation of improved images that specifically enhance the accuracy of external recognition modules, such as person detection or face authentication, by dynamically adjusting image processing parameters based on the accuracy of these modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025080559000001_ABST
    Figure 2025080559000001_ABST
Patent Text Reader

Abstract

To provide an image processing device capable of achieving image improvement processing with improved accuracy of an external module, a learning device, an image processing method, and a program.SOLUTION: An image processing device 1000 which improves an input image and performs recognition processing comprises: an image data acquisition unit 1001 which acquires an input image; an external module 1004 which performs recognition processing; an image improvement processing unit 1002 which generates input data input from the input image to the external module 1004; and a policy generation unit 1003 which generates parameters used by the image improvement processing unit 1002. The policy generation unit 1003 is optimized through reinforcement learning using the accuracy of the external module 1004.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, a learning apparatus, an image processing method, and a program.

Background Art

[0002] There are image quality improvement processes such as noise reduction processing for reducing noise included in an input image, super-resolution processing for enlarging the resolution of the input image, and blur removal processing for removing blur caused by movement or focus. Also, methods for realizing these processes by machine learning have been proposed.

[0003] Most of the techniques for image quality improvement processing by machine learning model the process of image quality degradation and generate artificial degraded images by simulating the degradation for the image before degradation. Then, supervised learning is performed with the degraded image as input data and the image before degradation as GT (Ground Truth, correct value), and the image quality improvement processing is performed using the model obtained by the learning. That is, one major objective of image quality improvement processing based on machine learning is to faithfully reconstruct the image before degradation from the degraded image.

[0004] On the other hand, due to the progress of machine learning technology, various image recognition technologies such as OCR for recognizing characters in an image, object detection for detecting target objects, and face authentication technology for authenticating a person based on a face are actually applied in software applications. OCR is an abbreviation for Optical Character Recognition.

[0005] Since the recognition accuracy of these image recognition technologies depends on the quality of the input image, in order to obtain a high-precision recognition processing result, a higher-quality input image is preferred, but it is not always the case that a faithful reproduction of the pre-deterioration image is suitable. For example, when an image taken with underexposure or overexposure has deteriorated due to noise, correcting the brightness to an appropriate level may result in better accuracy of the recognition result than faithfully reproducing the image taken with underexposure or overexposure through high-image-quality processing. That is, it is considered that tuning the high-image-quality processing based on the accuracy of the module performing the recognition processing can directly affect the improvement of the accuracy of the recognition processing.

[0006] Non-Patent Document 1 discloses a technique of connecting a network that performs a high-level vision task after a neural network that performs noise reduction processing and training the network for noise reduction processing. The high-level vision task is an image recognition task such as image classification and semantic region segmentation. In this training, by calculating and integrating both the loss for noise reduction processing and the loss of the high-level vision task, learning of noise reduction processing that improves the noise reduction performance and at the same time improves the accuracy of the recognition processing has been realized.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0008] Modules for performing image recognition may be provided as OSS or proprietary software and may be used to execute recognition processing via an API. OSS is an abbreviation for Open Source Software. API is an abbreviation for Application Programming Interface.

[0009] A neural network has a form in which layers that are differentiable functions are stacked. By performing forward propagation on the input data, calculating the error between the obtained inference result and GT using a loss function, backpropagating the calculated loss, calculating the correction amount of the parameters, and updating the parameters based on the correction amount, one iteration of learning is performed. That is, for the learning of a neural network, the network needs to be composed of differentiable layers.

[0010] Furthermore, the learning of a neural network not only needs to be differentiable in principle, but also needs to actually calculate both forward propagation and backpropagation. Therefore, it is usually executed using a learning framework of a single neural network. That is, in order to learn the high-quality processing optimized for the accuracy of a recognition module that executes a certain recognition task within the framework of Non-Patent Document 1, the recognition module needs to be implemented in the framework to be used.

[0011] However, in the case of a module provided as a binary file and performing recognition processing via an API, or implemented in another framework, the framework of Non-Patent Document 1 cannot be used unless it is implemented or transplanted into the framework to be used.

[0012] The present invention has been made to solve the above problems, and an object thereof is to provide an image improvement process that improves the accuracy of an external module.

Means for Solving the Problems

[0013] An image processing apparatus according to an embodiment of the present invention is an image processing apparatus that improves an input image and performs recognition processing, and includes: image data acquisition means for acquiring the input image; an external module that performs recognition processing; image improvement processing means for generating input data to be input to the external module from the input image; and policy generation means for generating parameters used by the image improvement processing means, wherein the policy generation means is optimized by reinforcement learning using the accuracy of the external module.

Effects of the Invention

[0014] According to the present invention, it is possible to provide an image improvement process that improves the accuracy of an external module.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0016] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. Note that the following embodiments do not limit the invention according to the claims, and not all combinations of the features described in the embodiments are essential for the solution means of the invention. In each figure, the same components may be denoted by the same reference numerals and the description may be omitted.

[0017] <Embodiment 1> In this embodiment, a case of image enhancement processing for improving the accuracy of an external module for person detection using a noisy image degraded by noise as an input will be described.

[0018] FIG. 1 is a diagram for explaining an example of a problem to be solved according to Embodiment 1 of the present invention. Image 101 is a noise-free image before being degraded by noise, in which a road, a person, and a car are shown. Person 102 is a person who is a detection target in the present embodiment. Image 103 is a noisy image in which degradation due to noise has occurred. Image 104 is an example of an enhanced image obtained by performing a process of improving the accuracy of person detection, which is an external module. Rectangle 105 is an example in which a bounding box output by an external module that performs person detection is drawn. Hereinafter, the external module that performs person detection will be appropriately referred to as a person detection module.

[0019] In the present embodiment, the noise is assumed to be noise generated in the process of converting photons detected by an image sensor into digital signals. As noise sources, there are photon shot noise, readout noise, dark current noise, quantization error, etc., but it is assumed that they are modeled. That is, it is assumed that a noisy image (for example, Image 103) with respect to a noise-free image (for example, Image 101) can be artificially generated by a noise model.

[0020] The process of estimating a noise-free image such as Image 101 from a noisy image such as Image 103 is called a noise reduction process. In the present embodiment, the main purpose is not to faithfully reproduce a noise-free image from a noisy image, but to estimate an enhanced image (for example, Image 104) that maximizes the accuracy when detecting the rectangle 105 by the person detection module.

[0021] FIG. 2 is a diagram showing the hardware configuration of the image processing apparatus according to Embodiment 1 of the present invention. The image processing apparatus 1000 includes a CPU 1201, a ROM 1202, a RAM 1203, an external memory 1204, an input unit 1205, a display unit 1206, an I / O 1207, and a system bus 1210. The CPU is an abbreviation for Central Processing Unit. The ROM is an abbreviation for Read Only Memory. The RAM is an abbreviation for Random Access Memory. The I / O is an abbreviation for Input / Output.

[0022] The CPU 1201 controls various devices connected to the system bus 1210. The ROM 1202 records the BIOS program and the boot program. The BIOS is an abbreviation for Basic Input Output System. The RAM 1203 is used as the main memory device of the CPU 201. The external memory 1204 stores the programs executed by the image processing apparatus 1000. The external memory 1204 also stores the image data and various parameters used when the image processing apparatus 1000 executes the program. The external memory 1204 may be an HDD or an SSD built into the image processing apparatus 1000. The HDD is an abbreviation for Hard Disk Drive. The SSD is an abbreviation for Solid State Drive. The image processing apparatus 1000 is an example of an information processing apparatus.

[0023] The input unit 1205 is a keyboard, a mouse, etc., and performs processing related to input of information and the like. The display unit 1206 outputs the calculation result of the image processing apparatus 1000 to the display device according to an instruction from the CPU 1201. The display device is, for example, a liquid crystal display device. The I / O 1207 is a communication interface and performs information communication via a network. The communication interface may be Ethernet (registered trademark), and the type such as USB or serial communication does not matter. The USB is an abbreviation for Universal Serial Bus.

[0024] FIG. 3 is a diagram showing the functional configuration at runtime of the image processing apparatus according to Embodiment 1 of the present invention. The image processing apparatus 1000 includes an image data acquisition unit 1001, an image improvement processing unit 1002, a policy generation unit 1003, and an external module 1004. Details of these functional configurations will be described with reference to FIG. 4 and the like.

[0025] FIG. 4 is a flowchart showing the processing at runtime of the image processing apparatus according to Embodiment 1 of the present invention. The processing at runtime will be described using this flowchart. This processing is realized by a program being executed by the CPU 1201.

[0026] In step S1001, the image data acquisition unit 1001 acquires image data. Step S1001 is an image data acquisition step. The image data acquired in step S1001 is the target data for image improvement processing and is a noisy image. The image data may be an image having 3 channels of RGB, a RAW image having 4 channels of RG1G2B, a grayscale image, or a hyperspectral image. Also, data other than images may be added to the input data. For example, in Non-Patent Document 2, a noise level map is estimated from a noisy image in the previous stage, the noisy image and the estimated noise level map are concatenated, and input to a noise reduction network in the subsequent stage to reconstruct a noise-reduced image. As in this example, information other than images may be input in association with the noisy image.

[0027] In step S1002, the image improvement processing unit 1002 creates an improved image from the image data acquired in the previous step. Step S1002 is an image improvement processing step. In the present embodiment, it is assumed that the image improvement processing unit 1002 is a single neural network (hereinafter referred to as an image improvement processing network). The configuration of the neural network assumed in the present embodiment is shown in FIG. 5. The image improvement processing unit 1002 generates input data to be input to the external module 1004. This input data is, for example, an image.

[0028] FIG. 5 is a diagram showing a neural network included in the image improvement processing unit 1002 according to Embodiment 1 of the present invention. A network 401 shown surrounded by a broken line is a part responsible for processing corresponding to the image improvement processing step (step S1002). The network 401 performs hierarchical encoding and decoding on the input image 402 using a convolutional layer, a pooling layer, or the like, and generates an improved image. The improved image 403 is an improved image generated by the network 401.

[0029] The network 401 can dynamically change parameters at runtime, In this embodiment, it is assumed that parameters 404, 405, and 406, which are parameters (weights and biases) of the convolutional layer of the network 401, can be changed. The weight may be something to multiply the parameter. The expression of the effects of the present invention is not limited to the position and number of parameters to be changed at runtime described in this embodiment. That is, for example, only the parameter 406 or a form of changing parameters other than the parameters 404, 405, and 406 may be used.

[0030] In the first time of step S1002, it is assumed that the parameters 404, 405, and 406 use the learned parameters as initial values. This learned parameter is a parameter pre-learned to perform noise reduction processing for faithfully reproducing a noise-free image from a noisy image.

[0031] In the loop L1001 of the flowchart in FIG. 4, after the first time, parameters estimated in the policy generation step (step S1003) described later are used.

[0032] Note that the image improvement processing step may have one or more outputs, and at least one of them may generate input data for the external module 1004. When the image improvement processing step has two or more outputs, at least one of them may be an image obtained by restoring the image before degradation when the input image is a degraded image.

[0033] In step S1003 of FIG. 4, the policy generation unit 1003 estimates the parameters of the image improvement process. Step S1003 is a policy generation step. In this embodiment, it is assumed that this policy generation unit 1003 is a single neural network (hereinafter referred to as a policy generation network). FIG. 6 shows an example of the network assumed here.

[0034] FIG. 6 is a diagram showing the neural network included in the policy generation unit 1003 according to Embodiment 1 of the present invention. The network 501 shown surrounded by a broken line is a neural network composed of a convolutional layer, a pooling layer, a fully connected layer, etc. It receives the improved image 502 as an input and outputs information for determining parameters to the output terminal 503. This network 501 is trained in the framework of reinforcement learning described later, and the output terminal 503 corresponds to an action (policy) in reinforcement learning. The network 501 has a fully connected layer 510 of the policy generation network. The policy generation unit 1003 is optimized by reinforcement learning using the accuracy of the external module 1004.

[0035] In reinforcement learning, during learning, an action is often sampled from a probability distribution, and during inference, an action showing the maximum probability is often selected. Here, it is also assumed that an action showing the maximum probability is selected. Further, the information for determining the output parameters may be the parameters (parameters 404, 405, 406 in FIG. 5) of the convolutional layer of the image improvement processing network (network 401 in FIG. 5) itself. Further, the information for determining the output parameters may be the update amount from the parameters of the current step in loop L1001. In this case, letting the parameters of the current step be xt and the output of the policy generation network at the current step be Δxt+1, the parameters xt+1 of the next step are determined by Equation 1.

Equation

[0036] The determined parameters can be deformed to be the kernels and biases of the convolutional layers (the convolutional layers using the parameters 404, 405, and 406 in FIG. 5) of the image improvement processing network (network 401 in FIG. 5) (output end 504 in FIG. 6).

[0037] As another form, a network such as the network 506 in FIG. 7 may be assumed. FIG. 7 is a diagram showing another example of the neural network included in the policy generation unit 1003, different from FIG. 6. Here, the parameter 507 in FIG. 7 is the parameter of the current step. This is input to the network 506 and integrated with the information derived from the input image by the internal process 508, and the parameter 509 of the next step is estimated.

[0038] In addition, the information for determining the parameters output by the policy generation network may be only the scalar values multiplied by the kernels of the convolutional layers of the image improvement processing network and the biases.

[0039] Note that the policy generation network also has an output end 505 different from the output end 504, and this is the output of the value function in the Actor-critic method which is one of the reinforcement learning methods. In the present embodiment, it will be described based on the Actor-critic method, but the manifestation of the effects of the present invention is not limited to this, and another reinforcement learning framework such as the Q-learning method may be used.

[0040] In the branch B1001 of FIG. 4, the CPU 1201 makes an end determination for the loop L1001. For the end determination, an output end for outputting a special token indicating the end may be provided in the policy generation network (network 501 in FIG. 6), and the determination may be made based on the output value thereof. Also, the end determination may be made based on the output value of the value prediction (the value of the output end 505 in FIG. 6). Further, a predetermined number of times may be determined in advance, and it may be determined to end when the loop of the predetermined number of times is completed.

[0041] In step S1004 of FIG. 4, the external module 1004 receives the improved image as input and executes recognition processing. Step S1004 is a recognition step. In the recognition step, a predetermined recognition result for the input image is output. In this embodiment, since person detection is assumed as the recognition processing, the coordinate values of the circumscribed rectangle indicating the person position, the reliability, etc. are output. Also, if a reliability threshold value has been determined in advance, the circumscribed rectangles below a predetermined threshold value are filtered out. The output result may be drawn on the image like rectangle 105 in FIG. 1, or may be used in some subsequent processing.

[0042] The above is the processing at runtime of this embodiment.

[0043] Next, the processing during learning will be described with reference to the functional configuration diagram of FIG. 8 and the flowchart of FIG. 9. In this embodiment, the processing during learning is performed by the learning device 3000. The learning device 3000 is an example of an information processing device. Since the hardware configuration of the learning device 3000 is the same as that of the image processing device 1000 in FIG. 2, detailed description thereof is omitted. When referring to the hardware configuration of the learning device 3000, the reference numerals in FIG. 2 are used. The learning device 3000 may be the same device as the image processing device 1000.

[0044] FIG. 8 is a diagram showing the functional configuration of the learning device according to Embodiment 1 of the present invention. The learning device 3000 includes a learning data storage device 3001, a learning data acquisition unit 3002, a degradation unit 3003, an image improvement processing unit 3004, a policy generation unit 3005, and an external module 3006. The learning device 3000 also has each functional configuration unit such as a reward calculation unit 3007, a loss calculation unit 3008, and a parameter update unit 3009, and a storage device. Details of each of these functions will be described with reference to FIG. 9 and the like.

[0045] FIG. 9 is a flowchart showing the processing of the learning device according to Embodiment 1 of the present invention. Using this flowchart, the processing during learning will be described. This processing is realized by a program being executed by the CPU 1201 of the learning device 3000.

[0046] In step S4001, the image improvement processing unit 3004 and the policy generation unit 3005 perform settings related to learning. Step S4001 is a setting step. In the present embodiment, the image improvement processing unit 3004 is realized by the neural network (network 401) shown in FIG. 5. Also in the present embodiment, the policy generation unit 3005 is realized by the neural network (network 501) shown in FIG. 6.

[0047] When initially learning the model of the neural network, initial values generated based on random numbers or the like are set for the parameters of each layer of the model. When additionally learning the learned parameters, the learned parameters are set. Regarding the image improvement processing unit 3004, it is assumed that learned parameters are set. At this time, it is assumed that the image improvement processing unit 3004 has been learned to perform noise reduction processing for faithfully reproducing a noise-free image from a noisy image, and the learned parameters have been obtained. That is, the learned parameters of the noise reduction processing are set for the parameters of network 401 in FIG. 5. Also, initial values are set for the parameters of the policy generation unit 3005.

[0048] Also in step S4001, the CPU 1201 performs settings for other hyperparameters related to learning. In the present embodiment, it is assumed that learning is performed by the stochastic gradient descent method. Therefore, as hyperparameters, the mini-batch size, the learning coefficient, the parameters of the solver of the stochastic gradient descent method, etc. are set.

[0049] The loop L4001 in FIG. 9 is a loop related to the iteration of the stochastic gradient descent method, and the processing from steps S4002 to S4009 is repeated N times. N may be a preset value. The loop L4001 may be terminated when the loss drops below a certain level.

[0050] In step S4002, the learning data acquisition unit 3002 acquires learning data from the learning data storage device 3001. Step S4002 is a learning data acquisition step. In the learning data storage device 3001, noise-free images, which are images with extremely little noise, and the GT of the coordinate values of the bounding rectangles of the people existing in this image are stored in association with each other. The images here are images suitable for the object of image improvement processing, such as RGB images, RAW images, grayscale images, hyperspectral images, etc. Also, the number of bounding rectangles is equal to the number of people existing in the image, and if there are no people in the image, there are no bounding rectangles. In this step, noise-free images are acquired in the amount of the mini-batch size set in the previous step.

[0051] In step S4003, the degradation unit 3003 adds artificial noise to the noise-free image. Step S4003 is a degradation step. As described above, it is assumed that the noise is modeled. In step S4003, noise is generated from the noise model to artificially create a noisy image.

[0052] Loop L4002 is a loop of time steps during one episode in reinforcement learning. The time length of one episode may be set in advance, such as 50 steps, or an end condition may be set based on the recognition accuracy described later, and the loop may be terminated when the end condition is satisfied. Loop L4002 repeats the processes of steps S4004 to S4007.

[0053] In step S4004, the image improvement processing unit 3004 performs image improvement processing. Step S4004 is an image improvement processing step. The image improvement processing is performed by the image improvement processing network (network 401 in FIG. 5) described at runtime. Also, the parameters (parameters 404, 405, 406 in FIG. 5) are set as the learned parameters of the noise reduction process as the initial values as described above.

[0054] In step S4005, the policy generation unit 3005 generates the parameters for the next step. Step S4005 is a policy generation step. The generation of the parameters is performed by the policy generation network (network 501 in FIG. 6, network 506 in FIG. 7) described at runtime.

[0055] In the processing during learning, the output terminal (output terminal 504 in FIG. 6, output terminal 509 in FIG. 7) outputs the parameters of the probability density function, and an action is sampled from the probability distribution based on the parameters. As described above, the action may be the parameter itself, but may also be an update amount from the parameters of the previous step, a scalar value multiplied by the kernel, a bias, etc. The parameters for the next step (parameters 404, 405, 406 in FIG. 5) are determined by any of these.

[0056] In step S4006, the external module 3006 performs recognition processing. Step S4006 is a recognition step. Here, the improved image generated in the image improvement processing step (step S4005) is input to the person detection module, and the coordinate values and reliability of the circumscribed rectangle of the person in the image are obtained. In step S4006, similar to the processing at runtime, if a reliability threshold has been determined in advance, circumscribed rectangles with a reliability below this threshold are filtered out.

[0057] In step S4007, the reward calculation unit 3007 calculates the reward. Step S4007 is a reward calculation step. Here, the reward is the detection accuracy between the coordinate values output in the recognition step (step S4006) and the GT. In the association between the GT of the circumscribed rectangle and the detection result, the IoU is calculated, and the success or failure of the detection is determined with an appropriate threshold such as 0.5. IoU is the abbreviation of Intersection over Union. And as the detection accuracy, Precision, Recall, F1-score, etc. are used appropriately according to the purpose. If you want to prevent over-detection, choose Precision; if you want to prevent under-detection, choose Recall; if you want to balance both and improve them better, choose the F1-score. This evaluation value is used as the reward in reinforcement learning.

[0058] In loop L4002, the above processing for one episode is executed until a predetermined end condition is satisfied.

[0059] In step S4008, the loss calculation unit 3008 calculates the loss. Step S4008 is a loss calculation step. The loss may be calculated based on the difference between the cumulative reward and the value prediction value, the entropy of the probability distribution generated by the policy generation network, etc., based on the Actor-critic method. Further, a loss based on image generation (image generation loss) may be integrated into this loss. Examples of the image generation loss include the L1 or L2 loss between the noise-free image and the improved image, and the Total Variation regularization of the improved image. In step S4008, these losses are integrated with appropriate weights to obtain the final loss.

[0060] In step S4009, the parameter update unit 3009 updates the parameters of the policy generation unit 3005. Step S4009 is a parameter update step. In this embodiment, based on the loss calculated in the loss calculation step (step S4008), the update amount of the parameters of each layer of the policy generation network is determined by the error backpropagation learning method, and the parameters are updated. The policy generation network is the network 501 in FIG. 6 or the network 506 in FIG. 7.

[0061] The above is the detailed processing during learning. Through these processes, the parameters of the policy generation unit 3005 used at runtime are learned.

[0062] This concludes the description of Embodiment 1. By executing as described above, an improved image that improves the accuracy of the external module for person detection can be generated from a noisy image degraded by noise.

[0063] In this embodiment, only the noise of the image sensor is assumed as the factor degrading the image. However, the present invention is not limited to only such degradation. For example, the present invention is applicable to degradation factors such as resolution degradation, blurring (optical blur, motion blur), and degradation due to haze in the atmosphere. However, in any case, it is necessary to model the degradation process and perform artificial degradation simulation on the pre-degraded image during learning.

[0064] <Modification Example of Embodiment 1> In Embodiment 1, the image improvement processing network (network 401 in FIG. 5) only generated an improved image. However, the present invention is not limited to this. In this modification example, the image improvement processing network performs image restoration that faithfully reproduces the pre-degraded image at the same time. An example of the image improvement processing network in this case is shown in FIG. 10.

[0065] FIG. 10 is a diagram showing a neural network included in the image improvement processing unit 1002 according to a modification example of Embodiment 1 of the present invention. The output terminal 407 of this network generates a restored image that faithfully reproduces the pre-degraded image, and the output terminal 408 generates an improved image that improves the accuracy of the external module 1004. At that time, it is assumed that the convolutional layer 409 in the former is learned only with a loss related to image reconstruction (L1 / L2 loss or TV loss). Also, the latter convolutional layer 410 can dynamically use the parameters output by the policy generation unit, similar to the parameters 404, 405, and 406 in FIG. 5.

[0066] <Embodiment 2> In Embodiment 1, one external module was used. However, the effectiveness of the present invention is not limited to this form. It is also possible to connect a plurality of external modules and perform adaptive image improvement for each module. In Embodiment 2, as an example, a case will be described in which a noisy image is input, and a face detection module and a face authentication module are connected to perform person identification.

[0067] Regarding the processing of Embodiment 2, it will be described with reference to the functional configuration diagram of FIG. 11 and the flowchart of FIG. 12. The image processing apparatus 2000 according to Embodiment 2 is an example of an information processing apparatus. Since the hardware configuration of the image processing apparatus 2000 is the same as that of the image processing apparatus 1000 in FIG. 2, detailed description thereof will be omitted. When referring to the hardware configuration of the image processing apparatus 2000, the reference numerals in FIG. 2 are used.

[0068] FIG. 11 is a diagram showing the functional configuration at runtime of the image processing apparatus according to Embodiment 2 of the present invention. The image processing apparatus 2000 includes an image data acquisition unit 2001, a first image improvement processing unit 2002, a first policy generation unit 2003, and a first external module 2004. The image processing apparatus 2000 also includes a second image improvement processing unit 2005, a second policy generation unit 2006, and a second external module 2007. Details of these functions will be described with reference to FIG. 12 and the like.

[0069] FIG. 12 is a flowchart showing the processing at runtime of the image processing apparatus according to Embodiment 2 of the present invention. Using this flowchart, the processing at runtime of Embodiment 2 will be described centering on the differences from Embodiment 1. This processing is realized by a program being executed by the CPU 1201 of the image processing apparatus 2000.

[0070] The image data acquisition step in step S2001 is the same processing as the image data acquisition step in Embodiment 1 (step S1001 in FIG. 4), so the description thereof will be omitted.

[0071] The processing of each of the first image improvement processing step (step S2002), the first policy generation step (step S2003), and the first recognition step (step S2004) is processing related to face detection. The first image improvement processing step is executed by the first image improvement processing unit 2002. The first policy generation step is executed by the first policy generation unit 2003. The first recognition step is executed by the first external module 2004.

[0072] In step S2002, the first image improvement processing unit 2002 performs image improvement processing to improve the accuracy of the face detection module. Also, in the first-time processing, the parameters of the noise reduction processing are used in the same manner as in Embodiment 1. In step S2003, the first policy generation unit 2003 generates parameters for the image improvement processing to improve the accuracy of the face detection module. In branch B2001, the CPU 1201 determines the end of loop L1001. In step S2004, the first external module 2004 performs face detection processing on the people in the image. The output of the face detection is the coordinate values of the rectangle indicating the position of the face and the confidence level. For details of other processing, it is the same as each processing in Embodiment 1.

[0073] Each processing of the second image improvement processing step (step S2005), the second policy generation step (step S2006), and the second recognition step (step S2007) is processing related to face authentication. The second image improvement processing step is executed by the second image improvement processing unit 2005. The second policy generation step is executed by the second policy generation unit 2006. The second recognition step is executed by the second external module 2007.

[0074] In step S2005, the second image improvement processing unit 2005 performs image improvement processing to improve the accuracy of the face recognition module. Also, in the first processing, parameters for noise reduction processing are used in the same manner as in the first image improvement processing step. In step S2006, the second policy generation unit 2006 generates parameters for image improvement processing to improve the accuracy of the face recognition module. In branch B2002, the CPU 1201 determines the end of loop L1002. In step S2007, the second external module 2007 performs face recognition processing on the person in the image. In face recognition, for the input face image, feature amounts of several hundreds of dimensions are extracted, and the similarity is calculated between this feature amount and the feature amount associated with the person ID registered in the database. In face recognition, person identification is performed based on this similarity. Therefore, in this embodiment, the face image is cropped from the coordinate values of the face rectangle which is the output of the first recognition step (step S2004) and the improved image which is the output of the second image improvement processing step (step S2005). In this embodiment, it is assumed that this cropped face image is input to the second external module 2007 which is the face recognition module. If there is no face rectangle, the process can be skipped. The above is the processing at runtime.

[0075] Regarding the processing during learning in Embodiment 2, face detection may be performed in the same manner as in Embodiment 1. Regarding face recognition in Embodiment 2, the evaluation value in the reward calculation step (step S4007 in FIG. 9) is different from that in Embodiment 1. In this embodiment, the false acceptance rate or the false rejection rate of the genuine user may be used. However, since a higher numerical value of the false acceptance rate or the false rejection rate of the genuine user represents worse performance, appropriate conversion (for example, the power of 1 - false acceptance rate, etc.) may be performed so that a higher numerical value represents better performance, and this may be used as the reward. Also, the false acceptance rate or the false rejection rate of the genuine user may be selected according to the purpose.

[0076] By executing as described above, it becomes possible to connect a plurality of external modules and perform adaptive image improvement for each module.

[0077] <Modification Example of Embodiment 2> In Embodiment 2, the face recognition module input a face image, converted it into a multi-dimensional feature amount, and calculated the similarity with the feature amount registered in the database. In a modification of Embodiment 2, the image improvement processing unit performs the conversion itself into this feature amount. That is, the neural network of the second image improvement processing unit 2005 has a structure that extracts feature amounts. Fig. 13 shows an example of this network structure.

[0078] Fig. 13 is a diagram showing the neural network of the second image improvement processing unit 2005 according to a modification of Embodiment 2 of the present invention. This network is pre-trained by noise reduction processing, and it is assumed that the output image 411 is a noise-reduced image. In the feature amount extraction in the second image improvement processing unit 2005, hierarchical intermediate features 412 are obtained and one-dimensionalized via GAP (data 413), and feature amounts 415 are extracted via the fully connected layer 414. GAP is an abbreviation for Global Average Pooling. In this modification, the second policy generation unit 2006 generates the parameters of the fully connected layer 414.

[0079] By implementing as described above, in a process such as collation using a feature amount as an input, an image improvement processing unit that improves the collation accuracy can be realized.

[0080] <Embodiment 3> In Embodiments 1 and 2, still images were assumed as input images, and cases were described in which the parameters of the image improvement processing were dynamically estimated multiple times with the same input image, and an image improvement processing was performed to better improve the result of the final recognition processing. Since the present invention is also applicable to videos, in this embodiment, a case in which a video is input and human detection is performed will be described.

[0081] Since the functional configuration in this embodiment is the same as that of the image processing apparatus 1000 in FIG. 3 of Embodiment 1, reference will be made to FIG. 3 and the description will be omitted. FIG. 14 is a flowchart showing the processing at the runtime of the image processing apparatus according to Embodiment 3 of the present invention. The difference from Embodiment 1 is that, within the loop L4001 related to time, the form is to execute from the image data acquisition step (step S4001) to the recognition step (step S4005). That is, Embodiment 3 is different from Embodiment 1 in that within the loop L4001, image data is acquired at each time and processed for different images.

[0082] Since the parameters of the image improvement process generated in the policy generation step (step S4004) are used in the next image improvement process, according to the mode of this embodiment, the image improvement process for the image at the next time is performed. For this reason, the parameters generated in the policy generation step (step S4004) are image improvement process parameters that improve the recognition accuracy for the image at the next time. In this embodiment, regarding the details of the processing of each of the other steps (steps S4001, S4002, S4003, S4004, and S4005), it is basically the same as that in Embodiment 1. The above is the processing at the runtime.

[0083] Subsequently, the processing during learning will be described. Since the functional configuration in the processing during learning is the same as that of the learning apparatus 3000 in FIG. 8 of Embodiment 1, reference will be made to FIG. 8 and the description will be omitted.

[0084] FIG. 15 is a flowchart showing the processing of the learning device according to Embodiment 3 of the present invention. The difference from Embodiment 1 is that the learning data acquisition step (Step S5002) and the degradation step (Step S5003) are inside the loop L5002 of the episode. That is, in Embodiment 3, as in the runtime, different images are acquired at each time, and adaptive image improvement processing is performed on the images. The parameters generated in the policy generation step (Step S5005) are used in the image improvement processing step (Step S5004) at the next time, and reinforcement learning is performed. Therefore, the policy generation unit 3005 of Embodiment 3 is learned to output the parameters of the image improvement processing for the image at the next time. From such a state, when the present invention is applied to a video, the policy generation unit 3005 is required to have a function of predicting the next period based on the current state. Therefore, a layer of a recurrent neural network such as LSTM may be introduced into the fully connected layer (the fully connected layer 510 in FIG. 6, etc.) of the policy generation network. LSTM is an abbreviation for Long-short-term-memory.

[0085] In this embodiment, regarding the details of the processing of each of the other steps (Steps S5001, S5002, S5003, S5004, S5005, S5006, S5007, S5008, and S5009), basically, it is the same as in Embodiment 1.

[0086] By executing as described above, it becomes possible to perform image improvement processing that takes a video as an input and adaptively improves the accuracy of the external module for each frame of the video.

[0087] (Other Embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0088] As described above, the preferred embodiments of the present invention have been explained. However, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.

[0089] The disclosure of this embodiment includes the following configurations. (Configuration 1) An image processing apparatus that improves an input image and performs recognition processing, image data acquisition means for acquiring the input image, an external module that performs recognition processing, image improvement processing means for generating input data to be input to the external module from the input image, policy generation means for generating parameters used by the image improvement processing means, and the policy generation means is optimized by reinforcement learning using the accuracy of the external module An image processing apparatus characterized by this. (Configuration 2) The input data is an image The image processing apparatus according to Configuration 1, characterized by this. (Configuration 3) The input image may be an image deteriorated due to some deterioration factor The image processing apparatus according to Configuration 1 or Configuration 2, characterized by this. (Configuration 4) The image improvement processing means has one or more outputs, at least one of which generates the input data, and when there are two or more outputs, at least one of them is an image obtained by restoring the image before deterioration when the input image is a deteriorated image The image processing apparatus according to any one of Configurations 1 to 3, characterized by this. (Configuration 5) The image improvement processing means is a neural network whose parameters can be dynamically changed The image processing apparatus according to any one of Configurations 1 to 4, characterized by this. (Configuration 6) The policy generation means receives the input data as an input, calculates information for determining parameters of a predetermined layer of the neural network of the image improvement processing means, and determines parameters based on the information. The image processing apparatus according to any one of Configurations 1 to 5, characterized in that. (Configuration 7) The image improvement processing means and the policy generation means are processed one or more times with respect to the input image. The parameters output by the policy generation means are used in the next processing of the image improvement processing means. The image processing apparatus according to any one of Configurations 1 to 6, characterized in that. (Configuration 8) The image improvement processing means and the policy generation means are processed one or more times with respect to the input image. The policy generation means calculates an update amount for the current parameters of a predetermined layer of the neural network of the image improvement processing means, and determines the next parameters by adding the update amount to the current parameters. The image processing apparatus according to any one of Configurations 1 to 7, characterized in that. (Configuration 9) The image improvement processing means and the policy generation means are processed one or more times with respect to the input image. The policy generation means is characterized by receiving the input data and further the current parameters of a predetermined layer of the neural network of the image improvement processing means as inputs, and outputting the next parameters. The image processing apparatus according to any one of Configurations 1 to 8, characterized in that. (Configuration 10) The image improvement processing means and the policy generation means are processed one or more times with respect to the input image. The policy generation means outputs the next parameters themselves of a predetermined layer of the neural network of the image improvement processing means. The image processing apparatus according to any one of Configurations 1 to 9, characterized in that. (Configuration 11) The policy generation means outputs weights to be multiplied by the parameters of a predetermined layer of the neural network of the image improvement processing means. The image processing apparatus according to any one of Configurations 1 to 10, characterized in that. (Configuration 12) The policy generation means outputs a bias of a predetermined layer of the neural network of the image improvement processing means. The image processing apparatus according to any one of Configurations 1 to 11, characterized in that. (Configuration 13) The image improvement processing means is a convolutional neural network, The policy generation means calculates information for determining the kernel of a predetermined convolutional layer of the convolutional neural network and determines the kernel based on the information. The image processing apparatus according to any one of Configurations 1 to 12, characterized in that. (Configuration 14) The policy generation means has a neural network, The external module is an independent module that cannot perform end-to-end error backpropagation learning with the neural network of the policy generation means. The image processing apparatus according to any one of Configurations 1 to 13, characterized in that. (Configuration 15) The input image is a frame of a video, The image data acquisition means sequentially acquires frames at a plurality of times, The image improvement processing means receives a frame at a certain time as an input and generates the input data at that time, The policy generation means generates parameters for the next time used in the image improvement processing means. The image processing apparatus according to any one of Configurations 1 to 14, characterized in that. (Configuration 16) The input data is a feature amount. The image processing apparatus according to any one of Configurations 1 to 15, characterized in that. (Configuration 17) The deterioration factors include noise generated during the process of imaging an image with an image sensor, degradation of resolution, blurring, and deterioration caused by imaging haze in the atmosphere. The image processing apparatus according to Configuration 3, characterized in that (Configuration 18) A learning apparatus that performs learning of a neural network included in the image processing apparatus according to Configuration 1, a learning data acquisition means for acquiring a pre-deterioration image and a correct value of recognition processing performed by a learning external module associated therewith; a deterioration means for deteriorating the pre-deterioration image used in learning to create a deteriorated image; a learning image improvement processing means for generating learning input data to be input to the external module from the deteriorated image; a learning policy generation means for generating parameters used in the learning image improvement processing means; the learning external module that receives the learning input data and performs recognition processing; a reward calculation means for calculating a reward from the correct value and the output of the learning external module; a loss calculation means for calculating a loss based on reinforcement learning from the reward and information calculated by the learning policy generation means; an update means for updating the parameters of the learning policy generation means from the loss; having A learning apparatus characterized by the above. (Method 1) An image processing method for improving an input image and performing recognition processing, an image data acquisition step of acquiring the input image; a recognition step of performing recognition processing; an image improvement processing step of generating input data to be input to the recognition step from the input image; a policy generation step of generating parameters used in the image improvement processing step; having The policy generation step is optimized by reinforcement learning using the accuracy of the recognition step. An image processing method characterized by the following. (Program 1) A program that improves an input image and performs recognition processing, causing a computer to image data acquisition means for acquiring the input image, authentication means for performing recognition processing, image improvement processing means for generating input data to be input to the external module from the input image, and policy generation means for generating parameters used by the image improvement processing means, functioning as wherein the policy generation means is optimized by reinforcement learning using the accuracy of the external module A program characterized by the above.

Explanation of Signs

[0090] 1000 Image processing apparatus 1001 Image data acquisition unit 1002 Image improvement acquisition unit 1003 Policy generation unit 1004 External module

Claims

1. An image processing apparatus that improves an input image and performs recognition processing, comprising: image data acquisition means for acquiring the input image; an external module that performs recognition processing; image improvement processing means for generating input data to be input to the external module from the input image; policy generation means for generating parameters used by the image improvement processing means; wherein the policy generation means is optimized by reinforcement learning using the accuracy of the external module An image processing apparatus characterized by the above.

2. The input data is an image The image processing apparatus according to claim 1, characterized in that.

3. The input image may be an image deteriorated by some deterioration factor The image processing apparatus according to claim 1, characterized in that.

4. The image improvement processing means has one or more outputs, at least one of which generates the input data, and when there are two or more outputs, at least one is an image obtained by restoring the image before deterioration when the input image is a deteriorated image The image processing apparatus according to claim 1, characterized in that.

5. The image improvement processing means is a neural network whose parameters can be dynamically changed The image processing apparatus according to claim 1, characterized in that.

6. The policy generation means receives the input data as an input, calculates information for determining parameters of a predetermined layer of the neural network of the image improvement processing means, and determines parameters based on the information The image processing apparatus according to claim 1, characterized in that.

7. The image improvement processing means and the policy generation means are processed one or more times for the input image, The parameters output by the policy generation means are used in the next processing of the image improvement processing means The image processing apparatus according to claim 1, characterized in that.

8. The image improvement processing means and the policy generation means are processed one or more times for the input image, The policy generation means calculates an update amount for the current parameters of a predetermined layer of the neural network of the image improvement processing means, and determines the next parameters by adding the update amount to the current parameters The image processing apparatus according to claim 1, characterized in that.

9. The image improvement processing means and the policy generation means are processed one or more times for the input image, The policy generation means receives, as inputs, the input data and further the current parameters of a predetermined layer of the neural network of the image improvement processing means, and outputs the next parameters. The image processing apparatus according to claim 1, characterized in that.

10. The image improvement processing means and the policy generation means are processed one or more times for the input image. The policy generation means outputs the next parameters themselves of a predetermined layer of the neural network of the image improvement processing means. The image processing apparatus according to claim 1, characterized in that.

11. The policy generation means outputs weights to be multiplied by the parameters of a predetermined layer of the neural network of the image improvement processing means. The image processing apparatus according to claim 1, characterized in that.

12. The policy generation means outputs the bias of a predetermined layer of the neural network of the image improvement processing means. The image processing apparatus according to claim 1, characterized in that.

13. The image improvement processing means is a convolutional neural network. The policy generation means calculates information for determining the kernel of a predetermined convolutional layer of the convolutional neural network, and determines the kernel based on the information. The image processing apparatus according to claim 1, characterized in that.

14. The policy generation means has a neural network. The external module is an independent module that cannot perform end-to-end error backpropagation learning with the neural network of the policy generation means. The image processing apparatus according to claim 1, characterized in that.

15. The input image is a frame of a video. The image data acquisition means sequentially acquires frames at a plurality of times. The image improvement processing means receives a frame at a certain time as an input, and generates the input data at that time. The policy generation means generates the parameters for the next time used in the image improvement processing means. The image processing apparatus according to claim 1, characterized in that

16. The input data is a feature amount. The image processing apparatus according to claim 1, characterized in that.

17. The degradation factor is noise generated in the process of imaging an image with an image sensor, degradation of resolution, blur, or degradation caused by imaging haze in the atmosphere. The image processing apparatus according to claim 3, characterized in that

18. A learning device for training a neural network included in the image processing apparatus according to claim 1, a learning data acquisition means for acquiring an image before degradation and a correct value of recognition processing performed by a learning external module associated therewith; a degradation means for degrading the image before degradation used in learning to create a degraded image; a learning image improvement processing means for generating learning input data to be input to the external module from the degraded image; a learning policy generation means for generating parameters used in the learning image improvement processing means; the learning external module that receives the learning input data and performs recognition processing; a reward calculation means for calculating a reward from the correct value and the output of the learning external module; a loss calculation means for calculating a loss based on reinforcement learning from the reward and information calculated by the learning policy generation means; an update means for updating parameters of the learning policy generation means from the loss; and characterized in that it has a learning device.

19. An image processing method for improving an input image and performing recognition processing, an image data acquisition step of acquiring the input image; a recognition step of performing recognition processing; an image improvement processing step of generating input data to be input to the recognition step from the input image; a policy generation step of generating parameters used in the image improvement processing step; and characterized in that the policy generation step is optimized by reinforcement learning using the accuracy of the recognition step. An image processing method characterized by the above.

20. A program for improving an input image and performing recognition processing, causing a computer to function as an image data acquisition means for acquiring the input image, an authentication means for performing recognition processing, an image improvement processing means for generating input data to be input to the external module from the input image, and a policy generation means for generating parameters used in the image improvement processing means, and characterized in that the policy generation means is optimized by reinforcement learning using the accuracy of the external module. A program characterized by the above.