Learning method, image recognition method, learning device, and image recognition system

By generating an image recognition model for machine learning by obtaining the difference image between a blurred image and an unblurred image, the problems of privacy protection and low image recognition accuracy in home or indoor environments are solved, and the learning efficiency and recognition accuracy are improved.

CN115843371BActive Publication Date: 2025-12-05PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180048827.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-16
Filing Date
2021-06-25
Publication Date
2025-12-05
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

In environments where privacy needs to be protected, such as homes or indoors, existing technologies struggle to effectively improve image recognition accuracy and machine learning efficiency because computational camera images are difficult for humans to visually recognize, making it difficult to assign labels to correct answers and reducing learning efficiency.

Method used

By acquiring the difference image between a blurred image captured by a first camera and an unblurred or less blurred image captured by a second camera, a blurred fourth image is generated for machine learning, and combined with the correct answer label, an image recognition model is created.

Benefits of technology

This approach achieves improved image recognition accuracy and machine learning efficiency while protecting privacy, reducing users' psychological burden and enhancing privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115843371B_ABST
    Figure CN115843371B_ABST
Patent Text Reader

Abstract

A learning device (20) acquires computational photography information about a computational photography camera (101) that captures an image with blur, acquires a normal image captured by a normal camera that captures an image without blur or with little blur and a correct answer label given to the normal image, generates an image with blur based on the computational photography information and the normal image, and creates an image recognition model for recognizing an image captured by the computational photography camera (101) by performing machine learning using the image with blur and the correct answer label.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an image recognition method and an image recognition system in an environment where privacy needs to be protected, such as in a home or indoors, and a learning method and a learning device for creating an image recognition model used in the image recognition. BACKGROUND

[0002] An image recognition system is disclosed in Patent Literature 1 below, which recognizes an object included in a computed photograph image captured by a light field camera or the like by inputting the computed photograph image to a recognizer that recognizes using a learned recognition model.

[0003] A computed photograph image is an image in which blurring is intentionally created so as to make it difficult for a person to visually recognize by causing a plurality of images having different viewpoints to overlap with each other or by not using a lens to make it difficult for a subject image to be in focus, and the like. Therefore, in order to construct an image recognition system in an environment where privacy needs to be protected, such as in a home or indoors, it is preferable to use a computed photograph image.

[0004] On the other hand, because a person has difficulty in visually recognizing a computed photograph image, in machine learning for creating a recognition model, it is difficult to give a correct answer label to a computed photograph image captured by a light field camera or the like. As a result, the learning efficiency is reduced.

[0005] Patent Literature 1 below does not take any countermeasures against this problem, and therefore it is desirable to improve the learning efficiency by implementing an effective countermeasure.

[0006] PRIOR ART DOCUMENTS

[0007] PATENT LITERATURE

[0008] Patent Literature 1: International Application Publication No. 2019 / 054092 SUMMARY

[0009] An object of the present application is to provide a technology that can both protect the privacy of a subject and improve the image recognition accuracy and the learning efficiency of machine learning in an image recognition system.

[0010] One embodiment of the present application relates to a learning method in which an information processing apparatus as a learning device acquires calculation imaging information related to a first camera that captures an image with blur, the calculation imaging information being a difference image between a first image captured by the first camera and a second image, the first image including a point light source in a light-on state, the second image including the point light source in a light-off state, acquires a third image captured by a second camera and a correct answer label given to the third image, the second camera capturing an image without blur or an image with less blur than the image captured by the first camera, generates a fourth image with blur based on the calculation imaging information and the third image, and creates an image recognition model for recognizing an image captured by the first camera by performing machine learning using the fourth image and the correct answer label. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a block diagram showing the configuration of an image recognition system according to the first embodiment.

[0012] Figure 2 is a flowchart showing the flow of main processing of the image recognition system.

[0013] Figure 3 is a diagram schematically showing the configuration of a multi-pinhole camera as one example of a calculation imaging camera, which is configured in a lensless manner.

[0014] Figure 4A is a diagram showing the positional relationship of the plurality of pinholes in the multi-pinhole camera.

[0015] Figure 4B is a diagram showing one example of a captured image captured by the multi-pinhole camera.

[0016] Figure 4C is a diagram showing one example of a captured image captured by the multi-pinhole camera.

[0017] Figure 5 is a flowchart showing the flow of main processing of the learning device.

[0018] Figure 6 is a block diagram showing the configuration of an image recognition system according to the second embodiment.

[0019] Figure 7 is a flowchart showing the flow of main processing of the image recognition system.

[0020] Figure 8A is a diagram for explaining the generation processing of a difference image.

[0021] Figure 8B is a diagram for explaining a generation process of a difference image.

[0022] Figure 8C is a diagram for explaining a generation process of a difference image.

[0023] Figure 9 is a flowchart showing a flow of main processing of a calculation imaging information acquisition section in a case where the LTM is used as calculation imaging information.

[0024] Figure 10 is a pattern diagram showing a plurality of PSFs.

[0025] Figure 11 is a pattern diagram showing a configuration of an image recognition system according to the third embodiment.

[0026] Figure 12 is a flowchart showing a flow of main processing of the image recognition system.

[0027] Figure 13 is a flowchart showing a flow of main processing of the image recognition system.

[0028] Figure 14 is a flowchart showing a flow of main processing of the image recognition system.

[0029] Figure 15 is a pattern diagram showing a configuration of an image recognition system according to the fourth embodiment.

[0030] Figure 16 is a flowchart showing a flow of main processing of the learning device.

[0031] Figure 17A is a pattern diagram showing a configuration of a multi-pinhole camera according to the modification.

[0032] Figure 17B is a pattern diagram showing a configuration of a multi-pinhole camera according to the modification.

[0033] Figure 17C is a pattern diagram showing a configuration of a multi-pinhole camera according to the modification.

[0034] Figure 17D is a pattern diagram showing a configuration of a multi-pinhole camera according to the modification.

[0035] Figure 18A is a pattern diagram showing a configuration of a multi-pinhole camera according to the modification.

[0036] Figure 18B is a pattern diagram showing a configuration of a multi-pinhole camera according to the modification.

[0037] Figure 18Cis a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0038] Figure 18D is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0039] Figure 19 is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0040] Figure 20 is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0041] Figure 21 is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0042] Figure 22A is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0043] Figure 22B is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0044] Figure 22C is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0045] Figure 22D is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0046] Figure 22E is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0047] Figure 22F is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0048] Figure 23A is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0049] Figure 23B is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains.

[0050] Figure 23C is a mode diagram showing the configuration of a multi-pinhole camera to which the modification pertains. DETAILED DESCRIPTION

[0051] (Basic knowledge of the invention)

[0052] Various recognition technologies such as action recognition of a person within an environment such as a home or a room or person recognition of an operator of a device are becoming increasingly important. In recent years, in order to recognize an object, a technique called deep learning is attracting attention. Deep learning refers to machine learning using a neural network of a multi-layer configuration, and by using a large amount of learning data, higher-accuracy recognition performance can be achieved compared to existing methods. In such object recognition, image information is particularly effective. Various methods have been proposed that can greatly improve existing object recognition capabilities by using a camera at an input device and performing deep learning using image information as input.

[0053] However, there is a problem that, when a camera is arranged within a home or the like, the photographed image can be leaked to the outside due to hacking or the like, which can infringe on privacy. Therefore, a countermeasure is needed that can protect the privacy of a subject even in a case where the photographed image is leaked to the outside.

[0054] A computational photograph photographed by a light field camera or the like is an image in which blurring is intentionally created so as to make it difficult for a person to visually recognize by overlapping a plurality of images different in viewpoint with each other or not using a lens to make it difficult to focus on a subject image or the like. Therefore, in order to construct an image recognition system in an environment in which privacy needs to be protected, particularly within a home or a room or the like, it is desirable to use a computational photograph.

[0055] In the image recognition system disclosed in the above-described Patent Document 1, an object region is photographed by a light field camera or the like, and a computational photograph acquired by the photographing is input to a recognizer. Thereby, the recognizer recognizes an object included in the computational photograph using a learned recognition model. In this way, by photographing an object region using a light field camera or the like that photographs a computational photograph, even in a case where the photographed image is leaked to the outside, since it is difficult for a person to visually recognize the computational photograph, the privacy of a subject can be protected.

[0056] In the image recognition system disclosed in the above-described Patent Document 1, a recognition model used by a recognizer is created by machine learning using a computational photograph photographed by a light field camera or the like as learning data. However, since it is difficult for a person to visually recognize a computational photograph, in machine learning for creating a recognition model, it is difficult to assign a correct correct answer label to a computational photograph photographed by a light field camera or the like. If an incorrect correct answer label is assigned to a computational photograph for learning, the learning efficiency of machine learning is reduced.

[0057] To solve this problem, the inventors of the present application have conceived the following scheme, that is, in the stage of accumulating learning data, instead of using an image having blur (hereinafter referred to as "blurred image") such as a calculation photographed image, an image having no blur (hereinafter referred to as "normal image") is used, and in the subsequent learning stage, machine learning is performed using a blurred image obtained by transforming a normal image based on calculation photographing information of the camera used. Thus, the present application has been conceived which can protect the privacy of the subject while improving the image recognition accuracy and the learning efficiency of machine learning.

[0058] Moreover, as another viewpoint of protecting privacy, how to reduce the psychological burden of the user photographed by the image recognition device is also important. By photographing a blurred image, it is possible to publicize that the privacy of the subject can be protected. However, in the case where the calculation photographing information is set in a field (a manufacturer's factory, etc.) unrelated to the user, the psychological burden of the user can increase because it can be suspected that the manufacturer can restore the blurred image to a normal image. On the other hand, considering that if the photographed user himself / herself can change the calculation photographing information, it is possible to reduce the psychological burden of the user, the present application has been conceived.

[0059] Next, each embodiment of the present application will be described.

[0060] One embodiment of the present application relates to a learning method in which an information processing device as a learning device acquires calculation photographing information related to a first camera that photographs an image having blur, the calculation photographing information being a difference image between a first image containing a point light source in a light-on state and a second image containing the point light source in a light-off state, acquires a third image photographed by a second camera that photographs an image having no blur or an image having less blur than the image photographed by the first camera, and a correct answer label given to the third image, generates a fourth image having blur based on the calculation photographing information and the third image, and creates an image recognition model for recognizing an image photographed by the first camera by performing machine learning using the fourth image and the correct answer label.

[0061] In the present application, the so-called "blur" means a state in which a person has difficulty in visually recognizing or simply a subject is in a state of being out of focus due to an influence of superimposing a plurality of images different in viewpoint on each other by photographing with an optical field camera or a lensless camera or the like or making a subject image difficult to focus by not using a lens. The so-called "image having blur" means an image in which a person has difficulty in visually recognizing or a subject is out of focus. The so-called "large blur" means a large degree of difficulty in visual recognition by a person or a large degree of being out of focus of a subject, and the so-called "small blur" means a small degree of difficulty or a small degree of being out of focus. The so-called "image having no blur" means an image in which a person easily visually recognizes or a subject is in focus.

[0062] According to this configuration, the subject region where the subject to be an image recognition target is photographed by the first camera that photographs an image having blur. Therefore, even in a case where the photographed image photographed by the first camera is leaked to the outside, since the image is difficult for a person to visually recognize, it is possible to protect the privacy of the subject. Moreover, the learning data, that is, the third image is photographed by the second camera that photographs an image having no blur or small blur. Since the image is easy for a person to visually recognize, it is possible to easily give a correct answer label to the third image. Further, the calculation imaging information related to the first camera is a difference image between a first image including a point light source in a light-on state and a second image including a point light source in a light-off state. Therefore, it is possible to correctly acquire the calculation imaging information related to the first camera actually used without being affected by a subject other than the point light source. As a result, it is possible to correctly generate the fourth image used at the time of machine learning based on the calculation imaging information and the third image. As a result, it is possible to protect the privacy of the subject and to improve the image recognition accuracy and the learning efficiency of machine learning.

[0063] In the above-described mode, the first camera can be one of an encoded aperture camera, a multi-pinhole camera, and an optical field camera. The encoded aperture camera includes a mask having a mask pattern with different transmittances for each region. In the multi-pinhole camera, a mask having a mask pattern with a plurality of pinholes is disposed on a light-receiving surface of an image sensor. The optical field camera acquires an optical field from a subject.

[0064] According to this configuration, by using one of an encoded aperture camera, a multi-pinhole camera, and an optical field camera as the first camera, it is possible to appropriately photograph an image having blur that is difficult for a person to visually recognize.

[0065] In the above-described mode, the first camera can not have an optical system that images light from a subject on an image sensor.

[0066] According to this configuration, since the first camera does not have an optical system that images light from the subject on the image sensor, blur can be intentionally made in the captured image captured by the first camera. As a result, it becomes more difficult to recognize the subject included in the captured image, and the effect of protecting the privacy of the subject can be further improved.

[0067] In the above-described mode, the mask can also be changed to another mask different from the mask pattern.

[0068] According to this configuration, since the calculation imaging information of the first camera can be changed by changing the mask, the calculation imaging information can be made different for each user, for example, by arbitrarily changing the mask by each user. As a result, it becomes more difficult for a third party to inversely transform the fourth image into the third image, and the effect of protecting the privacy of the subject can be further improved.

[0069] In the above-described mode, the calculation imaging information can also be one of a point spread function (PSF) and a light transmission matrix (LTM).

[0070] According to this configuration, by using one of the PSF and the LTM, the calculation imaging information about the first camera can be obtained easily and correctly.

[0071] In the above-described mode, the information processing apparatus can perform the lighting control of the point light source and the capturing control of capturing the first image by the first camera, and perform the extinguishing control of the point light source and the capturing control of capturing the second image by the first camera.

[0072] According to this configuration, by the information processing apparatus controlling the actions of the point light source and the first camera, the timing of lighting and extinguishing of the point light source can be synchronized with the timing of capturing by the first camera.

[0073] In the above-described mode, the information processing apparatus can perform the recapturing control of recapturing the first image and the second image by the first camera again in a case where the quality of the difference image is less than an allowable value.

[0074] According to this configuration, in a case where the quality of the difference image is less than an allowable value, by the information processing apparatus performing control to make the first camera recapture, a difference image in which the luminance value of the point light source is adjusted appropriately can be obtained. As a result, appropriate calculation imaging information about the first camera can be obtained.

[0075] In the above-described aspect, the information processing apparatus can correct at least one of an exposure time and a gain of the first camera in the rephotographing control so that a maximum luminance value is within a predetermined range for each of the first image and the second image.

[0076] According to this configuration, by correcting at least one of the exposure time and the gain of the first camera, a difference image in which the luminance value of the point light source is appropriately adjusted can be acquired through the rephotographing control.

[0077] Another aspect of the present technology relates to an image recognition method in an identification apparatus having an identification unit, in which an image captured by a first camera that captures an image having blur is input to the identification unit, the identification unit identifies the input image based on a learned image recognition model, and outputs an identification result of the identification unit, the image recognition model being created according to the above-described learning method.

[0078] According to this configuration, an object region in which a subject to be an image recognition target is captured by a first camera that captures an image having blur. Therefore, even in a case where the captured image captured by the first camera is leaked to the outside, since a person is difficult to visually recognize the image, it is possible to protect the privacy of the subject. Moreover, the learning data, i.e., the third image, is captured by a second camera that captures an image having no blur or small blur. Since a person is easy to visually recognize the image, it is possible to easily assign a correct correct answer label to the third image. Further, the calculation imaging information related to the first camera is a difference image between a first image including a point light source in an on state and a second image including a point light source in an off state. Therefore, it is possible to correctly acquire the calculation imaging information related to the actually used first camera without being affected by the subject other than the point light source. As a result, it is possible to correctly generate a fourth image used at the time of machine learning based on the calculation imaging information and the third image. As a result, it is possible to protect the privacy of the subject and to improve the image recognition accuracy and the learning efficiency of the machine learning.

[0079] Another aspect of the present application relates to a learning device including: an acquisition unit configured to acquire calculation imaging information related to a first camera that captures an image having blur, the calculation imaging information being a difference image between a first image captured by the first camera and a second image captured by the first camera, the first image including a point light source in a light-on state, the second image including the point light source in a light-off state; a storage unit configured to store a third image captured by a second camera that captures an image having no blur or less blur than the image captured by the first camera, and a correct answer label assigned to the third image; an image generation unit configured to generate a fourth image having blur based on the calculation imaging information acquired by the acquisition unit and the third image read from the storage unit; and a learning unit configured to create an image recognition model for recognizing an image captured by the first camera by performing machine learning using the fourth image generated by the image generation unit and the correct answer label read from the storage unit.

[0080] According to this configuration, an object region where a subject to be an image recognition target is present is captured by a first camera that captures an image having blur. Therefore, even in a case where a captured image captured by the first camera is leaked to the outside, the privacy of the subject can be protected because a person has difficulty visually recognizing the image. Also, the learning data, i.e., the third image, is captured by a second camera that captures an image having no blur or less blur. Because a person has an easy time visually recognizing the image, the correct correct answer label can be easily assigned to the third image. Further, the calculation imaging information related to the first camera is a difference image between the first image including the point light source in the light-on state and the second image including the point light source in the light-off state. Therefore, the calculation imaging information related to the first camera actually used can be correctly acquired without being affected by the subject other than the point light source. As a result, the image generation unit can correctly generate the fourth image used at the time of machine learning based on the calculation imaging information and the third image. As a result, the privacy of the subject can be protected and the image recognition accuracy and the learning efficiency of the machine learning can be improved.

[0081] Another aspect of the present application relates to an image recognition system including: an acquisition unit configured to acquire calculation imaging information related to a first camera that captures an image having blur, the calculation imaging information being a difference image between a first image captured by the first camera and a second image captured by the first camera, the first image including a point light source in a light-on state, the second image including the point light source in a light-off state; a storage unit configured to store a third image captured by a second camera and a correct answer label assigned to the third image, the second camera capturing an image having no blur or less blur than the image captured by the first camera; an image generation unit configured to generate a fourth image having blur based on the calculation imaging information acquired by the acquisition unit and the third image read from the storage unit; a learning unit configured to create an image recognition model by performing machine learning using the fourth image generated by the image generation unit and the correct answer label read from the storage unit; a recognition unit configured to recognize an image captured by the first camera based on the image recognition model created by the learning unit; and an output unit configured to output a recognition result of the recognition unit.

[0082] According to this configuration, an object region where a subject to be an image recognition target is present is captured by a first camera that captures an image having blur. Therefore, even in a case where a captured image captured by the first camera is leaked to the outside, the privacy of the subject can be protected because a person is less likely to visually recognize the image. Also, the learning data, i.e., the third image, is captured by a second camera that captures an image having no blur or less blur. Because a person is more likely to visually recognize the image, the correct correct answer label can be easily assigned to the third image. Further, the calculation imaging information related to the first camera is a difference image between a first image including a point light source in a light-on state and a second image including the point light source in a light-off state. Therefore, the calculation imaging information related to the first camera actually used can be correctly acquired without being affected by the subject other than the point light source. As a result, the image generation unit can correctly generate the fourth image used at the time of machine learning based on the calculation imaging information and the third image. As a result, the privacy of the subject can be protected and the image recognition accuracy and the learning efficiency of machine learning can be improved.

[0083] The present application can also be realized as a computer program for causing a computer to execute the features of the method, or as an apparatus or system that operates based on the computer program. Also, the computer program can be distributed through a non-transitory recording medium such as a CD-ROM, or through a communication network such as the Internet.

[0084] In addition, each of the embodiments described below is an example of the present application. The numerical values, shapes, constituent elements, steps, order of steps, and the like shown in the following embodiments are merely examples and are not intended to limit the present application. Furthermore, among the constituent elements in the following embodiments, those not recited in the independent claims representing the most general concepts are described as arbitrary constituent elements. Furthermore, the contents of all of the embodiments can be arbitrarily combined.

[0085] Embodiments of the present application will be described below with reference to the drawings. In addition, elements to which the same reference symbols are assigned in different drawings are the same or corresponding elements.

[0086] (First Embodiment)

[0087] Figure 1 is a block diagram showing the configuration of an image recognition system 10 according to the first embodiment of the present application. The image recognition system 10 includes a learning device 20 and a recognition device 30. The recognition device 30 has a computational imaging camera 101, a recognition section 106, and an output section 107. The recognition section 106 includes a processor such as a CPU and a memory such as a semiconductor memory. The output section 107 is a display device or a speaker or the like. Furthermore, the learning device 20 has a learning database 102, a computational imaging information acquisition section 103, a database correction section 104, and a learning section 105. The learning database 102 is a storage section such as an HDD, an SSD, or a semiconductor memory. The computational imaging information acquisition section 103, the database correction section 104, and the learning section 105 are processors such as CPUs.

[0088] Figure 2 is a flowchart showing the flow of the main processing of the image recognition system 10. This flowchart shows the flow of the recognition processing of an image performed by the recognition device 30. First, the computational imaging camera 101 captures an object region and inputs a computational imaging image obtained by the capturing to the recognition section 106 (step S101). Next, the recognition section 106 recognizes the computational imaging image using a learned image recognition model (step S102). This image recognition model is an image recognition model created by learning performed by the learning device 20. Next, the output section 107 outputs the result of the recognition performed by the recognition section 106. The details of the processing of each step will be described later.

[0089] The computational photography camera 101 differs from a normal camera that photographs a normal image without blur, and photographs a computational photography image with blur. The computational photography image is an image in which the subject cannot be recognized even when the photographed image is observed, but an image in which the subject can be recognized by the person or the recognition unit 106 can be generated by performing image processing on the photographed computational photography image.

[0090] Figure 3 is a diagram schematically showing the structure of a multi-pinhole camera 301 as an example of the computational photography camera 101 configured in a lensless manner. Figure 3 The multi-pinhole camera 301 shown has a multi-pinhole mask 301a and an image sensor 301b such as a CMOS. The multi-pinhole mask 301a is disposed at a distance from the light receiving surface of the image sensor 301b. The multi-pinhole mask 301a has a plurality of pinholes 301aa disposed at random or at equal intervals. The plurality of pinholes 301aa is also referred to as a multi-pinhole. The image sensor 301b acquires an image of a subject 302 through each pinhole 301aa. The image acquired through the pinhole is referred to as a pinhole image.

[0091] Since the pinhole images of the subject 302 differ depending on the positions and sizes of the pinholes 301aa, the image sensor 301b acquires an overlapping image in which a plurality of pinhole images slightly shifted and superimposed (multiple images). The positional relationship of the plurality of pinholes 301aa affects the positional relationship of the plurality of pinhole images projected on the image sensor 301b (i.e., the degree of superimposition of the multiple images), and the sizes of the pinholes 301aa affect the degree of blur of the pinhole images.

[0092] By using the multi-pinhole mask 301a, a plurality of pinhole images different in position and degree of blur can be superimposed and acquired. That is, a computational photography image with intentionally created multiple images and blur can be acquired. Therefore, the photographed image becomes an image in which the multiple images and blur are present, and an image in which the privacy of the subject 302 is protected by such blur can be acquired. Also, by changing the number, positions, and sizes of the pinholes, images different in blur manner can be acquired. That is, either the configuration in which the user can easily attach and detach the multi-pinhole mask 301a or the configuration in which a plurality of kinds of multi-pinhole masks 301a different in mask pattern are prepared in advance and the user can freely change the used multi-pinhole mask 301a can be adopted.

[0093] In addition, the change of the mask can be implemented by other various methods in addition to the replacement of the mask, for example:

[0094] • The user arbitrarily rotates the mask rotatably installed in front of the image sensor;

[0095] • The user makes an opening at an arbitrary position of a plate installed in front of the image sensor;

[0096] • The transmittance of each position within the mask is arbitrarily set by using a liquid crystal mask or the like that utilizes a spatial light modulator or the like.

[0097] • By using a mask formed of a stretchable material such as rubber, the position and size of the hole are changed by physically deforming the mask by applying an external force. The following describes these deformation examples in order.

[0098] (User arbitrarily rotates the mask)

[0099] Figure 17A to 17D is a diagram showing the structure of a multi-pinhole camera 301 in which the user arbitrarily rotates the mask. Figure 17A is a diagram showing an overview of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask, Figure 17B is a mode diagram showing a cross section thereof. The multi-pinhole camera 301 has a multi-pinhole mask 301a that can be rotated with respect to a housing 401, and a grip 402 is attached to the multi-pinhole mask 301a. The user can fix or rotate the multi-pinhole mask 301a with respect to the housing 401 by gripping and operating the grip 402. With respect to such a structure, a screw is provided to the grip 402, and the multi-pinhole mask 301a is fixed by tightening the screw and is rotated by loosening the screw. Figure 17C and Figure 17D is a mode diagram showing that the multi-pinhole mask 301a is rotated by 90 degrees when the grip 402 is rotated by 90 degrees. In this way, the multi-pinhole mask 301a can be rotated by the user operating the grip 402.

[0100] Furthermore, in the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask, as shown in Figure 17C , the multi-pinhole mask 301a can also be a pinhole arrangement that is asymmetric with respect to rotation. Thus, the user can realize various multi-pinhole patterns by rotating the mask.

[0101] Of course, the structure of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask can also be a structure that does not have the grip 402. Figure 18A , 18B is a mode diagram showing another configuration example of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask. Figure 18A is a diagram showing an overview of another configuration example of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask, Figure 18BA mode diagram showing its cross section. The multi-pinhole mask 301a is fixed to a lens barrel 411. Also, the image sensor 301b is provided to another lens barrel 412, and the lens barrel 411 and the lens barrel 412 are in a rotatable state by a screw structure. That is, the lens barrel 412 is located outside the lens barrel 411, and a male screw is engraved on the outside of the lens barrel 411 and a female screw is engraved on the inside of the lens barrel 412 at the joint thereof. Also, at the male screw of the lens barrel 411, a fixing member 413 is first installed, and then the lens barrel 412 is installed. The fixing member 413 is also engraved with a female screw as with the lens barrel 412. By adopting such a structure, when the lens barrel 411 is screwed into the lens barrel 412, the depth of screwing changes according to the position of the fixing member 413 screwed into the lens barrel 411, and the rotation angle of the multi-pinhole camera 301 can be changed.

[0102] Figure 18C 、 18D is a mode diagram showing that the depth of screwing changes according to the position of the fixing member 413 screwed into the lens barrel 411, and the rotation angle of the multi-pinhole camera 301 changes. Figure 18C is a mode diagram in the case where the fixing member 413 is screwed into the deep part of the lens barrel 411, Figure 18D is a mode diagram in the case where the fixing member 413 is screwed only to the middle of the lens barrel 411. As Figure 18C shown, in the case where the fixing member 413 is screwed into the deep part of the lens barrel 411, the lens barrel 412 can be screwed into the deep part of the lens barrel 411. On the other hand, as Figure 18D shown, in the case where the fixing member 413 is screwed only to the middle of the lens barrel 411, the lens barrel 412 can be screwed only to the middle of the lens barrel 411. Therefore, the depth of screwing changes according to the position of the fixing member 413 screwed into the lens barrel 411, and the rotation angle of the multi-pinhole mask 301a can be changed.

[0103] (User's opening of a hole in a mask)

[0104] Figure 19 is a mode diagram of the cross section of the multi-pinhole camera 301 in which the user opens a hole in an arbitrary part of the mask 301ab installed in front of the image sensor 301b. In Figure 19 , the same reference numerals are given to the same constituent elements as those of Fig. 17, and the description thereof is omitted. Initially, no pinhole exists in the mask 301ab. By letting the user open a plurality of holes in an arbitrary part of this mask 301ab using a needle or the like, a multi-pinhole mask of an arbitrary shape can be made.

[0105] (Modified example of arbitrarily setting the transmittance of each position in a mask using a spatial light modulator)

[0106] Figure 20is a mode diagram of a cross section of the multi-pinhole camera 301 configured by arbitrarily setting the transmittance of each position within the mask using the spatial light modulator 420. In Figure 20 , the same components as Figure 19 are given the same reference numerals and the description thereof is omitted. The spatial light modulator 420 is configured by liquid crystal or the like, and can change the transmittance of each pixel. This spatial light modulator 420 functions as a multi-pinhole mask. The change in transmittance can be controlled by omitting the spatial light modulator control section illustrated. Thus, by the user selecting an arbitrary pattern from among a plurality of transmittance patterns prepared in advance, various mask patterns (multi-pinhole patterns) can be realized.

[0107] (Deformation example in which the mask is deformed by applying an external force)

[0108] Figure 21 , 22A to 22F are mode diagrams of a cross section of the multi-pinhole camera 301 configured by deforming the mask by applying an external force. In Figure 21 , the same components as Figure 19 are given the same reference numerals and the description thereof is omitted. The multi-pinhole mask 301ac is configured by a plurality of masks 301al, 301a2, 301a3, each of which has a drive section (not illustrated) that independently applies an external force. Figure 22A to 22C is a mode diagram for explaining the three masks 301al, 301a2, 301a3 that configure the multi-pinhole mask 301ac. Here, each mask has a shape in which a fan shape and a circular ring are combined. Of course, this structure is only one example, and the shape is not limited to a fan shape, and the number of masks configuring is not limited to three. One or a plurality of pinholes are formed in each mask. In addition, a pinhole can not be formed in the mask. Here, two pinholes 301aal, 301aa2 are formed in the mask 301al, one pinhole 301aa3 is formed in the mask 301a2, and two pinholes 301aa4, 301aa5 are formed in the mask 301a3. By moving these three masks 301al to 301a3 by applying an external force, various multi-pinhole patterns can be produced.

[0109] Figure 22D to 22F indicates three multi-pinhole masks 301ac configured by the three masks 301al to 301a3. By moving each mask 301al to 301a3 in a different manner by omitting each drive section illustrated, a multi-pinhole mask 301ac having a different number of pinholes or a different pinhole position can be configured. Figure 22D , 22E is a mask having five pinholes, and in Figure 22F is a mask having four pinholes. Such a mask drive section can be realized using an ultrasonic motor or a linear motor that is widely used in autofocus or the like. Thus, by applying an external force, the number or position of the pinholes of the multi-pinhole mask 301ac can be changed.

[0110] Of course, the multi-pinhole mask can change the size of the pinholes in addition to the number or position of the pinholes. Figure 23A to 23C is a mode diagram illustrating the structure of a multi-pinhole mask 301ad of a multi-pinhole camera 301 configured to deform the mask by applying an external force. The multi-pinhole mask 301ad has a plurality of pinholes, is composed of a material having elasticity, and has four driving sections 421 to 424 that can independently control four corners. Of course, the number of driving sections need not necessarily be four. By driving each of the driving sections 421 to 424, the position or size of the pinholes of the multi-pinhole mask 301ad can be changed.

[0111] Figure 23B is a mode diagram illustrating a case where the driving sections 421 to 424 are driven in the same direction. In this diagram, the direction of the arrow shown by the driving section 421 to 424 indicates the driving direction of each driving section. In this case, the multi-pinhole mask 301ad moves in parallel with the driving direction of the driving section. On the other hand, Figure 23C is a mode diagram illustrating a case where the driving sections 421 to 424 are driven in a direction outward from the center portion of the multi-pinhole mask 301ad. In this case, since the multi-pinhole mask 301ad is stretched by elasticity, the size of the pinholes becomes large. Such driving sections 421 to 424 can be implemented using an ultrasonic motor or a linear motor that is widely used in autofocus and the like. Thus, by applying an external force, the position or size of the pinholes of the multi-pinhole mask 301ac can be changed.

[0112] Figure 4A is a diagram illustrating the positional relationship of the multi-pinhole camera 301, the plurality of pinholes 301aa. In this example, three pinholes 301aa are formed in a straight line. The interval between the left end pinhole 301aa and the central pinhole 301aa is set to L1, and the interval between the central pinhole 301aa and the right end pinhole 301aa is set to L2 (< L1).

[0113] Figure 4B and Figure 4C is a diagram illustrating one example of a captured image captured by the multi-pinhole camera 301. Figure 4B illustrates an example of a captured image in a case where the distance between the multi-pinhole camera 301 and the subject 302 is far and the subject image is small. Figure 4CAn example of a captured image in a case where the distance between the multi-pinhole camera 301 and the subject 302 is short and the subject image is large. By making the intervals L1, L2 different from each other, regardless of the distance between the multi-pinhole camera 301 and the subject 302, an overlapping image is captured by overlapping a plurality of images with different viewpoints, the overlapping image being an image in which a plurality of subject images coincide with each other in a manner that cannot be individually recognized.

[0114] As the computational camera 101, in addition to the multi-pinhole camera 301, a publicly known camera or the like shown below can also be used:

[0115] • an encoded aperture camera in which a mask having a mask pattern in which the transmittance differs for each region is arranged between the image sensor and the subject;

[0116] • a light field camera having a configuration in which a microlens array is arranged on the light receiving surface of the image sensor, for acquiring a light field;

[0117] • a compressive sensing camera that performs a weighted addition calculation on pixel information in the time space to capture an image.

[0118] Furthermore, it is desirable that the computational camera 101 does not have an optical system (lens, prism, mirror, or the like) for imaging light from the subject on the image sensor. By omitting the optical system, it is possible to achieve a small and lightweight camera, reduce costs, and improve design, and it is also possible to intentionally blur the captured image captured by the camera.

[0119] The recognition section 106 recognizes, using the image recognition model that is the result of learning by the learning device 20, the class information of the subject, such as a person (including the actions and expressions of the person), a car, a bicycle, or a signal, and the position information of each subject, from the image of the object region captured by the computational camera 101. In the learning for creating the image recognition model, machine learning such as deep learning using a multilayer neural network can be used.

[0120] The output section 107 outputs the result of the recognition by the recognition section 106. With regard to the output, an interface section can be provided that presents the recognition result to the user by means of an image, text, or sound, and a device control section can be provided that changes the control method in accordance with the recognition result.

[0121] The learning device 20 has a learning database 102, a computational camera information acquisition section 103, a database correction section 104, and a learning section 105. The learning device 20 performs learning for creating an image recognition model used by the recognition section 106 in correspondence with the computational camera information about the computational camera 101 actually used in the capturing of the object region.

[0122] Also, Figure 5 is a flowchart showing the flow of the main processing of the learning apparatus 20 of the image recognition system 10.

[0123] First, the calculation imaging information acquisition section 103 acquires calculation imaging information that is information indicating what kind of blurred image is imaged by the calculation imaging camera 101, the state of the blur (step S201). This can be acquired by the calculation imaging camera 101 having a transmission section and the calculation imaging information acquisition section 103 having a reception section, with a wire or wirelessly, or by the calculation imaging information acquisition section 103 having an interface and the user inputting the calculation imaging information to the calculation imaging information acquisition section 103.

[0124] As the calculation imaging information, for example, if the calculation imaging camera 101 is a multi-pinhole camera 301, a PSF (Point Spread Function) that indicates the state of the calculation imaging in two dimensions can be utilized. The PSF is a transfer function of a camera such as a multi-pinhole camera or an encoded aperture camera, and is expressed by the following relationship.

[0125] y = k * x

[0126] Here, y is a calculation imaging image having blur that is imaged by the multi-pinhole camera 301, k is the PSF, and x is a normal image that is imaged by a normal camera that does not blur the imaged scene. Also, * is a convolution operator.

[0127] Also, as the calculation imaging information, instead of the PSF, an LTM (Light Transport Matrix) that indicates calculation imaging information in four dimensions or more (two dimensions on the camera side and two dimensions or more on the object side) can be utilized. The LTM is a transfer function utilized by a light field camera.

[0128] For example, in the case where the calculation imaging camera 101 is the multi-pinhole camera 301, the PSF can be acquired by imaging a point light source with the multi-pinhole camera 301. This can be known from the fact that the PSF corresponds to the impulse response of the camera. That is, the imaged image of the point light source obtained by imaging the point light source with the multi-pinhole camera 301 itself is the PSF that is the calculation imaging information of the multi-pinhole camera 301. Here, as the imaged image of the point light source, it is preferable to use a difference image between when the point light is on and when the point light is off, for which a description will be given in the second embodiment to be described later.

[0129] Next, the database correction section 104 acquires normal images without blurring included in the learning database 102, and the learning section 105 acquires annotation information included in the learning database 102 (step S202).

[0130] Next, the database correction section 104 (image generation section) corrects the learning database 102 using the computational photography information acquired by the computational photography information acquisition section 103 (step S203). For example, in a case where the recognition section 106 recognizes the action of a person within an environment, the learning database 102 holds a plurality of normal images without blurring photographed by a normal camera and annotation information (correct answer label) given to each image, the annotation information indicating what action the person performed at which position in each image. In the case of using a normal camera, it is only necessary to give annotation information to an image photographed by the camera, however, in a case where a computational photography image is acquired by a multi-pinhole camera or a light field camera or the like, since it is not known what is photographed even by a person looking at the image, it is difficult to give annotation information. Moreover, even in a case where a learning process is performed on an image photographed by a normal camera that is significantly different from the computational photography camera 101, the recognition accuracy of the recognition section 106 does not become high. Here, by holding a database in which annotation information is given to an image photographed by a normal camera in advance as the learning database 102, and deforming only a photographed image in cooperation with the computational photography information of the computational photography camera 101, a learning dataset that matches the computational photography camera 101 is created, and the recognition accuracy is improved by performing a learning process. For this reason, the database correction section 104 calculates a corrected image y below using a PSF of the computational photography information acquired by the computational photography information acquisition section 103 with respect to a photographed image z photographed by a normal camera prepared in advance.

[0131] y = k * z

[0132] Here, k indicates a PSF of the computational photography information acquired by the computational photography information acquisition section 103, and * indicates a convolution operator.

[0133] The learning unit 105 performs a learning process using the corrected image thus calculated by the database correction unit 104 and the annotation information acquired from the learning database 102 (step S204). For example, in the case where the recognition unit 106 is constructed by a multi-layer neural network, the corrected image and the annotation information are used as training data (teacher data) to perform machine learning based on deep learning. As an algorithm for correcting the prediction error, a back propagation algorithm or the like can be employed. Thus, the learning unit 105 creates an image recognition model for the recognition unit 106 to recognize an image captured by the calculation camera 101. Since the corrected image becomes an image consistent with the calculation imaging information of the calculation camera 101, by such learning, learning suitable for the calculation camera 101 can be performed, and the recognition unit 106 can perform high-precision recognition processing.

[0134] According to the image recognition system 10 related to the present embodiment, the subject region where the subject 302 is present is imaged by the calculation camera 101 (first camera) that captures an image with blur, i.e., a calculation imaging image. Therefore, even in the case where the captured image of the calculation camera 101 is leaked to the outside, since it is difficult for a person to visually recognize the calculation imaging image, it is possible to protect the privacy of the subject 302. Moreover, the normal image (third image) accumulated in the learning database 102 is captured by a normal camera (second camera) that captures an image without blur (or an image with less blur than the calculation imaging image). Therefore, since it is easy for a person to visually recognize the image, it is possible to easily assign correct annotation information (correct answer label) to the normal image. As a result, it is possible to protect the privacy of the subject 302 and to improve the image recognition accuracy and the learning efficiency of machine learning.

[0135] Moreover, as the calculation camera 101, it is possible to appropriately capture an image with blur that is difficult for a person to visually recognize by using one of an encoded aperture camera, a multi-pinhole camera, and a light field camera.

[0136] Moreover, in the calculation camera 101, by omitting an optical system that images light from the subject 302 on the image sensor 301b, it is possible to intentionally create blur in the captured image of the calculation camera 101. As a result, since it is more difficult to recognize the subject 302 included in the captured image, it is possible to further improve the effect of protecting the privacy of the subject 302.

[0137] Moreover, in a case where the user can freely change the configuration of the multi-pin hole mask 301a used, since the calculation imaging information of the calculation imaging camera 101 also changes by changing the mask, for example, by each user arbitrarily changing the mask, it is possible to make the calculation imaging information different for each user. As a result, since it is difficult for a third party to perform the inverse transform from the correction image (fourth image) to the normal image (third image), it is possible to further improve the effect of protecting the privacy of the subject 302.

[0138] Moreover, by using one of the PSF and the LTM as the calculation imaging information, it is possible to simply and appropriately acquire the calculation imaging information on the calculation imaging camera 101.

[0139] Second Embodiment

[0140] Figure 6 is a mode diagram showing the configuration of the image recognition system 11 to which the second embodiment of the present application relates. In Figure 6 , the same reference numerals are given to the same constituent elements as Figure 1 , and the description thereof is omitted. The learning device 21 of the image recognition system 11 has the control section 108. Moreover, the image recognition system 11 has the light emitting section 109 existing in the subject region (environment) photographed by the calculation imaging camera 101. The light emitting section 109 is a light source regarded as a point light source existing in the environment, for example, is an LED mounted on an electrical device or an illumination LED. Moreover, it is also possible to function as the light emitting section 109 by only lighting and extinguishing the light point of only a part of a monitor such as an LED monitor. By controlling the light emitting section 109 and the calculation imaging camera 101 by the control section 108, the calculation imaging information acquisition section 103 acquires the calculation imaging information.

[0141] Moreover, Figure 7 is a flowchart showing the flow of the main processing of the image recognition system 11. The flowchart shows the flow of the processing of the calculation imaging information acquisition section 103 acquiring the calculation imaging information of the calculation imaging camera 101.

[0142] First, the control section 108 issues an instruction to light up the light emitting section 109 existing in the environment (step S111).

[0143] Next, the light emitting section 109 implements lighting in accordance with the instruction of the control section 108 (step S112).

[0144] Next, the control section 108 issues an instruction to cause the calculation imaging camera 101 to implement photographing (step S113). Thereby, the light emitting section 109 and the calculation imaging camera 101 can act while maintaining synchronization.

[0145] Next, the computational camera 101 performs a shooting operation according to the instructions of the control unit 108 (step S114). The captured image (first image) is input from the computational camera 101 to the computational camera information acquisition unit 103 and temporarily stored by the computational camera information acquisition unit 103.

[0146] Next, the control unit 108 sends an instruction to the light-emitting unit 109 to turn off (step S115).

[0147] Next, the light-emitting unit 109 extinguishes the light according to the instructions of the control unit 108 (step S116).

[0148] Next, the control unit 108 issues an instruction to the computational camera 101 to perform shooting (step S117).

[0149] Next, the computational camera 101 performs a shooting operation according to the instructions of the control unit 108 (step S118). The captured image (second image) is input from the computational camera 101 to the computational camera information acquisition unit 103.

[0150] Next, the camera information acquisition unit 103 calculates the difference image between the first image and the second image (step S119). By calculating the difference image between the first image of the light-emitting unit 109 when it is lit and the second image when it is turned off, the image of the light-emitting unit 109 in the lit state, i.e., the PSF, can be acquired without being affected by other subjects in the environment.

[0151] Next, the computational camera information acquisition unit 103 acquires the generated difference image as computational camera information of the computational camera 101 (step S120).

[0152] Using PSF as the information for image calculation, the camera 101 captures two images: one of the scene where the light source 109 is lit and the other of the scene where it is off. Ideally, the images captured when the light source is lit and when it is off should be captured with as little time difference as possible.

[0153] Figure 8A to Figure 8C This is an explanatory diagram used to illustrate the generation process of difference images. Figure 8A The image was captured by the computational camera 101 when the light-emitting part 109 was lit. It can be seen that the brightness value of the light-emitting part 109 is relatively high. Figure 8B The image was captured by the computational camera 101 when the light-emitting part 109 was off. It can be seen that the brightness value of the light-emitting part 109 is lower than when it is on. Figure 8C This refers to the image captured by the computational camera 101 when the light-emitting part 109 is lit. Figure 8A Subtract the image captured by the computational camera 101 when the light-emitting part 109 is turned off. Figure 8BThe difference image is obtained. Since the light emitting section 109 is shot as a point light source without being affected by the subject other than the light emitting section 109, it is possible to acquire the PSF.

[0154] Moreover, in the case of using the LTM as the calculation imaging information, it is also possible to use a plurality of light emitting sections 109 dispersedly arranged in the environment, acquire the PSFs at a plurality of positions, and use them as the LTM.

[0155] Figure 9 is a flowchart showing the flow of the main processing of the calculation imaging information acquisition section 103 in the case of using the LTM as the calculation imaging information. First, the PSF corresponding to each light emitting section 109 is acquired (step S301). This can be acquired using the difference image between when each light emitting section 109 is lit and when it is turned off as described above. By so doing, it is possible to acquire the PSF at a plurality of positions on the image. Figure 10 is a pattern diagram showing the plurality of PSFs thus acquired. In the case of this example, the PSF is acquired at six points on the image.

[0156] The calculation imaging information acquisition section 103 calculates the PSF of all the pixels of the image by performing interpolation processing on the plurality of PSFs thus acquired and uses it as the LTM (step S302). Such interpolation processing can be general image processing such as morphing. Moreover, the light emitting section 109 can also be the light of the user's smartphone or mobile phone. In this case, the user can also implement the lighting or turning off of the light emitting section 109 instead of the control section 108.

[0157] Moreover, in the case of using the LTM as the calculation imaging information, it is also possible to not arrange a plurality of light emitting sections 109, but to change the position of the light emitting section 109 by moving it using a smaller number of light emitting sections 109. For example, it is also possible to use the light of a smartphone or mobile phone as the light emitting section 109, and the user can implement the lighting and turning off while changing the position. Or, it is also possible to use an LED mounted on a mobile body such as a drone or a dust collector robot. Or, it is also possible to set the calculation imaging camera 101 on a mobile body or the like, or it is also possible to change the position of the light emitting section 109 on the calculation imaging image by changing the orientation or position by the user.

[0158] According to the image recognition system 11 related to the present embodiment, the computation photographing information related to the computation photographing camera 101 (first camera) is a difference image between the first image including the point light source in the light-on state and the second image including the point light source in the light-off state. Therefore, the computation photographing information related to the computation photographing camera 101 actually used can be correctly acquired without being affected by the subject other than the point light source. Thus, the correction image (fourth image) used in the machine learning can be correctly generated on the basis of the computation photographing information and the normal image (third image).

[0159] Further, by controlling the light emission section 109 and the operation of the computation photographing camera 101 by the control section 108 of the learning device 21, the timing of the light-on or light-off of the light emission section 109 can be correctly synchronized with the timing of the photographing by the computation photographing camera 101.

[0160] Third Embodiment

[0161] Figure 11 is a block diagram showing the configuration of the image recognition system 12 related to the third embodiment of the present application. In Figure 11 , the same reference numerals are given to the same constituent elements as those of Figure 6 The learning device 22 of the image recognition system 12 has a computation photographing information judging section 110. The computation photographing information judging section 110 judges the state of the quality of the computation photographing information acquired by the computation photographing information acquiring section 103. The learning device 22 switches the contents of the processing in accordance with the result of the judgment by the computation photographing information judging section 110.

[0162] Further, Figure 12 is a flowchart showing the flow of the main processing of the image recognition system 12. The flowchart shows the flow of the processing before and after the quality judgment processing by the computation photographing information judging section 110.

[0163] First, the computation photographing information acquiring section 103 generates a difference image between the first image when the light emission section 109 is lighted and the second image when the light emission section 109 is lighted off by the same method as that of step S119 Figure 7 ) of the above-described second embodiment (step S121).

[0164] Next, the calculation imaging information judging section 110 judges whether the quality of the difference image generated by the calculation imaging information acquisition section 103 is above the allowable value (step S122). Since the PSF needs not to image objects other than the point light source, it is possible to use the difference image between the on-time and the off-time. However, in the case where there is a change in the scene such as a person's motion amplitude being large or the brightness in the environment changing drastically between the on-time imaging and the off-time imaging, the change in the scene is also reflected in the difference image, and it is not possible to acquire the correct PSF. Here, the calculation imaging information judging section 110 counts the number of pixels having a brightness of a prescribed value or more in the difference image, and judges that the quality of the PSF is less than the allowable value in the case where the number of pixels is above a threshold value, and judges that the quality of the PSF is above the allowable value in the case where the number of pixels is less than the threshold value.

[0165] In the case where the calculation imaging information judging section 110 judges that the quality of the difference image is less than the allowable value (step S122: No), next, the control section 108 performs the following instruction: in order to perform imaging again, instructs the light emitting section 109 to emit light and to turn off, and instructs the calculation imaging camera 101 to image again (step S123). On the other hand, in the case where the calculation imaging information judging section 110 judges that the quality of the difference image is above the allowable value (step S122: Yes), next, the database correction section 104 corrects the learning database 102 using the calculation imaging information (PSF) acquired by the calculation imaging information acquisition section 103 as the difference image (step S124).

[0166] Here, as one of the causes of the quality degradation of the difference image, there can be an improper setting of the computational imaging camera 101. For example, in a case where the exposure time of the computational imaging camera 101 is too short or the gain of the signal amplification is too small, the image as a whole is darkened, and the luminance of the light emitting portion 109 is buried in noise. In contrast, in a case where the exposure time of the computational imaging camera 101 is too long or the gain of the signal amplification is too large, the luminance value of a high-luminance region within the image exceeds the upper limit value of the sensing range and is saturated, and the periphery of the light emitting portion 109 becomes a so-called whitening state. Here, the computational imaging information judging portion 110 confirms the maximum luminance value of each of the images at the time of lighting and the time of extinguishing of the light emitting portion 109, and in a case where the maximum luminance value exceeds the upper limit value or is less than the lower limit value (i.e., in a case where it is outside the prescribed range), it can be determined that the quality of the difference image is less than the allowable value. The computational imaging information judging portion 110 can determine whether the luminance of the light emitting portion 109 exceeds the sensing range and is saturated by judging whether the maximum luminance value of the image at the time of lighting of the light emitting portion 109 exceeds the upper limit value. Further, the computational imaging information judging portion 110 can determine whether the luminance of the light emitting portion 109 is buried in noise by judging whether the maximum luminance value of the image at the time of lighting of the light emitting portion 109 is less than the lower limit value. Further, in a case where it is determined that the luminance of the light emitting portion 109 is saturated or is buried in noise, the control portion 108 can also perform control to change the setting of the computational imaging camera 101 so that the maximum luminance value is within the prescribed range at the time of re-photographing.

[0167] Figure 13 is a flowchart showing the flow of the main processing of the image recognition system 12. This flowchart shows the flow of the processing before and after the quality judging processing by the computational imaging information judging portion 110.

[0168] First, the computational imaging information acquiring portion 103 acquires the first image photographed by the computational imaging camera 101 at the time of lighting of the light emitting portion 109 (step S131).

[0169] Next, the computational imaging information judging portion 110 judges whether the luminance of the first image acquired by the computational imaging information acquiring portion 103 is saturated by confirming whether the maximum luminance value of the image exceeds the upper limit value Thl (step S132).

[0170] When the maximum luminance value exceeds the upper limit value Thl (YES in step S132), that is, when the luminance of the image is in saturation, then the control section 108 instructs the calculation camera 101 to make the exposure time shorter and perform the rephotography (step S133). On the other hand, when the maximum luminance value is equal to or less than the upper limit value Thl (NO in step S132), then the calculation camera information judging section 110 judges whether the luminance of the light emitting section 109 is buried in noise by confirming whether the maximum luminance value of the first image acquired by the calculation camera information acquiring section 103 is less than the lower limit value Th2 (step S134).

[0171] When the maximum luminance value is less than the lower limit value Th2 (YES in step S134), that is, when the luminance of the light emitting section 109 is buried in noise, then the control section 108 instructs the calculation camera 101 to make the exposure time longer and perform the rephotography (step S135). On the other hand, when the maximum luminance value is equal to or more than the lower limit value Th2 (NO in step S134), then the calculation camera information judging section 110 judges that the quality of the first image acquired by the calculation camera information acquiring section 103 is high enough at the current exposure time. In this case, the control section 108 instructs the light emitting section 109 to be turned off, and also instructs the calculation camera 101 to perform the photography at the above-mentioned current exposure time. Thus, the calculation camera information acquiring section 103 acquires a second image of the light emitting section 109 when it is turned off (step S136). In addition, the control section 108 can control the exposure time of the calculation camera 101 so as to make the maximum luminance value be within the prescribed range with respect to the acquired second image as well as the above-mentioned first image.

[0172] Of course, the control section 108 can change the setting other than the exposure time of the calculation camera 101. For example, the gain can be changed.

[0173] Figure 14 This flowchart shows the flow of the main processing of the image recognition system 12. This flowchart shows the flow of the processing before and after the quality judging processing by the calculation camera information judging section 110.

[0174] In the judgment of step S132, when the maximum luminance value exceeds the upper limit value Thl, that is, when the luminance of the image is in saturation (YES in step S132), then the control section 108 instructs the calculation camera 101 to further reduce the gain and perform the rephotography (step S137).

[0175] In the judgment of step S134, when the maximum luminance value is less than the lower limit value Th2, that is, when the luminance of the light emitting section 109 is buried in noise (YES in step S134), then the control section 108 instructs the calculation camera 101 to further increase the gain and perform the rephotography (step S138).

[0176] Moreover, the control section 108 can not control the exposure time or the gain of the computational photography camera 101, but can control the luminance of the light emitting section 109. That is, in a case where it is judged by the computational photography information judging section 110 that the luminance of the light emitting section 109 is in saturation, the control section 108 controls the light emitting section 109 so as to reduce the luminance. On the contrary, in a case where it is judged by the computational photography information judging section 110 that the luminance of the light emitting section 109 is overwhelmed by noise, the control section 108 controls the light emitting section 109 so as to increase the luminance. By increasing the luminance of the light emitting section 109, the luminance difference from the noise is increased.

[0177] Moreover, the control section 108, in a case where it is judged by the computational photography information judging section 110 that the quality of the difference image is less than the allowable value, can also select another light emitting section existing in the subject region and instruct the other light emitting section to emit light and to be turned off. This is because, in a case of a light source having directivity, depending on the positional relationship between the computational photography camera 101 and the light emitting section 109, there is a case where the quality is reduced regardless of anything, and the above selection of the other light emitting section is more effective in such a case.

[0178] According to the image recognition system 12 related to the present embodiment, in a case where the quality of the difference image is less than the allowable value, the control section 108 can acquire a difference image in which the luminance value of the point light source is appropriately adjusted by controlling the computational photography camera 101 to take a photograph again. As a result, appropriate computational photography information about the computational photography camera 101 can be acquired.

[0179] Moreover, in the photographing control again, the control section 108 can acquire a difference image in which the luminance value of the point light source is appropriately adjusted by correcting at least one of the exposure time and the gain of the computational photography camera 101.

[0180] Fourth Embodiment

[0181] Figure 15 is a block diagram showing the configuration of the image recognition system 13 related to the fourth embodiment of the present application. In Figure 15 , the same components as those of the image recognition system 12 shown in Figure 1The same constituent elements are given the same reference symbols and the description thereof is omitted. The learning device 23 of the image recognition system 13 has a storage section 112 in which a plurality of image recognition models that have been learned are stored, and a model selection section 111 that selects one image recognition model from among the plurality of image recognition models. The learning device 23 of the image recognition system 13 does not cause the learning section 105 to learn the learning database 102 that has been corrected by the database correction section 104, but has the model selection section 111 that selects the best image recognition model that corresponds to the calculation imaging information of the calculation imaging camera 101 from among a plurality of image recognition models that have been learned in advance. For example, in the case where a plurality of multi-pin-hole masks 301a that differ in mask pattern are prepared in advance as described above, image recognition models that have been learned using captured images in the mounting state of each multi-pin-hole mask 301a are created in advance, and these plurality of image recognition models are stored in the storage section 112. The model selection section 111 selects one image recognition model that corresponds to the calculation imaging information of the calculation imaging camera 101 from among the plurality of image recognition models stored in the storage section 112.

[0182] Furthermore, Figure 16 is a flowchart that shows the flow of the main processing of the learning device 23 of the image recognition system 13. This flowchart shows the flow of the processing in which the model selection section 111 selects an image recognition model.

[0183] First, the calculation imaging information acquisition section 103 acquires the calculation imaging information of the calculation imaging camera 101 (step S201).

[0184] Next, the model selection section 111 selects one image recognition model that corresponds to the calculation imaging information acquired by the calculation imaging information acquisition section 103 from among the plurality of image recognition models stored in the storage section 112 (step S211). This can be achieved by preparing image recognition models that have been learned using various calculation imaging information in advance, and then selecting an image recognition model that has been learned using calculation imaging information that is closest to the calculation imaging information.

[0185] The image recognition model thus selected becomes an image recognition model that is suitable for the calculation imaging camera 101. The selected image recognition model is set as the image recognition model used by the recognition section 106. The recognition section 106 can perform a high-precision recognition process by using this image recognition model.

[0186] According to the image recognition system 13 related to the present embodiment, the learning device 23 selects one image recognition model that corresponds to the calculation imaging information of the calculation imaging camera 101 from among a plurality of image recognition models that have been learned. Thereby, because the learning device 23 does not need to perform learning again, it is possible to reduce the processing load of the learning device 23 and quickly start the use of the recognition device 30.

[0187] Industrial applicability

[0188] The learning method and the recognition method according to the present application are particularly useful for image recognition systems in environments where the privacy of the subjects needs to be protected.

Claims

1. A learning method, an information processing apparatus as a learning device acquires calculation imaging information related to a first camera that captures an image having blur, the calculation imaging information being a difference image between a first image captured by the first camera and a second image, the first image containing a point light source in an on state, the second image containing the point light source in an off state, the information processing apparatus acquires a third image captured by a second camera that captures an image having no blur or an image having less blur than the image captured by the first camera, and a correct answer label given to the third image, the information processing apparatus generates a fourth image having blur based on the calculation imaging information and the third image, the information processing apparatus creates an image recognition model for recognizing an image captured by the first camera by performing machine learning using the fourth image and the correct answer label. 2.The learning method according to claim 1, wherein the first camera is one of an encoded aperture camera, a multi-pinhole camera, and a light field camera, the encoded aperture camera has a mask having a mask pattern in which transmittance differs in each region, the multi-pinhole camera is a camera in which a mask having a mask pattern in which a plurality of pinholes are formed is disposed on a light receiving surface of an image sensor, and the light field camera acquires a light field from a subject. 3.The learning method according to claim 1 or 2, wherein the first camera has no optical system that images light from a subject on an image sensor. 4.The learning method according to claim 2, wherein the mask can be changed to another mask different from the mask pattern. 5.The learning method according to claim 1 or 2, wherein the calculation imaging information is one of a point spread function and a light transport matrix. 6.The learning method according to claim 1 or 2, wherein the information processing apparatus performs on state control of the point light source and performs imaging control of the first image captured by the first camera, and the information processing apparatus performs off state control of the point light source and performs imaging control of the second image captured by the first camera. 7.The learning method according to claim 6, wherein in a case where a quality of the difference image is less than an allowable value, the information processing apparatus performs re-imaging control of the first image and the second image captured again by the first camera. 8.The learning method according to claim 7, wherein the information processing apparatus corrects at least one of an exposure time and a gain of the first camera in the re-imaging control so that a maximum luminance value of each of the first image and the second image is within a prescribed range, respectively. 9.An image recognition method, in an recognition device having a recognition unit, an image captured by a first camera that captures an image having blur is input to the recognition unit, the recognition unit recognizes the input image based on a learned image recognition model, a recognition result of the recognition unit is output. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The image recognition model is an image recognition model created by the learning method according to any one of claims 1 to 8.

10. A learning device comprising: an acquisition unit configured to acquire calculation imaging information related to a first camera that captures an image having blur, the calculation imaging information being a difference image between a first image captured by the first camera and a second image, the first image including a point light source in a light-on state, the second image including the point light source in a light-off state; a storage unit configured to store a third image captured by a second camera and a correct answer label given to the third image, the second camera capturing an image having no blur or an image having less blur than the image captured by the first camera; an image generation unit configured to generate a fourth image having blur based on the calculation imaging information acquired by the acquisition unit and the third image read out from the storage unit; and a learning unit configured to create an image recognition model for recognizing an image captured by the first camera by performing machine learning using the fourth image generated by the image generation unit and the correct answer label read out from the storage unit.

11. An image recognition system comprising: an acquisition unit configured to acquire calculation imaging information related to a first camera that captures an image having blur, the calculation imaging information being a difference image between a first image captured by the first camera and a second image, the first image including a point light source in a light-on state, the second image including the point light source in a light-off state; a storage unit configured to store a third image captured by a second camera and a correct answer label given to the third image, the second camera capturing an image having no blur or an image having less blur than the image captured by the first camera; an image generation unit configured to generate a fourth image having blur based on the calculation imaging information acquired by the acquisition unit and the third image read out from the storage unit; a learning unit configured to create an image recognition model by performing machine learning using the fourth image generated by the image generation unit and the correct answer label read out from the storage unit; a recognition unit configured to recognize an image captured by the first camera based on the image recognition model created by the learning unit; and an output unit configured to output a recognition result of the recognition unit. ​ ​

Citation Information

Patent Citations

  • Image processing device, image processing method, processing device, processing method and program

    JP2020024612A

  • Image generation apparatus and method for generating image

    US20190279026A1