Learning method, image recognition method, learning device, and image recognition system
The learning method improves image recognition accuracy and learning efficiency in privacy-protected environments by generating image recognition models from blurred images created using calculation imaging information and non-blurred images from a second camera.
Patent Information
- Application Number
- JP2022536223
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-16
- Filing Date
- 2021-06-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-25
AI Technical Summary
Existing image recognition systems face challenges in improving image recognition accuracy and learning efficiency while protecting privacy, especially in environments where computational imaging images are difficult for humans to recognize.
A learning method that acquires calculation imaging information from a first camera capturing blurred images, generates a blurred image based on this information and a non-blurred image from a second camera, and creates an image recognition model through machine learning using these images.
This approach enhances image recognition accuracy and learning efficiency while ensuring privacy protection by utilizing blurred images that are difficult for humans to recognize.
Smart Images

Figure 0007684306000001 
Figure 0007684306000002 
Figure 0007684306000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image recognition method and an image recognition system in an environment where privacy protection is required, such as inside a home or indoors, and a learning method and a learning device for creating an image recognition model used for the image recognition.
Background Art
[0002] Patent Document 1 below discloses an image recognition system in which a computational imaging image captured by a light field camera or the like is input to a discriminator, and the discriminator uses a learned discrimination model to identify an object included in the computational imaging image.
[0003] A computational imaging image is an image in which a plurality of images with different viewpoints are superimposed, or an image in which a subject image is difficult to be focused due to the use of no lens or the like, and is an image in which visual recognition by a human is difficult due to intentionally created blurring. Therefore, it is suitable to use a computational imaging image for constructing an image recognition system in an environment where privacy protection is required, such as inside a home or indoors.
[0004] On the other hand, since a computational imaging image is difficult for visual recognition by a human, it is difficult to assign an accurate correct label to a computational imaging image captured by a light field camera or the like in machine learning for creating a discrimination model. As a result, the learning efficiency decreases.
[0005] According to Patent Document 1 below, since no countermeasure has been taken against this problem, it is desired to improve the learning efficiency by realizing an effective technical countermeasure.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
[0007] The present disclosure aims to provide a technology capable of improving image recognition accuracy and learning efficiency of machine learning while protecting the privacy of a subject in an image recognition system.
[0008] A learning method according to an aspect of the present disclosure is such that an information processing apparatus as a learning apparatus acquires calculation imaging information regarding a first camera that captures a blurred image, and the calculation imaging information is a difference image between a first image including a point light source in a lit state and a second image including the point light source in a turned-off state, which are captured by the first camera, and a third image captured by a second camera that captures a non-blurred image or an image with less blur than the first camera, and a correct label assigned to the third image are acquired, a fourth blurred image is generated based on the calculation imaging information and the third image, and an image recognition model for identifying an image captured by the first camera is created by performing machine learning using the fourth image and the correct label.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17A
Figure 17B
Figure 17C
Figure 17D
Figure 18A
Figure 18B
Figure 18C
Figure 18D
Figure 19
Figure 20
Figure 21
Figure 22A
Figure 22B
Figure 22C
Figure 22D
Figure 22E
Figure 22F
Figure 23A
Figure 23B
Figure 23C
Mode for Carrying Out the Invention
[0010] (Knowledge on which the present disclosure is based) In a home or indoors, various recognition technologies such as the recognition of the actions of people in the environment and the recognition of people as equipment operators are important. In recent years, for object recognition, a technology called deep learning has attracted attention. Deep learning is machine learning using a neural network with a multi-layer structure, and by using a large amount of learning data, it is possible to achieve higher-precision recognition performance compared to conventional methods. In such object recognition, image information is particularly effective. By using a camera as an input device and performing deep learning with the image information as input, various methods have been proposed to significantly improve the conventional object recognition ability.
[0011] However, placing a camera indoors or in a home has the problem that if the captured image leaks outside due to hacking or the like, privacy will be violated. Therefore, even if the captured image leaks outside, measures are necessary to protect the privacy of the subject.
[0012] The computational imaging image captured by a light field camera or the like is an image in which a plurality of images with different viewpoints are superimposed, or due to effects such as the subject image being difficult to focus because no lens is used, and the intentionally created blur makes it difficult for humans to visually recognize. Therefore, it is suitable to use the computational imaging image for constructing an image recognition system in an environment where privacy protection is required, such as particularly in a home or indoors.
[0013] In the image recognition system disclosed in Patent Document 1 above, a target area is photographed by a light field camera or the like, and the computational imaging image obtained by the photographing is input to an identifier. Thereby, the identifier uses a learned recognition model to identify the object included in the computational imaging image. In this way, by photographing the target area with a light field camera or the like that captures the computational imaging image, even if the captured image leaks outside, since the computational imaging image is difficult for humans to visually recognize, the privacy of the subject can be protected.
[0014] In the image recognition system disclosed in the above Patent Document 1, the recognition model used by the recognizer is created by performing machine learning using a computationally captured image captured by a light field camera or the like as learning data. However, since it is difficult for humans to visually recognize a computationally captured image, it is difficult to assign an accurate correct label to a computationally captured image captured by a light field camera or the like in the machine learning for creating the recognition model. If an incorrect correct label is assigned to the learning computational imaging image, the learning efficiency of the machine learning decreases.
[0015] In order to solve such problems, the inventor of the present invention uses a non-blurred image (hereinafter referred to as "normal image") instead of a blurred image (hereinafter referred to as "blurred image") such as a computationally captured image at the stage of accumulating learning data, and then, at the subsequent learning stage, performs machine learning using a blurred image obtained by converting a normal image based on the computational imaging information of the used camera. As a result, the inventor has obtained the knowledge that it is possible to improve the image recognition accuracy and the learning efficiency of machine learning while protecting the privacy of the subject, and has arrived at the present disclosure.
[0016] Also, from another perspective of privacy protection, it is also important to reduce the psychological burden on the user captured by the image recognition device. By capturing a blurred image, it is possible to appeal that the privacy of the subject is protected. However, if the computational imaging information is set in an area irrelevant to the user (such as the manufacturer's factory), there is a possibility that the psychological burden on the user may increase due to the suspicion that the manufacturer may be able to restore the normal image from the blurred image. On the other hand, the inventor considered that if the computational imaging information can be changed by the user himself / herself who is being captured, this psychological burden can be reduced, and thus arrived at the present disclosure.
[0017] Next, each aspect of the present disclosure will be described.
[0018] A learning method according to one aspect of the present disclosure is such that an information processing apparatus as a learning device acquires calculation imaging information regarding a first camera that captures a blurry image. The calculation imaging information is a difference image between a first image including a point light source in a lit state and a second image including the point light source in a turned-off state, which are captured by the first camera. The method also acquires a third image captured by a second camera that captures a non-blurry image or an image with less blur than the first camera, and a correct label assigned to the third image. Based on the calculation imaging information and the third image, a fourth blurry image is generated, and an image identification model for identifying an image captured by the first camera is created by performing machine learning using the fourth image and the correct label.
[0019] In the present disclosure, "blur" refers to a state in which a plurality of images with different viewpoints are superimposed when captured by a light field camera or a lensless camera, or a state in which it is difficult for a human to visually recognize due to the influence that the subject image is out of focus because no lens is used, or simply a state in which the subject is out of focus. A "blurry image" means an image that is difficult for a human to visually recognize or an image in which the subject is out of focus. "Large blur" means that the difficulty of visual recognition by a human is large or the degree of out-of-focus of the subject is large, and "small blur" means that the difficulty or the degree is small. A "non-blurry image" means an image that is easy for a human to visually recognize or an image in which the subject is in focus.
[0020] According to this configuration, the target area where the subject to be image-identified is located is imaged by the first camera that captures a blurry image. Therefore, even if the captured image by the first camera leaks outside, since the image is difficult for human visual recognition, the privacy of the subject can be protected. Also, the third image, which is learning data, is imaged by the second camera that captures a non-blurry or small image. Therefore, since the image is easy for human visual recognition, an accurate correct label can be easily assigned to the third image. Furthermore, the computational imaging information regarding the first camera is a difference image between a first image including a point light source in the lit state and a second image including the point light source in the off state. Therefore, the computational imaging information regarding the first camera actually used can be accurately acquired without being affected by subjects other than the point light source. Thereby, the fourth image used for machine learning can be accurately generated based on the computational imaging information and the third image. As a result, it becomes possible to improve the image identification accuracy and the learning efficiency of machine learning while protecting the privacy of the subject.
[0021] In the above aspect, the first camera may be any one of an encoded aperture camera provided with a mask having a mask pattern with different transmittance for each region, a multi-pinhole camera in which a mask having a mask pattern with a plurality of pinholes formed therein is disposed on the light-receiving surface of the image sensor, and a light field camera that acquires a light field from a subject.
[0022] According to this configuration, by using any one of an encoded aperture camera, a multi-pinhole camera, and a light field camera as the first camera, a blurry image that is difficult for human visual recognition can be appropriately imaged.
[0023] In the above aspect, the first camera may not have an optical system that forms an image of light from a subject on an image sensor.
[0024] According to this configuration, since the first camera does not have an optical system that forms an image of light from the subject on the image sensor, it is possible to intentionally create blur in the captured image by the first camera. As a result, it becomes more difficult to identify the subject included in the captured image, so that the effect of protecting the privacy of the subject can be further enhanced.
[0025] In the above aspect, it is preferable that the mask can be changed to another mask having a different mask pattern.
[0026] According to this configuration, since the calculated imaging information of the first camera also changes by changing the mask, for example, by each user arbitrarily changing the mask, the calculated imaging information can be made different for each user. As a result, it becomes difficult for a third party to perform inverse conversion from the fourth image to the third image, so that the effect of protecting the privacy of the subject can be further enhanced.
[0027] In the above aspect, the calculated imaging information may be either a Point Spread Function or a Light Transport Matrix.
[0028] According to this configuration, by using either the PSF or the LTM, it is possible to easily and appropriately acquire the calculated imaging information regarding the first camera.
[0029] In the above aspect, it is preferable that the information processing apparatus performs lighting control of the point light source and imaging control of the first image by the first camera, and performs extinguishing control of the point light source and imaging control of the second image by the first camera.
[0030] According to this configuration, by the information processing apparatus controlling the operations of the point light source and the first camera, it is possible to accurately synchronize the timing of turning on or off the point light source and the timing of imaging by the first camera.
[0031] In the above aspect, when the image quality of the differential image is less than the allowable value, the information processing apparatus may perform re - imaging control of the first image and the second image by the first camera.
[0032] According to this configuration, when the image quality of the differential image is less than the allowable value, by the information processing apparatus performing re - imaging control by the first camera, a differential image with the luminance value of the point light source appropriately adjusted can be obtained. As a result, it becomes possible to obtain appropriate calculation imaging information regarding the first camera.
[0033] In the above aspect, in the re - imaging control, the information processing apparatus may correct at least one of the exposure time and the gain of the first camera so that the maximum luminance value is within a predetermined range for each of the first image and the second image.
[0034] According to this configuration, by correcting at least one of the exposure time and the gain of the first camera, it becomes possible to obtain a differential image with the luminance value of the point light source appropriately adjusted by re - imaging control.
[0035] An image identification method according to an aspect of the present disclosure is an identification apparatus having an identification unit. An image captured by a first camera that captures a blurry image is input to the identification unit, the identification unit identifies the input image based on a learned image identification model, outputs the result of the identification by the identification unit, and the image identification model is an image identification model created by the learning method according to the above aspect.
[0036] According to this configuration, the target area where the subject to be image-identified is located is imaged by the first camera that captures a blurred image. Therefore, even if the captured image by the first camera leaks externally, since the image is difficult for human visual recognition, the privacy of the subject can be protected. Also, the third image which is learning data is imaged by the second camera that captures a non-blurred or slightly blurred image. Therefore, since the image is easy for human visual recognition, an accurate correct label can be easily assigned to the third image. Furthermore, the computational imaging information regarding the first camera is a difference image between a first image including a point light source in the lit state and a second image including the point light source in the off state. Therefore, the computational imaging information regarding the actually used first camera can be accurately acquired without being affected by subjects other than the point light source. Thereby, the fourth image used for machine learning can be accurately generated based on the computational imaging information and the third image. As a result, it becomes possible to improve the image identification accuracy and the learning efficiency of machine learning while protecting the privacy of the subject.
[0037] A learning device according to one aspect of the present disclosure includes an acquisition unit that acquires computational imaging information regarding a first camera that captures a blurred image, where the computational imaging information is a difference image between a first image including a point light source in the lit state and a second image including the point light source in the off state, both captured by the first camera, a storage unit that stores a third image captured by a second camera that captures a non-blurred image or an image with less blur than the first camera, and a correct label assigned to the third image, an image generation unit that generates a fourth blurred image based on the computational imaging information acquired by the acquisition unit and the third image read from the storage unit, and a learning unit that creates an image identification model for identifying an image captured by the first camera by performing machine learning using the fourth image generated by the image generation unit and the correct label read from the storage unit.
[0038] According to this configuration, the target area where the subject to be image-identified is located is imaged by the first camera that captures a blurry image. Therefore, even if the captured image by the first camera leaks outside, since the image is difficult for humans to visually recognize, the privacy of the subject can be protected. Also, the third image, which is learning data, is imaged by the second camera that captures a non-blurry or less blurry image. Therefore, since the image is easy for humans to visually recognize, an accurate correct label can be easily assigned to the third image. Furthermore, the computational imaging information regarding the first camera is a difference image between a first image including a point light source in the lit state and a second image including the point light source in the off state. Therefore, the computational imaging information regarding the actually used first camera can be accurately obtained without being affected by subjects other than the point light source. As a result, the image synthesis unit can accurately generate the fourth image used for machine learning based on the computational imaging information and the third image. Consequently, it becomes possible to improve the image identification accuracy and the learning efficiency of machine learning while protecting the privacy of the subject.
[0039] An image identification system according to an aspect of the present disclosure includes an acquisition unit that acquires computational imaging information regarding a first camera that captures a blurry image, where the computational imaging information is a difference image between a first image including a point light source in the lit state and a second image including the point light source in the off state, both of which are captured by the first camera, a storage unit that stores a third image captured by a second camera that captures a non-blurry image or an image with less blur than the first camera, and a correct label assigned to the third image, an image generation unit that generates a blurry fourth image based on the computational imaging information acquired by the acquisition unit and the third image read from the storage unit, a learning unit that creates an image identification model by performing machine learning using the fourth image generated by the image generation unit and the correct label read from the storage unit, an identification unit that identifies an image captured by the first camera based on the image identification model created by the learning unit, and an output unit that outputs an identification result by the identification unit.
[0040] According to this configuration, the target area where the subject to be image-identified is located is imaged by the first camera that captures a blurred image. Therefore, even if the captured image by the first camera leaks outside, since the image is difficult for human visual recognition, the privacy of the subject can be protected. Also, the third image, which is learning data, is imaged by the second camera that captures a non-blurred or small image. Therefore, since the image is easy for human visual recognition, an accurate correct label can be easily assigned to the third image. Furthermore, the computational imaging information regarding the first camera is a difference image between the first image including a point light source in the lit state and the second image including the point light source in the extinguished state. Therefore, the computational imaging information regarding the actually used first camera can be accurately obtained without being affected by subjects other than the point light source. As a result, the image synthesis unit can accurately generate the fourth image used for machine learning based on the computational imaging information and the third image. As a result, it becomes possible to improve the image identification accuracy and the learning efficiency of machine learning while protecting the privacy of the subject.
[0041] The present disclosure can be realized as a computer program for causing a computer to execute each characteristic configuration included in such a method, or can be realized as an apparatus or a system that operates based on this computer program. Also, it goes without saying that such a computer program can be distributed as a computer-readable non-volatile recording medium such as a CD-ROM, or can be distributed via a communication network such as the Internet.
[0042] Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Also, among the components in the following embodiments, components not described in the independent claims indicating the highest-level concept are described as optional components. In addition, in all embodiments, the respective contents can be combined with each other.
[0043] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that elements denoted by the same reference numerals in different drawings indicate the same or corresponding elements.
[0044] (First Embodiment) FIG. 1 is a schematic diagram showing the configuration of an image recognition system 10 according to the first embodiment of the present disclosure. The image recognition system 10 includes a learning device 20 and an identification device 30. The identification device 30 has a calculation imaging camera 101, an identification unit 106, and an output unit 107. The identification unit 106 includes a processor such as a CPU and a memory such as a semiconductor memory. The output unit 107 is a display device or a speaker, etc. The learning device 20 has a learning database 102, a calculation imaging information acquisition unit 103, a database correction unit 104, and a learning unit 105. The learning database 102 is a storage unit such as an HDD, an SSD, or a semiconductor memory. The calculation imaging information acquisition unit 103, the database correction unit 104, and the learning unit 105 are processors such as a CPU.
[0045] FIG. 2 is a flowchart showing the main processing procedure of the image identification system 10. The flowchart shows the flow of the image identification process by the identification device 30. First, the computational imaging camera 101 captures an image of the target area and inputs the computational imaging image obtained by the capture to the identification unit 106 (step S101). Next, the identification unit 106 identifies the computational imaging image using a learned image identification model (step S102). This image identification model is an image identification model created by learning by the learning device 20. Next, the output unit 107 outputs the result of the identification by the identification unit 106. Details of the processing of each step will be described later.
[0046] Unlike a normal camera that captures a normal image without blur, the computational imaging camera 101 captures a computational imaging image that is a blurred image. Although the computational imaging image is such that the subject cannot be recognized even by a person looking at the captured image due to intentionally created blur, an image that can be recognized by a person or identified by the identification unit 106 can be generated by performing image processing on the captured computational imaging image.
[0047] FIG. 3 is a diagram schematically showing the structure of a multi-pinhole camera 301 composed of a lensless as an example of the computational imaging camera 101. The multi-pinhole camera 301 shown in FIG. 3 has a multi-pinhole mask 301a and an image sensor 301b such as a CMOS. The multi-pinhole mask 301a is arranged at a certain distance from the light receiving surface of the image sensor 301b. The multi-pinhole mask 301a has a plurality of pinholes 301aa arranged randomly or at equal intervals. The plurality of pinholes 301aa are also called multi-pinholes. The image sensor 301b acquires an image of the subject 302 through each pinhole 301aa. The image acquired through the pinhole is called a pinhole image.
[0048] Since the pinhole images of the subject 302 vary depending on the positions and sizes of the respective pinholes 301aa, the image sensor 301b acquires a superimposed image in a state where a plurality of pinhole images are slightly shifted and overlapped (multiple images). The positional relationship of the plurality of pinholes 301aa affects the positional relationship of the plurality of pinhole images projected onto the image sensor 301b (that is, the degree of overlap of the multiple images), and the size of the pinhole 301aa affects the degree of blurring of the pinhole image.
[0049] By using the multi-pinhole mask 301a, it is possible to obtain a superimposed image of a plurality of pinhole images with different positions and degrees of blurring. That is, it is possible to obtain a computational imaging image in which multiple images and blurring are intentionally created. Therefore, the captured image becomes a multiple image and a blurred image, and an image in which the privacy of the subject 302 is protected by these blurs can be obtained. Also, by changing the number, position, and size of each pinhole, it is possible to obtain images with different blurring effects. That is, it may be configured such that the multi-pinhole mask 301a can be easily detached by the user, and a plurality of types of multi-pinhole masks 301a with different mask patterns are prepared in advance, and the user can freely exchange the multi-pinhole mask 301a to be used.
[0050] Note that such mask changes can be achieved not only by exchanging the mask, · by the user arbitrarily rotating a mask rotatably attached in front of the image sensor, · by the user making holes at arbitrary locations on a plate attached in front of the image sensor, · by using a liquid crystal mask or the like using a spatial light modulator or the like to arbitrarily set the transmittance of each position in the mask, · by forming the mask using an expandable material such as rubber and physically deforming the mask by applying an external force to change the position and size of the holes, and can be realized in various ways. Hereinafter, these modification examples will be described in order.
[0051] <Modification Example in Which the User Arbitrarily Rotates the Mask> Figs. 17A to 17D are schematic views showing the configuration of a multi-pinhole camera 301 in which the user can arbitrarily rotate the mask. Fig. 17A shows an overview of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask, and Fig. 17B shows a schematic cross-sectional view thereof. The multi-pinhole camera 301 has a multi-pinhole mask 301a that is rotatable with respect to its housing 401, and a gripping portion 402 is connected to the multi-pinhole mask 301a. The user can fix or rotate the multi-pinhole mask 301a with respect to the housing 401 by gripping and operating the gripping portion 402. Such a mechanism may be provided with a screw in the gripping portion 402, and the multi-pinhole mask 301a can be fixed by tightening the screw, and the multi-pinhole mask 301a can be rotated by loosening the screw. Figs. 17C and 17D show schematic views in which the multi-pinhole mask 301a rotates 90 degrees when the gripping portion 402 is rotated 90 degrees. In this way, by the user moving the gripping portion 402, the multi-pinhole mask 301a can be rotated.
[0052] Also, in the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask, the multi-pinhole mask 301a may have an asymmetric pinhole arrangement with respect to rotation, as shown in Fig. 17C. By doing so, it is possible to realize various multi-pinhole patterns by the user rotating the mask.
[0053] Of course, the configuration of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask may not have the grip portion 402. FIGS. 18A and 18B are schematic views showing another configuration example of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask. FIG. 18A shows an overview of another configuration example of the multi-pinhole camera 301 in which the user can arbitrarily rotate the mask, and FIG. 18B shows a schematic cross-sectional view thereof. The multi-pinhole mask 301a is fixed to the lens barrel 411. The image sensor 301b is installed in another lens barrel 412, and the lens barrel 411 and the lens barrel 412 are rotatable in a screw configuration. That is, there is a lens barrel 412 outside the lens barrel 411, and a male screw is cut on the outside of the lens barrel 411, which is the joint portion, and a female screw is cut on the inside of the lens barrel 412. First, a fixture 413 is attached to the male screw of the lens barrel 411, and then the lens barrel 412 is attached. The fixture 413 also has a female screw cut on it, similar to the lens barrel 412. With such a configuration, when the lens barrel 411 is screwed into the lens barrel 412, the depth of screwing changes depending on the screwing position of the fixture 413 into the lens barrel 411, and the rotation angle of the multi-pinhole camera 301 can be changed.
[0054] FIGS. 18C and 18D are schematic views showing that the depth of screwing changes depending on the screwing position of the fixture 413 into the lens barrel 411, and the rotation angle of the multi-pinhole camera 301 changes. FIG. 18C is a schematic view when the fixture 413 is screwed all the way into the lens barrel 411, and FIG. 18D is a schematic view when the fixture 413 is only screwed partway into the lens barrel 411. As shown in FIG. 18C, when the fixture 413 is screwed all the way into the lens barrel 411, the lens barrel 412 can be screwed all the way into the lens barrel 411. On the other hand, as shown in FIG. 18D, when the fixture 413 is only screwed partway into the lens barrel 411, the lens barrel 412 can only be screwed partway into the lens barrel 411. Therefore, the depth of screwing changes depending on the screwing position of the fixture 413 into the lens barrel 411, and the rotation angle of the multi-pinhole mask 301a can be changed.
[0055] <Modification example where the user makes a hole in the mask> FIG. 19 is a schematic cross-sectional view of a multi-pinhole camera 301 in which a user can make holes at any position on a mask 301ab attached in front of an image sensor 301b. In FIG. 19, the same components as those in FIG. 17 are denoted by the same reference numerals, and the description thereof is omitted. Initially, there are no pinholes in the mask 301ab. By making a plurality of holes at any position in this mask 301ab using a needle or the like by the user, a multi-pinhole mask of an arbitrary shape can be created.
[0056] <Modification example of arbitrarily setting the transmittance of each position in the mask using a spatial light modulator> FIG. 20 is a schematic cross-sectional view of a multi-pinhole camera 301 configured to arbitrarily set the transmittance of each position in a mask using a spatial light modulator 420. In FIG. 20, the same components as those in FIG. 19 are denoted by the same reference numerals, and the description thereof is omitted. The spatial light modulator 420 is composed of liquid crystal or the like and can change the transmittance for each pixel. This spatial light modulator 420 functions as a multi-pinhole mask. The change in the transmittance can be controlled by a spatial light modulator control unit (not shown). Therefore, by the user selecting an arbitrary pattern from a plurality of pre-prepared transmittance patterns, various mask patterns (multi-pinhole patterns) can be realized.
[0057] <Modification example of deforming the mask by applying an external force> Figs. 21, 22A to 22F are schematic cross-sectional views of a multi-pinhole camera 301 configured to deform a mask by applying an external force. In Fig. 21, the same reference numerals are given to the same components as in Fig. 19, and the description thereof is omitted. The multi-pinhole mask 301ac is composed of a plurality of masks 301a1, 301a2, 301a3, and each mask has a driving unit (not shown) for independently applying an external force. Figs. 22A to 22C are schematic views for explaining the three masks 301a1, 301a2, 301a3 constituting the multi-pinhole mask 301ac. Here, each mask has a shape in which a fan shape and an annular shape are combined. Of course, this configuration is an example, and the shape is not limited to a fan shape, and the number of components is not limited to three. One or more pinholes are formed in each mask. Note that the mask may not have a pinhole formed therein. Two pinholes 301aa1, 301aa2 are formed in the mask 301a1, one pinhole 301aa3 is formed in the mask 301a2, and two pinholes 301aa4, 301aa5 are formed in the mask 301a3. By moving these three masks 301a1 to 301a3 by applying an external force, various multi-pinhole patterns can be created.
[0058] Figs. 22D to 22F show three types of multi-pinhole masks 301ac composed of the three masks 301a1 to 301a3. By moving each of the driving units (not shown) of the masks 301a1 to 301a3 in different manners, masks having five pinholes are configured in Figs. 22D and 22E, and a mask having four pinholes is configured in Fig. 22F. Such a driving unit of the mask can be realized by using an ultrasonic motor or a linear motor that is widely used in autofocus and the like. In this way, the number and position of the pinholes in the multi-pinhole mask 301ac can be changed by applying an external force.
[0059] Of course, the multi-pinhole mask may be changed not only in the number and position of the pinholes but also in its size. FIGS. 23A to 23C are schematic diagrams for explaining the configuration of the multi-pinhole mask 301ad in the multi-pinhole camera 301 having a configuration in which the mask is deformed by the application of an external force. The multi-pinhole mask 301ad has a plurality of pinholes, is made of an elastic material, and has four drive units 421 to 424 that can be independently controlled at the four corners. Of course, the number of drive units does not have to be four. By moving each of the drive units 421 to 424, the position and size of the pinholes in the multi-pinhole mask 301ad can be changed.
[0060] FIG. 23B is a schematic diagram showing the state when the drive units 421 to 424 are moved in the same direction. In this figure, the directions of the arrows shown in the drive units 421 to 424 indicate the driving directions of the respective drive units. In this case, the multi-pinhole mask 301ad moves parallel to the driving direction of the drive units. On the other hand, FIG. 23C is a schematic diagram showing the state when the drive units 421 to 424 are moved in a direction outward from the central portion of the multi-pinhole mask 301ad. In this case, since the multi-pinhole mask 301ad is stretched according to its elasticity, the size of the pinholes becomes larger. Such drive units 421 to 424 can be realized by using ultrasonic motors or linear motors that are widely used in autofocus and the like. In this way, the position and size of the pinholes in the multi-pinhole mask 301ac can be changed by the application of an external force.
[0061] FIG. 4A is a diagram showing the positional relationship of a plurality of pinholes 301aa in the multi-pinhole camera 301. In this example, three pinholes 301aa arranged linearly are formed. The distance between the leftmost pinhole 301aa and the central pinhole 301aa is set to L1, and the distance between the central pinhole 301aa and the rightmost pinhole 301aa is set to L2 (<L1).
[0062] Figures 4B and 4C are diagrams showing an example of a captured image by the multi-pinhole camera 301. Figure 4B shows an example of a captured image when the distance between the multi-pinhole camera 301 and the subject 302 is relatively far and the subject image is small. Figure 4C shows an example of a captured image when the distance between the multi-pinhole camera 301 and the subject 302 is relatively close and the subject image is large. By varying the intervals L1 and L2, regardless of the distance between the multi-pinhole camera 301 and the subject 302, a superimposed image in a state where a plurality of subject images overlap in an indistinguishable manner is captured by superimposing a plurality of images with different viewpoints.
[0063] As the computational imaging camera 101, in addition to the multi-pinhole camera 301, · A coded aperture camera in which a mask having a mask pattern with different transmittances for each region is disposed between the image sensor and the subject, · A light field camera having a configuration in which a microlens array is disposed on the light-receiving surface of the image sensor and acquiring a light field, · A compressive sensing camera that captures an image by weighted addition of pixel information in space-time and other well-known cameras can also be used.
[0064] Also, in the computational imaging camera 101, it is desirable not to have an optical system (such as a lens, prism, mirror, etc.) for forming an image of light from the subject on the image sensor. By omitting the optical system, the camera can be made smaller and lighter, the cost can be reduced, and the design can be improved, and at the same time, intentional blurring can be created in the captured image by the camera.
[0065] The identification unit 106 uses the image identification model, which is the learning result of the learning device 20, to identify, with respect to the image of the target area captured by the computational imaging camera 101, the category information of the subjects such as the people (including their actions and expressions), automobiles, bicycles, or signals included in the image, and the position information of each subject. For the learning to create the image identification model, machine learning such as Deep Learning using a multi-layer neural network may be utilized.
[0066] The output unit 107 outputs the result identified by the identification unit 106. This may have an interface unit and present the identification result to the user by means of an image, text, or voice, etc., or may have a device control unit and change the control method according to the identification result.
[0067] The learning device 20 includes a learning database 102, a computational imaging information acquisition unit 103, a database modification unit 104, and a learning unit 105. The learning device 20 performs learning to create the image identification model used by the identification unit 106 in correspondence with the computational imaging information regarding the computational imaging camera 101 actually used for imaging the target area.
[0068] Also, FIG. 5 is a flowchart showing the main processing procedure of the learning device 20 of the image identification system 10.
[0069] First, the computational imaging information acquisition unit 103 acquires the computational imaging information, which is information representing the manner of blurring, regarding what kind of blurred image is captured by the computational imaging camera 101 (step S201). This may be such that the computational imaging camera 101 has a transmission unit and the computational imaging information acquisition unit 103 has a reception unit, and they exchange the computational imaging information by wire or wirelessly, or the computational imaging information acquisition unit 103 has an interface and the user inputs the computational imaging information to the computational imaging information acquisition unit 103.
[0070] As the computational imaging information, for example, if the computational imaging camera 101 is a multi-pinhole camera 301, a PSF (Point Spread Function) indicating the state of two-dimensional computational imaging may be used. The PSF is the transfer function of a camera such as a multi-pinhole camera or a coded aperture camera, and is expressed by the following relationship.
[0071] y = k * x
[0072] Here, y is a blurred computational imaging image captured by the multi-pinhole camera 301, k is the PSF, and x is a normal image captured by a normal camera without blur of the captured scene. Also, * is a convolution operator.
[0073] Also, as the computational imaging information, instead of the PSF, an LTM (Light Transport Matrix) indicating computational imaging information of four dimensions or more (two dimensions on the camera side and two dimensions or more on the subject side) may be used. The LTM is the transfer function used in a light field camera.
[0074] For example, when the computational imaging camera 101 is the multi-pinhole camera 301, the PSF can be obtained by photographing a point light source with the multi-pinhole camera 301. This can be understood because the PSF corresponds to the impulse response of the camera. That is, the captured image of the point light source itself obtained by imaging the point light source with the multi-pinhole camera 301 is the PSF as the computational imaging information of the multi-pinhole camera 301. Here, as the captured image of the point light source, it is desirable to use a difference image between when it is lit and when it is turned off, which will be described in the second embodiment below.
[0075] Next, the database correction unit 104 acquires a normal image without blur included in the learning database 102, and the learning unit 105 acquires annotation information included in the learning database 102 (step S202).
[0076] Next, the database correction unit 104 (image generation unit) corrects the learning database 102 using the computational imaging information acquired by the computational imaging information acquisition unit 103 (step S203). For example, when the identification unit 106 identifies the actions of a person in the environment, the learning database 102 holds a plurality of normal images taken with a normal camera without blur and annotation information (correct labels) assigned to each image indicating the position and action of the person in each image. When using a normal camera, annotation information may be assigned to the images taken with that camera. However, when acquiring computational imaging images such as a multi-pinhole camera or a light field camera, it is difficult to assign annotation information because it is not clear what is shown in the image even when a person looks at it. Also, even if learning processing is performed using images taken with a normal camera that is significantly different from the computational imaging camera 101, the identification accuracy of the identification unit 106 will not increase. Therefore, a database with annotation information pre-assigned to images taken with a normal camera is held as the learning database 102, and by deforming only the captured images according to the computational imaging information of the computational imaging camera 101, a learning dataset adapted to that computational imaging camera 101 is created, and the identification accuracy is improved by performing learning processing. For this purpose, the database correction unit 104 calculates the following corrected image y using the PSF, which is the computational imaging information acquired by the computational imaging information acquisition unit 103, for the pre-prepared captured image z with a normal camera.
[0077] y = k * z
[0078] Here, k represents the PSF, which is the computational imaging information acquired by the computational imaging information acquisition unit 103, and * represents the convolution operator.
[0079] The learning unit 105 performs learning processing (step S204) using the corrected image calculated by the database correction unit 104 and the annotation information acquired from the learning database 102 in this way. For example, when the identification unit 106 is constructed by a multi-layer neural network, machine learning by Deep Learning is performed using the corrected image and the annotation information as teacher data. As a correction algorithm for prediction error, the Back Propagation method or the like may be used. Thereby, the learning unit 105 creates an image identification model for the identification unit 106 to identify the image captured by the computational imaging camera 101. Since the corrected image is an image that matches the computational imaging information of the computational imaging camera 101, such learning enables learning adapted to the computational imaging camera 101, and the identification unit 106 can perform highly accurate identification processing.
[0080] According to the image identification system 10 according to the present embodiment, the target area where the subject 302 to be identified is located is imaged by a computational imaging camera 101 (first camera) that captures a computational imaging image that is a blurred image. Therefore, even if the captured image by the computational imaging camera 101 leaks outside, since the computational imaging image is difficult for human visual recognition, the privacy of the subject 302 can be protected. Further, the normal image (third image) stored in the learning database 102 is imaged by a normal camera (second camera) that captures a non-blurred image (or an image with less blur than the computational imaging image). Therefore, since the image is easy for human visual recognition, accurate annotation information (correct label) can be easily given to the normal image. As a result, it is possible to improve the image identification accuracy and the learning efficiency of machine learning while protecting the privacy of the subject 302.
[0081] Further, by using any one of an encoded aperture camera, a multi-pinhole camera, and a light field camera as the computational imaging camera 101, it is possible to appropriately capture a blurred image that is difficult for human visual recognition.
[0082] Further, in the computational imaging camera 101, by omitting the optical system that forms an image of the light from the subject 302 on the image sensor 301b, it is possible to intentionally create blurring in the captured image by the computational imaging camera 101. As a result, since it becomes more difficult to identify the subject 302 included in the captured image, it is possible to further enhance the privacy protection effect of the subject 302.
[0083] Also, when the multi-pinhole mask 301a to be used is configured to be freely changeable by the user, since the computational imaging information of the computational imaging camera 101 changes by changing the mask, for example, by each user arbitrarily changing the mask, the computational imaging information can be made different for each user. As a result, since it becomes difficult for a third party to perform inverse conversion from the corrected image (the fourth image) to the normal image (the third image), it is possible to further enhance the privacy protection effect of the subject 302.
[0084] Also, by using either the PSF or the LTM as the computational imaging information, it is possible to easily and appropriately acquire the computational imaging information regarding the computational imaging camera 101.
[0085] (Second Embodiment) FIG. 6 is a schematic diagram showing the configuration of an image identification system 11 according to the second embodiment of the present disclosure. In FIG. 6, the same components as those in FIG. 1 are denoted by the same reference numerals, and the description thereof is omitted. The learning device 21 of the image identification system 11 includes a control unit 108. Further, the image identification system 11 includes a light emitting unit 109 existing in the target area (environment) photographed by the computational imaging camera 101. The light emitting unit 109 is a light source that can be regarded as a point light source existing in the environment, and for example, is an LED mounted on an electric device or an LED for illumination. Also, it may function as the light emitting unit 109 by turning on and off only a part of the light of a monitor such as an LED monitor. By the control unit 108 controlling the light emitting unit 109 and the computational imaging camera 101, the computational imaging information acquisition unit 103 acquires the computational imaging information.
[0086] Further, FIG. 7 is a flowchart showing the procedure of the main processing of the image identification system 11. This flowchart shows the flow of the process in which the calculation imaging information acquisition unit 103 acquires the calculation imaging information of the calculation imaging camera 101.
[0087] First, the control unit 108 issues an instruction to turn on the light to the light emitting unit 109 existing in the environment (step S111).
[0088] Next, the light emitting unit 109 turns on the light according to the instruction of the control unit 108 (step S112).
[0089] Next, the control unit 108 issues an instruction to the calculation imaging camera 101 to perform imaging (step S113). Thereby, the light emitting unit 109 and the calculation imaging camera 101 can operate while being synchronized.
[0090] Next, the calculation imaging camera 101 performs imaging according to the instruction of the control unit 108 (step S114). The captured image (first image) is input from the calculation imaging camera 101 to the calculation imaging information acquisition unit 103 and temporarily held by the calculation imaging information acquisition unit 103.
[0091] Next, the control unit 108 issues an instruction to turn off the light to the light emitting unit 109 (step S115).
[0092] Next, the light emitting unit 109 turns off the light according to the instruction of the control unit 108 (step S116).
[0093] Next, the control unit 108 issues an instruction to the calculation imaging camera 101 to perform imaging (step S117).
[0094] Next, the calculation imaging camera 101 performs imaging according to the instruction of the control unit 108 (step S118). The captured image (second image) is input from the calculation imaging camera 101 to the calculation imaging information acquisition unit 103.
[0095] Next, the computational imaging information acquisition unit 103 creates a difference image between the first image and the second image (step S119). By obtaining the difference image between the first image when the light emitting unit 109 is lit and the second image when it is turned off in this way, it is possible to obtain a PSF that is an image only of the lit light emitting unit 109 without being affected by other subjects in the environment.
[0096] Next, the computational imaging information acquisition unit 103 acquires the created difference image as the computational imaging information of the computational imaging camera 101 (step S120).
[0097] When using the PSF as the computational imaging information in this way, the computational imaging camera 101 captures two images, one of a scene where the light emitting unit 109 is lit and the other of a scene where it is turned off. At this time, it is desirable to capture the lit image and the turned-off image with as little time difference as possible.
[0098] Figures 8A to 8C are diagrams for explaining the creation process of the difference image. Figure 8A is an image captured by the computational imaging camera 101 when the light emitting unit 109 is lit. It can be seen that the luminance value of the light emitting unit 109 is high. Figure 8B is an image captured by the computational imaging camera 101 when the light emitting unit 109 is turned off. It can be seen that the luminance value of the light emitting unit 109 is lower compared to when it is lit. Figure 8C shows a difference image obtained by subtracting Figure 8B, which is an image captured by the computational imaging camera 101 when the light emitting unit 109 is turned off, from Figure 8A, which is an image captured by the computational imaging camera 101 when the light emitting unit 109 is lit. Since it is not affected by subjects other than the light emitting unit 109 and only the light emitting unit 109, which is a point light source, is captured, it can be seen that the PSF has been obtained.
[0099] Also, when using the LTM as the computational imaging information, a plurality of light emitting units 109 distributed in the environment can be used to obtain the PSF at multiple positions, and this can be used as the LTM.
[0100] FIG. 9 is a flowchart showing the main processing procedure of the computational imaging information acquisition unit 103 when using LTM as computational imaging information. First, a PSF corresponding to each light emitting unit 109 is acquired (step S301). As described above, this may be acquired using the difference image between when each light emitting unit 109 is lit and when it is turned off. By doing so, PSFs at a plurality of positions on the image can be acquired. FIG. 10 shows a schematic diagram of a plurality of PSFs acquired in this way. In the case of this example, PSFs are acquired at 6 points on the image.
[0101] The computational imaging information acquisition unit 103 calculates the PSF at all pixels of the image by performing interpolation processing on the plurality of PSFs thus acquired, and sets it as LTM (step S302). Such interpolation processing may utilize general image processing such as morphing. Also, the light emitting unit 109 may be the light of the user's smartphone or mobile phone. In this case, the user may turn the light emitting unit 109 on and off instead of the control unit 108.
[0102] Also, when using LTM as computational imaging information, instead of arranging a plurality of light emitting units 109, a small number of light emitting units 109 may be used and the position of the light emitting unit 109 may be changed by movement. For example, the light of a smartphone or mobile phone may be used as the light emitting unit 109, and the user may turn it on and off while changing the location. Or, an LED mounted on a moving body such as a drone or a cleaning robot may be used. Or, the computational imaging camera 101 may be installed on a moving body or the like, or the position of the light emitting unit 109 on the computational imaging image may be changed by the user changing the orientation or position.
[0103] According to the image recognition system 11 according to this embodiment, the computational imaging information regarding the computational imaging camera 101 (first camera) is a difference image between a first image including a point light source in the lit state and a second image including a point light source in the off state. Therefore, the computational imaging information regarding the actually used computational imaging camera 101 can be accurately acquired without being affected by subjects other than the point light source. As a result, a corrected image (fourth image) used for machine learning can be accurately generated based on the computational imaging information and a normal image (third image).
[0104] Further, by the control unit 108 of the learning device 21 controlling the operations of the light emitting unit 109 and the computational imaging camera 101, the timing of turning on or off the light emitting unit 109 and the timing of imaging by the computational imaging camera 101 can be accurately synchronized.
[0105] (Third Embodiment) FIG. 11 is a schematic diagram showing the configuration of an image recognition system 12 according to the third embodiment of the present disclosure. In FIG. 11, the same reference numerals are given to the same components as in FIG. 6, and the description thereof is omitted. The learning device 22 of the image recognition system 12 has a computational imaging information determination unit 110. The computational imaging information determination unit 110 determines the image quality state of the computational imaging information acquired by the computational imaging information acquisition unit 103. The learning device 22 switches the content of the process according to the determination result of the computational imaging information determination unit 110.
[0106] Further, FIG. 12 is a flowchart showing the main processing procedure of the image recognition system 12. The flowchart shows the flow of processing before and after the image quality determination process by the computational imaging information determination unit 110.
[0107] First, the computational imaging information acquisition unit 103 creates a difference image between the first image when the light emitting unit 109 is lit and the second image when it is off by the same method as in step S119 (FIG. 7) of the second embodiment (step S121).
[0108] Next, the calculation imaging information determination unit 110 determines whether the image quality of the difference image created by the calculation imaging information acquisition unit 103 is equal to or higher than the allowable value (step S122). Since only a point light source other than the point light source should be reflected in the PSF, the difference image between when the light is on and when the light is off is used. However, if there is a change in the scene such that a person moves significantly or the brightness in the environment changes dramatically between the shooting when the light is on and the shooting when the light is off, the change will appear in the difference image, and it will be impossible to obtain an accurate PSF. Therefore, the calculation imaging information determination unit 110 counts the number of pixels having a luminance of a certain value or more in the difference image, and determines that the image quality of the PSF is less than the allowable value when the number of pixels is equal to or more than the threshold value, and determines that the image quality of the PSF is equal to or higher than the allowable value when the number of pixels is less than the threshold value.
[0109] When the calculation imaging information determination unit 110 determines that the image quality of the difference image is less than the allowable value (step S122: NO), next, the control unit 108 gives an instruction to the light emitting unit 109 to emit light and turn off the light, and an instruction to re-shoot the calculation imaging camera 101 in order to perform re-shooting (step S123). On the other hand, when the calculation imaging information determination unit 110 determines that the image quality of the difference image is equal to or higher than the allowable value (step S122: YES), next, the database correction unit 104 corrects the learning database 102 using the calculation imaging information (PSF) acquired by the calculation imaging information acquisition unit 103 as the difference image (step S124).
[0110] Here, it is considered that one of the reasons for the deterioration of the image quality of the differential image is that the settings of the computational imaging camera 101 are not appropriate. For example, when the exposure time of the computational imaging camera 101 is too short or the gain of signal amplification is too small, the image becomes dark overall, and the luminance of the light-emitting unit 109 is buried in noise. Conversely, when the exposure time of the computational imaging camera 101 is too long or the gain of signal amplification is too large, the luminance value of the high-luminance region in the image exceeds the upper limit value of the sensing range and saturates, and the periphery of the light-emitting unit 109 becomes a so-called white-out state. Therefore, the computational imaging information determination unit 110 checks the maximum luminance value of each image when the light-emitting unit 109 is lit and when it is turned off, and if it exceeds the upper limit value or is less than the lower limit value (that is, outside the predetermined range), it may be determined that the image quality of the differential image is less than the allowable value. By determining the image quality of the differential image based on whether the maximum luminance value of the image when the light-emitting unit 109 is lit exceeds the upper limit value, it is possible to determine whether the luminance of the light-emitting unit 109 exceeds the sensing range and saturates. Also, by determining the image quality of the differential image based on whether the maximum luminance value of the image when the light-emitting unit 109 is lit is less than the lower limit value, it is possible to determine whether the luminance of the light-emitting unit 109 is buried in noise. Further, when it is determined that the luminance of the light-emitting unit 109 is saturated or buried in noise, the control unit 108 may control to change the settings of the computational imaging camera 101 so that the maximum luminance value is within the above-mentioned predetermined range in the re-capture.
[0111] FIG. 13 is a flowchart showing the main processing procedure of the image identification system 12. The flowchart shows the flow of processing before and after the image quality determination process by the computational imaging information determination unit 110.
[0112] First, the computational imaging information acquisition unit 103 acquires a first image captured by the computational imaging camera 101 when the light-emitting unit 109 is lit (step S131).
[0113] Next, the calculation imaging information determination unit 110 determines whether the luminance of the image is saturated by checking whether the maximum luminance value of the first image acquired by the calculation imaging information acquisition unit 103 exceeds the upper limit value Th1 (step S132).
[0114] If the maximum luminance value exceeds the upper limit value Th1, that is, if the luminance of the image is saturated (step S132: YES), then the control unit 108 instructs the calculation imaging camera 101 to perform imaging again with a shorter exposure time (step S133). On the other hand, if the maximum luminance value is less than or equal to the upper limit value Th1 (step S132: NO), then the calculation imaging information determination unit 110 determines whether the luminance of the light emitting unit 109 is buried in noise by checking whether the maximum luminance value of the first image acquired by the calculation imaging information acquisition unit 103 is less than the lower limit value Th2 (step S134).
[0115] If the maximum luminance value is less than the lower limit value Th2, that is, if the luminance of the light emitting unit 109 is buried in noise (step S134: YES), then the control unit 108 instructs the calculation imaging camera 101 to perform imaging again with a longer exposure time (step S135). On the other hand, if the maximum luminance value is greater than or equal to the lower limit value Th2 (step S134: NO), then the calculation imaging information determination unit 110 determines that the image quality of the first image acquired by the calculation imaging information acquisition unit 103 is sufficiently high with the current exposure time. In this case, the control unit 108 instructs the light emitting unit 109 to turn off, and also instructs the calculation imaging camera 101 to perform imaging with the current exposure time. As a result, the calculation imaging information acquisition unit 103 acquires a second image when the light emitting unit 109 is turned off (step S136). Note that the control unit 108 may also control the exposure time of the calculation imaging camera 101 so that the maximum luminance value is within a predetermined range for the acquired second image as well as for the first image.
[0116] Of course, the control unit 108 may change settings other than the exposure time of the calculation imaging camera 101. For example, the gain may be changed.
[0117] FIG. 14 is a flowchart showing the main processing procedure of the image recognition system 12. The flowchart shows the processing flow before and after the image quality determination process by the calculation imaging information determination unit 110.
[0118] In the determination of step S132, when the maximum luminance value exceeds the upper limit value Th1, that is, when the luminance of the image is saturated (step S132: YES), next, the control unit 108 instructs the calculation imaging camera 101 to reduce the gain and perform imaging again (step S137).
[0119] In the determination of step S134, when the maximum luminance value is less than the lower limit value Th2, that is, when the luminance of the light emitting unit 109 is buried in noise (step S134: YES), next, the control unit 108 instructs the calculation imaging camera 101 to increase the gain and perform imaging again (step S138).
[0120] Further, the control unit 108 may control the luminance of the light emitting unit 109 instead of the exposure time or the gain of the calculation imaging camera 101. That is, when it is determined by the calculation imaging information determination unit 110 that the luminance of the light emitting unit 109 is saturated, the control unit 108 controls the light emitting unit 109 to lower the luminance. Conversely, when it is determined by the calculation imaging information determination unit 110 that the luminance of the light emitting unit 109 is buried in noise, the control unit 108 controls the light emitting unit 109 to increase the luminance. By increasing the luminance of the light emitting unit 109, the luminance difference from the noise is widened.
[0121] Further, when it is determined by the calculation imaging information determination unit 110 that the image quality of the difference image is less than the allowable value, the control unit 108 may select another light emitting unit existing in the target area and instruct the other light emitting unit to emit light and turn off the light. This is effective in cases where the image quality inevitably deteriorates depending on the positional relationship between the calculation imaging camera 101 and the light emitting unit 109 in the case of a light source having directivity.
[0122] According to the image recognition system 12 according to this embodiment, when the image quality of the difference image is less than the allowable value, the control unit 108 performs re-imaging control by the calculation imaging camera 101, so that a difference image with the luminance value of the point light source appropriately adjusted can be obtained. As a result, it becomes possible to acquire appropriate calculation imaging information regarding the calculation imaging camera 101.
[0123] Also, in the re-imaging control, by the control unit 108 modifying at least one of the exposure time and the gain of the calculation imaging camera 101, it becomes possible to acquire a difference image with the luminance value of the point light source appropriately adjusted.
[0124] (Fourth Embodiment) FIG. 15 is a schematic diagram showing the configuration of an image recognition system 13 according to the fourth embodiment of the present disclosure. In FIG. 15, the same reference numerals are given to the same components as in FIG. 1, and the description thereof is omitted. The learning device 23 of the image recognition system 13 includes a storage unit 112 in which a plurality of learned image recognition models are stored, and a model selection unit 111 that selects one image recognition model from among the plurality of image recognition models. The learning device 23 of the image recognition system 13 does not have the learning unit 105 learn the learning database 102 corrected by the database correction unit 104, but has a model selection unit 111, and selects the optimal image recognition model corresponding to the calculation imaging information of the calculation imaging camera 101 from among the plurality of image recognition models learned in advance. For example, when a plurality of types of multi-pinhole masks 301a with different mask patterns are prepared in advance as described above, image recognition models learned using the captured images in the mounted state of each multi-pinhole mask 301a are created in advance, and the plurality of image recognition models are stored in the storage unit 112. The model selection unit 111 selects one image recognition model corresponding to the calculation imaging information of the calculation imaging camera 101 from among the plurality of image recognition models stored in the storage unit 112.
[0125] FIG. 16 is a flowchart showing the main processing procedure of the learning device 23 of the image identification system 13. The flowchart shows the flow of the process in which the model selection unit 111 selects an image identification model.
[0126] First, the calculation imaging information acquisition unit 103 acquires the calculation imaging information of the calculation imaging camera 101 (step S201).
[0127] Next, the model selection unit 111 selects one image identification model corresponding to the calculation imaging information acquired by the calculation imaging information acquisition unit 103 from among a plurality of image identification models stored in the storage unit 112 (step S211). For this, image identification models learned in advance with various calculation imaging information may be prepared, and the image identification model learned with the calculation imaging information closest to the calculation imaging information may be selected.
[0128] The image identification model thus selected is an image identification model that is compatible with the calculation imaging camera 101. The selected image identification model is set in the identification unit 106 as the image identification model to be used by the identification unit 106. By using the image identification model, the identification unit 106 can perform highly accurate identification processing.
[0129] According to the image identification system 13 according to the present embodiment, the learning device 23 selects one image identification model corresponding to the calculation imaging information of the calculation imaging camera 101 from among a plurality of learned image identification models. Therefore, since it is not necessary for the learning device 23 to newly perform learning, the processing load of the learning device 23 can be reduced, and the operation of the identification device 30 can be started earlier.
Industrial Applicability
[0130] The learning method and the identification method according to the present disclosure are particularly useful for an image identification system in an environment where protection of the privacy of a subject is required.
Claims
1. An information processing apparatus as a learning device acquires calculation imaging information regarding a first camera that captures a blurry image, wherein the calculation imaging information is a difference image between a first image including a point light source in a lit state and a second image including the point light source in a turned-off state, both images being captured by the first camera, acquires a third image captured by a second camera that captures a non-blurry image or an image with less blur than the first camera, and a correct label assigned to the third image, generates a fourth blurry image based on the calculation imaging information and the third image, and creates an image identification model for identifying an image captured by the first camera by performing machine learning using the fourth image and the correct label. A learning method.
2. The first camera is any one of: an encoded aperture camera including a mask having a mask pattern with different transmittance for each region; a multi-pinhole camera in which a mask having a mask pattern with a plurality of pinholes formed therein is disposed on a light receiving surface of an image sensor; and a light field camera that acquires a light field from a subject, according to the learning method of Claim 1.
3. The first camera does not have an optical system for forming an image of light from a subject on an image sensor, according to the learning method of Claim 1 or 2.
4. The mask is changeable to another mask having a different mask pattern, according to the learning method of Claim 2.
5. The calculation imaging information is any one of a Point Spread Function and a Light Transport Matrix, according to the learning method of any one of Claims 1 to 4.
6. The information processing apparatus controls the lighting of the point light source and the imaging control of the first image by the first camera, and controls the turning-off of the point light source and the imaging control of the second image by the first camera, according to the learning method of any one of Claims 1 to 5.
7. The information processing apparatus performs re-imaging control of the first image and the second image by the first camera when the image quality based on the luminance value of the difference image is less than an allowable value, according to the learning method of Claim 6.
8. In the learning method according to claim 7, in the re-imaging control, for each of the first image and the second image, the information processing apparatus corrects at least one of the exposure time and the gain of the first camera so that the maximum luminance value is within a predetermined range.
9. In an identification apparatus having an identification unit, an image captured by a first camera that captures a blurred image is input to the identification unit, the identification unit identifies the input image based on a learned image identification model, the identification unit outputs the result of the identification, wherein the image identification model is an image identification model created by the learning method according to any one of claims 1 to 8. An image identification method.
10. an acquisition unit that acquires calculation imaging information regarding a first camera that captures a blurred image, wherein the calculation imaging information is a difference image between a first image including a point light source in a lit state and a second image including the point light source in a turned-off state, both captured by the first camera, a storage unit that stores a third image captured by a second camera that captures an image without blur or an image with less blur than the first camera, and a correct label assigned to the third image, an image generation unit that generates a fourth blurred image based on the calculation imaging information acquired by the acquisition unit and the third image read from the storage unit, a learning unit that creates an image identification model for identifying an image captured by the first camera by performing machine learning using the fourth image generated by the image generation unit and the correct label read from the storage unit, A learning apparatus comprising.
11. an acquisition unit that acquires calculation imaging information regarding a first camera that captures a blurred image, wherein the calculation imaging information is a difference image between a first image including a point light source in a lit state and a second image including the point light source in a turned-off state, both captured by the first camera, a storage unit that stores a third image captured by a second camera that captures an image without blur or an image with less blur than the first camera, and a correct label assigned to the third image, an image generation unit that generates a fourth blurred image based on the calculation imaging information acquired by the acquisition unit and the third image read from the storage unit, A learning unit that creates an image identification model by performing machine learning using the fourth image generated by the image generation unit and the correct label read from the storage unit; An identification unit that identifies an image captured by the first camera based on the image identification model created by the learning unit; An output unit that outputs an identification result by the identification unit; An image identification system comprising the above components.
Citation Information
Patent Citations
Imaging device, control method therefor, program, and storage medium
JP2019118098A
Model learning system, model learning method, program and storage medium
JP2020095428A
Image generation device and image generation method
WO2019054092A1