Image processing device, image processing method, and program

The described lensless camera system uses a mask and DNN-based signal processing to achieve accurate image recognition with privacy protection by optimizing mask patterns and reducing computational load, addressing privacy and efficiency issues in existing lensless camera technologies.

JP7782461B2Active Publication Date: 2025-12-09SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022568162
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-08
Filing Date
2021-11-24
Publication Date
2025-12-09
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

Existing lensless camera technologies lack effective privacy protection measures, and their mask patterns are susceptible to image reconstruction by third parties, leading to potential privacy breaches. Additionally, they do not optimize image quality and require unnecessary computational resources for image reconstruction.

Method used

The implementation of a mask that modulates incident light, an image sensor to capture the modulated image, and a signal processing unit that applies signal processing based on the mask pattern, combined with a DNN for image recognition, where the mask pattern is optimized to minimize computational load and ensure privacy by making reconstructed images unrecognizable.

Benefits of technology

This configuration enables highly accurate image recognition while protecting privacy by ensuring reconstructed images are not visually recognizable, reducing computational resources, and simplifying the device configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007782461000011
    Figure 0007782461000011
  • Figure 0007782461000012
    Figure 0007782461000012
  • Figure 0007782461000013
    Figure 0007782461000013
Patent Text Reader

Abstract

The present disclosure relates to an image processing apparatus, an image processing method, and a program with which it is possible to perform a highly precise image recognition process with a simple configuration, while respecting privacy concerns. A weight of the first layer of a DNN used for an image recognition process using a DNN is convolved with the mask of a lensless camera so that an image reconstructed on the basis of a captured image captured by an image sensor is provided by a processing result for the first layer of the DNN. In this way, the recovered image is provided by the processing result for the first layer of the DNN, the image being difficult to be recognized as a scene or an object when viewed by a person. Thus, the privacy of the image can be protected. The present disclosure may be applied to image recognition apparatuses using a lensless camera.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image processing device, an image processing method, and a program, and in particular to an image processing device, an image processing method, and a program that are simple in configuration and enable highly accurate image recognition processing while respecting privacy. [Background technology]

[0002] A lensless camera has been proposed that does not have optical blocks such as lenses, but instead captures images by placing a mask in front of the image sensor and modulating the incident light, and then reconstructs images at various focal lengths by applying signal processing to the captured image according to the mask pattern.

[0003] Lensless cameras are proposed for use in image recognition processing because they can reconstruct images of subjects at various distances from an image obtained in a single capture.

[0004] To improve the recognition accuracy of image recognition processing, it is essential to improve the image quality of a reconstructed image reconstructed from an image captured by a lensless camera.

[0005] Therefore, a technology has been proposed that improves the quality of the reconstructed image by dividing the mask into multiple regions and individually preparing a mask design and bandpass filter that matches the target wavelength range for each region (see Patent Document 1).

[0006] Furthermore, in the image recognition process, when an image is used for person authentication processing, consideration must be given to the privacy of the person in the captured image.

[0007] For this reason, for example, in the case of a multi-pinhole array, a technology has been proposed that takes advantage of the feature that the image observed directly on the image sensor is not an image formed by the subject, thereby realizing imaging that takes privacy into consideration (see non-patent document 1). [Prior art documents] [Patent documents]

[0008] [Patent Document 1] International Publication No. 2019 / 124106 [Non-patent literature]

[0009] [Non-Patent Document 1] Action Recognition from a Lensless Multi-Pinhole Camera / Satoshi Sato, Changxin Zhou, Pongsak Rasang, Ikunori Ishii, Ryota Fujimura, Takayoshi Yamashita, 23rd Symposium on Image Recognition and Understanding (MIRU2020) Summary of the Invention [Problem to be solved by the invention]

[0010] However, in the technology of Non-Patent Document 1, it cannot be said that active measures for privacy protection have been implemented in the pinhole array pattern (mask pattern), and if the pattern itself becomes known to a third party, it will be possible to reconstruct the image by calibration, so it cannot be said that this is a sufficient privacy measure.

[0011] Furthermore, in the technology of Patent Document 1, no special measures are taken for the boundary areas between adjacent sub-areas, or only a light-shielding wall is provided, which may result in a deterioration in image quality, as well as making manufacturing more difficult and expensive.

[0012] Furthermore, the mask pattern itself is intended to reduce the influence of diffraction effects due to differences in wavelength, and therefore has not been specifically optimized for subsequent recognition processing.

[0013] Furthermore, when attempting to achieve image recognition processing using a lensless camera, it is common to reconstruct an image at a preliminary stage of processing, even though a reconstructed image is not essential for image recognition processing, which requires extra computing resources and power for this processing.

[0014] The present disclosure has been made in consideration of such circumstances, and aims to realize highly accurate image recognition processing that takes privacy into consideration with a simpler configuration, particularly in image recognition processing using a lensless camera. [Means for solving the problem]

[0015] An image processing device and a program according to one aspect of the present disclosure include an image processing device and a program that include a mask that modulates incident light and transmits it, an image sensor that captures a modulated image based on the incident light modulated by the mask, and a signal processing unit that performs signal processing on the modulated image based on the mask pattern of the mask.

[0016] An image processing method according to one aspect of the present disclosure is an image processing method for an image processing device that includes a mask that modulates incident light and transmits it, an image sensor that captures a modulated image based on the incident light modulated by the mask, and a signal processing unit that applies signal processing to the modulated image based on a mask pattern of the mask, wherein the signal processing unit applies signal processing to the modulated image based on the mask pattern of the mask.

[0017] In one aspect of the present disclosure, a modulated image based on incident light modulated by a mask is captured, and signal processing based on a mask pattern of the mask that modulates the incident light and transmits it is performed on the modulated image. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of an image processing device that realizes image recognition processing using lensless imaging. [Figure 2] 1A and 1B are diagrams illustrating the imaging principle of a lensless camera. [Figure 3] FIG. 1 is a diagram illustrating an image processing device that realizes image recognition processing using lensless imaging according to the present disclosure. [Figure 4] FIG. 10 is a diagram illustrating that in image recognition processing using VGG16, processing for obtaining feature vectors is performed in parallel for multiple channels. [Figure 5] FIG. 1 is a diagram illustrating a configuration example of a first embodiment of an image processing device that realizes image recognition processing using lensless imaging according to the present disclosure. [Figure 6] 10A and 10B are diagrams illustrating an example of a mask pattern set for each sub-area. [Figure 7] FIG. 10 is a diagram illustrating that a margin area is set between adjacent sub-areas. [Figure 8] 10A and 10B are diagrams illustrating how to determine the width of a margin area between adjacent sub-areas. [Figure 9] 10A and 10B are diagrams illustrating examples of mask patterns when adjacent sub-areas overlap each other. [Figure 10] 6 is a flowchart illustrating a recognition process performed by the image processing device of FIG. 5. [Figure 11] This figure explains an example of a mask pattern formed by superimposing, for all channels, patterns obtained by convolving a basic pattern with low correlation for each channel with the weights of the first layer of the DNN. [Figure 12] FIG. 10 is a diagram illustrating a configuration example of a second embodiment of an image processing device that realizes image recognition processing using lensless imaging according to the present disclosure. [Figure 13] 13 is a flowchart illustrating image recognition processing by the image processing device of FIG. 12. [Figure 14] FIG. 10 is a diagram illustrating a first application example of a mask pattern. [Figure 15] FIG. 10 is a diagram illustrating a second application example of a mask pattern. [Figure 16] FIG. 1 is a diagram illustrating an example of the configuration of a general-purpose personal computer. DETAILED DESCRIPTION OF THE INVENTION

[0019] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0020] Hereinafter, embodiments of the present technology will be described in the following order. 1. Overview of image recognition using a lensless camera 2. First Embodiment 3. Second Embodiment 4. First application example 5. Second application example 6. Software implementation example

[0021] <<1. Overview of Image Recognition Using a Lensless Camera>> In describing the image processing device of the present disclosure that uses images captured by a lensless camera to achieve highly accurate image recognition processing while respecting privacy, an overview of image recognition processing using a lensless camera will first be described with reference to FIG. 1.

[0022] Fig. 1 shows an example of the configuration of an image processing device for explaining an overview of image recognition processing using a lensless camera. The image processing device 11 in Fig. 1 includes a mask 31, an image sensor 32, a reconstruction unit 33, and a recognition processing unit 34.

[0023] The mask 31, the image sensor 32, and the reconstruction unit 33 constitute an imaging device that uses lensless imaging, and functions as a so-called lensless camera.

[0024] The mask 31 is a plate-like structure made of a light-shielding material and provided in front of the image sensor 32. For example, the mask 31 may be configured to have a hole-like opening that transmits incident light, a transparent region where a focusing element such as a lens or an FZP (Fresnel Zone Plate) is provided, and a light-shielding non-transparent region other than the hole-like opening that transmits incident light. The mask 31 may further include an intermediate transparent region having a given transmittance (non-binary) that is intermediate between the opening and the light-shielding region, or may be configured by a diffraction grating or the like.

[0025] When the mask 31 receives light from the subject surface (the surface from which radiant light is emitted from a three-dimensional subject in reality) as incident light, it transmits the incident light through the transparent region, modulating the incident light from the subject surface as a whole and converting it into modulated light, and the converted modulated light is received by the image sensor 32 to capture an image.

[0026] The image sensor 32 is composed of a CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures an image consisting of modulated light obtained by modulating incident light from the subject surface using a mask 31, and outputs an image consisting of pixel-by-pixel signals to the reconstruction unit 33 as a captured image.

[0027] The mask 31 is large enough to encompass at least the entire surface of the image sensor 32, and is basically configured so that the image sensor 32 receives only modulated light that has been modulated by passing through the mask 31.

[0028] The transmission area formed on the mask 31 is at least equal to or larger than the pixel size of the image sensor 32 (size equal to or larger than the pixel size). A minute gap is provided between the image sensor 32 and the mask 31.

[0029] <Principles of lensless camera imaging> Here, we will explain the principle of imaging with a lensless camera. For example, as shown in the upper left of Figure 2, assume that incident light from point light sources PA, PB, and PC on the subject plane passes through mask 31 and is received as light rays with light intensities a, b, and c at positions Pa, Pb, and Pc on image sensor 32, respectively.

[0030] 2, the detection sensitivity of each pixel has directivity according to the angle of incidence as a result of the incident light being modulated by the transmission areas set in the mask 31. Giving the detection sensitivity of each pixel incident angle directivity here means that the light receiving sensitivity characteristics according to the angle of incidence of incident light vary depending on the area on the image sensor 32.

[0031] That is, assuming that the light sources constituting the subject plane are point light sources, light rays of the same light intensity emitted from the same point light source are incident on the image sensor 32, but the angle of incidence changes for each region on the imaging surface of the image sensor 32 due to modulation by the mask 31. The mask 31 changes the angle of incidence of the incident light depending on the region on the image sensor 32, thereby providing light receiving sensitivity characteristics, i.e., incident angle directivity. Therefore, even light rays of the same light intensity are detected with different sensitivities for each region on the image sensor 32 by the mask 31 provided in front of the imaging surface of the image sensor 32, and detection signals of different detection signal levels are detected for each region.

[0032] More specifically, as shown in the upper right part of Fig. 2, the detection signal levels DA, DB, and DC of pixels at positions Pa, Pb, and Pc on the image sensor 32 are respectively expressed by the following equations (1) to (3). Note that the equations (1) to (3) in Fig. 2 are upside down relative to the positions Pa, Pb, and Pc on the image sensor 32 in Fig. 2.

[0033]

number

number

number

[0034] Here, α1 is a coefficient for the detection signal level a that is set according to the angle of incidence of a ray from a point light source PA on the object plane to be restored at the position Pa on the image sensor 32.

[0035] Also, β1 is a coefficient for the detection signal level b that is set according to the incident angle of the light ray from the point light source PB on the object plane to be restored at the position Pa on the image sensor 32.

[0036] Furthermore, γ1 is a coefficient for the detection signal level c that is set according to the incident angle of the light ray from the point light source PC on the object plane to be restored at the position Pa on the image sensor 32.

[0037] Therefore, (α1×a) of the detection signal level DA indicates the detection signal level due to the light beam from the point light source PA at the position Pa.

[0038] Furthermore, (β1×b) of the detection signal level DA indicates the detection signal level due to the light beam from the point light source PB at the position Pa.

[0039] Furthermore, (γ1×c) of the detection signal level DA indicates the detection signal level due to the light beam from the point light source PC at the position Pa.

[0040] Therefore, the detection signal level DA is expressed as a composite value obtained by multiplying each component of the point light sources PA, PB, and PC at position Pa by the respective coefficients α1, β1, and γ1. Hereinafter, the coefficients α1, β1, and γ1 will be collectively referred to as the coefficient set.

[0041] Similarly, for the detection signal level DB at point light source Pb, the coefficient set α2, β2, γ2 corresponds to the coefficient set α1, β1, γ1 for the detection signal level DA at point light source PA. Also, for the detection signal level DC at point light source Pc, the coefficient set α3, β3, γ3 corresponds to the coefficient set α1, β1, γ1 for the detection signal level DA at point light source Pa.

[0042] However, the detection signal levels of the pixels at positions Pa, Pb, and Pc are values ​​expressed as the sum of the products of the light intensities a, b, and c of the light rays emitted from the point light sources PA, PB, and PC, respectively, and the coefficients. Therefore, these detection signal levels are a mixture of the light intensities a, b, and c of the light rays emitted from the point light sources PA, PB, and PC, and are therefore different from the image of the subject that is formed.

[0043] That is, by forming simultaneous equations using the coefficient set α1, β1, γ1, coefficient set α2, β2, γ2, coefficient set α3, β3, γ3, and the detection signal levels DA, DB, DC, and solving the light intensities a, b, c, the pixel values ​​at each position Pa, Pb, Pc are obtained as shown in the lower right of Figure 2. This reconstructs and restores the restored image (final image), which is a collection of pixel values.

[0044] Furthermore, when the distance between the image sensor 32 shown in the upper left of Figure 2 and the subject plane changes, the coefficient sets α1, β1, γ1, α2, β2, γ2, and α3, β3, γ3 will change, respectively. By changing these coefficient sets, it is possible to reconstruct restored images (final images) of subject planes at various distances.

[0045] Therefore, by changing the coefficient set to correspond to various distances in one imaging operation, it is possible to reconstruct images of object planes at various distances from the imaging position.

[0046] As a result, in imaging using a lensless camera realized by the mask 31, image sensor 32, and reconstruction unit 33 in Figure 1, there is no need to be aware of the phenomenon of so-called defocus, which occurs when imaging with a general imaging device that uses a lens, and as long as the image is captured so that the subject to be captured is included in the field of view, images of the subject plane at various distances can be reconstructed after imaging by changing the coefficient set according to the distance.

[0047] The detection signal level shown in the upper right of Fig. 2 is not a detection signal level corresponding to an image in which the image of the subject is formed, and is therefore not a pixel value but a mere observation value, resulting in an image made up of observation values. Also, the detection signal level shown in the lower right of Fig. 2 is made up of signal values ​​for each pixel corresponding to an image in which the image of the subject is formed, and is the value of each pixel in the restored image (final image), so is a pixel value.

[0048] With this configuration, the mask 31, the image sensor 32, and the reconstruction unit 33 can function as a so-called lensless camera. As a result, an imaging lens is not an essential component, so it is possible to reduce the height of the imaging device, that is, the thickness in the direction of light incidence in the configuration that realizes the imaging function. Also, by changing the coefficient set in various ways, it is possible to reconstruct and restore final images (reconstructed images) on object planes at various distances.

[0049] Hereinafter, an image captured by the image sensor 32 before reconstruction will be simply referred to as a captured image, and an image reconstructed and restored by signal processing of the captured image will be referred to as a final image (restored image). Therefore, by changing the above-mentioned coefficient set in various ways from a single captured image, images on the object plane at various distances can be reconstructed as final images.

[0050] Returning now to the explanation of Fig. 1, the reconstruction unit 33 is provided with the coefficient set described above, and uses the coefficient set according to the distance from the imaging position to the object plane to reconstruct a final image (reconstructed image) based on the captured image captured by the image sensor 32, and outputs the reconstructed image to the recognition processing unit 34.

[0051] The recognition processing unit 34 performs image recognition processing using a DNN (Deep Neural Network) based on the final image supplied from the reconstruction unit 33, and outputs the recognition result.

[0052] More specifically, the recognition processing unit 34 includes a DNN first layer processing unit 51-1, a DNN second layer processing unit 51-2, . . . a DNN nth layer processing unit 51-n, and a recognition unit 52.

[0053] The DNN first layer processing unit 51-1, the DNN second layer processing unit 51-2, ..., the DNN nth layer processing unit 51-n each perform convolution processing for each of the n layers that make up the DNN and output the results to the subsequent stage, and the DNN nth layer processing unit 51-n, which has performed processing on the nth layer, which is the final layer, outputs the convolution result to the recognition unit 52.

[0054] The recognition unit 52 recognizes an object based on the convolution results of the nth layer supplied from the DNN nth layer processing unit 51-n, and outputs the recognition results.

[0055] (Recognition processing by the recognition processing unit) Here, an input image when the distance from the imaging position to the subject is a predetermined distance is defined as input image X, and a captured image, which is an image modulated by mask 31 of mask pattern A and observed by image sensor 32 while ignoring the effects of diffraction, noise, etc., is expressed as A*X. Note that * represents convolution.

[0056] Furthermore, when the restoration matrix consisting of the above-mentioned coefficient set for reconstructing the captured image A*X modulated by the mask 31 of the mask pattern A into a final image (reconstructed image) X' is defined as G, the reconstructed final image (reconstructed image) X' is expressed by the following equation (4).

[0057]

number

[0058] When the DNN first layer processing unit 51-1 performs a convolution operation with weight P1 on the final image (restored image) X', the processing result of the DNN first layer processing unit 51-1, which is the processing result of the DNN first layer, is expressed by the following equation (5).

[0059]

number

[0060] Similarly, in the DNN second layer processing unit 51-2, a convolution operation with weight P2 is performed on the processing result of the DNN first layer processing unit 51-1, in the DNN third layer processing unit 51-3, a convolution operation with weight P3 is performed on the processing result of the DNN first layer processing unit 51-2, and thereafter in the DNN nth layer processing unit 51-n, a convolution operation with weight Pn is performed on the processing result of the DNN (n-1)th layer processing unit 51-(n-1).

[0061] Finally, as the processing result of the DNN nth layer processing unit 51-n, the processing result P1*P2*···Pn*X' is output to the recognition unit 52, and the recognition unit 52 performs object recognition processing based on the processing result P1*P2*···Pn*X'.

[0062] It is assumed here that the mask pattern of the mask 31 is a MURA (Modified Uniformly Redundant Arrays) pattern, but other mask patterns may also be used, such as a URA (Uniformly Redundant Arrays) pattern. These patterns have a repeating structure called a cyclic coded mask, and are often used in the world of coded aperture imaging, which uses a mask to reconstruct a scene.

[0063] For details on the MURA pattern, see E.E. Fenimore and T.M. Cannon, “Coded aperture imaging with uniformly redundant arrays,” Applied Optics, vol. 17, no. 3, pp. 337-347, 1978.

[0064] For details of the URA pattern, see S.R. Gottesman and E. Fenimore, “New family of binary arrays for coded aperture imaging,” Applied Optics, vol. 28, no. 20, pp. 4344-4352, 1989.

[0065] As described above, in the object recognition process using the lensless camera, a final image (restored image) X' corresponding to the input image X is reconstructed in the stage preceding the recognition processing unit 34, and then the recognition process is performed in the recognition processing unit 34.

[0066] As described above, the final image (restored image) is obtained by restoring the captured image A*X modulated by the mask 31 using the restoration matrix G to an image that can be visually recognized by humans.

[0067] However, in the processing of the recognition processing unit 34, the image that is input to the recognition processing does not necessarily have to be an image that can be visually recognized by humans, as long as it contains information necessary for the recognition processing.

[0068] Therefore, in the recognition process using a lensless camera, the process of restoring the image to a final image (restored image) by the reconstructor 33 cannot be said to be essential.

[0069] Furthermore, when the recognition processing by the recognition processing unit 34 is, for example, a person's facial image such as face authentication, the input image will inevitably be a person's facial image, and as the final image (restored image) is reconstructed, consideration must be given to privacy protection.

[0070] For this reason, facial images that become the final reconstructed images (restored images) must be discarded after being used in the recognition process, or must be specially managed to prevent them from being exposed to third parties in some way to protect privacy.

[0071] Therefore, in the present disclosure, the mask pattern constituting the mask 31 is obtained by convolving the above-mentioned mask pattern A with a weight P1 used in the convolution operation of the DNN first layer processing unit 51-1, and in the reconstruction unit 33, an image P1*X' obtained by convolving the weight P1 with the input image X is reconstructed as the final image (reconstructed image).

[0072] With this configuration, the convolution calculation of the weight P1 by the DNN first layer processing unit 51-1 in the recognition processing unit 34 can be replaced with optical processing using the mask pattern of the mask 31.

[0073] As a result, it is possible to simplify the device configuration of the recognition processing unit 34, and along with the simplification of the device configuration, it is possible to reduce the processing load on the image processing device 101 as a whole.

[0074] Furthermore, the image reconstructed by the reconstruction unit 33 is an image P1*X' obtained by convolving a weight P1 with a final image (reconstructed image) X' corresponding to the input image X, making it difficult for a human to visually recognize the content, but it can be an image that contains information essential for the recognition process. Even if the final image and mask pattern A are stolen and the corresponding restoration matrix G is identified, the reconstructed image contains information essential for the recognition process, but is not an image that a human can visually recognize as a scene; in other words, it is two-dimensional information that contains only information essential for the recognition process.

[0075] As a result, privacy protection is possible for the reconstructed image, and even if the reconstructed image is stolen, it cannot be visually recognized, so there is no need for special management to prevent it from being exposed to third parties.

[0076] <<2. First Embodiment>> Next, a configuration example of the first embodiment of the image processing device of the present disclosure will be described with reference to the block diagram of FIG.

[0077] The image processing device 101 in FIG. 3 includes a mask 131, an image sensor 132, a reconstruction unit 133, and a recognition processing unit .

[0078] The image sensor 132 and the reconstruction unit 133 have the same configuration as the image sensor 32 and the reconstruction unit 33 in FIG. 1, respectively.

[0079] The basic function of the mask 131 is the same as that of the mask 31 in Fig. 1, but the mask pattern is obtained by convolving the mask pattern A in Fig. 1 with a weight P1 used in the convolution operation of the DNN first layer processing unit 51-1 in Fig. 1. Hereinafter, the mask pattern of the mask 131 will be expressed as a mask pattern A*P1.

[0080] The basic functions of the recognition processing unit 134 are similar to those of the recognition processing unit 34 in FIG. 1, but the recognition processing unit 134 does not have a configuration corresponding to the DNN first layer processing unit 51-1 in FIG. 1 that realizes the processing of the first layer of the DNN.

[0081] That is, the recognition processing unit 134 includes a DNN second layer processing unit 151-2 to a DNN n-th layer processing unit 151-n, and a recognition unit 152. The functions of the DNN second layer processing unit 151-2 to the DNN n-th layer processing unit 151-n and the recognition unit 152 are similar to those of the DNN second layer processing unit 51-2 to the DNN n-th layer processing unit 51-n and the recognition unit 52 in FIG.

[0082] With this configuration, modulation is applied to the input image X by the mask 131, and the captured image captured by the image sensor 132 is expressed as A*P1*X.

[0083] As a result, the reconstruction unit 133 multiplies the captured image A*P1*X by the restoration matrix G, and outputs an image as shown in the following equation (6) to the recognition processing unit 134 as a reconstructed image.

[0084]

number

[0085] That is, the final image (restored image) reconstructed in the reconstruction unit 133 expressed by equation (6) is an image P1*X' obtained by convolving a weight P1 with the final image (restored image) X' corresponding to the input image X, which is the processing result in the DNN first layer processing unit 51-1.

[0086] As a result, in the image processing device 101 of FIG. 3, the convolution calculation of the weight P1 by the DNN first layer processing unit 51-1 in the recognition processing unit 34 of FIG. 1 is replaced with optical processing using the mask pattern of the mask 131.

[0087] As a result, it is possible to simplify the device configuration of the recognition processing unit 134, and the processing load can be reduced along with the simplification of the device configuration.

[0088] Furthermore, the image reconstructed by the reconstruction unit 133 is an image P1*X' obtained by convolving a weight P1 with the final image X' corresponding to the input image X, and the information necessary for the recognition process is recorded as a captured image by the image sensor 132.

[0089] As a result, it is possible to protect privacy, and even if the final image is stolen, since it cannot be visually recognized, there is no need for special management to protect privacy, such as managing it so that it is not exposed to third parties.

[0090] <Multiple channels> Incidentally, in the image processing device 101 that realizes the above-mentioned recognition processing, the processing of the first layer of the DNN multi-layer hierarchical processing is generally performed in parallel by multiple channels, and subsequent object recognition is generally performed based on multiple feature vectors obtained in each channel.

[0091] More specifically, for example, when realizing object recognition processing using VGG16, which is used for image recognition tasks, the processing of the first layer of the DNN is performed using multiple channels, as shown in FIG. 4 .

[0092] As shown in Figure 4, when the input image Pin is a 224x224 resolution image with 3 RGB channels, the first layer processing outputs a 64-channel convolution result (224x224x64) (convolutional ReLU (Rectified Linear Unit)).

[0093] Using this result, Maxpooling reduces the layer to a 112x112 resolution, and then 128 channels of convolution are performed on the 112x112 resolution (112x112x128).

[0094] Next, using this result, the layer is reduced to 56x56 resolution using Maxpooling, and then 256 channels of convolution are performed on the 56x56 resolution (56x56x256).

[0095] Furthermore, using this result, the layer is reduced to 28x28 resolution using Maxpooling, and subsequently, 512-channel convolution is performed on the 28x28 resolution (28x28x512).

[0096] Furthermore, using this result, the layer is reduced to 14x14 resolution using Maxpooling, and subsequently, 512-channel convolution is performed on the 14x14 resolution (14x14x512).

[0097] In addition, using this result, the layer is reduced to 7x7 resolution using Maxpooling, and subsequently, 512 channels of convolution are performed on the 7x7 resolution (7x7x512).

[0098] Then, based on the convolution result of 512 channels at 7x7 resolution (7x7x512), the feature vector judgment result (fully connected+ReLU) (1x1x4096) is generated.

[0099] Then, a probability function (softmax) based on the feature vector determination result (fully connected + ReLU) (1 x 1 x 4096) is output as the recognition result.

[0100] <Multi-channel image recognition device> Therefore, the image recognition device of the present disclosure also needs to have a multi-channel configuration according to the number of channels in the first layer.

[0101] That is, the image processing device corresponding to multiple channels of the present disclosure is configured as shown in, for example, the image processing device 111 of FIG. 5, in which the image processing device 101 described with reference to FIG. 3 is made into a multi-channel device.

[0102] 5 includes a mask 131′, an image sensor 132, a division unit 161, Ch1 reconstruction units 162-1 to Chm reconstruction units 162-m, and a recognition processing unit 134′. Note that the following description will be given on the assumption that the number of channels in the first layer is m.

[0103] The mask 131' has the same basic function as the mask 131, but is configured as a mask pattern B divided into sub-areas equal to the number of channels.

[0104] That is, the mask 131' is divided into sub-areas SA1 to SAm corresponding to the number of channels=m, for example, as shown in FIG.

[0105] The mask pattern B that covers the entirety of each subarea SA1 to SAm consists of mask pattern B (=As*Ps11+As*Ps21+···As*Psm1) in which each of the mask patterns As, which is the basic pattern in each subarea of ​​channels 1 to m, is convolved with weight Psm1 used for the convolution calculation in the DNN first layer processing unit 51-1 in Figure 1.

[0106] In addition, the "+" in As*Ps11+As*Ps21+···As*Psm1 representing mask pattern B means that it is placed at a different position within a two-dimensional plane consisting of the incident surface of the incident light on mask 131', and is not intended to be superimposed with respect to the incident direction of the incident light.

[0107] The mask pattern As for each sub-area, which is the basic pattern, is a mask pattern used in so-called lensless cameras, such as a MURA pattern or an M sequence, and is a binary pattern.

[0108] That is, in FIG. 6, as an example, the mask pattern of the sub-area SA1 is shown to be made up of As*Ps11, and the mask pattern of the sub-area SA2 is shown to be made up of As*Ps21.

[0109] <Gap between adjacent sub-areas (if margin area is provided)> Between the sub-areas, margin areas BL are formed, which are light-shielding areas indicated by white areas in FIG. 6, in order to suppress interference of images between adjacent areas.

[0110] 7 is a side cross-sectional view of the mask 131′ and the image sensor 132 viewed from a direction perpendicular to the incident direction of the incident light. In FIG. 7, the downward direction in the drawing is the incident direction of the incident light, and the range of width 2×wb between the sub-areas SA1 and SA2 on the mask 131′ is the margin region BL.

[0111] As shown in the upper part of Figure 7, since there is a field of view for each pixel on the imaging surface of image sensor 132, the effective area on image sensor 132 where an image can be captured for sub-area SA1 of mask 131' is designated as effective area EA1, and the effective area on image sensor 132 where an image can be captured for sub-area SA2 is designated as effective area EA2.

[0112] Therefore, an area Dz1 between the effective areas EA1 and EA2 becomes a dead zone where imaging is not possible.

[0113] The width wb of the margin area BL is, for example, as shown in Figure 8, when the FOV (Field of View) of each pixel of the image sensor 132 is 2θ, the size required to prevent the mask pattern of an adjacent sub-area from entering that range.

[0114] In a typical lensless camera, the relationship between the mask size wm (the distance from the center of a rectangular sub-area in mask 131′ to the edge of the sub-area) and the sensor size ws (the distance from the center of the effective area of ​​a rectangular sub-area in image sensor 132 to the edge of the effective area) is wm > ws.

[0115] Therefore, it is sufficient to consider the relationship between the FOV of the pixels on the edge of the effective area of ​​the image sensor 132 and the sub-area of ​​the mask 131′ as shown in Figure 8, and the minimum required size wb of the margin area is expressed by the following equation (7).

[0116]

number

[0117] In the above, we have described an example in which a margin area having a width of 2×wb is set as a light-shielding area between subareas on the mask 131′. However, since the margin area need only be wider than the width 2×wb, a margin area BL′ having a width wc may be set between the width wb, as shown in the lower part of FIG. 7. In this case, a dead zone that cannot be imaged is set between the effective areas EA1 and EA2, consisting of an area Dz1′ wider than the area Dz1 in the upper part of FIG. 7. Furthermore, a light-shielding wall perpendicular to the incident direction of the incident light may be formed at the edge of each subarea to suppress interference between adjacent subareas. This reduces the margin area and increases the effective area of ​​the image sensor 132.

[0118] <Gap between adjacent sub-areas (when configured by overlapping)> Furthermore, an example has been described in which margin regions BL are provided between the sub-areas described with reference to Figures 6 to 8. However, in such a configuration, a dead zone is created on the image sensor 132, reducing the amount of information in the input image X and resulting in a reduced S / N ratio.

[0119] Therefore, by devising the mask pattern As, which is the basic pattern used to generate the mask pattern for each sub-area, it is possible to reduce the influence of interference between adjacent sub-areas and make it easier to separate channels by signal processing.

[0120] For example, by creating a mask pattern As that serves as a basic pattern for each channel, i.e., for each subarea, using pseudo-random signals with low correlation obtained from M sequences or Gold codes, even if signals from adjacent subareas are mixed into the observed values, the correlation with the basic pattern for each channel can be reduced, thereby minimizing the impact of interference between adjacent subareas. Ideally, there would be no correlation between the basic patterns for each channel and adjacent subareas, but the pseudo-random signals are used to minimize the correlation as much as possible.

[0121] By applying this principle, sub-areas can be set on the mask 131′ so as to prevent the dead zones that occur on the image sensor 132 described with reference to FIG. 7 from occurring. In this case, the sub-areas may be in complete contact with each other.

[0122] Furthermore, since there is low correlation between the mask patterns As that are the basic patterns for each sub-area, it is possible to arrange the mask patterns of adjacent sub-areas so that they overlap, as shown in Figure 9, as long as the effective areas on the image sensor 132 do not overlap.

[0123] In FIG. 9, effective areas EA1 and EA2 in which images of adjacent sub-areas are captured are set on the image sensor 132, but they overlap with the corresponding sub-areas SA11 and SA12 on the mask 131'.

[0124] Even with this configuration, if the correlation between the basic mask patterns of the sub-areas SA11 and SA12 is low, for example, if they are orthogonal to each other, there will be no mutual interference in the images captured in the effective areas EA1 and EA2 on the image sensor 132.

[0125] With this configuration, no dead zone is set on the image sensor 132, so it is possible to suppress a reduction in the amount of information in the input image X and improve the S / N ratio.

[0126] 5, the image sensor 132 captures a captured image B*X based on incident light modulated by a mask 131′ having a mask pattern B for the input image X, and outputs the captured image B*X to the division unit 161.

[0127] The dividing unit 161 divides the captured image B*X into sub-areas SA1 to SAm, and outputs the divided images to the Ch1 reconstructing units 162-1 to Chm reconstructing units 162-m of the corresponding channels.

[0128] That is, the division unit 161 outputs, from the captured image B*X, the captured image As*Ps11*X of the sub-area SA1 corresponding to channel 1 to the Ch1 reconstruction unit 162-1, outputs, from the captured image B*X, the captured image As*Ps21*X of the sub-area SA2 corresponding to channel 2 to the Ch2 reconstruction unit 162-2, and... outputs, from the captured image B*X, the captured image As*Psm1*X of the sub-area SAm corresponding to channel m to the Chm reconstruction unit 162-m.

[0129] The Ch1 reconstruction units 162-1 to Chm reconstruction units 162-m reconstruct images by multiplying the captured images As*Ps11*X to As*Psm1*X of the sub-areas SA1 to SAm of each channel by the restoration matrix Gs corresponding to the mask pattern As of the low-resolution image corresponding to the sub-area, and output the reconstructed images to the recognition processing unit 134' for each channel.

[0130] That is, the Ch1 reconstruction unit 162-1 of channel 1 reconstructs the image Ps11*X' processed by the DNN first layer processing unit 52-1 in Figure 1 by multiplying the captured image As*Ps11*X of subarea SA1 by the reconstruction matrix Gs corresponding to the mask pattern As, and outputs it to the Ch1 second layer processing unit 151-1-2, Ch2 second layer processing unit 151-2-2, ..., Chm2 second layer processing unit 151-m2-2 of the recognition processing unit 134'.

[0131] Similarly, the Ch2 reconstruction unit 162-2 of channel 2 multiplies the captured image As*Ps21*X of subarea SA2 by the restoration matrix Gs corresponding to the mask pattern As to restore the image Ps21*X' processed by the DNN first layer processing unit 52-1 in Figure 1, and outputs it to the Ch1 second layer processing unit 151-1-2, Ch2 second layer processing unit 151-2-2, ..., Chm2 second layer processing unit 151-m2-2 of the recognition processing unit 134'.

[0132] Furthermore, the Chm reconstruction unit 162-m of channel m reconstructs the image Psm1*X' processed by the DNN first layer processing unit 52-1 in Figure 1 by multiplying the captured image As*Psm1*X of the subarea SAm by the restoration matrix Gs corresponding to the mask pattern As, and outputs it to the Ch1 second layer processing unit 151-1-2, Ch2 second layer processing unit 151-2-2, ..., Chm2 second layer processing unit 151-m2-2 of the recognition processing unit 134'.

[0133] The recognition processing unit 134' includes Ch1 second layer processing units 151-1-2 to Chm2 second layer processing units 151-m2-2, Ch1 third layer processing units 151-1-3 to Chm3 third layer processing units 151-m3-3, Ch1 n-th layer processing units 151-1-n to Chm n n-th layer processing unit 151-m n -n is provided.

[0134] That is, the recognition processing unit 134′ has a configuration corresponding to the DNN second layer processing unit 151-2 to the DNN n-th layer processing unit 151-n in the recognition processing unit 134 of FIG. 3 for a plurality of channels.

[0135] More specifically, Ch1 second layer processing units 151-1-2 to Chm2 second layer processing units 151-m2-2, Ch1 third layer processing units 151-1-3 to Chm3 third layer processing units 151-m3-3, Ch1 nth layer processing units 151-1-n to Chm n n-th layer processing unit 151-m n-n sequentially convolves the weights with the reconstructed image Ps11*X' of channel 1, the reconstructed image Ps21*X' of channel 2, ..., and the reconstructed image Psm1*X' of channel m to calculate feature vectors and outputs them to the recognition unit 152'.

[0136] The recognition unit 152' is a final stage of the recognition processing unit 134', which includes Ch1 n-th layer processing units 151-1-n to Chm n n-th layer processing unit 151-m n Based on the feature vector, which is the result of the convolution operation supplied from -n, the object in the subject of the input image X is recognized.

[0137] That is, the Ch1 reconstruction units 162-1 to Chm reconstruction units 162-m multiply the captured images As*Ps11*X to As*Psm1*X of the subareas SA1 to SAm of each channel by the restoration matrix Gs corresponding to the mask pattern As, and output the processing results by the DNN first layer processing unit 52-1 in Figure 1, i.e., the images Ps11*X' to Psm1*X', which are the processing results of the DNN first layer, to the recognition processing unit 134'.

[0138] Therefore, the reconstructed images Ps11*X' to Psm1*X' are output as the processing results of the first layer of DNN processing in each channel on low-resolution restored images obtained by dividing the image sensor 132 into sub-areas for each channel.

[0139] Therefore, the restored images reconstructed in each channel contain the information necessary for recognition processing, but are output to the recognition processing unit 134' in an image state that makes it difficult for humans to visually recognize the scene or object, thereby protecting privacy. Also, because privacy is protected for the restored images, there is no need for special management that takes privacy into consideration.

[0140] In addition, since the processing of the first layer of the DNN in each channel is transferred to optical processing in the mask 131', there is no need for a configuration corresponding to the DNN first layer processing unit 52-1 in Figure 1 for each channel, which makes it possible to simplify the device configuration.

[0141] Furthermore, since a configuration corresponding to the DNN first layer processing unit 52-1 is not required, the processing load on the image processing device 111 can be reduced.

[0142] <Recognition processing by the image processing device in Figure 5> Next, the image recognition process performed by the image processing device 111 in FIG. 5 will be described with reference to the flowchart in FIG.

[0143] In step S11, the image sensor 132 photoelectrically converts the input image X, which is light from the subject modulated by the mask 131' of the mask pattern B in which the weights related to the processing of the first layer of the DNN for each channel are convolved with the mask pattern As, which is the basic pattern for each sub-area, to generate the captured image B*X and output it to the division unit 161.

[0144] In step S12, the division unit 161 divides the captured image B*X into restored images As*Ps11*X to As*Psm1*X for each sub-area corresponding to each channel, and outputs them to the Ch1 reconstruction units 162-1 to Chm reconstruction units 162-m of the corresponding channels.

[0145] In step S13, the Ch1 reconstruction units 162-1 to 162-m multiply the images As*Ps11*X to As*Psm1*X by the restoration matrix Gs for each channel to obtain restored images consisting of the processing results of the first layer of the DNN, and output them to the Ch1 second layer processing units 151-1-2 to Chm2 second layer processing units 151-m2-2 of the recognition processing unit 134'.

[0146] In step S14, the Ch1 second layer processing unit 151-1-2 to the Chm2 second layer processing unit 151-m2-2, the Ch1 third layer processing unit 151-1-3 to the Chm3 third layer processing unit 151-m3-3, ..., the Ch1 n-th layer processing unit 151-1-n to the Chm n n-th layer processing unit 151-m n -n executes the processing of the second layer and subsequent layers of the DNN for each channel, obtains a feature vector, and outputs it to the recognition unit 152′.

[0147] In step S15, the recognition unit 152' executes recognition processing of the object based on whether or not the information on the feature vector of each channel matches the feature vector of the object, and outputs the recognition result.

[0148] Through the above processing, the configuration functions as a lensless camera, and the restored images reconstructed in each channel are output to the recognition processing unit 134' in a state where they are difficult to recognize as objects with the human eye. As a result, privacy protection can be applied to the restored images, and special management with consideration for privacy protection is not required when managing the restored images.

[0149] In addition, since the processing of the first layer of the DNN in Figure 1 for each channel is transferred to optical processing in mask 131', there is no need for a configuration to realize the processing of the first layer of the DNN in Figure 1 for each channel, which makes it possible to simplify the device configuration.

[0150] Furthermore, since the processing of the first layer of the DNN is eliminated, the processing load on the image processing device 111 can be reduced.

[0151] In the above, we have described an example in which the division unit 161 divides the captured image captured by the image sensor 132 into channels using a mask pattern in which sub-areas for each channel are provided in the mask 131' and weights related to the processing of the first layer of DNN are convolved, and the divided images are output to the Ch1 reconstruction unit 162-1 to Chm reconstruction unit 162-m of each channel.

[0152] However, the splitting unit 161 may be omitted by having the Ch1 reconstruction unit 162-1 to Chm reconstruction unit 162-m of each channel extract and process the information of the sub-area required for processing its own channel from the captured image B*X.

[0153] (About learning mask patterns) In generating mask patterns to ensure privacy protection, the recognition processing network is trained to generate mask patterns that simultaneously satisfy two different criteria.

[0154] Here, the mask pattern is generated from the weights used in the processing of the first layer of the DNN, and only the weights of the first layer of the DNN are the target of optimization using two different types of indicators.

[0155] The first index is a loss function that directly defines the performance of the recognition process, which is the degree of deviation of the ID classification and recognition results from the correct data. The first layer and the second and subsequent layers of the recognition process are optimized to minimize this loss function.

[0156] The second index is an index for ensuring privacy protection, that is, an index for ensuring that the final reconstructed image (restored image) is an image that is difficult for humans to recognize as a scene or object by visual inspection. To achieve this, a network is separately trained to perform image restoration processing, which restores the modulated image generated when the original image is modulated by the mask pattern, assuming that the mask pattern itself is known. This image restoration processing network is trained to minimize the difference between the original image and the modulated image.

[0157] On the other hand, by evaluating the difference between the image output by the image restoration processing network and the original image using indices that express image similarity such as PSNR, MSE, SSIM, and VGG16, and training the weights in the first layer of the DNN so that the evaluation value is low, it is possible to create a system that can perform recognition processing while ensuring privacy protection.

[0158] In other words, in the case of facial images, the weights of the first layer of the DNN are determined through learning to create a mask pattern that leaves only the information necessary for recognition processing from the original facial image, and performs modulation processing to minimize the information necessary to restore the facial image.

[0159] By reflecting the weights of the first layer of the DNN in the mask pattern, the captured image obtained by modulating the input image with the mask pattern contains only the information required for the subsequent recognition process. As a result, while maintaining sufficient accuracy in the recognition process, the final image (reconstructed image) is so different that it is difficult for the human eye to recognize it as the input image, i.e., an image with high privacy protection.

[0160] As an evaluation index for privacy protection, the recognition rate when a recognizer for normal images is run on the restored image output by the image restoration processing network can be used. In this case, by training the weights of the first layer of the DNN so that the recognition rate is low, a system that can perform recognition processing while ensuring privacy protection can be created.

[0161] <<3. Second Embodiment>> In the above, we have described an example in which the mask 131 is divided into sub-areas, a channel is assigned to each sub-area, a mask pattern is generated by convolving weights related to the corresponding DNN first layer processing, the captured image captured by the image sensor 132 is divided into sub-areas, and the processing results of the DNN first layer for each channel are obtained as a reconstructed image.

[0162] However, the mask patterns that serve as the basic patterns in mask 131 may be ones that have low correlation with each other for each channel, and patterns in which weights convolved in the processing of the first layer of DNN are superimposed on all basic patterns may be used.

[0163] That is, for example, by defining the mask pattern Ki, which is the basic pattern set for each channel, as in the following equation (8), the correlation between the basic patterns of each channel becomes the lowest.

[0164]

number

[0165] where I is the identity matrix, i.e., the mask patterns Ki are uncorrelated and become identity matrices when multiplied by themselves.

[0166] Using the mask pattern Ki as a basic pattern set for each channel, the mask pattern C is set as expressed by the following equation (9).

[0167]

number

[0168] Here, Pi1 is the weight used in the convolution operation of the first layer of the DNN for channel i.

[0169] That is, the mask pattern C expressed by equation (9) can be represented schematically as shown in FIG.

[0170] That is, the basic patterns of each of the m channels are set to Ki, which has low correlation with each other, and the mask patterns Ki*Pi1 of each channel, which are convolved with the weight Pi1 convolved in the first layer processing of the DNN for each channel, are superimposed on m channels to form the mask pattern C. Note that the "+" shown in Figure 11 represents the superposition of mask patterns.

[0171] By modulating the input image X using the mask pattern C configured in this way and using the mask pattern Ki, which is the basic pattern of each channel, as a restoration matrix for the captured image C*X, as expressed in the following equation (10), it is possible to obtain the processing results of the first layer of the DNN for each channel (extract them from the captured image C*X).

[0172]

number

[0173] As a result, in the image processing device 111 of Figure 5, when extracting feature vectors, recognition processing was performed using feature vectors obtained on a sub-area basis, but by using a mask pattern C convolved with each basic pattern and weight Pi1 convolved in the processing of the first layer of DNN for each channel in this way, it is possible to obtain feature vectors using a high-resolution restored image using the entire image captured by the image sensor 132.

[0174] FIG. 12 shows an example of the configuration of an image processing device 111 that realizes image recognition processing using a mask made up of the mask pattern C described with reference to FIG.

[0175] In the image processing device 111 in FIG. 12, components having the same functions as those in the image processing device 111 in FIG. 5 are given the same reference numerals, and the description thereof will be omitted as appropriate.

[0176] That is, the image processing device 111 of Figure 12 differs from the image processing device 111 of Figure 5 in that a mask 131'' and Ch1 reconstruction units 171-1 to 171-m are provided instead of the mask 131', the division unit 161, and the Ch1 reconstruction units 162-1 to 162-m.

[0177] The mask 131'' is composed of a mask pattern C for each of the above-mentioned channels, which is convolved with a basic pattern having low correlation and a weight Pi1 that is convolved in the processing of the first layer of the DNN for each channel.

[0178] The Ch1 reconstruction unit 171-1 to Chm reconstruction unit 171-m use the mask pattern Ki, which is the basic pattern in the mask 131'', as a restoration matrix in each channel to obtain (extract) the processing result by multiplying the reconstructed image X' in each channel from the captured image C*X by the weight Pi1 of the first layer of the DNN, and output the processing result to the subsequent recognition processing unit 134'.

[0179] Note that the components of the recognition processing unit 134' in FIG. 12 are similar to the components of the recognition processing unit 134' in FIG. 5, but differ in that the components of the recognition processing unit 134' in FIG. 5 process information in sub-area units assigned to each channel, whereas the components of the recognition processing unit 134' in FIG. 12 process information of the entire area imaged by the image sensor 132 in each channel.

[0180] <Recognition processing by the image processing device in Figure 12> Next, the recognition processing by the image processing device 111 in FIG. 12 will be described with reference to the flowchart in FIG.

[0181] In step S31, the image sensor 132 photoelectrically converts the input image X, which is light from a subject modulated by the mask 131'' of the superimposed mask pattern C, into a mask pattern obtained by convolving a mask pattern Ki, which is a basic pattern for each channel, with a weight Pi1 related to the processing of the first layer of the DNN, to generate a captured image C*X, and outputs the image to the Ch1 reconstruction unit 171-1 to the Chm reconstruction unit 171-m.

[0182] In step S32, the Ch1 reconstruction unit 171-1 to the Chm reconstruction unit 171-m multiply the captured image C*X by a mask pattern Ki, which is the basic pattern of each channel, to obtain a restored image consisting of the processing results of the DNN first layer in each channel, and output it to the Ch1 second layer processing unit 151-1-2 to the Chm2 second layer processing unit 151-m2-2.

[0183] In step S33, the Ch1 second layer processing unit 151-1-2 to the Chm2 second layer processing unit 151-m2-2, the Ch1 third layer processing unit 151-1-3 to the Chm3 third layer processing unit 151-m3-3, ..., the Ch1 nth layer processing unit 151-1-n to Chm n n-th layer processing unit 151-m n -n executes the processing of the second layer and subsequent layers of the DNN for each channel, obtains a feature vector, and outputs it to the recognition unit 152′.

[0184] In step S34, the recognition unit 152' executes recognition processing of the object based on whether or not the information on the feature vector of each channel matches the feature vector of the object, and outputs the recognition result.

[0185] Through the above processing, the restored images reconstructed in each channel are output to the recognition processing unit 134' in an image state that makes it difficult for humans to recognize objects visually, thereby realizing privacy protection and eliminating the need for special privacy-conscious management of the restored images.

[0186] In addition, since the processing by the DNN first layer processing unit 52-1 in Figure 1 for each channel is diverted to optical processing in the mask 131'', processing of the first layer of DNN for each channel is no longer necessary, which makes it possible to simplify the device configuration.

[0187] Furthermore, since the processing of the first layer of the DNN is no longer necessary, the processing load on the image processing device 111 can be reduced.

[0188] Furthermore, since the feature vector of each channel can be obtained using the entire image captured by the image sensor 132, it is possible to obtain the feature vector of each channel with higher accuracy than when the feature vector is obtained in sub-area units.

[0189] <<4. First Application Example>> As a method for performing different convolutions of multiple channels in parallel, division into subareas and superposition of basic patterns with low correlation with each other may be combined.

[0190] For example, as shown in FIG. 14, when a mask 131''' is provided, the entire mask 131''' is divided into a plurality of subareas MSA1 to MSAx, and each of the subareas MSA1 to MSAx uses superimposed basic patterns that have low correlation with each other.

[0191] 14, the area is divided into subareas MSA1 to MSAx, and in subarea MSA1, a mask pattern is formed by superimposing a mask pattern in which a weight P11 used in the first layer processing of the DNN is convolved with a basic pattern Ka1 in channel 1 (Ch1), a mask pattern in which a weight P21 used in the first layer processing of the DNN is convolved with a basic pattern Ka2 in channel 2 (Ch2), and a mask pattern in which a weight Pt1 used in the first layer processing of the DNN is convolved with a basic pattern Kat in channel t (Cht). Here, the basic patterns Ka1 to Kat have low correlation with each other.

[0192] In addition, in sub-area MSA2, a mask pattern is formed in which a weight P(t+1)1 used in the first layer processing of the DNN is convolved with a basic pattern Kb1 in channel (t+1) (Ch(t+1)), a mask pattern in which a weight P(t+2)1 used in the first layer processing of the DNN is convolved with a basic pattern Kb2 in channel (t+2) (Ch(t+2)), and a mask pattern in which a weight P(t+u)1 used in the first layer processing of the DNN is convolved with a basic pattern Kbu in channel (t+u) (Ch(t+u)). Here, the basic patterns Kb1 to Kbu have low correlation with each other.

[0193] Such a mask pattern makes it possible to achieve a balance between the trade-off between the resolution of each channel and the number of channels when dividing the subareas, and the trade-off between the number of channels and the S / N ratio of the signal extracted by signal processing when superimposing basic patterns that have low correlation with each other.

[0194] The configuration of the image processing device 111 when using such a mask pattern is a combination of a configuration corresponding to the division unit 161 required when using the above-mentioned mask using the sub-area, and a configuration corresponding to a reconstruction unit for each channel that extracts the processing results of the first layer of DNN for each channel from an image captured using a mask in which a mask pattern in which a basic pattern with low correlation for each channel and a weight used for processing the first layer of DNN are convolved is superimposed. The individual components used in the combination are the same as the above-mentioned components, so their description will be omitted.

[0195] <<5. Second Application Example>> When dividing the entire mask into sub-areas, it is not necessary for each sub-area to have the same size and shape.

[0196] For example, as shown in the mask pattern of mask 131'''' in FIG. 15, subarea BSA may be arranged in the upper left of the drawing, and to the right of it, four subareas SSA1 to SSA4 may be arranged that are similar in shape to subarea BSA but are 1 / 4 the size, i.e., have 1 / 4 the resolution, and further, a subarea LSA that is long in the horizontal direction and has a different shape may be arranged in the lower part of the drawing. Although not shown, a subarea that is long in the vertical direction may also be arranged in the same way as subarea LSA that is long in the horizontal direction.

[0197] By changing the size, shape, and resolution of the sub-areas in this way, it is possible to assign a resolution according to the importance of the channel, or to prioritize a higher or lower S / N ratio.

[0198] <<6. Example of execution by software>> The above-described series of processes can be executed by hardware, but can also be executed by software. When the series of processes are executed by software, the programs constituting the software are installed from a recording medium into a computer incorporated in dedicated hardware, or into, for example, a general-purpose computer that can execute various functions by installing various programs.

[0199] 16 shows an example of the configuration of a general-purpose computer. This personal computer has a built-in CPU (Central Processing Unit) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.

[0200] Connected to the input / output interface 1005 are an input unit 1006 including input devices such as a keyboard and a mouse through which a user inputs operation commands, an output unit 1007 that outputs a processing operation screen and images of processing results to a display device, a storage unit 1008 including a hard disk drive or the like that stores programs and various data, and a communication unit 1009 including a LAN (Local Area Network) adapter or the like that executes communication processing via a network typified by the Internet. Also connected to the input / output interface 1005 is a drive 1010 that reads and writes data from / to removable storage media 1011 such as a magnetic disk (including a flexible disk), an optical disk (including a CD-ROM (Compact Disc-Read Only Memory) and a DVD (Digital Versatile Disc)), a magneto-optical disk (including an MD (Mini Disc)), or a semiconductor memory.

[0201] The CPU 1001 executes various processes in accordance with a program stored in a ROM 1002 or a program read from a removable storage medium 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, installed in a storage unit 1008, and loaded from the storage unit 1008 into a RAM 1003. The RAM 1003 also stores data necessary for the CPU 1001 to execute various processes as appropriate.

[0202] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0203] The program executed by the computer (CPU 1001) can be provided by being recorded on a removable storage medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0204] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting a removable storage medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.

[0205] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0206] 16 realizes the functions of the division unit 161, Ch1 reconstruction unit 162-1 to Chm reconstruction unit 162-m, and recognition processing unit 134' in FIG. 5, and the functions of the Ch1 reconstruction unit 171-1 to Chm reconstruction unit 171-m, and recognition processing unit 134' in FIG. 12.

[0207] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0208] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.

[0209] For example, the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.

[0210] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0211] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0212] The present disclosure can also be configured as follows.

[0213] <1> A mask that modulates the incident light and transmits it; an image sensor that captures a modulated image based on the incident light modulated by the mask; a signal processing unit that performs signal processing based on a mask pattern of the mask on the modulated image; An image processing device comprising: <2> the signal processing unit performs signal processing on the modulated image based on the mask pattern in a plurality of channels; the mask pattern is set for each of the channels, The pattern set for each channel is a pattern obtained by weighting and adding a weight used for signal processing for each channel to a binary pattern that is a basic pattern set for each channel. <1> The image processing device according to claim 1. <3> the incident light is reflected light reflected from a subject, the signal processing is image recognition processing of the subject using a DNN (Deep Neural Network), The weights are the weights of the first layer of the DNN. <2> The image processing device according to claim 1. <4> The weights of the first layer of the DNN are trained so that the recognition accuracy of the image recognition process is higher and the similarity between the original image captured by the image sensor without modulation by the mask and the restored image obtained by restoring the modulated image using the weights of the first layer of the DNN is lower. <3> The image processing device according to claim 1. <5> The mask is divided into sub-areas equal to the number of channels; The mask pattern is configured by arranging a pattern for each channel in each of the sub-areas. <2> The image processing device according to claim 1. <6> A margin area made of a light-shielding area is provided between the sub-areas. <5> The image processing device according to claim 1. <7> The width of the margin area is set based on the FOV (Field of View) of the pixels of the image sensor. <6> The image processing device according to claim 1. <8> Light-shielding walls are provided between the plurality of sub-areas. <5> The image processing device according to claim 1. <9> The sub-areas are identical in size and shape. <5> The image processing device according to claim 1. <10> The sizes and shapes of the sub-areas are non-identical. <5> The image processing device according to claim 1. <11> The basic patterns set for each of the channels, which constitute the patterns arranged in each of the sub-areas, are set to be uncorrelated with each other or to have a correlation lower than a predetermined correlation. <5> The image processing device according to claim 1. <12> The basic patterns set for each of the channels, which constitute the patterns arranged in each of the sub-areas, are set to be uncorrelated with each other or to have a correlation lower than a predetermined correlation using a pseudo-random signal. <5> The image processing device according to claim 1. <13> The mask pattern is a combination of patterns set for each channel. <2> The image processing device according to claim 1. <14> The basic patterns set for each channel are set to be uncorrelated with each other or to have a correlation lower than a predetermined correlation. <13> The image processing device according to claim 1. <15> The basic patterns set for each channel are set to be uncorrelated with each other or lower than a predetermined correlation using pseudo-random signals. <14> The image processing device according to claim 1. <16> The mask is divided into sub-areas equal to the number of channels; A pattern of the plurality of channels is arranged in each of the sub-areas, The pattern for each sub-area is a combination of patterns corresponding to a plurality of channels. <2> The image processing device according to claim 1. <17> The mask pattern is composed of transmission, light blocking, and any intermediate value. <1> ~ <16> 10. The image processing device according to claim 9, wherein <18> The mask pattern is formed by a diffraction grating. <1> ~ <16> 10. The image processing device according to claim 9, wherein <19> A mask that modulates the incident light and transmits it; an image sensor that captures a modulated image based on the incident light modulated by the mask; a signal processing unit that applies signal processing to the modulated image based on a mask pattern of the mask; An image processing method for an image processing apparatus comprising: The signal processing unit performs signal processing based on a mask pattern of the mask on the modulated image. Image processing methods. <20> A mask that modulates the incident light and transmits it; an image sensor that captures an image as a modulated image based on the incident light modulated by the mask; a signal processing unit that performs signal processing based on a mask pattern of the mask on the modulated image; A program that makes it work. [Explanation of symbols]

[0214] 101, 111 image processing device, 131, 131' to 131'''' mask, 132 image sensor, 133 reconstruction unit, 134, 134' recognition processing unit, 151-2 to 151-n, 151-1-2 to 151-1-n, 151-2-2 to 151-2-n, 151-m-2 to 151-mn DNN second layer processing unit to DNN nth layer processing unit, 152, 152' recognition unit, 171-1 to 171-m Ch1 reconstruction unit to Chm reconstruction unit

Claims

1. A mask that modulates the incident light and transmits it; an image sensor that captures a modulated image based on the incident light modulated by the mask; a signal processing unit that performs signal processing on the modulated image based on a mask pattern of the mask, the signal processing unit performs signal processing on the modulated image based on the mask pattern in a plurality of channels; the mask pattern is made up of a plurality of patterns set for each of the channels, The plurality of patterns set for each channel are patterns obtained by weighting and adding weights used in signal processing for each channel to binary patterns that are basic patterns set for each channel. Image processing device.

2. the incident light is reflected light reflected from a subject, the signal processing is image recognition processing of the subject using a DNN (Deep Neural Network), The weights are the weights of the first layer of the DNN. The image processing device according to claim 1 .

3. The weights of the first layer of the DNN are trained so that the recognition accuracy of the image recognition process is higher and the similarity between the original image captured by the image sensor without modulation by the mask and the restored image obtained by restoring the modulated image using the weights of the first layer of the DNN is lower. The image processing device according to claim 2 .

4. The mask is divided into sub-areas equal to the number of channels; The mask pattern is configured by arranging a pattern for each channel in each of the sub-areas. The image processing device according to claim 1 .

5. A margin area made of a light-shielding area is provided between the sub-areas. The image processing device according to claim 4 .

6. The width of the margin area is set based on the field of view (FOV) of the image sensor pixels. The image processing device according to claim 5 .

7. Light-shielding walls are provided between the plurality of sub-areas. The image processing device according to claim 4 .

8. The sub-areas are identical in size and shape. The image processing device according to claim 4 .

9. The sizes and shapes of the sub-areas are non-identical. The image processing device according to claim 4 .

10. The basic patterns set for each of the channels, which constitute the patterns arranged in each of the sub-areas, are set to be uncorrelated with each other or to have a correlation lower than a predetermined correlation. The image processing device according to claim 4 .

11. The basic patterns set for each of the channels, which constitute the patterns arranged in each of the sub-areas, are set to be uncorrelated with each other or to have a correlation lower than a predetermined correlation using a pseudo-random signal. The image processing device according to claim 4 .

12. The mask pattern is a combination of patterns set for each channel. The image processing device according to claim 1 .

13. The basic patterns set for each channel are set to be uncorrelated with each other or to have a correlation lower than a predetermined correlation. The image processing device according to claim 12.

14. The basic patterns set for each channel are set to be uncorrelated with each other or lower than a predetermined correlation using pseudo-random signals. The image processing device according to claim 13.

15. The mask is divided into sub-areas equal to the number of channels; A pattern of the plurality of channels is arranged in each of the sub-areas, The pattern for each sub-area is a combination of patterns corresponding to a plurality of channels. The image processing device according to claim 1 .

16. The mask pattern is composed of transmission, light blocking, and any intermediate value. The image processing device according to claim 1 .

17. The mask pattern is formed by a diffraction grating. The image processing device according to claim 1 .

18. A mask that modulates the incident light and transmits it; an image sensor that captures a modulated image based on the incident light modulated by the mask; an image processing method of an image processing apparatus including a signal processing unit that applies signal processing based on a mask pattern of the mask to the modulated image, the signal processing unit performs signal processing on the modulated image based on the mask pattern in a plurality of channels; the mask pattern is made up of a plurality of patterns set for each of the channels, The plurality of patterns set for each channel are patterns obtained by weighting and adding weights used in signal processing for each channel to binary patterns that are basic patterns set for each channel. Image processing methods.

19. A mask that modulates the incident light and transmits it; an image sensor that captures a modulated image based on the incident light modulated by the mask; a signal processing unit that performs signal processing based on a mask pattern of the mask on the modulated image; the signal processing unit performs signal processing on the modulated image based on the mask pattern in a plurality of channels; the mask pattern is made up of a plurality of patterns set for each of the channels, The plurality of patterns set for each channel are patterns obtained by weighting and adding weights used in signal processing for each channel to binary patterns that are basic patterns set for each channel. program.

Citation Information

Patent Citations

  • Image processing method, image processing device, imaging device, and image processing program

    JP2019003609A

  • Imaging device

    WO2017145348A1

  • Imaging device, imaging method, and imaging element

    WO2019124106A1