Imaging System
The camera system optically modulates light rays to protect privacy by making images unrecognizable, while enabling recognition through a planar modulation element and imaging element, addressing privacy concerns in imaging technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2026-03-11
AI Technical Summary
Existing imaging technologies face challenges in protecting privacy while enabling recognition of individual information, as captured images can be hacked or leaked, raising concerns about unauthorized access and privacy violations.
A camera system with a planar modulation element and imaging element that optically modulates light rays from multiple directions, destroying spatial projection information and preserving essential recognition information, combined with a recognition unit that performs recognition without restoring the image to a retinal format.
The system allows for recognition of individual information while ensuring privacy by making captured images unrecognizable to humans and secure from leaks, maintaining privacy even if data is intercepted.
Smart Images

Figure 0007828107000010 
Figure 0007828107000011 
Figure 0007828107000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to a camera technology capable of protecting privacy by capturing an image of a subject by modulating the image to a level that makes the individual unrecognizable. [Background technology]
[0002] Traditionally, cameras capture optical images by projecting a retinal image, i.e., a human-readable image of focused light, through a lens onto an image sensor, measuring the received light intensity at each pixel of the image sensor, and then digitizing the optical image. The captured image data is typically read out in a raster scan sequence, maintaining spatial relationships, and then transferred, for example, via the Internet and saved as a data file. If information is hacked or leaked during transfer or storage, the content can be easily observed. Today, image privacy issues caused by such data leaks and unilateral disclosure by third parties are becoming increasingly serious. For example, there have been cases where camera-equipped IoT glasses were banned in restaurants and their release was discontinued, and cases where third parties requested the deletion of images uploaded to social media.
[0003] Furthermore, lensless cameras or flat cameras have been proposed in recent years (see, for example, Patent Document 1). This type of camera uses a plate-shaped modulator that modulates transmitted light instead of a lens, thereby achieving a thinner imaging device. The imaging device includes a modulator that modulates light intensity using a first pattern formed in a concentric circle, an image sensor that converts the light image transmitted through the modulator into image data, and an image processing unit that performs cross-correlation calculations between the image data output from the image sensor and pattern data representing a second pattern, thereby enabling the restoration of a subject image. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-61109 Summary of the Invention [Problem to be solved by the invention]
[0005] While surveillance cameras and other devices are used to identify individuals for crime prevention purposes, many of today's smartphones, smart speakers, and IoT devices are not necessarily intended for individual identification or evidence recording. Instead, they are used as sensors and monitors for recognizing individuals' facial expressions, gestures, and behavior. Drone and autonomous driving cameras are also sensors for environmental recognition and obstacle detection, and do not necessarily record the privacy of people captured in their images. While there are uses for cameras that are not originally intended to record or store private information, the very use of cameras raises concerns about privacy violations, limiting their use. This creates a dilemma, preventing cameras from being used solely as sensors for gesture recognition, and preventing the development of applications and services for the coming IoT and Society 5.0 era. A commonly proposed solution to this problem involves capturing images, encoding them on the edge device, transferring them, and then decrypting them on the server device before recognition. However, even with this approach, concerns remain about the risk of unencoded and decoded images being leaked due to hacking or information leaks.
[0006] Furthermore, in the imaging device described in Patent Document 1, the data acquired by the image sensor is image information that can be restored, and therefore there is a risk that the information may be hacked or leaked by a third party and made public, and there is no consideration for privacy protection.
[0007] The present invention has been made in view of the above, and aims to provide a camera and imaging system that enable recognition (identification) of attached information of an individual while protecting the privacy of the individual subject. [Means for solving the problem]
[0008] The camera of the present invention includes a planar imaging element having an array of multiple pixels each made of a photosensitive element, and a planar modulation element arranged in front of the imaging element and having a pattern formed thereon for modulating incident light, the pattern including an array of multiple light-transmitting portions that guide light rays from multiple directions from a subject to one pixel.
[0009] According to the present invention, light rays from a subject are optically modulated by a modulation element and then captured by an imaging element. Spatial projection information, such as an optical retinal image, is destroyed in the captured image, but essential information necessary for recognition can be preserved. This makes it difficult to visually recognize the content from recorded or leaked data sequences, thereby protecting privacy. [Effects of the Invention]
[0010] According to the present invention, it is possible to take a photograph in which individual recognition of the subject is impossible, while the recognition of the individual's attached information for the intended purpose is possible, thereby protecting privacy. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a schematic side view, partially in cross section, showing the configuration of an imaging system according to the present invention; [Figure 2] FIG. 2 is a diagram showing the relationship between the pattern of the modulation element and the pixels of the imaging element. [Figure 3] This figure explains the relationship between the presence or absence and type of modulation element and the captured image. (A) is a lensless case, (B) is a case where a pinhole-shaped hole (pinhole) is drilled, and (C) is a case where a mask with multiple or different sized light-transmitting portions is interposed on the surface. [Figure 4] 10A and 10B are diagrams showing other patterns of modulation elements, in which (A) is a mask in which light-transmitting portions of different sizes are formed, and (B) is a mask in which light-transmitting portions are formed randomly or in a sparsely and densely distributed pattern. [Figure 5] FIG. 10 shows another embodiment of a modulation element. [Figure 6]Lensless imaging diagrams for small and large distances between the image and the encoding plane. (A) is the small distance case, and (B) is the large distance case. [Figure 7] FIG. 1 is a system diagram illustrating visual privacy protection for face recognition via lensless imaging. [Figure 8] In the figure showing a visual comparison of different imaging systems, the scale above the measurements and patterns indicates the ratio between blur and exposure, and the dimension of the fixed pattern is m=32×32. [Figure 9] , This figure shows the confusion matrix of LwoC-woRec for multiple Hi and Ri, where each value (i,j) indicates the Top1 accuracy (%) of the jth recognition function Rj for the input lensless measurement of the ith coding pattern Hi. [Figure 10] This figure shows the Top1 accuracy (%) of VGG-Face2 for coded image size n = 63 × 63 (Fig. 10(A)) and 127 × 127 (Fig. 10(B)) when the coded pattern size m = 32 × 32. [Figure 11] Fig. 1 illustrates the learning patterns and measurement of human visual privacy protection with various weights, where n = 63 × 63 and m = 32 × 32. [Figure 12] This figure shows the confusion matrix between the training pattern Hi and the recognition function Ri in VGG-Face2, with 10 classes, n = 63 × 63, b = 32 × 32. [Figure 13] FIG. 10 illustrates another embodiment of a hardware implementation for lensless imaging. [Figure 14] FIG. 1 shows the displayed image, the actual coded pattern on the spatial light modulator, and the actual captured measurements (rescaled to maximum and minimum values for better visual quality). DETAILED DESCRIPTION OF THE INVENTION
[0012] FIG. 1 is a schematic side view, partially in cross section, illustrating the configuration of an imaging system 1 according to the present invention. In FIG. 1, the imaging system 1 includes a camera 11 and a recognition unit 12. The camera 11 includes, from the front in the optical axis direction, a mask 2, which is an embodiment of a modulation element, an imaging element 4, a thin bonding layer 3 that optically bonds the mask 2 to the imaging element 4, and a readout unit 5 that reads out image data captured by the imaging element 4 from each pixel. For ease of explanation, the camera 11 is shown exaggerated in size relative to the subject P. The bonding layer 3 may be an adhesive layer alone in an integrated configuration, or a physical connection structure may also be employed.
[0013] Camera 11 is a digital camera equipped with an image sensor 4. Image sensor 4 is typically configured with a large number of pixels 42 arranged in a matrix on the front surface of a rectangular, plate-like (planar) main body 41. Each pixel 42 is a tiny photosensitive element such as a CCD, and generates a voltage signal corresponding to the brightness of received light.
[0014] The mask 2 is a sheet-like or thin plate-like body having a size corresponding to the imaging element 4. The mask 2 has light-blocking properties, and has light-transmitting sections 21 consisting of a plurality of holes or light-transmitting areas formed at appropriate locations on its surface. Conversely, the mask 2 may have light-transmitting properties, but have light-blocking treatment applied to areas of its surface other than the light-transmitting sections 21.
[0015] 2 is a diagram showing the arrangement relationship between the pattern (modulation pattern) of the light-transmitting portions 21 of the mask 2 and the pixels 42 of the image sensor 4. The light-transmitting portions 21 are preferably provided corresponding to the pixels 42, and are formed at a predetermined pitch in at least one direction of the rows and columns of the pixels 42.
[0016] The size of the light-transmitting portions 21 does not need to be uniform, and it is preferable that all or some of them are equal to or larger than the size of the pixels 42. In FIGS. 1 and 2, the size of the light-transmitting portions 21 is several times the size of the pixels 42, but it may be several tens to several hundreds of times larger. By including the large-sized light-transmitting portions 21 of the mask 2 as described above, light rays L1, L2 (or other light rays) incident from multiple directions on the subject P are made incident on the same pixel 42. In this way, light rays from multiple directions are made incident on the same pixel 42, i.e., modulated, that is, spatial projection information is optically destroyed without forming a retinal image, and imaging is performed, thereby reducing the level to such an extent that personal recognition of the subject P cannot be reproduced from the captured image itself.
[0017] FIG. 3 is a diagram illustrating the relationship between the presence or absence and type of mask 2 and the captured image. FIG. 3(A) shows a lensless system in which an image of a subject is captured by the image sensor 4 without the use of a mask 20A. In FIG. 3(A), light rays from all directions of the subject are equally incident on all pixels 42, resulting in a uniform, completely meaningless captured image. On the other hand, if a single pinhole-shaped hole (pinhole) is drilled in the mask 20B as in FIG. 3(B), only incident light from one direction of the subject passes through the pinhole and measures different intensities, resulting in a perfectly formed image similar to a retinal image, as in a normal photograph.
[0018] FIG. 3(C) shows an arrangement in which a mask 2a having a plurality of light-transmitting portions 21a or light-transmitting portions 21a of different sizes on the surface is interposed, and light rays that have passed through a plurality of light-transmitting portions 21a are combined and directed to each of pixel 421 and pixel 422, and light rays that have passed through the same light-transmitting portion 21a are combined and directed to both pixel 421 and pixel 422, thereby performing imaging.
[0019] In the above case of Figure 3(A), privacy can be completely protected because visual information is completely lost, but since no information remains, that is, all pixels are integrated (averaged) in the same way and are inseparable, it becomes impossible to identify what is shown in the subject image, such as when recognizing the image.In contrast, in the case of Figure 3(B), the subject image itself is shown, so there is no loss of data, but it is actually more vulnerable to privacy.
[0020] On the other hand, as in the case of Figure 3(C), if the captured image is put into this intermediate state using a mask 2a, it is possible to make it visually unrecognizable what is being shown in the captured image. Therefore, in this case, even if the captured image itself is hacked or leaked, the incomprehensible state is maintained, and even if the captured image and mask information are stolen and image processing is performed for reproduction, reproduction to the level of personal recognition is not possible, so privacy protection is still ensured. For example, when put into the intermediate state as in Figure 3(C), as can be seen from the captured image G, even if the shading pattern is measured and additional information such as the subject's location information can be recognized, the subject itself cannot be reproduced.
[0021] 4 and 5 are diagrams showing other aspects of the modulation element, with Fig. 4 showing another mask pattern and Fig. 5 showing another embodiment. Fig. 4(A) shows mask 2b in which light-transmitting portions 21b and 22b of different sizes are formed, and Fig. 4(B) shows mask 2c in which light-transmitting portions 21c are formed randomly or in a sparsely and densely distributed pattern. The shape of the light-transmitting portions may be rectangular (including slit-shaped), polygonal, or circular.
[0022] FIG. 5 shows a light-transmitting thin plate-like body 2d in place of a mask 2 as an example of a modulation element. The plate-like body 2d may be in the form of a sheet. At least one of the front and back surfaces of the plate-like body 2d is formed into an uneven rough surface 21d (corresponding to a light-transmitting portion). The unevenness of the rough surface 21d may include a minute convex or concave lens shape. The size in the surface direction of the uneven surface forming the rough surface 21d may be a size corresponding to the size of the pixel 42 or several to a hundred times that size. The uneven surface forming the rough surface 21d corresponds to the light-transmitting portion.
[0023] Plate-shaped body 2d is not a condensing lens that enables regular light collection, but rather refracts light rays L11, L12, and L13 from multiple directions, for example, within plate-shaped body 2d, and directs them in irregular directions, as shown in Figure 5. In other words, rough surface 21d causes transmitted light rays L11, L12, and L13 to be incident on pixels 42 that are in a non-corresponding positional relationship, such as incident on the same pixel 42 or on a different pixel 42 that is skipped. As a result, the spatial projection information of the image from the subject is optically destroyed, and the captured image becomes meaningless information that makes it impossible to recognize the individual.
[0024] Returning to FIG. 1 , the readout unit 5 outputs a voltage signal (measurement signal) generated by each pixel 42 of the image sensor 4. The readout unit 5 reads out the signals of each pixel 42 in a predetermined order along the array direction, for example, in accordance with raster scanning. Furthermore, when reading out signals from the image sensor 4, the readout unit 5 may read out the signals in a random order or by adding together the signals of multiple pixels and then electronically encrypting them, thereby outputting an image that is difficult for a human to understand. This image is effectively recognized (determined) by machine learning using a recognition unit 12 having parameters suitable for, for example, determining the gender of the subject. The recognition unit 12 may be integral or semi-integrated with the camera 11, or may be connected by wire, wirelessly, or via an internet connection in a remote location (e.g., a monitoring room).
[0025] The recognition unit 12 performs recognition (determination) on input image information using parameters acquired through machine learning, and outputs the results. The recognition unit 12 effectively performs recognition (determination) specialized for a specific intended use. The parameters stored in the parameter storage unit 121 of the recognition unit 12 are modeled through machine learning. As machine learning, at least one of the so-called supervised learning, unsupervised learning, reinforcement learning, and deep learning learning methods is adopted.
[0026] Machine learning has an input layer, an output layer, and at least one hidden layer between them that simulate (model) a neuron network, and each layer has a structure in which multiple nodes are connected by edges. Parameters refer to the weight values of each edge in each layer. For example, in supervised learning, when recognizing (determining) the gender of a subject from images captured by camera 11, each image of multiple subjects captured by camera 11 is input to the input layer of the simulated network, and a corresponding answer (label) is presented. The weight values are updated during feedback and learning is performed. By performing such learning on a large number of subjects, the feature values of each subject are reflected in the parameters, improving the accuracy of the determination.
[0027] It is also preferable to simultaneously train the signals read out from the modulation element and image sensor 4 and the recognition unit 12 using, for example, a deep learning framework. In this case, it is preferable to train the deep learning using an adversarial learning framework so that the captured image is as visually meaningless as possible, thereby making it possible to capture images that are incomprehensible to humans and even incapable of personal recognition by the recognition unit 12 without compromising the recognition function. In this way, the set of camera 11 and recognition unit 12 is designed by optimizing the parameters of the recognition unit 12, which is software, and the hardware design, which is the pattern of the modulation element 2, in relation to each other within a machine learning framework.
[0028] In this way, by essentially designing a modulation pattern in such a way that a ray of light that has passed through one light-transmitting section is incident on multiple pixels, or in such a way that each ray of light that has passed through multiple light-transmitting sections is incident on one pixel, it is possible to produce a modulation element that makes it impossible to recognize an individual, while making it possible to recognize the individual's associated information.
[0029] The present invention also includes the following aspects.
[0030] (1) The present camera 11 can also be constructed by placing the present modulation element on either the front or rear surface of the photographic lens of a normal camera. In this case, the modulation element should be designed to modulate the optical image taking into account the imaging performance of the photographic lens.
[0031] (2) Specific intended uses of the imaging system 1 include gender determination, age determination, gestures (actions), personal ID, and various other types of additional information that do not lead to the identification of the subject. The determination results can be notified by further providing a display, speaker, etc. that displays the determination results from the recognition unit 12. The imaging system 1 can also be applied to non-human animals and other individuals. Therefore, the imaging system 1 can be applied not only as a portable system but also as a stationary system.
[0032] (3) The modulation pattern on the surface of the modulation element may be irregular, or it is preferable from a production standpoint to arrange one or more types of modulation patterns repeatedly in at least one of the vertical and horizontal directions for each certain size. The size of the split modulation pattern depends on the recognition application, but may be a size corresponding to an area of several tens to several hundreds of pixels 42, for example, an array area of 100 x 100 pixels, or smaller or larger, in relation to the number of pixels 42. Furthermore, it may also be possible to form adjacent pinholes as part of the modulation element pattern, as shown in Figure 3(B), and direct light rays passing through both pinholes to the same pixel.
[0033] (4) Instead of a fixed type, the mask 2 can be made of a material that changes the modulation pattern, such as a liquid crystal display (LCD) panel. By changing the modulation pattern, it can be switched by an electrical signal to a preset pattern depending on the application, and can also be switched over time for the same application, which further improves privacy performance in either case.
[0034] We then present our experiments, which (A) model lensless acquisition and evaluate various lensless imaging methods, (B) demonstrate visual privacy preservation through custom loss functions for human and machine vision and methods for training unique pairs of coded patterns and recognition features, (C) demonstrate our experiments along with hardware realization, and (D) discuss our experimental conclusions.
[0035] (A) Safe lensless imaging We first provide a background on lensless imaging for visual privacy protection and imaging systems for face recognition.
[0036] (1) Coded lensless image Lensless imaging is a novel technique for capturing images without complex lens systems. A coded pattern is used to modulate incident light at a single or multiple pixels. The latter approach is more common because it allows for single-shot image capture without modifying the pattern. Lensless imaging is illustrated in Figure 6 for short (A) and long (B) distances d1 between the image and coded plane. Given a scene x and a coded pattern H, the lensless measurement y is expressed as (Equation 1):
[0037]
number
[0038] where * is the convolution operator and η is additive noise. As the distance d1 decreases, the camera can be thin, like a FlatCam (i.e., a camera that can capture images without a lens), but the angle of the incident light ray is also limited by the field of view of the pixels on the sensor 4. As the distance increases, the field of view is defined by the entrance pupil of the camera, the diameter of the mask 2. For the same resolution as the binary pattern H and a large kernel size, increasing the distance d1 blurs the image and improves visual privacy protection. Therefore, a large distance d1 is adopted. The binary pattern H is learned by modeling the coded imaging as a binary convolution.
[0039] (2) Lensless imaging system for face recognition The imaging system 1 shown in Figure 7 uses a lensless camera 11 equipped with a mask 2 and a sensor 4 to capture images and send them without reconstruction to a recognition unit 12 based on ResNet18 (a convolutional neural network with a depth of 18 layers).
[0040] First, we evaluated imaging scenarios including conventional coded imaging (using fixed and trained patterns). For fixed lensless imaging, we used a pinhole, a defocused pattern, and a random pattern without reconstruction (Rand-woRec). For trained lensless imaging, patterns were trained without constraints and without training (LwoC-woRec). The reconstruction network is described below.
[0041] [Table 1]
[0042] Table 1 shows the Top1 accuracy (%) of various sampling methods using ResNet18. For LwC-MSE, α = 10 -8 , and for LwC-TV, α = 10 -6 Note that Top1 accuracy (%) is an expression of the recognition rate, and refers to the recognition rate of the first candidate. As shown in the results in Table 1, either conventional imaging or pinhole imaging achieves the highest accuracy. Defocusing and randomly coded imaging result in a loss of accuracy of 20% to 40%.
[0043] As shown in Figure 7, it is best for the recognition result b to be accurate, but at the same time, it is required that the captured image y be blurred (unintelligible to humans). Simply optimizing to improve the recognition rate results in a non-blurred captured image y (pinehole has good performance in Table 1), but blurring y results in a trade-off where the recognition rate drops (defocus and random have poor performance in Table 1). This method solves this trade-off by simultaneously optimizing the mask 2 pattern (for blurring) and the recognition unit 12. LwC-TV can sometimes perform well despite the blurred image, or even better than pinehole. In other words, it achieves pattern generation that is intelligible to machines even if not recognizable to humans.
[0044] Figure 8 also shows a visual comparison of various imaging systems. The scale above the measurement and pattern indicates the ratio of blur to exposure. Figure 8 shows that traditional pinhole imaging reveals image details, while defocused and random pattern imaging does not. Therefore, there is a trade-off between accuracy and visual privacy. While training patterns significantly improve recognition accuracy with a loss of approximately 5% compared to pinhole and traditional imaging, they do not guarantee visually secure measurements. As shown in Figure 8, when the coded ratio r is small (i.e., r = 1 / 16), LwoC-woRec reveals the subject's identity. Therefore, a method for controlling the trade-off between accuracy and privacy is desirable.
[0045] (B) Safe learning lensless imaging (1) Protecting privacy from human vision To make it impossible to identify people from lensless images, we wanted to train an encoding pattern so that the captured image would be the same as the image captured with defocused patterns while maintaining high recognition performance. To achieve this, we maximize the blur of the captured image by minimizing the mean squared error (MSE) in (Equation 2).
[0046]
number
[0047] where l m denotes a matrix of all 1s. This is the coded pattern for defocused imaging. Conversely, as shown in Figure 8, the training pattern may converge to smaller local regions (or smaller variations). Thus, measurements are convolved from smaller regions of the image, revealing more information. As a result, we maximize the total variation (TV) of the coded pattern as shown in Equation 3.
[0048]
number
[0049] Here, Δ x and Δ y represent the horizontal and vertical gradient operators, respectively. When using the TV loss, the training patterns need to be more diverse than when using the MSE loss.
[0050] (2) Protecting privacy from machine vision In security applications, the pair of pattern Hi and recognition function Ri must be unique. In other words, a correct {Ri, Hi} indicates high recognition function, while a mismatched {Ri, Hj} indicates low recognition function. To illustrate this more clearly, the pattern Hi and recognition function Ri function function like a key. Accuracy is high only when the key Hi and keyhole Ri match, and low accuracy when they do not. Even if a key Hi and keyhole Ri are public keys, if i is unknown, even if an image captured with Hi is intercepted, the corresponding Ri cannot be identified, making it impossible to directly intercept information. This can be applied, for example, by temporally varying Hi on an LCD panel and synchronizing Ri on the server side with this, as in an ATM (Automatic Teller Machine) code table, further enhancing security.
[0051] Optimizing as in (B).(1) above generates a pattern that is imperceptible to humans but easy for machines to understand. In other words, it is possible that the image will be easily detectable by any learning device (for example, as an extreme example, a mask with horizontal stripes for person A and vertical stripes for person B). To prevent this, the condition in (Equation 4) below is added, and an image encoded with a certain pattern Hi is optimized so that it can be distinguished only by Ri that is also optimized at the same time, and is difficult to distinguish with other Ri, making it impossible for a recognition function Rj that is unaware of pattern Hi to identify it. In other words, this realizes the generation of a mask 2 pattern that simultaneously achieves recognition rate, blurring, and machine privacy (making it difficult to discern the correlation between changes in the captured image and the label).
[0052] For example, a plurality of types of patterns Hi and recognition functions Ri optimized for each type of pattern are stored (prepared) in advance as combinations in a storage unit (not shown), for example, a storage unit in the recognition unit 12, and a control unit (not shown, including the recognition unit 12) stores and controls this combination information. When the recognition unit 12 or the control unit (not shown) selects mask 2 of pattern Hi during a certain photograph, the recognition function Ri as a set is selected and applied to the recognition process instead of the incompatible recognition function Rj, thereby executing the recognition process in the intended, i.e., optimized, state. In this way, application like a code table can further enhance security.
[0053] However, the aforementioned method only protects privacy from human vision, and training multiple examples will generate similar pattern-recognition feature pairs. This can be seen in Figure 9, which shows high accuracy along the diagonal. Note that Figure 9 also shows the confusion matrix of LwoC-woRec for multiple Hi and Ri, where each value (i,j) indicates the Top 1 accuracy of the jth recognition feature Rj for the input lensless measurement of the ith coding pattern Hi. Privacy preservation in machine vision is required to train unique pairs {Ri,Hi}. regIf represents the cross-entropy loss function of the input x and the label b, it is easy to reduce the accuracy of mismatched pairs by (Equation 4).
[0054]
number
[0055] (Equation 4) requires extensive computation with multiple inferences of Ri as the number of unique pairs M increases. Finally, the training loss is a combination of the visual privacy-preserving losses of human vision and machine vision, as expressed in (Equation 5).
[0056]
number
[0057] For new pairs of coded patterns H and R, a more complex loss is added.
[0058] (C) Experimental results of simulation data (1) Dataset and training (1-1) Dataset Here, we present the main results for the VGG-Face2 dataset (pre-trained model). We also conducted additional experiments on the adjusted Microsoft® Celeb (MS-Celeb) and CASIA datasets. For all datasets, we selected the 10 classes with the largest number of images and divided them into training and test sets in a 95:5 ratio. Random cropping and vertical flipping were employed to capture the data.
[0059] (1-2) Training Here, we used ResNet18 for face recognition. The network was trained using a stochastic gradient descent optimizer. The mini-batch size was 128. Three settings were used: image size n = {63 × 63, 127 × 127}, and coded pattern size m = {32 × 32, 64 × 64}. The coded ratio was defined as r = n / m, and the aperture ratio is expressed as the total number of "1" elements in the pattern relative to the entire pattern area. After training, the network with the highest Top1 test accuracy was selected as the final solution. The weighting factors α and β were set to 10 -2 From 10 -8 We tested various combinations up to For reconstruction, we used 17 residual blocks to learn the residuals between clean and captured images from the Div2K (training and test images) dataset.
[0060] (2) Human visual privacy performance Evaluating visual privacy is extremely difficult due to a lack of research on methods for measuring the human eye's ability to recognize objects. Generally, blurred images make it difficult for humans to recognize subjects. Therefore, we employed a non-reference blur metric to evaluate visual privacy quality. As shown in Table 1, all training pattern schemes produced high recognition accuracy with a loss of less than 5% compared to traditional pinhole imaging. Furthermore, reconstruction is not necessary for recognition, but it does reduce accuracy. It should be noted that better reconstruction methods can improve accuracy. However, these methods require fixed coding patterns, making them unsuitable for our method. Conversely, reconstruction midway through the process may increase security risks. Furthermore, recent studies have suggested that direct recognition outperforms initial reconstruction.
[0061] From Figures 10(A) and (B), it is easy to observe that the MSE loss provides a trade-off between defocused imaging and unconstrained imaging (LwoC-woRec), while the TV loss has a trade-off between Rand-woRec and LwoC-woRec. The smaller the weight, the closer the result to the unconstrained result. As the curve moves to the upper right, the TV loss provides slightly better results than the MSE loss. Note that in Figures 10(A) and (B), the mask patterns are the same (32x32), but the image sizes are different (the amount of information varies depending on the number of pixels), resulting in different recognition rates. Because (B) has a higher resolution than (A), it has a higher recognition rate even with the same amount of optical blur.
[0062] The effect of the weighting coefficient is shown in Figure 11. The smaller the weight, the smaller the aperture ratio and the higher the accuracy, but more information is revealed. Visually, both the MSE and TV loss functions can ensure visual privacy at the expense of accuracy. Conversely, reducing the aperture ratio reduces optical efficiency. Although this effect is not considered in our simulations, it significantly affects the recognition accuracy from actual measurements.
[0063] The results of this experiment show that the weighting coefficient α is 10 -4 ~10 -6 , and the MSE loss is 10 -6 ~10 -8 Based on this experiment, we recommend α=10 for a good trade-off between performance and privacy. -4 We choose the TV loss of α=10 for higher accuracy. -5 was selected.
[0064] (3) Machine visual privacy and security performance For security applications, we define two objective scores for the confusion matrix of patterns and recognition features: self-accuracy and mutual-accuracy. Self-accuracy is shown in (Equation 6) and is defined as the average of the diagonal of the confusion matrix. It is the average accuracy using the correct pair H and R.
[0065]
number
[0066] The cross-accuracy is the average accuracy of the off-diagonal lines of the confusion matrix, and represents the performance when mismatched pairs of training patterns and recognition functions are used. Generally, a high self-accuracy and a low cross-accuracy are desirable. The larger the difference between the performance of the self-accuracy and the cross-accuracy, the better. The confusion matrices of various methods are shown in Figure 12.
[0067] Table 2 shows the Top1 accuracy (%) of various sampling methods using ResNet18, with α = 10 for LwC-MSE. -8 , and for LwC-TV, α = 10 -6 , and for LwC-TV-Reg, α=10 -4 , β=10 -6 is.
[0068] [Table 2]
[0069] As the results in Table 2 show, without constraints, LwoC-woRec achieves the highest self-accuracy, but also high inter-accuracy. The loss of human vision due to MSE and TV improves visual privacy for human vision, but does not help protect against machine vision. Therefore, high average (70%) and maximum (80%) inter-accuracy values are reported. Conversely, L, as shown in (Equation 4), reg mvThe ML loss for visual privacy preservation in machine vision helps reduce mutual accuracy while maintaining high accuracy. The ML loss is effective up to M=3, with a 40% accuracy gap between self-accuracy and mutual accuracy, compared to 18% for LwoC, 4% for LwC, and 12% for LwC-TV. Unfortunately, as the number of unique pairs in M increases, the effectiveness of the ML loss decreases as mutual accuracy increases. One reason is that the training framework is sequential, making it more difficult to train new unique pairs. However, accuracy is also significantly affected by the hyperparameters α and β, which have not yet been optimized.
[0070] (4) Experimental results using real data (Hardware Realization) To verify the proposed method, we implemented a prototype imaging system as shown in Figure 13. This camera consists of a monochrome imaging sensor 4 (Grasshoper3 model GS3-U3-41C6M-C, 2048 × 2048) and a mask 2B. The mask 2B consists of a spatial light modulator 20B (SLM; LC 2012, 1024 × 768) and polarizers 20f and 20b placed before and after the spatial light modulator 20B. The relative angle between the two polarizers is adjusted to modulate the intensity of the incident light. The distance between the sensor 4 and the code surface of the mask 2B is approximately 17 mm. A monitor (Plasma display) for displaying the image is installed approximately 1 m away from the SLM.
[0071] The coded patterns were rescaled from 32 × 32 to 716 × 716 and zero-padded to make the SLM size 1024 × 768. Five different coded patterns were evaluated for Mask 2, as shown in Figure 14. To compensate for differences in aperture ratio, the shutter time was manually selected. The facial test image was also rescaled and calibrated on the display screen so that it appeared centered on the image sensor. However, there was still interreflection between the image sensor and the SLM. Therefore, images captured with the SLM aperture close to the center were used for correction. Furthermore, to reduce the effects of noise and reduced light efficiency, an average of 10 times the captured measurements were used as input for the recognition function.
[0072] First, measurements were taken with various patterns, as shown in Figure 14, captured in 16-bit grayscale. Unlike simulations, in real imaging scenarios, pinhole imaging results in very low quality due to the extremely low amount of light. Visible images can also be observed in the captures. Similar to simulations, no privacy information was observed from measurements with defocus and random patterns (50% exposure). Furthermore, without constraints, the learning pattern LwoC revealed more information than the TV loss constraint.
[0073] For face recognition applications, we selected subsets of the highest resolution images—70 and 20—from the CASIA train and test sets, respectively, to capture real-world lensless measurements. Prior to face recognition, the captured images were normalized and further cropped to 80% of the central face region. A background image was captured for each image using an all-zero mask. Light leakage was corrected by subtracting the background image. The final training images were resized to 128x128 for training. Furthermore, the simulated resNet18 model was retrained using the above real-world image data to refine the model to match real-world images.
[0074] Although it achieves high performance in simulations, pinhole imaging performs poorly on real datasets due to inefficient light capture at low coded ratios. Pinhole images are noisier than other images, limiting performance. Pinhole images also contain many details, with a small blur score of 0.140. Defocused imaging results in poorer recognition performance. The captured images have a small blur score due to the lack of information. Random masks also perform slightly better, but are still worse than the trained masks of LwoC and LwC-TV.
[0075] [Table 3]
[0076] Table 3 shows the Top 1 accuracy (%) using the selected CASIA10 surface dataset. Table 3 also shows the experimental results for real images, which, like the simulation, demonstrate that the performance of the proposed LowC-TV is sufficiently high even when the image is blurred, i.e., apparent privacy is protected. It also demonstrates that the reduction in image contrast in real implementations can be improved by using background subtraction (subtracting the brightness value of an image without anything in it from the captured image).
[0077] (D) Conclusion and Discussion We have proposed a trained lensless imaging system to protect visual privacy from both human vision models and target machine vision models. To protect visual privacy from human vision, we maximize measurement blurring using MSE and maximize the variation of the training patterns using TV loss. Through experiments, we confirmed that our method can address the trade-off between visual privacy protection and recognition accuracy in lensless imaging. Although accuracy is slightly reduced, this method can sufficiently protect visual privacy. Furthermore, we used recognition loss to protect visual privacy from machine vision models. A sequential training framework is presented to enable security applications by training multiple unique pairs of coded patterns and deep learning-based recognition functions.
[0078] Here, we are based on the simple hypothesis that the less blurry an image is, the less likely humans are to recognize an object. However, the threshold blur metric for object recognition is not clear and depends on the coded ratio. Meanwhile, blind image deblurring techniques can be used to reconstruct the original image. Further research on subjective quality assessment and the impact of learned kernels is encouraged.
[0079] Our sequential training method was able to learn unique pairs of coded patterns and recognition features. However, the framework is limited in the number of unique pairs (i.e., key space) it can handle. How to handle the case of a large number of unique pairs of H and R (i.e., an increase in M) remains an open question. Furthermore, techniques involving adversarial samples can be further integrated to provide a better training method. Unlike previous techniques that used fixed patterns, we learn coded patterns to achieve higher recognition accuracy. However, the system was trained only on simulated data.
[0080] As described above, the camera of the present invention includes a planar imaging element having an array of multiple pixels each made of a photosensitive element, and a planar modulation element arranged in front of the imaging element and having a pattern formed thereon for modulating incident light, the pattern including an array of multiple light-transmitting portions that guide light rays from a subject coming from multiple directions to one pixel.
[0081] According to the present invention, light rays from a subject are optically modulated by a modulation element and then captured by an imaging element. Spatial projection information, such as an optical retinal image, is destroyed in the captured image, but essential information necessary for recognition can be preserved. This makes it difficult to visually recognize the content from a recorded or leaked data stream, thereby protecting privacy.
[0082] Preferably, the light-transmitting portion includes a portion that guides light rays from multiple directions to multiple pixels. With this configuration, by dispersing the light rays that have passed through one light-transmitting portion, spatial projection information is further destroyed, thereby protecting privacy.
[0083] Preferably, the light-transmitting portion is a mask having holes drilled in the surface thereof for blocking light. According to this configuration, the modulation element can be easily produced by drilling the holes.
[0084] Furthermore, it is preferable that the hole is larger than the size of the pixel. With this configuration, a plurality of light rays can be transmitted, and spatial projection information is destroyed accordingly.
[0085] Furthermore, it is preferable that the light-transmitting portion is a light-transmitting plate-like body having an uneven surface. With this configuration, it is possible to produce a modulation element by, for example, surface processing of a light-transmitting member other than a mask.
[0086] Furthermore, the imaging system according to the present invention preferably includes a readout unit that reads out the image of the subject captured by the camera, and a recognizer that performs predetermined recognition of the subject's attached information from the readout image. According to the present invention, the recognition is performed directly by the recognizer without restoring it to a retinal image, which has the advantage of not involving any visually understandable image at all, thereby protecting privacy.
[0087] Furthermore, it is preferable that the modulation element and the recognizer are optimized in terms of both the degree of blurring of the image of the subject captured through the pattern of the modulation element and the recognition rate of the recognizer. With this configuration, it is possible to simultaneously achieve the best processing conditions for the blurring of the captured image through the pattern and the recognition rate of the recognition unit.
[0088] It is also preferable to have a storage unit that stores combinations of multiple types of patterns Hi (i=1, 2, ...) and recognition functions Ri optimized for each type of pattern, and a control unit that selects the combination of pattern and recognition function (Hi, Ri) when capturing an image. This configuration can be applied like a so-called code table, further enhancing security. [Explanation of symbols]
[0089] 1. Imaging System 11 Camera 12 Recognition unit (recognizer) 2, 2a, 2b, 2c, 2B Mask (modulation element) 20B Spatial Light Modulator (Modulation Element) 21,21c, 21b, 22b, 21c Transparent part 2d Plate (modulator) 21d Rough surface (partially transparent) 4. Image sensor 42 pixels
Claims
1. A lensless imaging system including a camera, a readout unit, and a recognition unit, The camera is a planar imaging element in which a plurality of pixels each made of a photosensitive element is arranged; a planar modulation element disposed in front of the imaging element and having a pattern formed thereon for modulating incident light; the pattern includes an array of light-transmitting portions each of which guides light rays from a plurality of directions from an object to one pixel; the reading unit reads out an image of a subject captured by the camera; the recognition unit performs a predetermined determination on the attached information of the subject from the read captured image; An imaging system characterized in that the modulation element and the recognition unit are optimized in terms of both the degree of blurring of the image of the subject captured through the pattern of the modulation element and the recognition rate of the judgment made by the recognition unit.
2. The recognition unit includes: a storage unit that stores in advance combinations of a plurality of types of patterns Hi (i = 1, 2, ...) and recognition functions Ri optimized for each type of pattern; 2. The imaging system according to claim 1, further comprising a control unit for selecting a combination pattern and a recognition function (Hi, Ri) when imaging.
3. 3. The imaging system according to claim 1, wherein the pattern includes a pattern that guides light rays that have passed through a common light-transmitting portion to a plurality of pixels.
4. 4. The imaging system according to claim 1, wherein the light-transmitting portion is a mask surface for blocking light and holes are formed therein.
5. 4. The imaging system according to claim 1, wherein the light-transmitting portion is a light-transmitting plate-like body having an uneven surface.
Citation Information
Patent Citations
Imaging apparatus and imaging method
JP2018061109A