Image acquisition device, method, and program

The image acquisition device modulates light to obscure specific attributes like facial features, addressing the issue of sensitive information capture and leakage by selectively omitting unwanted content while maintaining image quality.

JP7759006B2Active Publication Date: 2025-10-23NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024534815
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-10-23
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Existing image acquisition technologies fail to effectively prevent the capture and potential leakage of sensitive information, such as personal information or other unwanted content, during photography.

Method used

An image acquisition device that modulates light using spatially spreading patterns, selectively generates and applies modulation patterns to obscure specific attributes like facial features, and reconstructs images to omit unwanted information while preserving other details.

Benefits of technology

Enables the capture of images where sensitive information is omitted, ensuring privacy protection without compromising the quality of other image content, and securing against information leakage even if data is compromised.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007759006000001
    Figure 0007759006000001
  • Figure 0007759006000002
    Figure 0007759006000002
  • Figure 0007759006000003
    Figure 0007759006000003
Patent Text Reader

Abstract

According to the present invention, a modulation unit modulates light from a subject while switching between a plurality of modulation patterns that spread spatially. An observation unit acquires an observation signal for each modulation pattern by observing the light modulated by the modulation unit. A modulation pattern generation unit uses the modulation patterns up to the C-th modulation pattern and the observation signals corresponding to these intensity modulation patterns to generate the (C+1)-th modulation pattern such that the information amount for a region having a predetermined attribute is smaller than the information amount for the other region. An image reconstruction unit uses the (C+1)-th and succeeding modulation patterns and the observation signals corresponding to these modulation patterns to acquire a reconstructed image that represents the subject.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image acquisition device, method, and program. [Background technology]

[0002] With the recent spread of smartphones, it has become possible for anyone to easily take photos anywhere. However, there are cases where objects that should not be captured in a photo, such as the faces of third parties or the text on street signs, may appear in the photo from the perspective of protecting personal information and preventing information leaks. Therefore, there is a technology that processes the image data of a photo after it has been taken to remove objects that should not be captured (see, for example, Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Hakon Hukkelas, Rudolf Mester, and Frank Lindseth. Deepprivacy: A generative adversarial network for face anonymization. In International Symposium on Visual Computing, pp. 565-578. Springer, 2019. Summary of the Invention [Problem to be solved by the invention]

[0004] However, such technology may not necessarily protect personal information or prevent information leaks because the image data of photographs that show things that should not be captured is acquired by the photographing device. In other words, such technology may not be able to prevent information that should be prevented from leaking from being leaked through photography. Furthermore, the problem of information leaking through photography is not limited to personal information and information leaks, but is a common problem that occurs with any information that should be prevented from leaking. On the other hand, it is preferable to faithfully capture image data of things that do not pose a problem if captured. In view of the above circumstances, an object of the present invention is to provide a technique that can acquire an image in which information of things that should not be captured is selectively omitted during photography. [Means for solving the problem]

[0005] One aspect of the present invention is an image acquisition device having a modulation unit that modulates light from a subject by switching between multiple spatially spreading modulation patterns; an observation unit that acquires an observation signal for each modulation pattern by observing the light modulated by the modulation unit; a modulation pattern generation unit that uses the modulation patterns up to the Cth modulation pattern and the observation signal corresponding to the intensity modulation pattern to generate the (C+1)th modulation pattern so that information of a region having a predetermined attribute in an image reconstructed from the (C+1)th observation signal is made less clear than information of other regions; and an image reconstruction unit that acquires a reconstructed image representing the subject using the modulation patterns from the (C+1)th modulation pattern and the observation signal corresponding to the modulation pattern.

[0006] One aspect of the present invention is an image acquisition method comprising: a modulation step of modulating light from a subject while switching between multiple spatially spreading modulation patterns; an observation step of observing the modulated light to obtain an observation signal for each modulation pattern; a modulation pattern generation step of using the modulation patterns up to the Cth modulation pattern and the observation signal corresponding to the intensity modulation pattern to generate the (C+1)th modulation pattern so that information of a region having a predetermined attribute in an image reconstructed from the (C+1)th observation signal is made less clear than information of other regions; and an image reconstruction step of obtaining a reconstructed image representing the subject using the modulation patterns from the (C+1)th modulation pattern onwards and the observation signal corresponding to the modulation pattern.

[0007] One aspect of the present invention is a program for causing a computer to function as the image acquisition device according to the above aspect. [Effects of the Invention]

[0008] According to the above aspect, it is possible to obtain an image in which information of things that should not be captured is selectively omitted. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 10 is a diagram illustrating an example of the relationship between the number of elements of an observation signal and reconstructed image data. [Figure 2] 1 is a diagram illustrating an example of the configuration of an image acquisition device according to a first embodiment. [Figure 3] 1 is a schematic configuration diagram of an image acquisition device according to a first embodiment. [Figure 4] 10A and 10B are diagrams illustrating an example in which information about an area in which a face is captured is blocked in the first embodiment. [Figure 5] 2A to 2C are diagrams illustrating examples of images acquired by the image acquisition device according to the first embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of the configuration of an image acquisition device according to a second embodiment. [Figure 7] FIG. 1 illustrates a computer configuration according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] In the following embodiment, an example will be described in which an image acquisition device acquires a single monochrome image by observing visible light. Note that in other embodiments, the invention is not limited to monochrome images, and can also be applied to color images and multispectral images. Furthermore, in other embodiments, the invention can also be applied to non-visible electromagnetic waves such as near-infrared light instead of visible light.

[0011] <Image acquisition model> First, a model of image acquisition in the embodiment will be described. The model of image acquisition in the embodiment acquires an observation signal (y) that is the result of modulating light from the object O (unknown image x of the object O) with an observation matrix (Φ) based on the following formula (1), and then reconstructs an image of the object O from the observation matrix and the observation signal through calculation processing, thereby acquiring reconstructed image data (x') of the object O. Note that the unknown signal x is virtual two-dimensional image data that would be obtained if the object O were photographed with N pixels. However, since the object O is not actually photographed with N pixels, the unknown signal x does not actually exist.

[0012] y=Φx (1) However, in equation (1), y∈R M , x∈R N , Φ∈R M×N In equation (1), the unknown image x is an N-dimensional vector obtained by rearranging two-dimensional image data with N pixels into one dimension. The observed signal y is an M-dimensional vector. Furthermore, all elements of the observed signal y are not acquired simultaneously, but are acquired sequentially. Note that the elements of the observed signal y do not have to be acquired completely sequentially. For example, an operation of simultaneously acquiring four elements of the observed signal y may be repeated sequentially. The elements of the observed signal y are acquired at a sufficiently high speed, and the light (x) from the subject O can be considered to be constant during this time.

[0013] The observation matrix Φ uses values generated from random numbers of a normal distribution. When M ≪ N, the reconstructed image data x' can only be roughly restored, and it is known that when M / N is about 1 / 4, the reconstructed image data x' can be restored almost without degradation (x' ≒ x). Figure 1 is a diagram showing an example of the relationship between the number of elements of the observed signal and the reconstructed image data. Also, based on Equation (1), a method of obtaining an observed signal y with fewer elements than the unknown image x, that is, when M < N, and obtaining a reconstructed image x' by restoring the unknown image x through computational processing is also called compressive sensing. However, in the image acquisition method in the embodiment, it is not necessarily required that M < N.

[0014] The image acquisition device according to the embodiment grasps the content of the subject O roughly by generating a provisional reconstruction from the observation matrix and the observation signal up to a certain point in an image acquisition method that acquires the light from the subject O while modulating it and sequentially acquires the modulated components, and devises the subsequent modulation method to block the acquisition of detailed information of a specific target. That is, by controlling the observation matrix Φ according to the sequentially acquired observation signal y, it is realized that the image information of a specific target is not acquired. Therefore, if it conforms to the above model (acquiring the light from the subject O while modulating it and sequentially acquiring the modulated components), it can be used without depending on a specific observation device. For example, single pixel imaging and ghost imaging conform to the above model. Note that in the image acquisition method in the embodiment, whether M < N (compressive sensing) or M ≧ N, the image information of a specific target can be selectively omitted. Also, in the following embodiments, the attribute of the target not to be acquired, that is, the detailed information of a specific target is a feature that can identify an individual's face. That is, the following embodiments realize an image acquisition device that does not capture features of a face that can identify an individual. Note that the attribute of the target not to be acquired can also be applied to, for example, characters other than faces.

[0015] <First Embodiment> The image acquisition device 1 according to the first embodiment acquires an image by single pixel imaging, in which information about facial features that can identify an individual is selectively omitted. FIG. 2 is a diagram showing an example of single-pixel imaging. In single-pixel imaging, an image is acquired by repeating the process of intensity modulation by the spatial light modulator 12 and then observing the luminance of the incident light by the photoelectric converter 13 M times while changing the modulation pattern. In this case, the i-th modulation pattern is represented by the i-th row vector (Φ i ∈R N ) corresponds to the observation matrix Φ. Each modulation pattern by the spatial light modulator 12 is a pattern with N elements distributed on a two-dimensional plane, and the modulation pattern on the two-dimensional plane is rearranged into one dimension, that is, an N-dimensional vector is the row vector of the observation matrix Φ. The value (observation value) acquired by the photoelectric converter 13 at the i-th time is expressed as y i The i-th observation y i is the observed result of modulating the light from the object O with the i-th modulation pattern, and y i =Φ i x. In other words, the modulation pattern Φ i and the observed signal y i The relationship between these corresponds to equation (1). The image acquisition device 1 according to the first embodiment reconstructs an image of the object O from M modulation patterns (observation matrix Φ) and M corresponding acquired values ​​(observation signals y). Note that in other embodiments, the modulation pattern may be any pattern that is spatially distributed, such as a pattern distributed on a one-dimensional straight line or a pattern distributed on a curved surface.

[0016] A first embodiment of an image capture device 1 for blocking acquisition of facial feature information that can identify an individual using single-pixel imaging will now be described with reference to Fig. 3. Fig. 3 is a schematic diagram of the image capture device 1 according to the first embodiment. The image acquisition device 1 includes an observation device 10 and a control device 30. The observation device 10 and the control device 30 alternately repeat M times the acquisition of observation values ​​based on a modulation pattern by the observation device 10 and the generation of a modulation pattern based on the observation values ​​by the control device 30. The control device 30 then outputs a reconstructed image x' as an acquired image based on the observation values ​​and modulation patterns up to the Mth time. M may be a predetermined number of times or may be adaptively determined based on a predetermined condition. The predetermined condition is, for example, determined based on a provisional reconstructed image that is calculated sequentially. The control device 30 outputs the provisional reconstructed image after the Mth repetition as the output of the image acquisition device 1.

[0017] The observation device 10 includes a housing 11, a spatial light modulator 12, and a photoelectric converter 13. The spatial light modulator 12 is provided in an opening of the housing 11. The photoelectric converter 13 is provided inside the housing 11 so as to face the spatial light modulator 12. Thus, the photoelectric converter 13 is configured to observe the light that has passed through the spatial light modulator 12.

[0018] The spatial light modulator 12 is configured to modulate the light beam in accordance with a modulation pattern (Φ i ) to intensity-modulate the light incident from the object O. The spatial light modulator 12 uses, for example, a given modulation pattern (Φ i ) and may be, for example, a transmissive spatial light modulation device configured with a transmissive liquid crystal microdisplay.

[0019] The photoelectric converter 13 converts the intensity of the received light into an electrical signal, digitizes it, and outputs it. The photoelectric converter 13 is composed of, for example, a photomultiplier tube or a photodiode. The value acquired by the photoelectric converter 13 is the i-th element of the observation signal y (y i ), which are called observations.

[0020] The control device 30 includes an observation signal storage unit 31, an image reconstruction unit 32, a modulation pattern generation unit 33, and an observation matrix storage unit 34. The control device 30 generates a modulation pattern Φ i and the determined modulation pattern Φi is output to the observation device 10.

[0021] The observation signal storage unit 31 receives the observation value (y i ) and records them with an index indicating the order of acquisition. The observed signal storage unit 31 outputs an observed signal (y) whose elements are the observed values ​​recorded up to that point.

[0022] The image reconstruction unit 32 performs image reconstruction processing from the observation matrix Φ recorded in the observation matrix storage unit 34 and the observed signal y recorded in the observed signal storage unit 31. The image reconstruction unit 32 performs image reconstruction processing from, for example, the pseudo-inverse matrix Φ of the observation matrix Φ + and calculate the pseudo-inverse matrix Φ + and the observed signal y, the reconstructed image x' can be calculated. Note that in other embodiments, the image reconstructing unit 32 may calculate the reconstructed image x' using other methods. The image reconstructing unit 32 outputs the reconstructed image x' generated up to the M-1th time as the provisional reconstructed image. When the number i of observed values ​​acquired up to that point (when the i-th element has been acquired) is sufficiently smaller than N, this provisional reconstructed image is expected to be an inaccurate image, i.e., the facial features shown in the provisional reconstructed image are expected to be unclear. On the other hand, the image reconstructing unit 32 sets the reconstructed image x' generated for the Mth time as output data of the image acquisition device 1.

[0023] The modulation pattern generator 33 generates a modulation pattern from the tentatively reconstructed image acquired by the image reconstructor 32. Specifically, the modulation pattern generator 33 generates a modulation pattern Φ to be used for the (i+1)th element by using the tentatively reconstructed image generated from the observed signal acquired up to the i-th element. i+1 The modulation pattern generation unit 33 generates the (i+1)th modulation pattern so that, in a tentative reconstructed image reconstructed from the (i+1)th observed signals, information on an area where facial features exist becomes unclear compared to information on other areas. A specific method for generating the modulation pattern by the modulation pattern generation unit 33 will be described later.

[0024] The observation matrix memory unit 34 records the modulation pattern Φ generated by the modulation pattern generation unit 33 while attaching an index indicating the generation order. The observation matrix memory unit 34 outputs a modulation matrix Φ having the modulation patterns Φ i recorded so far as elements. i as an element.

[0025] Hereinafter, according to the first embodiment, the reason why a function capable of accurately acquiring information on other parts in addition to acquiring only face information can be realized will be explained. When i≪N, it is expected that the modulation pattern generation unit 33 is given a provisional reconstruction image that can only restore rough features. In this case, it is impossible to determine in which region of the provisional reconstruction image the face exists. At this time, the modulation pattern generation unit 33 generates random numbers with a normal distribution in all regions and outputs them as the modulation pattern Φ i+1 . Also, when i = 0 when the provisional reconstruction image does not exist yet, the modulation pattern generation unit 33 similarly generates random numbers with a normal distribution and outputs them as the first modulation pattern Φ1.

[0026] If the modulation pattern generation unit 33 continues to output a modulation pattern that is a random number with a normal distribution, according to the compressive sensing principle, every time a new y i is acquired, the provisional reconstruction image gradually becomes more faithful to the subject. If, hypothetically, the modulation pattern that is a random number with a normal distribution is continuously output until about i = M = N / 4, as in general single pixel imaging, all of the subject may be accurately acquired, and there is a possibility that an image including privacy information such as face features may be acquired as the reconstructed image.

[0027] It can be expected that the provisional reconstruction image at a certain point C where i is sufficiently smaller than N can determine the approximate position of the object in the image, but cannot determine the details. That is, at a certain i (C≦i<M), it is expected that a provisional reconstruction image with a fidelity such that the position of the face can be known is input to the modulation pattern generation unit 33 while facial features that can identify an individual do not appear. At this time, the modulation pattern Φ i+1In this example, if all elements corresponding to the areas of facial features necessary for identifying an individual are set to 0, the observation device 10 will no longer acquire facial features that can identify an individual from the (i+1)th onward. Fig. 4 is a diagram showing an example of blocking information about areas that include a face in the first embodiment. In Fig. 4, the element values ​​(modulation intensity) of areas of the subject O through which light from an individual's face O2 passes are set to 0, and a modulation pattern is used in which the element values ​​of other areas, including a non-face object O1, are determined by normal distribution, so that the spatial light modulator 12 acquires information about the object O1 while blocking acquisition of information about the individual's face O2.

[0028] In subsequent provisionally reconstructed images, the fidelity of the blocked area remains unchanged from time C, while the fidelity of other areas gradually improves. By repeating this process up to the Mth time, the image capture device 1 can capture an image in which facial feature information is selectively omitted and information in other areas is faithfully reproduced. In this way, the image acquisition device 1 can realize the function of not only acquiring facial information but also accurately acquiring other parts by devising a modulation pattern using a provisional reconstructed image at a point C less than M times.

[0029] In other words, the image acquisition device 1 according to the first embodiment has a feedback structure that determines subsequent modulation patterns based on observed values ​​up to a certain point in time. This feedback structure uses observed values ​​up to a certain point in time to identify areas containing objects with attributes that should not be captured, and subsequently uses a modulation pattern that does not capture details of those areas, thereby realizing a function of not capturing details of those areas that should not be captured. In this way, this feedback structure enables the image acquisition device 1 to selectively acquire images that distinguish between attributes that should be captured and attributes that should not be captured. For example, the image acquisition device 1 can avoid capturing facial details and faithfully capture everything else. Note that in other embodiments, the image acquisition device 1 can similarly avoid capturing only text such as license plates. Furthermore, unlike encryption or obfuscation technology, even if all acquired data were leaked, data with attributes that should not be captured cannot be restored by anyone, eliminating the risk of leakage.

[0030] In other words, encryption and obfuscation technologies make it difficult to read existing original data that includes information that should not be copied, and therefore, if the original data is leaked, the information that should not be copied will be read. In contrast, with the image acquisition device according to the first embodiment, even if the observed signal y stored in the observed signal storage unit 31 and the observation matrix Φ stored in the observation matrix storage unit 34 are leaked and a third party attempts to reconstruct an image based on these, the information in the area of ​​the reconstructed image where the information that should not be copied exists will be missing, and the information that should not be copied will not be read.

[0031] 5 is a diagram showing an example of an image acquired by the image acquisition device 1 according to the first embodiment. The image shown in FIG. 5 is an image reconstructed from the observed signal y and the observation matrix Φ. In the image shown in FIG. 5, facial features cannot be recognized, but other information (such as the background, the person's clothing, and their posture) is faithfully reproduced. In this way, the image acquisition device 1 according to the first embodiment can selectively acquire information by separating what is acceptable to capture (what should be captured) from what is not. The image acquisition device 1 controls information about what is not acceptable to capture so that it is not observed. Therefore, even if data is leaked, it is difficult to restore information about what is not acceptable to capture, making the image highly secure. On the other hand, since the image data other than the face is faithfully acquired, it can be used for various image processing applications such as general object recognition and posture estimation. In other words, the image acquisition device 1 according to the first embodiment can achieve privacy protection without sacrificing convenience.

[0032] An example of the modulation pattern generating unit 33 will be described below.

[0033] <First example of modulation pattern generation unit 33> The modulation pattern generation unit 33 according to the first example generates the following modulation pattern by executing a random number generation process, a detection process, and a mask process. The modulation pattern generation unit 33 first generates a base pattern using random numbers with a Gaussian distribution (random number generation process). The base pattern is a pattern in which the modulation intensities of all regions are determined by the random numbers. Next, the modulation pattern generation unit 33 detects a region in the provisional reconstructed image that contains an object that should not be captured, i.e., a face region (detection process). The face region may be detected by, for example, a pattern matching process or a process using a trained model for detecting face regions. The method for extracting the face region is not limited to a specific method. The modulation pattern generation unit 33 estimates the positions of features (e.g., eyes, nose, mouth, etc.) that represent individual characteristics from the detected face region. Next, the modulation pattern generation unit 33 generates a modulation pattern by replacing the values ​​of elements at the estimated positions of the features in the base pattern generated by the random numbers with zero (mask process). This is expected to prevent further acquisition of information about objects that satisfy the occlusion condition.

[0034] It should be noted that the masking process by the modulation pattern generating unit 33 is not limited to completely blocking the acquisition of information by setting all elements of the modulation pattern of the face region to 0, but may be configured so that the amount of information about the modulation intensity of the face region is smaller than the amount of information about the modulation intensity of other regions. For example, the modulation pattern generating unit 33 may smooth the base pattern of the face region, or may acquire only the average value of the pixels in the face region by setting all modulation patterns in the face region to the same value, or may acquire only a specific spatial frequency component of the face region by making the modulation pattern of the face region a striped pattern of a specific frequency. Furthermore, the amount of information on the modulation intensity is not limited to the amount of information on the spatial axis, but may be the amount of information on the time axis. For example, the modulation pattern generation unit 33 may match the pattern of the face region in the next modulation pattern with the current modulation pattern.

[0035] <Second Example of Modulation Pattern Generator 33> The modulation pattern generation unit 33 according to the second example generates a modulation pattern using a learned model (pattern generation model) that is learned to take a provisional reconstruction image as an input and output a modulation pattern by, for example, a machine learning method. The learning of the parameters of the pattern generation model is performed in the following procedure. First, a large amount of image data is prepared as image data of a simulated subject O (hereinafter referred to as simulated unknown image data). The learning device generates an L-dimensional observation signal using the simulated unknown image data and L base patterns (1 < L < M) generated based on random numbers. Note that L may be determined based on uniform random numbers or may be determined according to normal random numbers.

[0036] The base pattern according to the second example may be generated according to random numbers of a normal distribution or may be generated by the pattern generation model during learning. When generating the base pattern by the pattern generation model during learning, the learning device can generate the base pattern by the following recursive procedure. First, the learning device first generates the first base pattern based on random numbers. The learning device generates the first base pattern by inputting the provisional reconstruction image obtained from the first base pattern that has already been generated into the pattern generation model during learning. The learning device generates the third base pattern by inputting the provisional reconstruction image obtained from the two (first and second) base patterns that have already been generated into the pattern generation model during learning. By repeating this, the learning device can generate the L-th base pattern by inputting the provisional reconstruction image obtained from the L - 1 base patterns that have already been generated into the pattern generation model during learning.

[0037] The learning device calculates the Lth tentative reconstructed image data based on an L-dimensional observed signal and an observation matrix consisting of L base patterns. Next, the learning device obtains the L+1th modulation pattern by inputting the Lth tentative reconstructed image data into the pattern generation model being trained. The learning device calculates observed values ​​using simulated unknown image data and the modulation pattern generated by the pattern generation model. The learning device generates an L+1-dimensional observed signal using the L-dimensional observed signal and the newly calculated observed values. The learning device calculates the L+1th tentative reconstructed image data based on the L+1-dimensional observed signal and an observation matrix consisting of L base patterns and modulation patterns. The learning device calculates a loss function from the simulated unknown image data, the Lth tentative reconstructed image data, and the L+1th tentative reconstructed image data using the loss function. The learning device updates the parameters of the pattern generation model so as to reduce the loss function. The learning device learns the parameters of the pattern generation model by repeating this procedure for a predetermined number of epochs.

[0038] The loss function includes two terms: a faithfulness term and a concealment term. The faithfulness term is a term that reduces the difference between the simulated unknown image data and the (L+1)th reconstructed image data. In other words, the higher the fidelity of the (L+1)th reconstructed image data, the smaller the value of the faithfulness term. The faithfulness term may be, for example, the norm of the difference between the simulated unknown image data and the (L+1)th reconstructed image data. The concealment term is a term that increases the difference in feature space when the facial regions included in the simulated unknown image data and the facial regions included in the (L+1)th tentative reconstructed image data are converted into facial feature vectors using a predetermined feature extractor (e.g., the intermediate layer output of a face recognition model). Increasing the difference in feature space means that the faces are not determined to be the same person. In other words, the more blurred the facial features in the (L+1)th reconstructed image data, the smaller the value of the concealment term. The concealment term may be, for example, the norm of the difference between the facial feature vector of the simulated unknown image data and the facial feature vector of the (L+1)th reconstructed image data multiplied by -1. In addition, the occlusion term may be obtained by inputting the face area contained in the simulated unknown image data and the face area contained in the L+1th provisional reconstructed image data into a model that calculates the probability that the person appearing in two images is the same person, and multiplying the probability by -1.

[0039] Furthermore, in the above example, the learning device calculates the loss function using the (L+1)th reconstructed image data for generating the (L+1)th modulation pattern, but this is not limited to this. Alternatively, the learning device may acquire reconstructed image data up to the (L+K)th (K is an appropriate integer) reconstructed image data beyond the (L+1)th, and calculate the loss function using the (L+K)th reconstructed image data. In other words, the learning device may calculate the loss function for the (L+1)th modulation pattern so as to evaluate the blurring of facial features in the reconstructed image data obtained after the (L+1)th, and the fidelity of the entire image. By expressing the loss function as a combination of the occlusion term and the faithful term, the learning device can learn the parameters of the pattern generation model so as to generate a modulation pattern that can faithfully acquire other information while missing information in the face region in the image of subject O. The machine learning method is, for example, deep learning.

[0040] In this way, the modulation pattern generation unit 33 according to the second example generates a modulation pattern by inputting the provisional reconstructed image into a trained pattern generation model trained by executing a machine learning process. In this case as well, the modulation pattern generated by the modulation pattern generation unit 33 is configured so that the amount of information on the modulation intensity of the face region is smaller than the amount of information on the modulation intensity of other regions.

[0041] <Third Example of Modulation Pattern Generator 33> In the second example, the learning device generates "simulated acquired data" using a specific reconstruction method (similar to that used by the image acquisition device 1) and evaluates the concealment performance based on the provisionally reconstructed image. On the other hand, if the observed signals and observation matrix are leaked and a third party (hereinafter referred to as an attacker) attempts to obtain facial information by performing image reconstruction from the observed signals and observation matrix, it is unknown what image reconstruction method the attacker will use. Because the reconstruction method used may differ, the concealment performance expected during actual shooting and training may not necessarily match the concealment performance achieved by the attacker when reconstructing images. In particular, if the attacker uses a reconstruction method that provides better facial feature restoration, concerns about concealment performance arise. Therefore, to improve concealment performance, the learning device reconstructs images using a machine learning model (reconstruction model) and trains the reconstruction model and the pattern generation model using opposing loss functions. That is, the loss function of the reconstruction model has a faithful term similar to that of the pattern generation model and a term that reduces the difference between the face region in the simulated unknown image data and the face region in the (L+1)th reconstructed image data when viewed in feature space, i.e., a term that is opposing the concealment term of the pattern generation model. As a result, the parameters of the reconstruction model are trained to reconstruct facial information as faithfully as possible, while the parameters of the pattern generation model are trained so that concealment can be ensured even with the provisionally reconstructed image data generated by that reconstruction model. Because the reconstruction model is trained to reconstruct facial information as faithfully as possible, it is unlikely that an attacker will be able to restore facial information more faithfully than this, and it is expected that security will be stronger than in the second example. Note that the reconstruction model may be any learning model.

[0042] In this way, the modulation pattern generation unit 33 according to the third example generates a modulation pattern by inputting a provisional reconstructed image into a pattern generation model that has been trained adversarially against the reconstruction model. In this case, too, the modulation pattern generated by the modulation pattern generation unit 33 is configured so that the amount of information about the modulation intensity of the facial region is smaller than the amount of information about the modulation intensity of other regions. Note that the image reconstruction unit 32 of the image acquisition device 1 generates reconstructed image data without using a trained reconstruction model. This is because the image acquisition device 1 is not required to faithfully reproduce facial information.

[0043] <Second embodiment> The image acquisition device 1 according to the second embodiment includes: Ghost imaging allows for the acquisition of images that selectively omit information about facial features that can identify an individual. FIG. 6 is a diagram showing an example of the configuration of an image acquisition device 1 according to the second embodiment. The spatial light modulator 12 according to the second embodiment receives a modulation pattern (Φ i ) and irradiates the light modulated by the spatial light modulator 12 onto the object O. The spatial light modulator 12 may be, for example, a projector. The photoelectric converter 13 according to the second embodiment is provided at a position where it can receive the modulated light reflected by the object O. The configuration of the control device 30 is the same as that of the first embodiment.

[0044] In this way, the image acquisition device 1 according to the second embodiment can acquire an image in which information about facial features that can identify an individual is selectively omitted by ghost imaging.

[0045] <Other embodiments> Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. For example, the image acquisition device 1 according to the embodiment described above generates modulation patterns that successively reduce the amount of information in the face area after the start of acquisition of observation values, but is not limited to this. For example, in another embodiment, the number C of modulations at which the face area can be detected may be specified in advance, and the modulation pattern generation unit 33 may generate modulation patterns up to the Cth time based on random numbers regardless of whether the face area can be detected, and generate modulation patterns from the C+1th time onwards so that the amount of information in the face area is reduced.

[0046] Furthermore, the modulation pattern generating unit 33 according to the embodiment described above generates a modulation pattern Φ based on the provisional reconstructed image x′ generated by the image reconstructing unit 32. i For example, in another embodiment, the modulation pattern generator 33 generates a modulation pattern Φ based on the observed signal y and the observation matrix Φ. i That is, the modulation pattern generation unit 33 may generate the modulation pattern Φ directly from the observed signal y and the observation matrix Φ without calculating the tentative reconstructed image x' based on the observed signal y and the observation matrix Φ. i For example, the pattern generation models according to the second and third examples may take the observed signal y and the observation matrix Φ as input and generate the modulation pattern Φ i In this case, the image reconstructing unit 32 may calculate the reconstructed image x' from the observed signal y and the observation matrix Φ only the Mth time.

[0047] <Computer configuration> FIG. 7 is a diagram illustrating a computer configuration according to at least one embodiment. The control device 30 or the learning device may be configured by a computer 50 including a processor 51, a memory 52, an auxiliary storage device 53, an interface 54, etc., which are connected by a bus, as shown in Fig. 7. The computer 50 functions as the control device 30 including an observation signal storage unit 31, an image reconstruction unit 32, a modulation pattern generation unit 33, and an observation matrix storage unit 34 by executing an image acquisition program. The computer 50 functions as the learning device by executing a learning program. Examples of the processor 51 include a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), and a microprocessor. The image acquisition program or the learning program may be recorded on a computer-readable recording medium. The computer-readable recording medium may be the auxiliary storage device 53 or an external storage device connected via the interface 54. Examples of the computer-readable recording medium include storage devices such as magnetic disks, magneto-optical disks, optical disks, and semiconductor memories. The image acquisition program may be transmitted via a telecommunications line. Note that all or part of the functions of the image acquisition device or the learning device may be implemented using a custom LSI (Large Scale Integrated Circuit), such as an ASIC (Application Specific Integrated Circuit) or a PLD (Programmable Logic Device). Examples of PLDs include PAL (Programmable Array Logic), GAL (Generic Array Logic), CPLD (Complex Programmable Logic Device), and FPGA (Field Programmable Gate Array). Such integrated circuits are also included in the processor 51. [Explanation of symbols]

[0048] 1...Image acquisition device 10...Observation device 11...Housing 12...Spatial light modulator 13...Photoelectric converter 30...Control device 31...Observation signal storage unit 32...Image reconstruction unit 33...Modulation pattern generation unit 34...Observation matrix storage unit

Claims

1. a modulation unit that modulates light from a subject by switching between multiple spatially expanding modulation patterns; an observation unit that acquires an observation signal for each modulation pattern by observing the light modulated by the modulation unit; a modulation pattern generation unit that generates the (C+1)th modulation pattern using the modulation patterns up to the Cth modulation pattern and the observed signals corresponding to the modulation patterns so that information of a region having a predetermined attribute in an image reconstructed from the (C+1)th observed signal becomes less clear than information of other regions; an image reconstruction unit that acquires a reconstructed image representing the subject using the modulation patterns from (C+1) onward and the observation signals corresponding to the modulation patterns; An image acquisition device having:

2. The modulation pattern generation unit generates the (C+1)th modulation pattern from the reconstructed image generated from the modulation patterns up to the Cth time and the observed signal, using a trained pattern generation model that has been trained to receive a reconstructed image as an input and output a modulation pattern. The image acquisition device of claim 1 .

3. The pattern generation model is calculating first reconstructed image data from the training image data; inputting the first reconstructed image data into the pattern generation model to obtain a tentative modulation pattern; calculating an observation value using the learning image data and the tentative modulation pattern; calculating second reconstructed image data using the tentative modulation pattern and a learning observation signal including an observation value obtained in a calculation process of the first reconstructed image data and the observation value calculated using the modulation pattern; updating parameters of the pattern generation model so as to reduce a difference between the learning image data and the second reconstructed image data and to increase a difference between a feature amount of the predetermined attribute included in the learning image data and a feature amount of the predetermined attribute included in the second reconstructed image data; Learned through learning methods including The image acquisition device of claim 2 .

4. In the step of calculating the second reconstructed image data in the learning method, the second reconstructed image data is calculated using a reconstruction model that receives the learning observation signal and the tentative modulation pattern as inputs and outputs reconstructed image data; The learning method further includes a step of updating parameters of the reconstruction model so as to reduce a difference between the learning image data and the second reconstructed image data and to reduce a difference between a feature amount of the predetermined attribute included in the learning image data and a feature amount of the predetermined attribute included in the second reconstructed image data. Contains The image acquisition device of claim 3 .

5. The modulation pattern generation unit A spatially spreading base pattern is generated based on random numbers, a region having the predetermined attribute is detected from a reconstructed image generated from the modulation patterns up to the Cth time and the observed signal, and the amount of information of the detected region in the base pattern is reduced to generate the (C+1)th modulation pattern. The image acquisition device of claim 1 .

6. The modulation pattern generation unit generates the modulation pattern every time the observation unit acquires the observation signal. The image acquisition device according to any one of claims 1 to 5.

7. a modulation step of modulating light from a subject while switching between a plurality of spatially spreading modulation patterns; an observation step of observing the modulated light to obtain an observation signal for each modulation pattern; a modulation pattern generation step of generating the (C+1)th modulation pattern using the modulation patterns up to the Cth modulation pattern and the observed signals corresponding to the modulation patterns so that information of a region having a predetermined attribute in an image reconstructed from the (C+1)th observed signal is less clear than information of other regions; an image reconstructing step of acquiring a reconstructed image representing the subject using the modulation patterns from (C+1) onward and the observation signals corresponding to the modulation patterns; An image acquisition method comprising:

8. A program for causing a computer to function as the image acquisition device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Privacy-protecting biometric feature recognition method and privacy-protecting biometric feature recognition device

    CN114140823A

  • Monitoring system

    WO2021157721A1