Recognition system, recognition method, collection device, and program
The recognition system addresses privacy and efficiency challenges in face recognition by anonymizing features through random unitary transformations and dictionary matrix learning, achieving effective face identification with a 71.3% discrimination rate.
Patent Information
- Application Number
- JP2024123208
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Existing image recognition systems face challenges in efficiently performing data analysis while protecting privacy, particularly in face recognition, and require significant user effort in preparing learning input images.
A recognition system that extracts facial regions, applies random unitary transformations to anonymize features, learns a dictionary matrix, and estimates observer identity using anonymized data, leveraging edge and cloud computing to maintain privacy.
Enables efficient face recognition with privacy protection by processing anonymized data, achieving a 71.3% discrimination rate in identifying similar faces while maintaining privacy.
Smart Images

Figure 2026021937000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a recognition system, a recognition method, a collection device, and a program. [Background technology]
[0002] Image recognition, such as face recognition and person matching, has a wide range of practical applications, including security, and has been the subject of active research for a long time. By using edge cloud computing, such image recognition systems can offload processing that is difficult to perform on the edge side to the cloud, which is expected to improve computational efficiency.
[0003] One method for performing data analysis and signal processing while protecting privacy is private sparse modeling based on random unitary transformations (Non-Patent Document 1).
[0004] There is a method for learning by concealing pixel values of face image data (Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Takayuki Nakachi, Yukihiro Bando, "Sparse Data Modeling in the Privacy-Preserving Domain," IEICE Fundamentals Review Vol.16 No.2 [Non-patent document 2] Wang, et.al, “Edge and cloud-aided secure sparse representation for face recognition”, EUSIPCO, 2019. Summary of the Invention [Problem to be solved by the invention]
[0006] Non-Patent Document 1 can achieve both privacy protection and data analysis, but does not disclose any specific applications of the data analysis.
[0007] In Non-Patent Document 2, as shown in Fig. 2-3, one learning input image contains only the facial region of one person to be authenticated, which makes learning easy, but imposes a heavy burden on the user in preparing the learning input image.
[0008] The present disclosure has been made in consideration of the above circumstances, and an object of the present disclosure is to provide a technology that can easily recognize a face image while protecting privacy. [Means for solving the problem]
[0009] A recognition system according to one aspect of the present disclosure includes a calculation unit that extracts facial regions of a plurality of recognition target persons from training image data and calculates training features for each of the plurality of facial regions; a transformation unit that converts each of the plurality of training features into a plurality of anonymized training features using a random unitary transformation; a learning unit that learns a dictionary matrix using the plurality of anonymized training features; and an estimation unit that outputs an estimation result of an observer of observed image data using the dictionary matrix.
[0010] A recognition method according to one aspect of the present disclosure includes a computer extracting facial regions of a plurality of recognition target persons from training image data, calculating training features for each of the plurality of facial regions, converting each of the plurality of training features into a plurality of anonymized training features using a random unitary transformation, learning a dictionary matrix using the plurality of anonymized training features, and outputting an estimation result of the observer of the observed image data using the dictionary matrix.
[0011] A collection device according to one aspect of the present disclosure includes a calculation unit that extracts facial regions of a plurality of recognition target persons from training image data and calculates training features for each of the plurality of facial regions, and a conversion unit that converts each of the plurality of training features into a plurality of concealed training features using a random unitary transformation.
[0012] One aspect of the present disclosure is a program that causes a computer to function as the collection device or the recognition device. [Effects of the Invention]
[0013] According to the present disclosure, it is possible to provide a technology that can recognize a face image while protecting privacy. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram illustrating the system configuration of a recognition system according to the present disclosure. [Figure 2] FIG. 2 is a functional block diagram of the collection device. [Figure 3] FIG. 3 is a functional block diagram of the recognition device. [Figure 4] FIG. 4 is a diagram illustrating the relationship between the observed signal, the dictionary matrix, and the sparse coefficients. [Figure 5] FIG. 5 is a diagram illustrating an example of a method for calculating a sparse model. [Figure 6] FIG. 6 is a sequence diagram illustrating an example of processing in the recognition system. [Figure 7] FIG. 7 is a diagram illustrating the hardware configuration of a computer used in the collection device or recognition device. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the description of the drawings, the same parts are designated by the same reference numerals and the description thereof will be omitted.
[0016] (Recognition System) The recognition system 1 according to the present disclosure identifies a recognition target person that has a high degree of similarity to a facial area included in a newly acquired image from among facial areas of multiple recognition targets that have been previously learned. The recognition system 1 is used, for example, to determine whether a person who enters a surveillance area where a surveillance camera is installed is a pre-registered recognition target person.
[0017] As shown in FIG. 1, the recognition system 1 includes an imaging device 10, a collection device 20, and a recognition device 30.
[0018] The image capturing device 10 captures image data including one or more human face regions. The image capturing device 10 is, for example, a surveillance camera, and captures images of people entering and leaving a monitored area. The recognition system 1 may include one or more image capturing devices 10. The number of image capturing devices 10 is not important.
[0019] The image capturing device 10 captures training image data 21 used during training and observed image data 22 used during estimation. The training image data 21 includes face regions of one or more recognition targets. The recognition targets are, for example, people who are permitted to enter and exit the area to be monitored. The observed image data 22 includes face regions of one or more observers. The observers are people who have entered and exited the area to be monitored, and it is determined whether or not they are recognition targets who have been permitted to enter and exit in advance.
[0020] The collection device 20 collects learning image data 21 and observed image data 22 from the image capture device 10 and performs processing to transmit them to the recognition device 30. Specifically, the collection device 20 extracts a face region included in the image data, conceals the feature amount of the extracted face region, and transmits them to the recognition device 30.
[0021] The recognition device 30 learns a dictionary matrix using feature amounts of the facial region of the recognition target person extracted and anonymized from the training image data 21. The recognition device 30 also identifies a recognition target person having a high degree of similarity to the facial region of the observer using feature amounts of the facial region of the observer extracted and anonymized from the observed image data 22, and outputs the identifier of the identified recognition target person as an estimation result. The estimation result output by the recognition device 30 is transmitted to the collection device 20, for example, and displayed.
[0022] The image capturing device 10 and the collection device 20 are installed near a monitored area. The image capturing device 10 and the collection device 20 are connected using a communication network provided within a limited range, such as a local area network (LAN) or short-range wireless communication. The collection device 20 realizes edge computing.
[0023] The recognition device 30 is installed at a location away from the monitored area, such as a data center on a cloud. The recognition device 30 is, for example, a cloud computer. The recognition device 30 and the collection device 20 are bidirectionally connected using a communication network that is provided over a wide area, such as a commercial or public network.
[0024] In the present disclosure, a first path from the image capture device 10 to the collection device 20 is shorter than a second path from the image capture device 10 to the recognition device 30. The image capture device 10 is arranged to connect to the recognition device 30 via the collection device 20.
[0025] In the recognition system 1 according to the present disclosure, the collection device 20 does not transmit the image data itself to the recognition device 30, but transmits the image data in a concealed form. The recognition system 1 can learn a dictionary matrix and estimate the observer without transmitting private information, such as the face of the person to be recognized or the observer, over a widely available communication network.
[0026] (collection device) A collection device 20 according to the present disclosure will be described with reference to Fig. 2. The collection device 20 includes training image data 21, observed image data 22, training feature values PL, and observed feature values PO, as well as functions of a calculation unit 26 and a conversion unit 27. Each piece of data is stored in a storage device such as a memory 902 or a storage 903. Each function is implemented in a CPU 901.
[0027] The training image data 21 and the observed image data 22 are image data collected from the image capture device 10. The training image data 21 includes face regions of one or more recognition targets. The training image data 21 is teacher data used for learning the dictionary matrix in the recognition device 30. The observed image data 22 includes face regions of one or more observers. The observed image data 22 is used to estimate recognition targets similar to the observer using the dictionary matrix.
[0028] The collection device 20 processes the training image data 21 during learning of the dictionary matrix, and processes the observed image data 22 during estimation using the dictionary matrix. The collection device 20 performs similar processing during learning and estimation.
[0029] First, the learning process will be described.
[0030] The calculation unit 26 extracts face regions of each of the multiple recognition targets from the training image data 21 and calculates feature amounts for each of the multiple face regions as training feature amounts PL. Here, the calculation unit 26 preferably outputs feature amounts that are robust to the influence of the shooting environment, such as lighting and facial orientation, rather than pixel values of the face regions. The calculation unit 26 calculates the feature amounts using general-purpose software such as dlib (https: / / dlib.net / face_recognition.py.html).
[0031] The conversion unit 27 converts the training feature PL into the concealed training feature PLe using a random unitary transformation. The random unitary transformation is disclosed in, for example, Non-Patent Document 1.
[0032] When the training image data 21 includes face regions of multiple recognition targets, the collection device 20 calculates a training feature PL for each of the multiple recognition targets. The collection device 20 conceals each training feature PL and calculates each concealed training feature PLe. Furthermore, when there are multiple training image data 21, the collection device 20 calculates a concealed training feature PLe for each training image data 21 and each recognition target.
[0033] The collection device 20 transmits the anonymized learning features PLe to the recognition device 30.
[0034] Next, the estimation process will be described.
[0035] The calculation unit 26 extracts the face region of the observer from the observed image data 22 and calculates the feature amount of the face region as the observed feature amount P0. As in the learning process, the calculation unit 26 preferably calculates feature amounts that are robust to the influence of the shooting environment using dlib or the like.
[0036] The transform unit 27 transforms the observed feature PO into the concealed observed feature POe using a random unitary transform. The random unitary transform is disclosed in, for example, Non-Patent Document 1.
[0037] When the observed image data 22 includes face regions of multiple observers, the calculation unit 26 of the collection device 20 extracts the multiple face regions from the observed image data 22 and calculates an observed feature value PO for each of the multiple face regions. The conversion unit 27 of the collection device 20 converts each of the multiple observed feature values PO into multiple concealed observed feature values POe using a random unitary transformation. Furthermore, when there are multiple observed image data 22, the conversion unit 27 calculates an anonymized observed feature value POe for each observed image data 22 and each observer.
[0038] The collection device 20 transmits the anonymized observed feature POe to the recognition device 30.
[0039] In general, it is difficult to extract a face region from anonymized image data. Therefore, before transmitting the data to the cloud, the collection device 20 calculates features of the face region and transmits the anonymized features to the recognition device 30. This allows the recognition device 30 to process the anonymized data, thereby protecting privacy.
[0040] (recognition device) A recognition device 30 according to the present disclosure will be described with reference to FIG. 3. The recognition device 30 includes data such as anonymized training features PLe, anonymized observation features POe, dictionary matrix data 31, and estimation results 32, as well as functions of a learning unit 36, an estimation unit 37, and an output unit 38. Each piece of data is stored in a storage device such as a memory 902 or a storage 903. Each function is implemented in a CPU 901. Note that the process of learning and estimating a dictionary matrix is described in detail in, for example, Non-Patent Document 1 or Non-Patent Document 2.
[0041] The anonymized learning feature PLe and the anonymized observation feature POe are data acquired from the collection device 20, as described with reference to FIG.
[0042] The dictionary matrix data 31 is data that specifies the dictionary matrix learned by the learning unit 36.
[0043] The estimation result 32 is data resulting from estimation by the estimation unit 37. In the present disclosure, the estimation result 32 includes, for example, identifiers of a plurality of recognition targets who have a high degree of similarity to the observer.
[0044] The recognition device 30 processes anonymized training features PLe for multiple recognition targets during dictionary matrix training. The recognition device 30 processes anonymized observed features POe for one or more observers during estimation using the dictionary matrix.
[0045] First, the learning process will be described.
[0046] During learning, the recognition device 30 acquires anonymized training features PLe for a plurality of, preferably all, recognition targets from the collection device 20. The learning unit 36 learns a dictionary matrix using the plurality of anonymized training features PLe. The learning unit 36 outputs dictionary matrix data 31 for identifying the learned dictionary matrix. An example of an algorithm for learning a dictionary matrix using the anonymized training features PLe is the K-SVD (K-Singular Value Decomposition) algorithm.
[0047] A dictionary matrix is a matrix in which basic patterns (atoms) are used as column vectors and column vectors of each pattern are arranged. In the present disclosure, a basic pattern is sparsely coded data of anonymized learning features PLe for one target person. Here, the sparsely coded data is data that expresses essential features of the anonymized learning features PLe for one target person. The dictionary matrix has data of basic patterns for each target person in the column direction.
[0048] Next, the estimation process will be described.
[0049] The estimation unit 37 uses the dictionary matrix learned by the learning unit 36 to output an estimation result 32 of the observer of the observed image data 22. Specifically, the estimation unit 37 estimates sparse coefficients from the dictionary matrix and the anonymized observation feature POe. The estimation unit 37 outputs the estimation result 32 of the observer from the similarity between the estimated sparse coefficients and the dictionary matrix. Here, the estimation result 32 includes identifiers of the observer and multiple recognition targets whose facial regions have a high similarity.
[0050] First, the estimation unit 37 calculates sparse coefficients for the anonymized observed feature POe. An algorithm for calculating sparse coefficients is the OMP (Orthogonal Matching Pursuit) algorithm.
[0051] When the observed image data 22 includes multiple face regions, the estimation unit 37 of the recognition device 30 receives multiple anonymized observed features POe from the collection device 20. The estimation unit 37 calculates a similarity for each of the multiple anonymized observed features POe and outputs an estimation result 32.
[0052] If the anonymized observed feature POe is y, its sparse coefficient is x, and the dictionary matrix is D, then a linear model is established as shown in Figure 4. In Figure 4, R represents a set of real numbers, M is the number of rows in the dictionary matrix, and K is the number of columns.
[0053] Using the calculated sparse coefficients, the observer's features are modeled by the sparse model shown in Fig. 5. The estimation unit 37 uses the sparse model to calculate the similarity between each column of the dictionary matrix and the features of each recognition target. In Fig. 5, by applying the constraint that the number of non-zero components of x is less than a threshold ε, a sparse solution is estimated in which elements less than the threshold ε are non-zero and other elements are zero. The kNN (k-nearest neighbor) algorithm is an algorithm for calculating the similarity.
[0054] The estimation unit 37 identifies the identifiers of one or more recognition targets with high similarity and outputs them as the estimation result 32.
[0055] The output unit 38 transmits the estimation result 32 to a predetermined destination, such as the collection device 20.
[0056] (Recognition method) The recognition method according to the present disclosure will be described with reference to Fig. 6. In Fig. 6, steps S101 to S105 are processes during learning, and steps S111 to S117 are processes during estimation.
[0057] In step S101, the collection device 20 collects training image data 21 of the person to be recognized from the photographing device 10.
[0058] In step S102, the collection device 20 calculates a training feature PL of the face region of each recognition target person in the training image data 21. In step S103, the collection device 20 conceals the training feature PL calculated in step S102. In step S104, the collection device 20 transmits the training feature PLe concealed in step S103 to the recognition device 30.
[0059] In step S105, the recognition device 30 learns a dictionary matrix using the anonymized training feature PLe received in step S104.
[0060] In step S111, the collecting device 20 collects the observation image data 22 of the observer from the photographing device 10.
[0061] In step S112, the collection device 20 calculates an observed feature value PO of the observer's face region in the observed image data 22. In step S113, the collection device 20 conceals the observed feature value PO calculated in step S112. In step S114, the collection device 20 transmits the observed feature value POe concealed in step S113 to the recognition device 30.
[0062] In step S115, the recognition device 30 estimates sparse coefficients using the anonymized observation features POe received in step S114 and the dictionary matrix learned in step S105. In step S116, the recognition device 30 estimates a recognition target person similar to the observer using the sparse coefficients and dictionary matrix estimated in step S115.
[0063] In step S117, the recognition device 30 transmits to the collection device 20 an estimation result 32 including the identifier of the person to be recognized estimated in step S116.
[0064] The results of verification of the recognition system 1 according to the present disclosure will be described. In the verification, the recognition system 1 prepares three sets of training image data 21 for each of the 77 recognition targets and learns a dictionary matrix. The recognition system 1 uses the dictionary matrix to calculate the similarity between one of the recognition targets as an observer and the observed image data 22 of that observer, and identifies the identifier of the recognition target of the training image data 21 with the highest similarity. The recognition system 1 determines whether the identifier of the identified recognition target matches the identifier of the observer.
[0065] When multiple observers were tested, the discrimination rate was 71.3%. Of the 665 test data items, there were 474 correct answers and 191 incorrect answers.
[0066] According to the recognition system 1 of the present disclosure, when multiple people appear in a captured image, it is possible to identify which of the recognition targets each person is in the image while keeping the image concealed. As a result, the recognition system 1 can be used to identify a person to be tracked from a security camera, or to compare the person with a person registered in advance to enter a facility.
[0067] The estimation result 32 output by the recognition system 1 includes identifiers of multiple recognition targets who have a high degree of similarity to the observer. The recognition system 1 can be applied to a system that determines whether an observer who enters a predetermined area where a surveillance camera is installed is a pre-registered recognition target. Even if the recognition system 1 cannot confirm a match between the observer and the recognition target, it can estimate multiple recognition targets who are similar to the observer, making it possible to narrow down the recognition targets.
[0068] The recognition system 1 according to the present disclosure can recognize face images while protecting privacy.
[0069] The collection device 20 and the recognition device 30 according to the present disclosure described above each use a general-purpose computer system including, for example, a CPU (Central Processing Unit, processor) 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906. In this computer system, the CPU 901 executes a program loaded on the memory 902, thereby realizing each function of the collection device 20 and the recognition device 30.
[0070] The collection device 20 and the recognition device 30 may each be implemented by a single computer or by multiple computers. Furthermore, the collection device 20 and the recognition device 30 may each be a virtual machine implemented on a computer.
[0071] The programs of the collection device 20 and the recognition device 30 can be stored in a computer-readable recording medium such as a HDD, an SSD, a Universal Serial Bus (USB) memory, a Compact Disc (CD), or a Digital Versatile Disc (DVD), or can be distributed via a network. The computer-readable recording medium is, for example, a non-transitory recording medium.
[0072] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the present disclosure. [Explanation of symbols]
[0073] 1 Recognition System 10 Imaging equipment 20 Collection Device 21 Learning image data 22 Observation image data 26 Calculation section 27 Conversion unit 30 recognition device 31 Dictionary matrix data 32 Estimation results 36 Learning Department 37 Estimation part 38 Output section 901 CPU 902 memory 903 Storage 904 Communication equipment 905 Input Device 906 Output Device PL learning feature PO Observation feature PLe concealed learning features POe Concealed observation features
Claims
1. a calculation unit that extracts face regions of a plurality of recognition target persons from the training image data and calculates training features for each of the plurality of face regions; a transformation unit that transforms each of the plurality of training features into a plurality of concealed training features using a random unitary transformation; a learning unit that learns a dictionary matrix using the plurality of anonymized learning features; an estimation unit that uses the dictionary matrix to output an estimation result of the observer of the observed image data; A recognition system comprising:
2. the calculation unit extracts a face region of the observer from the observed image data and calculates an observation feature amount of the face region; the transforming unit transforms the observed feature into a concealed observed feature using the random unitary transformation; The estimation unit estimates sparse coefficients from the dictionary matrix and the anonymized observed features, and outputs an estimation result of the observer from a similarity between the estimated sparse coefficients and the dictionary matrix. The recognition system of claim 1 .
3. The estimation result includes identifiers of a plurality of recognition targets having high similarity. The recognition system of claim 2 .
4. the calculation unit extracts a plurality of face regions from the observed image data and calculates an observed feature amount for each of the plurality of face regions; the transforming unit transforms each of a plurality of observed features into a plurality of concealed observed features using the random unitary transformation; The estimation unit outputs the estimation result for each of the plurality of anonymized observed features. The recognition system of claim 2 .
5. A first path from an image capturing device that captures the learning image data to a collection device that includes the calculation unit and the conversion unit is shorter than a second path from the image capturing device to a recognition device that includes the learning unit. The recognition system of claim 1 .
6. The computer extracting face regions of each of a plurality of recognition target persons from the training image data, and calculating training features for each of the plurality of face regions; Transforming each of the plurality of training features into a plurality of anonymized training features using a random unitary transformation; learning a dictionary matrix using the plurality of anonymized training features; Using the dictionary matrix, an estimation result of the observer of the observed image data is output. Recognition method.
7. a calculation unit that extracts face regions of a plurality of recognition target persons from the training image data and calculates training features for each of the plurality of face regions; A transformation unit that transforms each of the plurality of training features into a plurality of concealed training features using a random unitary transformation. A collecting device comprising:
8. A program that causes a computer to function as the collection device according to claim 7.