The determination method, the non-transitory computer-readable storage medium storing the determination program, and the information processing device
By generating differential image data and combining spatial and frequency region information, the problem of insufficient accuracy in deformable image recognition is solved, achieving efficient deformable image recognition and reducing cost and time consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2021-01-27
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies lack sufficient accuracy in determining whether a facial image is distorted, and may allow unauthorized users to successfully verify identity through distortion attacks.
By generating differential image data, spatial region information is used to detect signal anomalies in the deformation processing, and frequency region information is used for re-determination when no judgment can be made. Deep learning methods are combined to improve the judgment accuracy.
It improves the accuracy of deformed image identification, reduces costs and shortens processing time, and avoids the use of specific equipment.
Smart Images

Figure CN116724332B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a determination method, a non-transitory computer-readable storage medium storing a determination program, and an information processing apparatus. Background Technology
[0002] In facial authentication, there are instances of improper behavior known as distorted attacks. To address this, techniques for determining whether facial image data is distorted have been disclosed (e.g., see Patent Document 1).
[0003] Patent Document 1: Japanese Patent Publication No. 2020-525947
[0004] However, there are cases where sufficient accuracy cannot be obtained when determining whether facial image data is a deformed image. Summary of the Invention
[0005] In one aspect, the object of the present invention is to provide a determination method, a determination procedure, and an information processing apparatus that can improve the accuracy of determining whether facial image data is a deformed image.
[0006] In one aspect, the determination method involves a computer performing the following processes: when facial image data has been acquired, generating facial image data with noise removed using a specific algorithm based on the facial image data; generating differential image data between the acquired facial image data and the generated facial image data; determining whether the acquired facial image data is a composite image based on information contained in the differential image data; and if the acquired facial image data is not determined to be a composite image, determining whether the acquired facial image data is a composite image based on information contained in frequency data generated from the differential image data.
[0007] It can improve the accuracy of judgment. Attached Figure Description
[0008] Figure 1 This is an example of a deformed image.
[0009] Figure 2 Figures (a) to (i) are used to illustrate the principles of the embodiments.
[0010] Figure 3 (a) is a functional block diagram illustrating the overall structure of an information processing device. Figure 3 (b) is a block diagram illustrating the hardware structure.
[0011] Figure 4 This is a flowchart illustrating the feature extraction process.
[0012] Figure 5This is a flowchart illustrating the learning process.
[0013] Figure 6 This is a graph illustrating the training data.
[0014] Figure 7 This is a flowchart illustrating the decision-making process. Detailed Implementation
[0015] Facial recognition is a technology that uses facial features to verify identity. In facial recognition, when verification is required, facial feature data acquired by sensors is compared with pre-registered facial feature data. Identity verification is performed by determining whether the similarity exceeds a threshold. Facial recognition is used in passport control, ID card registration, and access control systems.
[0016] In facial recognition, there are instances of misconduct known as morphing attacks. In morphing attacks, such as... Figure 1 As illustrated, a composite image is generated by combining a source image of a first person and a target image of a second person using a morphing technique. This composite image is called a morphing image. The source image of the first person is not a facial image with modifications applied, but rather an actual facial image of the first person. Similarly, the target image of the second person is not a facial image with modifications applied, but rather an actual facial image of the second person.
[0017] Deformed images include fusion images (so-called clipped images) obtained by partially combining two facial images. Clipped images are, for example, images obtained by cutting out and replacing the eyes and mouth with the eyes and mouth of another person.
[0018] The deformed image includes an interpolated image obtained by interpolating two facial images. An interpolated image is obtained by interpolating two facial images. For example, an interpolated image is obtained by averaging two facial images.
[0019] The distorted image simultaneously possesses the facial features of both a first person and a second person. Therefore, if the distorted image is pre-registered as registered facial feature data, it is possible for both the first and second persons to successfully verify their identity. For example, when creating a passport, if a distorted image composed of the facial image data of the first person (source image) and the facial image data of the second person (target image) is registered, then both the first and second persons can use the same passport.
[0020] Here, methods for detecting warp attacks are investigated. For example, techniques using the mode noise (PRNU: Photo Response Non-Uniformity) of image sensors can be considered. Specifically, the energy properties of spatial and spectral features extracted from PRNU modes across image units can be considered. However, this method is costly and complex due to the need for specialized sensors.
[0021] Next, we consider using a deep learning-based technique for residual noise. Specifically, this technique uses residual noise, obtained by differencing the facial image data with the noise-removed facial image data, to determine whether the facial image data being analyzed is a deformed image. However, this technique uses only spatial region information, raising concerns that it may only be able to determine deformed images of specific types.
[0022] In this regard, the following embodiments will describe an information processing apparatus, a determination method, and a determination procedure that can suppress costs and improve the determination accuracy of deformed images.
[0023] Example 1
[0024] First, the principle of this embodiment will be explained.
[0025] As mentioned above, distorted images include interpolated images, clipped images, etc. They are synthesized from multiple facial image data of different people. Each facial image data was acquired, for example, by a different camera. In this case, noise of varying intensity remains in each facial image data. Even if each facial image data was acquired by the same camera, different intensities of noise still remain in each facial image data due to differences in the timing, environment, etc., of acquiring the facial image data.
[0026] To address this, differential image data is acquired between the deformed image and the deformed image after noise removal processing. This differential image data retains some of the removed noise components. Therefore, the differential image data represents residual noise. This residual noise contains noise of varying intensities, resulting in discontinuous noise levels. Therefore, signal anomalies caused by the deformation processing can be detected.
[0027] For example, in a clipped image, traces of direct pixel manipulation are left within the image plane. Therefore, by analyzing spatial region information representing residual noise in a defined spatial region, the edges of the clipped portions in the clipped image can be detected.
[0028] However, for interpolated images without clipped edges, it is difficult to determine whether they are deformed images even by analyzing spatial region information.
[0029] For this interpolated image, there are patterns in the frequency region that are difficult to represent in the spatial region information. For example, there are peaks, lines, and other processing traces that appear in the frequency region expressed on a logarithmic amplitude scale but are not represented in the spatial region information. Therefore, by analyzing the information in this frequency region, signal anomalies caused by deformation processing can be detected, thereby improving the accuracy of determining whether the facial image data of the target is a deformed image.
[0030] Figure 2 (a) is a graph illustrating the actual facial image data of the first person. Figure 2 (b) is a diagram illustrating the actual facial image data of the second person. No traces of distortion processing are left in the residual noise of these facial image data. Therefore, as... Figure 2 As illustrated in (c), even when the residual noise is converted to the frequency space, no signal anomalies caused by the deformation process are present as a characteristic.
[0031] Figure 2 (d) is a diagram illustrating a clipping image that partially combines facial image data of a first person and facial image data of a second person. Figure 2 (e) represents the spatial region information when the residual noise of the clipped image is represented by a defined spatial region. In this spatial region information, the edges of the clipped portions can be easily detected as signal anomalies. Figure 2 (f) is a graph illustrating frequency region information when the residual noise of a clipped image is converted into frequency space. For example... Figure 2 As illustrated in (f), with Figure 2 Compared to case (c), as a signal anomaly, it is more likely to exhibit characteristics such as a longitudinal line passing through the center.
[0032] Figure 2 (g) is an example of an interpolated image obtained by averaging the facial image data of the first person and the facial image data of the second person. Figure 2 (h) represents the spatial region information of the interpolated image under the condition that the residual noise is represented by a specified spatial region. Since the interpolated image does not contain edges, it is difficult to show the characteristics of signal anomalies in this spatial region information. Figure 2 (i) is a graph illustrating frequency region information when the residual noise of the interpolated image is converted into frequency space. For example... Figure 2 As illustrated by the arrow in (i), peaks, lines, etc., that are not present in spatial region information appear as features of signal anomalies.
[0033] Based on the above, regarding facial image data designated as the object, even if it is not determined to be a deformed image based on spatial region information, it can be re-determined using frequency region information, thereby improving the accuracy of deformed image determination. Furthermore, since no additional special equipment is required, costs can be suppressed.
[0034] The following is a detailed description of this embodiment.
[0035] Figure 3 (a) is a functional block diagram illustrating the overall structure of the information processing device 100. For example... Figure 3 As illustrated in (a), the information processing apparatus 100 includes a feature extraction processing unit 10, a learning processing unit 20, a determination processing unit 30, and an output processing unit 40. The feature extraction processing unit 10 includes a face image acquisition unit 11, a color space conversion unit 12, a noise filter unit 13, a difference image generation unit 14, a first feature extraction unit 15, a second feature extraction unit 16, a feature score calculation unit 17, a determination unit 18, and an output unit 19. The learning processing unit 20 includes a training data storage unit 21, a training data acquisition unit 22, a training data classification unit 23, and a model creation unit 24. The determination processing unit 30 includes a face image acquisition unit 31 and a determination unit 32.
[0036] Figure 3 (b) is a block diagram illustrating the hardware structure of the feature extraction processing unit 10, the learning processing unit 20, the decision processing unit 30, and the output processing unit 40. Figure 3 As illustrated in (b), the information processing device 100 includes a CPU 101, RAM 102, storage device 103, display device 104, interface 105, etc.
[0037] The CPU (Central Processing Unit) 101 is a central processing unit. The CPU 101 contains one or more cores. The RAM (Random Access Memory) 102 is volatile memory that temporarily stores programs executed by the CPU 101, data processed by the CPU 101, etc. The storage device 103 is a non-volatile storage device. For example, the storage device 103 can be a ROM (Read Only Memory), a solid-state drive (SSD) such as flash memory, or a hard disk driven by a hard disk drive. The storage device 103 stores the determination program involved in this embodiment. The display device 104 is a display device such as a liquid crystal display. The interface 105 is an interface device for connecting to external devices. For example, facial image data can be acquired from an external device via the interface 105. By executing the determination program through the CPU 101, the feature extraction processing unit 10, the learning processing unit 20, the determination processing unit 30, and the output processing unit 40 of the information processing apparatus 100 are realized. Furthermore, dedicated circuits or other hardware can be used as the feature extraction processing unit 10, the learning processing unit 20, the decision processing unit 30, and the output processing unit 40.
[0038] (Feature extraction processing)
[0039] Figure 4 This is a flowchart illustrating the feature extraction process performed by the feature extraction processing unit 10. For example... Figure 4 As illustrated, the facial image acquisition unit 11 acquires facial image data (step S1).
[0040] Next, the color space conversion unit 12 converts the color space of the facial image data acquired in step S1 into a specified color space (step S2). For example, the color space conversion unit 12 converts the facial image data into the HSV color space, which consists of three components: hue, saturation chroma, and value brightness. In the HSV color space, brightness or image intensity can be separated from chroma or color information.
[0041] Next, the noise filter unit 13 performs noise removal processing on the facial image data obtained in step S2 to generate noise-removed facial image data (step S3). In step S3, the noise filter unit 13 performs noise removal processing based on a specific algorithm. In the noise removal processing, known techniques for removing image noise can be used.
[0042] Next, the differential image generation unit 14 generates differential image data between the facial image data obtained in step S2 and the facial image obtained in step S3 (step S4). By generating the differential image data, residual noise remaining in the facial image obtained in step S1 can be obtained. Alternatively, without performing the processing in step S2, differential image data between the facial image data obtained in step S1 and the facial image data after noise removal processing can be generated.
[0043] Next, the first feature extraction unit 15 extracts features of signal anomalies caused by deformation processing from the spatial region information of the differential image data as first features (step S5). For example, the first feature extraction unit 15 extracts vector values of spatial region information such as LBP (Local Binary Pattern) and CoHOG (Co-occurrence Histograms of Oriented Gradients) as features of signal anomalies. For example, the first feature extraction unit 15 can extract vector values of spatial regions by using statistics obtained by comparing the pixel value of the pixel of interest with the pixel values of the pixels surrounding the pixel of interest. Furthermore, the first feature extraction unit 15 can also extract vector values of spatial regions by using deep learning. The first feature extraction unit 15 can also use feature quantities represented by numerical expressions.
[0044] Next, the feature score calculation unit 17 calculates the accuracy of the facial image data obtained in step S1 as deformed image data as a feature score based on the first feature extracted in step S5 (step S6). For example, statistics of the vector values obtained in step S5 can be used as feature scores.
[0045] Next, the determination unit 18 determines whether the feature score calculated in step S5 is higher than a threshold (step S7). This threshold can be determined in advance, for example, based on the variation value when feature scores are calculated from spatial regions for multiple facial image data.
[0046] If the determination in step S7 is "yes", the output unit 19 outputs the feature score calculated in step S6 (step S8).
[0047] If the result is "no" in step S7, the second feature extraction unit 16 generates frequency information of the frequency region based on the differential image data (step S9). For example, the second feature extraction unit 16 can generate frequency information of the frequency region by performing a digital Fourier transform on the differential image data.
[0048] Next, the second feature extraction unit 16 extracts features of the signal anomaly caused by the deformation processing as second features based on the frequency information generated in step S9 (step S10). For example, the second feature extraction unit 16 extracts vector values of frequency regions such as the gray-level co-occurrence matrix (GLCM) as features of the signal anomaly. For example, the second feature extraction unit 16 can extract vector values of frequency regions by using statistics obtained by comparing the pixel value of the pixel of interest with the pixel values of the pixels surrounding the pixel of interest. Furthermore, the second feature extraction unit 16 can also extract vector values of frequency regions by using deep learning. The second feature extraction unit 16 can also use feature quantities expressed numerically.
[0049] Next, the feature score calculation unit 17 calculates the accuracy of the facial image data obtained in step S1 as deformed image data as a feature score based on the frequency features extracted in step S10 (step S11). For example, statistics of the vector values obtained in step S10 can be used as feature scores.
[0050] Next, the determination unit 18 determines whether the feature score calculated in step S11 is higher than a threshold (step S12). The threshold in step S12 can be determined in advance, for example, based on the variation value of the feature score calculated from the frequency region for multiple facial image data.
[0051] If the determination in step S12 is "yes", the output processing unit 40 outputs the feature score calculated in step S11 (step S8).
[0052] If the determination in step S12 is "no", the feature score calculation unit 17 calculates the accuracy of the facial image data obtained in step S1 as deformed image data as a feature score based on the first feature extracted in step S5 and the second feature extracted in step S10 (step S13). Then, the output processing unit 40 outputs the feature score calculated in step S13 (step S8).
[0053] (Learning Processing)
[0054] Figure 5 This is a flowchart illustrating the learning process performed by the learning processing unit 20. For example... Figure 5 As illustrated, the training data acquisition unit 22 acquires the training data stored in the training data storage unit 21 (step S21). The training data is training data for facial images, including unmodified actual facial image data and deformed image data. Figure 6As illustrated, each training data point is associated with either an identifier representing actual facial image data (Bonafide) or an identifier representing a deformed image (Morphing). This training data is pre-created by the user and others and stored in the training data storage unit 21. The training data acquired in step S21 is then processed... Figure 4 Feature extraction processing.
[0055] Next, the training data classification unit 23 classifies the feature scores output by the output unit 19 for each training data into actual facial image data and deformed image data based on the identifiers stored in the training data storage unit 21 (step S22).
[0056] Next, the model creation unit 24 creates an evaluation model based on the classification results of step S22 (step S23). For example, it creates a classification model by drawing a separating hyperplane (boundary plane) based on the spatial relationship between the feature score distribution of each training data and the identifier. A classification model can be created through the above processing.
[0057] (Decision Processing)
[0058] Figure 7 This is a flowchart illustrating the decision processing performed by the decision processing unit 30. For example... Figure 7 As illustrated, the facial image acquisition unit 31 acquires facial image data (step S31). The facial image data acquired in step S31 is then processed... Figure 4 Feature extraction processing. Since the facial image data is facial image data used for passport making, etc., it is input from an external device via interface 105.
[0059] Next, the determination unit 32 determines whether the feature score output by the output unit 19 is an actual image or a deformed image using the classification model created by the model creation unit 24 (step S32). The determination result of step S32 is output by the output processing unit 40. The determination result output by the output processing unit 40 is displayed on the display device 104.
[0060] According to this embodiment, firstly, spatial region information of residual noise is used to determine whether facial image data is a deformed image. By using spatial region information of residual noise, discontinuities in noise intensity can be easily detected. If the facial image data is not determined to be a deformed image in the determination using spatial region information, frequency region information is used to re-determine whether the facial image data is a deformed image. By using frequency region information, the determination accuracy is improved. In addition, since special image sensors are not required, costs can be reduced. Furthermore, since frequency region information is used to re-determine whether the image is a deformed image in the determination using spatial region information, the processing time can be shortened compared to the case where both spatial region information and frequency region information are used for determination.
[0061] In the above examples, the noise filter unit 13 is an example of a facial image data generation unit that generates facial image data with noise removed using a specific algorithm based on the acquired facial image data. The difference image generation unit 14 is an example of a difference image data generation unit that generates difference image data between the acquired facial image data and the generated facial image data. The determination unit 18 is an example of a first determination unit that determines whether the acquired facial image data is a composite image based on information contained in the difference image data. The determination unit 18 is also an example of a second determination unit that determines whether the acquired facial image data is a composite image based on information contained in the frequency data. The determination processing unit 30 is an example of a determination processing unit that further determines whether the facial image data, which has been determined to be a composite image, is a composite image based on a classification model obtained by machine learning using training data of multiple facial image data. The second feature extraction unit 16 is an example of a frequency data generation unit that generates the frequency data based on the difference image data using digital Fourier transform.
[0062] The embodiments of the present invention have been described in detail above, but the present invention is not limited to the specific embodiments involved. Various modifications and alterations can be made within the scope of the spirit of the present invention as described in the claims.
[0063] Explanation of reference numerals in the attached figures
[0064] 10 Feature Extraction Processing Unit
[0065] 11 Facial Image Acquisition Unit
[0066] 12 Color Space Conversion Department
[0067] 13 Noise Filter Section
[0068] 14 Differential Image Generation Unit
[0069] 15 First Feature Extraction Unit
[0070] 16 Second Feature Extraction Unit
[0071] 17 Characteristic Score Calculation Section
[0072] 18 Judgment Department
[0073] 19 Output Section
[0074] 20 Learning Processing Department
[0075] 21 Training Data Storage Department
[0076] 22 Training Data Acquisition Department
[0077] 23 Training Data Classification Department
[0078] 24 Model Making Department
[0079] 30 Judgment and Processing Department
[0080] 31 Facial Image Acquisition Unit
[0081] 32 Judgment Department
[0082] 40 Output Processing Unit
[0083] 100 Information Processing Device
Claims
1. A determination method, characterized in that, The computer performs the following processing: Having acquired facial image data, noise-removed facial image data is generated based on this data using a specific algorithm. Generate differential image data between the acquired facial image data and the generated facial image data. The spatial region information contained in the differential image data is used to determine whether the acquired facial image data is a synthetic image. If the acquired facial image data is not determined to be a synthetic image, the determination of whether the acquired facial image data is a synthetic image is based on the information contained in the frequency data generated from the differential image data.
2. The determination method according to claim 1, characterized in that, When determining whether the acquired facial image data is a synthetic image based on the information contained in the differential image data, the determination is made by detecting discontinuities in noise intensity.
3. The determination method according to claim 1, characterized in that, The above computer performs the following processing: A classification model obtained by machine learning using training data from multiple facial images is used to further determine whether the aforementioned facial image data, which has already been determined to be a synthetic image, is a synthetic image.
4. The determination method according to claim 1, characterized in that, The color space of the facial image data obtained above is the HSV color space.
5. The determination method according to claim 1, characterized in that, The frequency data is generated from the differential image data using digital Fourier transform.
6. The determination method according to claim 1, characterized in that, The information contained in the aforementioned differential image data is a statistical measure obtained by comparing the pixel value of the pixel of interest with the pixel values of the pixels surrounding the pixel of interest.
7. The determination method according to claim 1, characterized in that, The information contained in the aforementioned differential image data is used as a feature quantity expressed numerically.
8. A non-transitory computer-readable storage medium, characterized in that, A decision program is stored that causes the computer to perform the following processes: Having acquired facial image data, noise-removed facial image data is generated based on the aforementioned facial image data using a specific algorithm. Generate differential image data between the acquired facial image data and the generated facial image data; Based on the spatial region information contained in the aforementioned differential image data, it is determined whether the acquired facial image data is a synthetic image; and If the acquired facial image data is not determined to be a synthetic image, the determination of whether the acquired facial image data is a synthetic image is based on the information contained in the frequency data generated from the differential image data.
9. The non-transitory computer-readable storage medium according to claim 8, characterized in that, When determining whether the acquired facial image data is a synthetic image based on the information contained in the differential image data, the determination is made by detecting discontinuities in noise intensity.
10. The non-transitory computer-readable storage medium according to claim 8, characterized in that, The computer described above shall perform the following processing: A classification model obtained by machine learning using training data from multiple facial images is used to further determine whether the aforementioned facial image data, which has already been determined to be a synthetic image, is a synthetic image.
11. The non-transitory computer-readable storage medium according to claim 8, characterized in that, The color space of the facial image data obtained above is the HSV color space.
12. The non-transitory computer-readable storage medium according to claim 8, characterized in that, The frequency data is generated from the differential image data using digital Fourier transform.
13. The non-transitory computer-readable storage medium according to claim 8, characterized in that, The information contained in the aforementioned differential image data is a statistical measure obtained by comparing the pixel value of the pixel of interest with the pixel values of the pixels surrounding the pixel of interest.
14. The non-transitory computer-readable storage medium according to claim 8, characterized in that, The information contained in the aforementioned differential image data is used as a feature quantity expressed numerically.
15. An information processing device, characterized in that, have: The facial image data generation unit, upon acquiring facial image data, generates facial image data that has had noise removed using a specific algorithm based on the aforementioned facial image data; The differential image data generation unit generates differential image data between the acquired facial image data and the generated facial image data; The first determination unit determines whether the acquired facial image data is a synthetic image based on the spatial region information contained in the differential image data. as well as The second determination unit determines whether the acquired facial image data is a synthetic image based on information contained in the frequency data generated from the differential image data if it does not determine that the acquired facial image data is a synthetic image.
16. The information processing apparatus according to claim 15, characterized in that, When the first determination unit determines whether the acquired facial image data is a synthetic image based on the information contained in the differential image data, it makes the determination by detecting discontinuities in noise intensity.
17. The information processing apparatus according to claim 15, characterized in that, The system includes a determination processing unit that uses a classification model obtained by machine learning from training data of multiple facial image data to further determine whether the facial image data, which has already been determined to be a synthetic image, is a synthetic image.
18. The information processing apparatus according to claim 15, characterized in that, The color space of the facial image data obtained above is the HSV color space.
19. The information processing apparatus according to claim 15, characterized in that, It has a frequency data generation unit that generates the frequency data based on the differential image data by digital Fourier transform.
20. The information processing apparatus according to claim 15, characterized in that, The information contained in the aforementioned differential image data is a statistical measure obtained by comparing the pixel value of the pixel of interest with the pixel values of the pixels surrounding the pixel of interest.
21. The information processing apparatus according to claim 15, characterized in that, The first determination unit mentioned above uses feature quantities expressed numerically as information contained in the differential image data.
Citation Information
Patent Citations
Method and device for training neural network for image identification
CN106485192A
Detecting manipulated images
JP2020525947A