A security testing method and system for face liveness detection model based on structured light
Through the structured light depth imaging principle, the existing 3D live detection methods are solved, and the low-cost and fast security testing of multi-scene face authentication system is achieved.
Patent Information
- Application Number
- CN202310542078.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-05-15
AI Technical Summary
The existing 3D live detection methods are costly and cannot be reused, and it is difficult to conduct safety testing of equipment in large-scale and multi-scenarios.
The structured light depth imaging principle is adopted to forge non-existent depth images for security testing of 3D face authentication system. Security testing is carried out by obtaining template infrared speckle images, extracting facial areas, reconstructing three-dimensional structures, modulating depth information and projecting images.
It provides a low-cost, migable 3D face live detection method, which is suitable for all intelligent systems deploying structured light depth cameras, with fast test speed and wide application range.
Smart Images

Figure CN116612374B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular to a method and system for security testing of a face liveness detection model based on structured light. Background Art
[0002] Facial recognition has become a mainstream biometric authentication technology, widely used in areas such as mobile phone unlocking, venue access control, and financial payments. However, facial recognition is vulnerable to replay attacks, such as using photos or videos to forge faces. Therefore, using 3D liveness detection for facial authentication has become a popular facial authentication solution for many smart devices. 3D liveness detection technology uses a depth camera to obtain three-dimensional information about a target and use it to determine whether the target is alive. Depth cameras include structured light cameras, Time of Flight cameras, and LiDAR. Due to cost and imaging performance considerations, the mainstream depth camera currently uses structured light cameras.
[0003] Currently, some replay attack security testing methods for 3D liveness detection exist, such as using 3D printed masks and head models. However, these methods are expensive and non-reusable, and cannot be used to test the security of devices in a wide range of scenarios. Therefore, a portable, low-cost 3D face liveness detection security testing method for depth sensors is urgently needed. Summary of the Invention
[0004] In response to the above-mentioned problems in the prior art, the present invention provides a method and system for security testing of a facial liveness detection model based on structured light deep forging. Based on the principle of structured light depth imaging, the method can forge non-existent depth images for security testing of the liveness detection module in a 3D face authentication system. Compared with 3D printed masks, head models, etc., this method is universal for smart devices using structured light depth cameras and has a low migration cost.
[0005] The technical solution of the present invention is:
[0006] In a first aspect, the present invention proposes a method for security testing of a face liveness detection model based on structured light, comprising the following steps:
[0007] Step 1: Obtain a template infrared speckle image of a structured light depth camera in a face authentication system under test;
[0008] Step 2: Extract the facial region from the two-dimensional image containing the target face, obtain the target face RGB image, reconstruct the three-dimensional structure of the face, and generate the corresponding face depth image;
[0009] Step 3: Modulate the infrared speckle image with depth information based on the template infrared speckle image and the depth information in the face depth image;
[0010] Step 4: Align the target face RGB image with the infrared speckle image with depth information and project them together;
[0011] Step 5: The face recognition system under test collects the projected image and transmits it to the face liveness detection model. If it passes the face recognition, the face recognition system is unsafe; otherwise, it is safe, and the security test is completed.
[0012] Furthermore, the step 2 includes:
[0013] 2.1) Obtain a face detection bounding box in the 2D image containing the target face, extract the facial region using a feature point extraction method, remove background elements, and then crop it into a fixed-size RGB image of the target face;
[0014] 2.2) A UNet model is used to generate a facial depth image. The UNet model uses ResNet-50 as an encoder to reduce the target face RGB image to an embedded feature map. A decoder consisting of 5 transposed convolutional layers and 10 convolutional layers is then used to reconstruct the embedded feature map into a fixed-size facial depth image, where each pixel represents the absolute value of the depth.
[0015] Furthermore, the step 3 includes:
[0016] 3.1) Using non-maximum suppression method to filter the template infrared speckle image;
[0017] 3.2) The depth information in the face depth image is modulated into the filtered template infrared speckle image. The modulation process is expressed as:
[0018] S=T+Φ(D)
[0019] Where Φ(.) represents the mapping function that converts depth information into scattering displacement, S represents the infrared speckle image with depth information, and T represents the coordinate set of each pixel in the filtered template infrared speckle image.
[0020] Furthermore, the mapping function for converting depth information into scattering displacement is:
[0021]
[0022] Among them, Φ(.) represents the mapping function, k p represents the number of pixels within 1 mm physical length in the infrared speckle image delivery submodule, L represents the baseline distance between the structured light depth camera and the infrared speckle image delivery submodule, and f p represents the focal length of the infrared speckle image projection submodule, d refRepresents the reference depth, and d represents the depth value of the pixel in the face depth image.
[0023] Furthermore, before aligning the target face RGB image with the infrared speckle image with depth information in step 4, a projection correction step is also included, specifically:
[0024] Perform perspective transformation on infrared speckle images with depth information:
[0025]
[0026] Where (x,y) is the point in the infrared speckle image captured by the structured light depth camera, is the perspective transformation matrix, is a point in the infrared speckle image after perspective transformation, is the perspective transformation parameter.
[0027] The calculation formula of the infrared speckle image after projection correction is as follows:
[0028]
[0029] Among them, (x′, y′) is the pixel point in the infrared speckle image after projection correction.
[0030] In a second aspect, the present invention proposes a structured light-based face liveness detection model security testing system, which is characterized by comprising:
[0031] An image acquisition module, which is used to obtain a template infrared speckle image of a structured light depth camera in a face authentication system under test;
[0032] The depth estimation module is used to extract the facial area from the two-dimensional image containing the target face, obtain the target face RGB image and reconstruct the three-dimensional structure of the face to generate the corresponding face depth image;
[0033] A deep fake module is used to modulate an infrared speckle image with depth information based on the template infrared speckle image and the depth information in the face depth image;
[0034] The image projection module is used to align the target face RGB image with the infrared speckle image with depth information, and project the two. The projected image is collected by the face recognition system under test and transmitted to the face liveness detection model. If it passes the face recognition, the face recognition system is unsafe. Otherwise, it is safe, and the security test is completed.
[0035] The beneficial effects of the present invention are:
[0036] 1. This paper reverse-engineers the structured light depth camera and proposes a security testing method from the perspective of the depth sensor. The method constructs a triangular relationship between a dot projector, an infrared camera, and a depth plane. By constructing a "depth information-speckle offset" mapping function, the infrared speckle pattern is depth-modulated. A specific infrared speckle pattern is projected through an infrared projector to interfere with the depth measurement of the structured light depth camera. This method can generate facial images of arbitrary depth to perform security testing on the liveness detection module in the 3D face authentication system.
[0037] 2. The testing method of this invention is highly transferable and applicable to all intelligent systems that deploy structured light depth cameras. It can also be applied to the testing of other image recognition systems that use depth images for detection. Compared with other methods such as 3D-printed masks and head models, this method offers low cost, rapid testing speed, and wide application range. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is an architecture diagram of a face liveness detection system based on structured light deepfake proposed by the present invention;
[0039] Figure 2 This is a flow chart of the method for face liveness detection based on structured light deepfake proposed in the present invention;
[0040] Figure 3 Schematic diagram of the method for acquiring infrared speckle images of the camera template under test proposed in the present invention;
[0041] Figure 4 This is a schematic diagram of the filtering algorithm effect proposed by the present invention;
[0042] Figure 5 This is a schematic diagram of imaging modeling of the structured light depth camera proposed in the present invention;
[0043] Figure 6 This is a schematic diagram of the depth information mapping function calculation proposed by the present invention;
[0044] Figure 7 Schematic diagram of the security testing method for forged depth speckle projection and structured light depth camera proposed in the present invention; DETAILED DESCRIPTION
[0045] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are for illustrative purposes only, and those skilled in the art will readily appreciate other obvious variations. The basic principles of the present invention defined in the following description may be applied to other embodiments, variations, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the present invention.
[0046] The accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0047] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to actual circumstances.
[0048] The present invention provides a human face liveness detection security testing system based on structured light deep fake. Figure 1 The system architecture diagram of the exemplary embodiment is shown in FIG. Figure 1 As shown, the system architecture includes an image acquisition module, a depth estimation module, a deep fake module, an image projection module, and a face recognition system under test. The image acquisition module is used to collect the infrared scattered point image emitted by the infrared scattered point projection module. It can be a device capable of receiving and imaging infrared light with wavelengths including but not limited to 780nm to 2526nm, including but not limited to mobile phones without infrared filters, cameras, or night vision devices.
[0049] Image Acquisition Module Depth Estimation Module and Deep Fake Module The image acquisition module can be a background server that provides image analysis and processing;
[0050] The image projection module consists of two parts: a visible light image projection submodule and an infrared speckle image projection submodule. The visible light image projection submodule can be a high-precision color projection device, while the infrared image playback submodule is a projection device capable of projecting infrared images, including but not limited to wavelengths between 780nm and 2526nm, such as a laser array or infrared projector.
[0051] The 3D face recognition system is the tested system of the present invention, which can be any system or device that includes a 3D liveness detection algorithm model, such as a smart phone, an access control system, a face-swiping payment terminal, etc. The structured light depth camera in the tested face recognition system can emit infrared speckles arranged in a fixed pattern.
[0052] The image acquisition module and the image projection module can communicate with the deepfake module and the depth estimation module respectively via wired or wireless means. The image acquisition module transmits the collected template infrared speckle pattern to the deepfake module via wired or wireless means. The depth estimation module generates an RGB image and a facial depth image of the target face based on the two-dimensional image containing the target face, and transmits them to the deepfake module. The deepfake module generates a modulated infrared speckle image and transmits it to the image projection module for correction of projection distortion and alignment. The infrared speckle image projection submodule projects the modulated infrared speckle image, and the visible light image projection submodule projects the aligned RGB image of the target face. In one embodiment, the image acquisition module is composed of an infrared camera. The depth projection module is composed of a high-precision color projector and an infrared projector, and can transmit data with the deepfake module.
[0053] The following describes a facial liveness detection security testing method based on structured light deepfakes according to this exemplary embodiment. Application scenarios of this method include, but are not limited to, downloading a two-dimensional image containing a target face from a social network, extracting the tester's facial region using the face extraction method proposed in this example, eliminating the influence of background elements, and obtaining an RGB image of the target face. A convolutional neural network (CNN) model is then used to reconstruct the three-dimensional structure of the facial region and generate a corresponding facial depth image. The depth information of the facial depth image is modulated into the template infrared speckle image of the camera being tested and projection distortion is corrected. The image is then aligned with the face position of the target face RGB image, and the aligned infrared speckle image and target face RGB image are projected separately and input into the 3D face authentication system under test to complete the security test.
[0054] Figure 2 An exemplary process of a structured light-based face liveness detection model security testing method is shown, including:
[0055] Step 1: The image acquisition module obtains the template infrared speckle image of the structured light depth camera in the face authentication system under test and transmits it to the deep fake module.
[0056] Step 2: The depth estimation module extracts the facial area from the two-dimensional image containing the target face, obtains the target face RGB image and reconstructs the three-dimensional structure of the face, generates the corresponding face depth image, and transmits it to the deep fake module.
[0057] In step 3, the deep fake module uses the depth information in the received template infrared speckle image and the face depth image to modulate and obtain an infrared speckle image with depth information.
[0058] In step 4, the depth projection module aligns the target face RGB image with the infrared speckle image with depth information and projects them.
[0059] Step 5: The 3D face recognition system under test collects the projected image and transmits it to the face liveness detection model. If it passes the face recognition, the face recognition system is unsafe; otherwise, it is safe, and the security test is completed.
[0060] Based on the above method, the present invention conducts a security test on the face liveness detection system based on structured light depth. If the face recognition system can pass the test directly, it means that there are security defects in the system, and its liveness detection model needs to be strengthened in a targeted manner to improve the security of the face recognition system.
[0061] Below Figure 2 Each step is described in detail.
[0062] In step 1, if Figure 3 As shown, an infrared camera collects the template infrared speckle image of the structured light depth camera in the face authentication system under test, and the template infrared speckle image is transmitted to the deep forgery module through the data interface for subsequent forgery processing.
[0063] In step 2, the depth estimation module extracts the facial region from the 2D image containing the target face, obtains the target face RGB image, and estimates the depth information of the facial region. The specific method is as follows:
[0064] 1) To accurately capture the face region, the face detection bounding box should be large enough to encompass the entire facial contour; the face should be centered within the bounding box, minimizing background artifacts. This paper uses MTCNN to obtain the face detection bounding box, employs the dlib68 feature point extraction method to extract the primary facial region, and removes background artifacts. This image is then cropped and resized to 224×224 as the target face RGB image, the default input size for the depth estimation model described below.
[0065] 2) A pixel-to-pixel method based on a convolutional neural network (CNN) is used to estimate depth information from the cropped RGB image of the target face. In this example, UNet is used as the basic model architecture, and ResNet-50 is used as the encoder to reduce the 224×224 pre-processed face image to a 7×7 embedding feature map. A decoder consisting of 5 transposed convolutional layers and 10 convolutional layers is used to reconstruct the feature map into a 224×224 face depth image, where each pixel represents the absolute value of the depth.
[0066] In step 3, after receiving the template infrared speckle image and the face depth image, the deep fake module modulates the depth information in the face depth image into the template infrared speckle image to form an infrared speckle image with depth information as a fake image. The specific method is as follows:
[0067] 1) Template speckle extraction. Since different structured light depth cameras use different template infrared speckle images, it is first necessary to obtain the template infrared speckle image of the structured light depth camera in the face authentication system under test. The original infrared image is usually affected by various noises, resulting in inaccurate extracted template infrared speckle images, thereby introducing additional errors in deepfakes. Based on the characteristic that each infrared scattered point in the template infrared speckle image is sparsely distributed on the image, this paper proposes a filtering algorithm based on non-maximum suppression. Its main algorithm is as follows:
[0068] The present invention uses non-maximum suppression (NMS) to eliminate background noise. NMS is a mathematical method that selects the maximum value in an array and suppresses other values. For example, if the array A is {A1, A2, ..., An}, and the maximum value of A is Amax, then:
[0069] NMS(A)={0,0,...,A max ,...,0}
[0070] After the above filtering algorithm, the results are as follows Figure 4 As shown, (a) is before filtering, and (b) is after filtering.
[0071] 2) Scatter mapping function modeling. In order to deceive the structured light depth camera, the present invention first reversely analyzes the structured light imaging process. The imaging modeling of the structured light depth camera is as follows: Figure 5 As shown, the depth measurement can be obtained based on the trigonometric function relationship as follows:
[0072]
[0073] Among them, the structured light depth camera uses a reference depth of d ref , the scattered point displacement is Δx c , target depth d t , f c is the focal length of the infrared camera.
[0074] In order to determine the depth value of the scattering point in the filtered template infrared speckle image, the present invention extracts the coordinates of each scattering point in the filtered template into a set T = {(x1, y1)).....(x n ,y n )}, and based on the coordinates in the set T, extract the corresponding depth information from the face depth map to construct the set D = {d1, d2, ..., d n}.
[0075] The depth information is modulated into the filtered template infrared speckle image. The modulation process can be expressed as:
[0076] S=T+Φ(D)
[0077] Where Φ(.) is the mapping function that converts depth information into scattering displacement, and S is the infrared speckle image with depth information.
[0078] The structured light depth camera uses a reference depth d ref and displacement Δx c To calculate the target depth d t .like Figure 6 As shown, the depth value is d t The scattered displacement Δx c The calculation formula is:
[0079]
[0080] Among them, k c is the number of pixels within 1 mm of physical length in the structured light depth camera, L is the baseline distance between the structured light depth camera and the infrared speckle image projection submodule, and f is the focal length of the structured light depth camera. According to the principle of similar triangles, the displacement Δx in the camera c can be converted into a displacement Δx in the projector p , the calculation formula is:
[0081]
[0082] Among them, f p is the focal length of the infrared speckle image projection submodule, k p It is the number of pixels within 1 mm physical length in the infrared speckle image projection submodule.
[0083] In summary, the depth value d t and the scattering point displacement Δx p The mapping function between can be expressed as:
[0084]
[0085] Therefore, Φ(.) can be expressed as:
[0086]
[0087] In step 4, if Figure 7As shown, this example uses an infrared dot matrix projector and a projection reflector to project an infrared speckle image with depth information. Specifically, the target face RGB image is first aligned with the infrared speckle image with depth information. The infrared dot matrix projector then projects the infrared speckle image with depth information onto the projection reflector. The projection reflector is placed within the shooting range of the structured light depth camera in the face recognition system under test. During projection, there may be a physical distance L between the projector and the structured light depth camera in the face recognition system under test, resulting in projection distortion of the captured scattered image. Before alignment, the present invention uses perspective transformation to perform projection correction. The perspective transformation function is as follows:
[0088]
[0089] Where (x,y) is the point in the original scattered image captured by the structured light depth camera, is the perspective transformation matrix, is a point in the perspective-transformed scattered image, is the perspective transformation parameter.
[0090] The calculation formula of the projected corrected scattering image is as follows:
[0091]
[0092] where (x′, y′) is the point in the projected rectified scatter image.
[0093] In step 5, the 3D face authentication system under test collects the projected image and transmits it to the face liveness detection model for security testing.
[0094] Specifically, if the face liveness detection model outputs "live", it can be considered that the 3D face recognition system has a security threat; if the face liveness detection model outputs "not live", it means that the face recognition system under test is safe in resisting deep fake attacks.
[0095] It should be noted that although the image acquisition and processing and depth generation modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0096] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented as the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software, which may be collectively referred to herein as a "circuit", "module" or "system". Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and implementation are intended to be exemplary only, and the true scope and spirit of the present invention are indicated by the claims.
[0097] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings and that various modifications and variations can be made without departing from the scope thereof, which is limited only by the appended claims.
Claims
1. A security testing method for a face liveness detection model based on structured light, characterized in that: The following steps are involved: Step 1: Obtain a template infrared speckle image of a structured light depth camera in a face authentication system under test; Step 2: Extract the facial region from the two-dimensional image containing the target face, obtain the target face RGB image, reconstruct the three-dimensional structure of the face, and generate the corresponding face depth image; Step 3: Modulate the infrared speckle image with depth information based on the template infrared speckle image and the depth information in the face depth image; Step 4: Align the target face RGB image with the infrared speckle image with depth information and project them together; Step 5: The face recognition system under test collects the projected image and transmits it to the face liveness detection model. If it passes the face recognition, the face recognition system is unsafe; otherwise, it is safe, and the security test is completed.
2. The method for security testing of a face liveness detection model based on structured light according to claim 1, characterized in that: The step 2 includes: 2.1) Obtain a face detection bounding box in the 2D image containing the target face, extract the facial region using a feature point extraction method, remove background elements, and then crop it into a fixed-size RGB image of the target face; 2.2) A UNet model is used to generate a facial depth image. The UNet model uses ResNet-50 as an encoder to reduce the target face RGB image to an embedded feature map. A decoder consisting of 5 transposed convolutional layers and 10 convolutional layers is then used to reconstruct the embedded feature map into a fixed-size facial depth image, where each pixel represents the absolute value of the depth.
3. The method for security testing of a face liveness detection model based on structured light according to claim 1, characterized in that: The step 3 includes: 3.1) Using non-maximum suppression method to filter the template infrared speckle image; 3.2) The depth information in the face depth image is modulated into the filtered template infrared speckle image. The modulation process is expressed as: S=T+Φ(D) Where Φ(.) represents the mapping function that converts depth information into scattering displacement, S represents the infrared speckle image with depth information, and T represents the coordinate set of each pixel in the filtered template infrared speckle image.
4. The method for security testing of a face liveness detection model based on structured light according to claim 3, characterized in that: The mapping function for converting depth information into scattering displacement is: Among them, Φ(.) represents the mapping function, k p represents the number of pixels within 1 mm physical length in the infrared speckle image delivery submodule, L represents the baseline distance between the structured light depth camera and the infrared speckle image delivery submodule, and f p represents the focal length of the infrared speckle image projection submodule, d ref Represents the reference depth, and d represents the depth value of the pixel in the face depth image.
5. The method for security testing of a face liveness detection model based on structured light according to claim 1, characterized in that: Before aligning the target face RGB image with the infrared speckle image with depth information in step 4, a projection correction step is also included, specifically: Perform perspective transformation on infrared speckle images with depth information: Where (x,y) is the point in the infrared speckle image captured by the structured light depth camera, is the perspective transformation matrix, is a point in the infrared speckle image after perspective transformation, is the perspective transformation parameter; The calculation formula of the infrared speckle image after projection correction is as follows: Among them, (x′, y′) is the pixel point in the infrared speckle image after projection correction.
6. A face liveness detection model security testing system based on structured light, characterized by: include: An image acquisition module, which is used to obtain a template infrared speckle image of a structured light depth camera in a face authentication system under test; The depth estimation module is used to extract the facial area from the two-dimensional image containing the target face, obtain the target face RGB image and reconstruct the three-dimensional structure of the face to generate the corresponding face depth image; A deep fake module is used to modulate an infrared speckle image with depth information based on the template infrared speckle image and the depth information in the face depth image; The image projection module is used to align the target face RGB image with the infrared speckle image with depth information, and project the two. The projected image is collected by the face recognition system under test and transmitted to the face liveness detection model. If it passes the face recognition, the face recognition system is unsafe. Otherwise, it is safe, and the security test is completed.
7. The structured light-based face liveness detection model security testing system according to claim 6, characterized in that: The image projection module includes: Visible light image projection submodule, which is used to project the target face RGB image; The infrared speckle image projection submodule is used to project infrared speckle images with depth information.
8. The structured light-based facial liveness detection model security testing system according to claim 6, wherein the face recognition system under test includes a face liveness detection model, and the structured light depth camera in the face recognition system under test is capable of emitting infrared speckles arranged in a fixed pattern.
Citation Information
Patent Citations
Face recognition security defense method and system based on multi-modal visual information
CN114419710A
Depth information playback method and system based on infrared dot matrix
CN114612542A