Facial presentation attack detection method and system for double-recording link

Through the deep learning model, the RGB image data and dense point cloud data in the insurance dual recording process are voxelized and convolutional, which solves the problem of difficult to supervise and detect audio-visual data forgery behavior in the prior art, and achieves high accuracy and robust facial presentation attack detection.

CN120014716APending Publication Date: 2025-05-16CHINA LIFE INSURANCE CO LTD HEBEI BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411922375.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively supervise and detect audio and video data forgery in the insurance dual recording process, especially offline fraud, lacks effective tracking technical means, and traditional live detection methods based on plane 2D images are greatly affected by ambient light and cannot effectively identify and present attacks.

Method used

The deep learning model is adopted to collect RGB image data and dense point cloud data, and then voxelize it and perform convolution processing. The 3DCNN model is used to extract the features of voxel data, and then face presentation attack detection is performed.

Benefits of technology

It improves the accuracy of facial attack detection, enhances the robustness and adaptability of the detection system, and can maintain high detection performance under light changes, angle changes or facial occlusion, automatically determines the detection results, and whether there is a risk of facial attack.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014716A_ABST
    Figure CN120014716A_ABST
Patent Text Reader

Abstract

The invention relates to a face presentation attack detection method and system for a double-recording link. The method comprises the following steps: acquiring video data; voxelizing the video data to obtain voxel data; performing convolution on the voxel data through a deep learning model to obtain a probability result; and performing face presentation attack detection according to a confidence threshold and the probability result to obtain a detection result. According to the invention, the accuracy of video face presentation attack detection can be improved, and the security of a double-recording link is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security technology, and in particular relates to a facial presentation attack detection method and system for a dual recording link. Background Art

[0002] In insurance sales, the double recording process refers to an effective regulatory measure taken before the contract is signed to protect the insurance consumer's right to know. The specific implementation method is: during the insurance sales process, especially the reading of key information such as the terms, responsibilities, exclusions, fees, risks, etc. of the insurance product, simultaneous audio and video recording is carried out.

[0003] Existing regulatory agencies face a huge insurance market and a massive amount of double-recording files. It is impossible to manually review each double-recording document throughout the process, which makes it difficult to supervise the falsification of audio and video materials, and difficult to warn and prevent risks. In particular, there is a lack of effective tracking technology for offline fraud. To ensure that insurance double-recording meets the requirements and avoid irregular double-recording, insurance institutions need to conduct manual quality inspections on a large number of double-recording files. However, due to the possibility of adverse selection in technology and human factors, there is a possibility of partial falsification of double-recording, and there is currently a lack of effective tracking technology. In addition, the current traditional liveness detection method based on flat 2D images is greatly affected by ambient lighting, and when attacked by presentation, it can only read color information and cannot effectively identify attacks such as the other screen displaying the pre-recorded video. Summary of the invention

[0004] In view of the above deficiencies in the prior art, an object of the present invention is to provide a facial presentation attack detection method and system for dual recording.

[0005] The present invention provides a facial presentation attack detection method for a dual recording link, comprising:

[0006] S1: Collect and record data;

[0007] S2: voxelizing the recorded data to obtain voxel data;

[0008] S3: Convolving the voxel data through a deep learning model to obtain a probability result;

[0009] S4: Perform facial presentation attack detection according to the confidence threshold and the probability result to obtain a detection result.

[0010] According to a facial presentation attack detection method for a dual recording link provided by the present invention, the recorded data in step S1 includes RGB image data and dense point cloud data.

[0011] According to a facial presentation attack detection method for a dual recording link provided by the present invention, the detection model in step S3 is a 3DCNN model.

[0012] According to a facial presentation attack detection method for a dual recording link provided by the present invention, step S3 further comprises:

[0013] S31: inputting the voxel data into a deep learning model;

[0014] S32: extracting features from the voxel data using the deep learning model to obtain output features;

[0015] S33: Perform tensor output according to the output features to obtain a probability result.

[0016] According to a facial presentation attack detection method for a dual recording link provided by the present invention, step S32 further comprises:

[0017] S321: performing convolution on the voxel data through a first convolution module to obtain convolution features;

[0018] S322: Pooling the convolutional features to obtain pooled features;

[0019] S323: extracting features from the pooled features through a shared fully connected layer to obtain point cloud features;

[0020] S324: Processing the point cloud features through a Sigmoid activation function to obtain channel attention;

[0021] S325: Weighting the channel attentions channel by channel to obtain an attention cycle unit;

[0022] S326: Perform feature extraction through the attention cycle unit to obtain output features.

[0023] According to a facial presentation attack detection method for a dual recording link provided by the present invention, the first convolution module in step S321 includes a ReLU layer and multiple 3D convolution layers.

[0024] According to a facial presentation attack detection method for a dual recording link provided by the present invention, in step S322, the pooling of the convolutional features includes global maximum pooling and global average pooling.

[0025] According to a facial presentation attack detection method for a dual recording link provided by the present invention, step S33 further comprises:

[0026] S331: Input the output features into a second convolution module to obtain dimension reduction features;

[0027] S332: Perform adaptive pooling on the dimension reduction features to obtain an output tensor;

[0028] S333: Process the output tensor through a softmax activation function to obtain a probability result.

[0029] According to a facial presentation attack detection method for a dual recording link provided by the present invention, step S4 specifically includes:

[0030] When the probability result is greater than or equal to a preset confidence threshold, the detection result is that there is a risk of facial presentation attack;

[0031] When the probability result is less than a preset confidence threshold, the detection result is that there is no risk of facial presentation attack.

[0032] The present invention also provides a facial presentation attack detection system for dual recording, which is used to execute a facial presentation attack detection method for dual recording as described in any one of the above items, including:

[0033] A collection module, used for collecting and recording data;

[0034] A voxel module, used for voxelizing the recorded data to obtain voxel data;

[0035] A convolution module, wherein the convolution module is configured as a deep learning model and is used to convolve the voxel data to obtain a probability result;

[0036] The determination module is used to perform facial presentation attack detection according to the confidence threshold and the probability result to obtain a detection result.

[0037] The present invention provides a facial presentation attack detection method and system for a dual recording link. By collecting RGB image data and dense point cloud data as recording data, and using a 3DCNN model to perform deep convolution processing on voxel data, facial features can be captured and analyzed more comprehensively, thereby effectively distinguishing the difference between a real face and an attacker, and improving the accuracy of facial presentation attack detection. The method of the present invention not only takes into account two-dimensional image information, but also combines three-dimensional point cloud data, which is helpful for coping with various complex scenes and lighting conditions, and enhancing the robustness and adaptability of the detection system, especially in the case of lighting changes, angle changes or facial occlusion, etc., it can still maintain a high detection performance. In addition, the deep learning model designed by the present invention can efficiently extract key features from voxel data, providing strong support for subsequent probability calculations. According to a preset confidence threshold, the present invention can automatically judge the detection result, that is, whether there is a risk of facial presentation attack. This flexible detection mechanism enables the system to adjust the threshold according to actual needs to achieve the best detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings are only used to illustrate specific embodiments and are not considered to limit the present invention. In the entire drawings, the same reference symbols represent the same components. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0039] Figure 1 A schematic flow chart of a facial presentation attack detection method for a dual recording link provided by an embodiment of the present invention;

[0040] Figure 2 A schematic diagram of the structure of a facial presentation attack detection system for a dual recording session provided by an embodiment of the present invention.

[0041] Reference numerals:

[0042] 100, acquisition module; 200, voxel module; 300, convolution module; 400, determination module. DETAILED DESCRIPTION

[0043] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work should fall within the scope of protection of the present invention.

[0044] Furthermore, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts disclosed in the present invention.

[0045] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the orientation or position relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. The terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0046] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of methods and systems consistent with some aspects of the present invention as detailed in the appended claims.

[0047] In order to better understand the present invention, the terms used in the embodiments of the present invention are first explained below.

[0048] Double recording: simultaneous audio and video recording of the insurance sales process.

[0049] Voxelization: Voxelize the entire point cloud into a three-dimensional voxel grid, using an approximation of each voxel grid to represent the voxel block, thereby simplifying calculations, accelerating processing, reducing noise, and facilitating subsequent three-dimensional data processing and analysis.

[0050] 3DCNN: A deep learning model that processes three-dimensional input data and can better capture feature information in the spatial dimension.

[0051] Facial presentation attack: refers to the use of various tools and techniques to simulate or forge the facial features of legitimate users in order to deceive biometric systems, especially face recognition systems, and thus obtain unauthorized access or authentication.

[0052] The embodiments of the present invention are described below with reference to the accompanying drawings.

[0053] like Figure 1 As shown, the present invention provides a facial presentation attack detection method for a dual recording link, comprising:

[0054] S1: Collect and record data.

[0055] The recorded data in step S1 includes RGB image data and dense point cloud data.

[0056] The purpose of step S1 is to synchronously obtain pixel images and infrared dense point clouds from the mobile camera device. RGB image data is color image information captured by the camera device. RGB image data can intuitively reflect the facial expressions, movements and surrounding environment of both parties to the transaction. It is an important basis for judging whether the two parties are truly involved and whether there is fraud. The dense point cloud data is three-dimensional spatial information captured by infrared or depth sensors. Compared with RGB images, point cloud data provides more spatial information and can construct a three-dimensional model of the two parties to the transaction, so as to more accurately judge their position and posture. In the double recording link, point cloud data helps to identify whether there is a facial presentation attack (such as using photos, videos, etc. to deceive).

[0057] S2: voxelizing the recorded data to obtain voxel data.

[0058] Furthermore, this step is to voxelize the dense point cloud data obtained in the above step S1. Voxelization is the process of converting three-dimensional space data into voxels (volume pixels). Voxels are the smallest units in three-dimensional space, similar to pixels in two-dimensional images. In the dual recording link, the purpose of voxelizing the recorded data is to convert the dense point cloud data into a format that is easier to process and analyze.

[0059] The purpose of voxelization in step S2 is, firstly, to simplify data processing: voxelization simplifies complex three-dimensional spatial data into a regular grid structure, making subsequent data processing and analysis more efficient; secondly, it also improves detection accuracy: by dividing the three-dimensional space into smaller voxel units, abnormal behaviors such as facial attacks can be captured and identified more finely; in addition, it also enhances robustness: voxelized data is highly robust to interference factors such as lighting and occlusion, and can maintain high detection accuracy in complex environments.

[0060] S3: Convolve the voxel data through a deep learning model to obtain a probability result.

[0061] Wherein, the detection model in step S3 is a 3DCNN model.

[0062] Furthermore, the working principle of 3DCNN is to extract spatial and temporal features of video data through layer-by-layer convolution and pooling operations. In the convolution layer, the 3D convolution kernel performs local convolution operations on the input data to extract features. The obtained features are activated by nonlinear activation functions (such as ReLU) in subsequent layers and downsampled by pooling layers to reduce the dimension and amount of calculation of the data. Finally, in the fully connected layer, the extracted features are used for classification or regression tasks.

[0063] Wherein, step S3 further comprises:

[0064] S31: Input the voxel data into a deep learning model.

[0065] S32: Extract features from the voxel data using the deep learning model to obtain output features.

[0066] Step S32 further includes:

[0067] S321: Convolve the voxel data through a first convolution module to obtain convolution features.

[0068] Among them, the first convolution module in step S321 includes a ReLU layer and multiple 3D convolution layers.

[0069] Furthermore, in step S321, the voxelized data is first sent to a combination of several 3D convolutional layers and leaky ReLU layers to process the 3D spatial data and extract spatial features. The leaky ReLU (Rectified Linear Unit) is a nonlinear activation function that allows small negative gradient values ​​to pass through, avoiding completely "dead" neurons and helping to solve the gradient vanishing problem.

[0070] S322: Pooling the convolutional features to obtain pooled features.

[0071] In step S322, the pooling of the convolutional features includes global maximum pooling and global average pooling.

[0072] In step S322, the feature layer data in step S321 is input into the spatial attention unit, and global average pooling (AvgPool) and global maximum pooling (MaxPool) are performed respectively. The spatial attention unit is designed to enhance the model's attention to important spatial areas in the input data. Global average pooling and global maximum pooling respectively extract average features and maximum features from the entire feature map, which helps the model understand global context information.

[0073] S323: Extracting features from the pooled features through a shared fully connected layer to obtain point cloud features.

[0074] The features obtained in step S322 enter the shared fully connected layer (Shared MLP) for processing in step S323 to further obtain features from the unordered point cloud data. MLP (Multi-layer Perceptron) is a fully connected neural network used to further process the features extracted from the previous steps. Since point cloud data is usually unordered, using shared MLP can ensure that the model is insensitive to the arrangement of the input data.

[0075] S324: Process the point cloud features through a Sigmoid activation function to obtain channel attention.

[0076] In step S324, the output is converted into a value between 0 and 1 mainly through the Sigmoid function. These values ​​can be interpreted as weights, that is, the channel attention map is obtained, that is, the weight of each channel of the input feature layer is obtained, indicating the importance of each channel.

[0077] S325: Weight the channel attentions channel by channel to obtain an attention cycle unit.

[0078] S326: Perform feature extraction through the attention cycle unit to obtain output features.

[0079] In steps S325 to S326, the weights obtained in step S324 are mainly used to reweight each channel of the input feature map, that is, the weights are re-weighted channel by channel through multiplication to the input feature layer, completing the attention unit cycle, enhancing important features, and suppressing unimportant features.

[0080] S33: Perform tensor output according to the output features to obtain a probability result.

[0081] Wherein, step S33 further comprises:

[0082] S331: Input the output features into a second convolution module to obtain dimension reduction features.

[0083] Furthermore, in step S331, the final result output by the attention unit is input into a smaller 3D convolution layer for dimensionality reduction to reduce the number of model parameters, aiming to reduce the computational complexity and overfitting risk of the model by reducing the size and number of feature maps.

[0084] S332: Perform adaptive pooling on the dimension reduction features to obtain an output tensor.

[0085] In step S332, an adaptive average pooling layer is used to ensure that the output tensor has a fixed size, that is, the parameters are input into the adaptive average pooling layer to adjust the output tensor shape, which is necessary for subsequent fully connected layers or classification tasks.

[0086] S333: Process the output tensor through a softmax activation function to obtain a probability result.

[0087] Finally, in step S333, the output result is normalized using the softmax normalization function. The Softmax function converts the output into a probability distribution so that the sum of all output values ​​is 1, which is then used as the output layer for subsequent multi-classification problems.

[0088] S4: Perform facial presentation attack detection according to the confidence threshold and the probability result to obtain a detection result.

[0089] In the aforementioned steps S1 to S3, the probability result of facial presentation attack detection on the current recorded data has been obtained from the detection model. The probability result is usually a value between 0 and 1, indicating the possibility that the current data is judged to have a facial presentation attack risk.

[0090] Then, the present invention compares this probability result with a preset confidence threshold in step S4. If the probability result is greater than or equal to the threshold, it is considered that the current data has a risk of facial presentation attack; if the probability result is less than the threshold, it is considered that the current data does not have a risk of facial presentation attack.

[0091] Finally, the final detection result is output based on the comparison result. The result is a binary judgment, such as "there is a risk" or "there is no risk". Of course, it can also be provided as a more detailed report including probability results, such as "the detection probability is 0.75, and there is a risk of facial presentation attack."

[0092] Wherein, step S4 specifically includes:

[0093] When the probability result is greater than or equal to a preset confidence threshold, the detection result is that there is a risk of facial presentation attack;

[0094] When the probability result is less than a preset confidence threshold, the detection result is that there is no risk of facial presentation attack.

[0095] In a specific embodiment, the initial setting is to set the confidence threshold to 0.5, that is, when the probability result output by the detection model is greater than or equal to 0.5, the system considers that there is a risk of facial presentation attack; otherwise, it considers that there is no risk.

[0096] In practical applications, it is necessary to adjust the threshold according to the actual performance of the detection system and the specific requirements of the application scenario. For example, in this embodiment, it is a double recording link in the insurance industry. In identity authentication in financial transactions, the application scenario is very sensitive to false alarms, so the threshold needs to be set higher to reduce the false alarm rate. In normal contract signing or other security monitoring anomaly detection, the application scenario is very sensitive to missed alarms, so the threshold needs to be set lower to increase the recall rate.

[0097] like Figure 2 As shown, the present invention also provides a facial presentation attack detection system for a dual recording link, comprising:

[0098] The acquisition module 100 is used to acquire the recorded data;

[0099] A voxel module 200, used for voxelizing the recorded data to obtain voxel data;

[0100] A convolution module 300, the convolution module is configured as a deep learning model, and is used to convolve the voxel data to obtain a probability result;

[0101] The determination module 400 is used to perform facial presentation attack detection according to the confidence threshold and the probability result to obtain a detection result.

[0102] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0103] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0104] The present invention provides a facial presentation attack detection method and system for a dual recording link, which simultaneously adopts optical RGB image color information and dense point cloud depth information, significantly improves the three-dimensional judgment of the screen in front of the camera and the real person, while retaining the recognition ability of masks and wrapped printing paper, reducing the influence of ambient light, and can improve the accuracy of detection of presentation attacks, improve the ability to control the risk of forged customer image in the dual recording link of insurance sales, and improve the security of the dual-path link.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed in the present invention should be covered within the protection scope of the present invention.

Claims

1. A facial presentation attack detection method for dual recording, characterized in that: include: S1: Collect and record data; S2: voxelizing the recorded data to obtain voxel data; S3: Convolving the voxel data through a deep learning model to obtain a probability result; S4: Perform facial presentation attack detection according to the confidence threshold and the probability result to obtain a detection result.

2. A facial presentation attack detection method for dual recording according to claim 1, characterized in that: The recorded data in step S1 includes RGB image data and dense point cloud data.

3. A facial presentation attack detection method for dual recording according to claim 1, characterized in that: The detection model in step S3 is a 3DCNN model.

4. A facial presentation attack detection method for dual recording according to claim 1, characterized in that: Step S3 further comprises: S31: inputting the voxel data into a deep learning model; S32: extracting features from the voxel data using the deep learning model to obtain output features; S33: Perform tensor output according to the output features to obtain a probability result.

5. A facial presentation attack detection method for dual recording according to claim 4, characterized in that: Step S32 further includes: S321: performing convolution on the voxel data through a first convolution module to obtain convolution features; S322: Pooling the convolutional features to obtain pooled features; S323: extracting features from the pooled features through a shared fully connected layer to obtain point cloud features; S324: Processing the point cloud features through a Sigmoid activation function to obtain channel attention; S325: Weighting the channel attentions channel by channel to obtain an attention cycle unit; S326: Perform feature extraction through the attention cycle unit to obtain output features.

6. A facial presentation attack detection method for dual recording according to claim 5, characterized in that: The first convolution module in step S321 includes a ReLU layer and multiple 3D convolution layers.

7. A facial presentation attack detection method for dual recording according to claim 5, characterized in that: In step S322, the pooling performed on the convolutional features includes global maximum pooling and global average pooling.

8. A facial presentation attack detection method for dual recording according to claim 4, characterized in that: Step S33 further comprises: S331: Input the output features into a second convolution module to obtain dimension reduction features; S332: Perform adaptive pooling on the dimension reduction features to obtain an output tensor; S333: Process the output tensor through a softmax activation function to obtain a probability result.

9. A facial presentation attack detection method for dual recording according to claim 1, characterized in that: Step S4 specifically includes: When the probability result is greater than or equal to a preset confidence threshold, the detection result is that there is a risk of facial presentation attack; When the probability result is less than a preset confidence threshold, the detection result is that there is no risk of facial presentation attack.

10. A facial presentation attack detection system for dual recording, used to execute a facial presentation attack detection method for dual recording as claimed in any one of claims 1 to 9, characterized in that: include: A collection module, used for collecting and recording data; A voxel module, used for voxelizing the recorded data to obtain voxel data; A convolution module, wherein the convolution module is configured as a deep learning model and is used to convolve the voxel data to obtain a probability result; The determination module is used to perform facial presentation attack detection according to the confidence threshold and the probability result to obtain a detection result.