Remote identity verification method based on hand-face interaction deformation judgment
By introducing a method of hand-face interaction deformation judgment in facial recognition identity verification, using deformation model and 3D point cloud technology, the problem of difficult to judge non-rigid deformation of faces in the existing technology is solved, and the defense ability of AIGC synthesis attacks is significantly improved.
Patent Information
- Application Number
- CN202510199360.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
The existing identity verification method based on face interaction fails to effectively consider the deformation caused by non-rigid changes in the face when interacting with the finger, which makes it easy to be simulated by AIGC technology.
A remote identity verification method based on hand-face interaction deformation judgment is adopted. By randomly selecting preset actions, using a mobile phone camera to collect images, combining the face key point model and gesture key point model, we can judge whether the deformation of face and hand interactions meets expectations through the deformation model. The deformation model processes images through convolutional neural networks, estimates deformation information, and generates 3D point cloud data to correct and reconstruct 3D shapes.
Effectively judge the deformation of face and hand interaction, improve the defense ability of synthetic attacks, and reduce the possibility of black-produced attackers attacking by injecting prefabricated videos.
Smart Images

Figure CN119992629A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of face recognition technology, and in particular to a remote identity verification method based on hand-face interactive deformation judgment. Background Art
[0002] In the field of face liveness, determining whether the user's face is a real person is an important issue. With the rapid development of AIGC technology, synthesized face videos are becoming more and more realistic. Not only can they synthesize actions such as "shaking head" and "opening mouth" in real time, but they can also synthesize faces under different light conditions in real time according to needs. This makes the commonly used motion liveness and light liveness technologies more difficult to defend against such synthetic attacks. Existing identity verification methods based on face interaction mainly rely on the positional relationship between the face and fingers for judgment, which is achieved through face key points and finger key points. However, this type of method does not take into account the deformation caused by non-rigid changes in the face when the face interacts with the fingers, and is therefore easily simulated by AIGC technology. Summary of the invention
[0003] In order to help solve the above technical problems, the present application provides a remote identity verification method based on hand-face interactive deformation judgment, which adopts the following technical solutions: A remote identity verification method based on hand-face interactive deformation judgment, wherein the method comprises: Step S1: randomly select a preset action and send it, display a schematic diagram corresponding to the preset action on the front-end interface of the mobile phone, and require the user to interact with the hand and face according to the schematic diagram; Step S2: Capture several frames of continuous images through the mobile phone camera and transmit them to the backend server; Step S3: The background server uses the face key point model and the hand gesture key point model to obtain preliminary key points, and determines whether they meet the positional relationship of the issued action. If not, a failure result is immediately returned to the user. If they meet, the deformation model is used to determine whether the deformation of the face and hand interaction meets expectations. If not, a failure result is returned to the user. The deformation model determines whether the deformation of the face and hand interaction meets expectations in the following ways: Step S31: the input image is processed by a convolutional neural network, and the convolutional neural network extracts key features of the input image; Step S32: estimating deformation information in the image through the deformation model to obtain an A-dimensional offset, where the A-dimensional offset represents an adjustment amount in the 3D model space, and is used to correct and reconstruct the 3D shape; Step S33: Generate corresponding point clouds by combining the predefined Flame face model and MANO hand model with the offset of the A dimension. For the face, a B-dimensional point cloud is generated, which indicates the position of each point on the face surface in the 3D space. For the hand, a C-dimensional point cloud is generated, which indicates the position of each point on the hand in the 3D space. Step S34: The deformation model outputs point cloud data containing estimated face and hand parameters.
[0004] Preferably, step S1 includes: the preset action includes a single-finger pressing face preset action, a double-finger pinching face preset action and a hand clenching fist pressing face preset action, the single-finger pressing face preset action acts on the single-finger pressing face area, the single-finger pressing face area includes the upper left cheek, the middle left cheek, the lower left cheek, the upper right cheek, the middle right cheek, and the lower right cheek, the double-finger pinching face preset action acts on the double-finger pinching face area, the double-finger pinching face area includes the left cheek, the right cheek and the chin, and the hand clenching fist pressing face preset action acts on the hand clenching fist pressing face area, and the hand clenching fist pressing face area includes the left cheek and the right cheek.
[0005] Preferably, the step S3 comprises: judging the overlapping area by using a 256-point key point model of the face and a 25-point key point model of the gesture in combination with a preset action.
[0006] Preferably, A is 5023, B is 5023×3, and C is 778×3.
[0007] To summarize, this application determines whether it is a real person by judging the deformation caused by the interaction between the face and the hands, thereby solving the problem that the current AIGC model has poor synthesis effect for non-rigid deformations, and greatly improving the defense capability against synthetic attacks. This application enhances the randomness of the identity verification process by randomly issuing action groups, and reduces the possibility of black market attackers launching attacks by injecting a pre-made video. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A schematic diagram of an embodiment of a preset action of the present application; Figure 2 A schematic block diagram of an embodiment of a deformation model of the present application. DETAILED DESCRIPTION
[0009] The present application is further described below in conjunction with the accompanying drawings, and the structure and principle of the present application are very clear to people in the field. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0010] Figure 1 is a schematic diagram of an embodiment of the preset action of this application, Figure 2 A schematic block diagram of an embodiment of a deformation model of the present application.
[0011] Combination Figure 1 and Figure 2 The remote identity verification method based on hand-face interaction deformation judgment of the present application may include: Step S1: Randomly select a preset action and send it, display a schematic diagram corresponding to the preset action on the front-end interface of the mobile phone, and require the user to interact with the hand and face according to the schematic diagram. The preset actions include a single-finger face pressing preset action, a double-finger face pinching preset action, and a hand fist pressing face preset action. The single-finger face pressing preset action acts on the single-finger face pressing area, and the single-finger face pressing area includes the upper left cheek, the middle left cheek, the lower left cheek, the upper right cheek, the middle right cheek, and the lower right cheek. The double-finger face pinching preset action acts on the double-finger face pinching area, and the double-finger face pinching area includes the left cheek, the right cheek, and the chin. The hand fist pressing face preset action acts on the hand fist pressing face area, and the hand fist pressing face area includes the left cheek and the right cheek.
[0012] In step S1, by randomly selecting preset actions, the difficulty of prediction and simulation by attackers can be increased, thereby improving the security of the system. Displaying a schematic diagram corresponding to the preset action on the front-end interface of the mobile phone can intuitively guide the user to complete the action and ensure the accuracy and consistency of the action. Step S1 increases the system's requirements for the accuracy and consistency of real user actions, reducing the possibility of successful simulation by attackers.
[0013] Step S2: Capture several frames of continuous images through the mobile phone camera and transmit them to the background server.
[0014] Step S3: Use the universal 256-point key point model of the face and the 25-point key point model of the gesture to judge the overlapping area in combination with the preset action. Take the action of pressing the left middle cheek with a single finger as an example. It is necessary to judge whether the coordinates of the fingertip position of the index finger are in the left cheek position area, and judge whether the distance of the 2D coordinate point is within the set threshold. The background server uses the above-mentioned face key point model and gesture key point model to obtain preliminary key points, and judge whether they meet the position relationship of the issued action. If not, a failure result is immediately returned to the user. If it meets, the deformation model is used to judge whether the deformation of the face and hand interaction meets the expectations. If not, a failure result is returned to the user. The deformation model judges whether the deformation of the face and hand interaction meets the expectations in the following ways: Step S31: The input image is processed by a convolutional neural network, and the convolutional neural network extracts key features of the input image.
[0015] Step S32: estimate the deformation information in the image through the deformation model to obtain an A-dimensional offset, where the A-dimensional offset represents the adjustment amount in the 3D model space and is used to correct and reconstruct the 3D shape.
[0016] Step S33: Generate corresponding point clouds through the predefined Flame face model and MANO hand model combined with the A-dimensional offset. For the face, a B-dimensional point cloud is generated, which represents the position of each point on the face surface in the 3D space. For the hand, a C-dimensional point cloud is generated, which represents the position of each point on the hand in the 3D space.
[0017] Step S34: The deformation model outputs point cloud data containing estimated face and hand parameters.
[0018] In this embodiment, A is 5023, B is 5023×3, and C is 778×3. In actual use, it is only necessary to obtain the 5023-dimensional offset through model estimation, and to determine whether the offset is consistent with the issued action in combination with the issued action.
Claims
1. A remote identity verification method based on hand-face interactive deformation judgment, characterized in that: The method comprises: Step S1: randomly select a preset action and send it, display a schematic diagram corresponding to the preset action on the front-end interface of the mobile phone, and require the user to interact with the hand and face according to the schematic diagram; Step S2: Capture several frames of continuous images through the mobile phone camera and transmit them to the backend server; Step S3: The background server uses the face key point model and the hand gesture key point model to obtain preliminary key points, and determines whether they meet the positional relationship of the issued action. If not, a failure result is immediately returned to the user. If they meet, the deformation model is used to determine whether the deformation of the face and hand interaction meets expectations. If not, a failure result is returned to the user. The deformation model determines whether the deformation of the face and hand interaction meets expectations in the following ways: Step S31: the input image is processed by a convolutional neural network, and the convolutional neural network extracts key features of the input image; Step S32: estimating deformation information in the image through the deformation model to obtain an A-dimensional offset, where the A-dimensional offset represents an adjustment amount in the 3D model space, and is used to correct and reconstruct the 3D shape; Step S33: Generate corresponding point clouds by combining the predefined Flame face model and MANO hand model with the offset of the A dimension. For the face, a B-dimensional point cloud is generated, which indicates the position of each point on the face surface in the 3D space. For the hand, a C-dimensional point cloud is generated, which indicates the position of each point on the hand in the 3D space. Step S34: The deformation model outputs point cloud data containing estimated face and hand parameters.
2. The remote identity verification method based on hand-face interactive deformation judgment according to claim 1 is characterized in that: The step S1 includes: the preset actions include a single-finger pressing face preset action, a double-finger pinching face preset action and a hand clenching fist pressing face preset action, the single-finger pressing face preset action acts on the single-finger pressing face area, the single-finger pressing face area includes the upper left cheek, the middle left cheek, the lower left cheek, the upper right cheek, the middle right cheek, and the lower right cheek, the double-finger pinching face preset action acts on the double-finger pinching face area, the double-finger pinching face area includes the left cheek, the right cheek and the chin, and the hand clenching fist pressing face preset action acts on the hand clenching fist pressing face area, the hand clenching fist pressing face area includes the left cheek and the right cheek.
3. The remote identity verification method based on hand-face interactive deformation judgment according to claim 1 is characterized in that: The step S3 includes: judging the overlapping area by using the 256-point key point model of the face and the 25-point key point model of the gesture in combination with the preset action.
4. The remote identity verification method based on hand-face interactive deformation judgment according to claim 1 is characterized in that: A is 5023, B is 5023×3, and C is 778×3.