Face image processing method, device, computer equipment and storage medium

By generating and projecting expressive three-dimensional facial images onto a model, the method addresses inefficiencies in facial recognition verification, reducing costs and improving system performance.

CN111553284BActive Publication Date: 2025-07-15WUHAN UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010357461.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-29
Publication Date
2025-07-15
Estimated Expiration
2040-04-29

AI Technical Summary

Technical Problem

The existing facial recognition verification method requires a large number of volunteers or customized 3D silicone masks, which are cost-effective and inefficient.

Method used

By acquiring the first face image of the first user, generating a three-dimensional projected image, extracting the expression features in the second user's face video, fusing it to the three-dimensional projected image for expression reconstruction, generating a synthetic face video and projecting it to a three-dimensional solid model for playback, and for face recognition.

Benefits of technology

It effectively improves the verification efficiency of face recognition, reduces costs, and improves the security and accuracy of face recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111553284B_ABST
    Figure CN111553284B_ABST
Patent Text Reader

Abstract

The present application relates to a face image processing method, apparatus, computer device, and storage medium. The method includes: obtaining a first face image of a first user; generating a corresponding three-dimensional projection image based on the face features of the first face image; obtaining a second face video of a second user, and extracting face expression features based on each second face image in the second face video; respectively fusing the extracted face expression features into the three-dimensional projection image for expression reconstruction to obtain a synthesized face video; projecting the synthesized face video onto a three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition. By using this method, the verification efficiency of face recognition can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a face image processing method, apparatus, computer device, and storage medium. Background Art

[0002] Face recognition is a biometric identification technology that performs identity recognition based on human facial feature information. With the rapid development of computer technologies, face recognition technology is increasingly used for identity verification in more and more application scenarios. To verify the recognition effect of a face recognition system, face recognition can be performed through real test users or 3D mask models. Currently, these verification methods usually require recruiting a large number of volunteer users or customizing a large number of 3D silicone masks, which consume a large amount of resource costs, have a high cost, and a low verification efficiency. Summary of the Invention

[0003] Based on this, to address the above technical problems, it is necessary to provide a face image processing method, apparatus, computer device, and storage medium that can effectively improve the verification efficiency of face recognition.

[0004] A face image processing method, the method comprising:

[0005] Obtain a first face image of a first user;

[0006] Generate a corresponding three-dimensional projection image based on the facial features of the first face image;

[0007] Obtain a second face video of a second user, and extract facial expression features based on each second face image in the second face video;

[0008] Fuse each of the extracted facial expression features into the three-dimensional projection image for expression reconstruction to obtain a synthesized face video;

[0009] Project the synthesized face video onto a three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition.

[0010] A face image processing apparatus, the apparatus comprising:

[0011] An image acquisition module, configured to obtain a first face image of a first user;

[0012] An image conversion module, configured to generate a corresponding three-dimensional projection image based on the facial features of the first face image;

[0013] An expression extraction module, configured to obtain a second face video of a second user, and extract facial expression features based on each second face image in the second face video;

[0014] An expression reconstruction module, configured to respectively fuse the obtained face expression features into the three-dimensional projection image for expression reconstruction, so as to obtain a synthesized face video;

[0015] A video projection module, configured to project the synthesized face video onto a three-dimensional solid model for playing, and the projected three-dimensional solid model is used for face recognition.

[0016] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0017] Obtain a first face image of a first user;

[0018] Generate a corresponding three-dimensional projection image based on the face features of the first face image;

[0019] Obtain a second face video of a second user, and extract face expression features based on each second face image in the second face video;

[0020] Respectively fuse the obtained face expression features into the three-dimensional projection image for expression reconstruction, so as to obtain a synthesized face video;

[0021] Project the synthesized face video onto a three-dimensional solid model for playing, and the projected three-dimensional solid model is used for face recognition.

[0022] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0023] Obtain a first face image of a first user;

[0024] Generate a corresponding three-dimensional projection image based on the face features of the first face image;

[0025] Obtain a second face video of a second user, and extract face expression features based on each second face image in the second face video;

[0026] Respectively fuse the obtained face expression features into the three-dimensional projection image for expression reconstruction, so as to obtain a synthesized face video;

[0027] Project the synthesized face video onto a three-dimensional solid model for playing, and the projected three-dimensional solid model is used for face recognition.

[0028] The above-mentioned face image processing method, device, computer equipment and storage medium, after obtaining the first face image of the first user, generate a corresponding three-dimensional projection image based on the face features of the first face image, thereby being able to effectively generate a face image for projection onto a three-dimensional solid model and displaying the effect of a real face. By obtaining the second face video of the second user, extracting face expression features from each second face image in the second face video, and then respectively fusing the extracted face expression features into the three-dimensional projection image for expression reconstruction, it is possible to accurately and effectively perform expression reconstruction on the expression features in the three-dimensional projection image, and effectively generate a synthetic face video with high authenticity that includes an expression response. Project the synthetic face video onto the three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition. By performing face recognition on the projection of the synthetic face video that fuses the face of the first user and the expression of the second user on the three-dimensional solid model, the effect of the face recognition system can be effectively verified. By generating a synthetic face video with high accuracy and authenticity and projecting it onto a reusable three-dimensional solid model for face recognition, the cost of verification can be effectively saved, and the efficiency of the face recognition system can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a diagram of the application environment of the face image processing method in an embodiment;

[0030] Figure 2 It is a schematic flowchart of the face image processing method in an embodiment;

[0031] Figure 3 It is a schematic flowchart of a person performing expression reconstruction in an embodiment;

[0032] Figure 4 It is a schematic diagram of projecting the synthetic face video onto the three-dimensional solid model in an embodiment;

[0033] Figure 5 It is a schematic flowchart of the face image processing method in another embodiment;

[0034] Figure 6 It is a schematic flowchart of the face image processing method in yet another embodiment;

[0035] Figure 7 It is a schematic flowchart of the face image processing method in still another embodiment;

[0036] Figure 8 It is a schematic flowchart of the face image processing method in a specific embodiment;

[0037] Figure 9 It is a block diagram of the structure of the face image processing device in an embodiment;

[0038] Figure 10 It is the internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0040] The solution provided by the embodiment of the present application relates to technologies such as biometric recognition, computer vision, and image processing based on artificial intelligence. Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results in theory, technology, and application systems. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0041] Computer Vision Technology (CV): Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further performing graphic processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0042] The face image processing method provided by the present application can be applied to a terminal or a server. It can be understood that it can also be applied to a system including a terminal and a server and be implemented through the interaction between the terminal and the server. Refer to Figure 1 , through the interaction between the terminal and the server, it can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The first face image is sent to the server 104. The server 104 obtains the first face image of the first user and generates a corresponding three-dimensional projection image based on the face features of the first face image; the server 104 obtains the second face video of the second user collected by the terminal 102 and extracts face expression features from each second face image in the second face video; the extracted face expression features are respectively fused into the three-dimensional projection image for expression reconstruction, and a synthesized face video is obtained and sent to the terminal 102. The terminal 102 projects the synthesized face video onto a three-dimensional solid model for playback, and the projected three-dimensional solid model is used for face recognition. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0043] In one embodiment, as Figure 2 shown, a face image processing method is provided. Taking the terminal in Figure 1 as an example, the method includes the following steps:

[0044] S202, obtain the first face image of the first user.

[0045] Among them, a face image refers to an image including a user's face. The face image can be extracted from an image or video stream including the user's face to obtain the face image corresponding to the face area.

[0046] The terminal can pre-obtain an image or video including the first user's face from the local database and extract the first face image corresponding to the face area. The terminal can also crawl an image or video including the first user's face from the network and extract the first face image corresponding to the face area as the first face image to be processed.

[0047] In one of the embodiments, the terminal can also directly collect an image or video including the first user's face and extract the first face image corresponding to the first user from the collected image or video.

[0048] After the terminal obtains the initial face image including the first user's face, since there may be various noises and random interferences in the obtained original image, it is necessary to perform image preprocessing such as gray correction and noise filtering on the directly obtained initial face image. Specifically, the terminal first detects the face area in the initial face image and preprocesses the image based on the face area detection result. For a face image, its preprocessing process may include light compensation, gray transformation, histogram equalization, normalization, geometric correction, filtering, and sharpening of the initial face image. After preprocessing the initial face image, the first face image of the first user is obtained.

[0049] S204. Generate a corresponding three-dimensional projection image based on the face features of the first face image.

[0050] Among them, the first face image is a two-dimensional image, and a two-dimensional image refers to a planar image that does not contain depth information. Three-dimensional (i.e., 3D) refers to a spatial system formed by adding a direction vector to the planar two-dimensional system. A three-dimensional projection image can represent an image of a three-dimensional model projected onto a two-dimensional (i.e., 2D) plane, and a corresponding three-dimensional image is displayed after being projected onto the three-dimensional solid model.

[0051] The first face image includes face feature points such as eyebrows, eyes, nose, lips, and chin. The face feature points are mainly distributed at face contours such as the brow bone, nose bridge, eye contour, lip contour, and jaw line, which can reflect the contour and expression of the face. After the terminal obtains the first face image of the first user, it first extracts the feature points from the first face image, estimates the face pose based on the feature points, and obtains the face features of the first face image based on these extracted feature points and the face pose.

[0052] The face image obtained by the terminal is a two-dimensional image. When projecting the image onto a three-dimensional solid model for use, it is necessary to establish a mapping relationship between the 2D feature points and 3D feature points of the face. The terminal can match and map the two-dimensional coordinate information of the feature points in the first face image with the feature points in the preset three-dimensional face model respectively to obtain the corresponding three-dimensional coordinate information.

[0053] Specifically, after the terminal obtains the face features of the first face image, it can perform face modeling through three-dimensional model mapping based on the obtained two-dimensional face feature points to obtain the three-dimensional coordinate information corresponding to the face features. For example, a 3D deformation model can be used for face modeling. The terminal corrects the face pose according to the face orientation to obtain the face features with normalized expressions. The terminal then generates a three-dimensional projection image based on the face features corrected by three-dimensional deformation mapping, and the obtained three-dimensional projection image can be directly used for projection onto the three-dimensional solid model.

[0054] In one embodiment, the above-mentioned face image processing method further includes: identifying occluded feature points in the first face image; calculating the yaw angle based on the feature points in the face features; and correcting the occluded feature points based on the yaw angle to obtain the face features of the first face image.

[0055] When the face image is not a frontal face image, some face feature points may be occluded, and thus these occluded face feature points need to be restored. The terminal can detect whether there are occluded feature points in the first face image. When the occluded feature points in the first face image are identified, the occluded feature points need to be corrected and restored. Assuming that the face is approximately regarded as a cylinder, when the expression and pose of the face change, the positions of the feature points will move along some parallel circular arcs on the surface of the cylinder, that is, they will move along the intersection line of the plane parallel to the bottom surface passing through the feature point and the cylinder surface. Based on this principle, the occluded feature points can be restored.

[0056] Specifically, after the terminal extracts the feature points in the first face image and identifies the occluded feature points in the first face image, it calculates the yaw angle based on the feature points in the face features. Specifically, it can be based on obtaining the feature points that will not be occluded during the face pose transformation (such as the tip of the nose, the center of the eyebrows, etc.). Taking the positions of these feature points on the frontal face as the starting points and the positions on the deflected face as the end points, the yaw angle of the face is determined based on the central angle corresponding to the connected arc. Then, the occluded feature points are corrected according to the yaw angle and the unoccluded feature points, so as to accurately and effectively obtain the face features of the first face image.

[0057] S206, obtaining the second face video of the second user, and extracting face expression features based on each second face image in the second face video.

[0058] Herein, the second user refers to another user different from the first user. A video is a sequence of images that change continuously at a rate exceeding a preset number of frames per second. The second face video refers to a video including the face region of the second user, and the second face video includes a series of consecutive frames of images including faces. Specifically, the second face video can be a video of the face corresponding to the expression response collected according to a specified instruction. For example, during the face recognition process, the user is required to make corresponding responses according to the instructions of the identity verification system, and record the response video for further processing.

[0059] The terminal extracts the face expression features in the second face video by obtaining each frame of the second face image in the second face video and performing expression recognition and expression feature extraction on each frame of the second face image.

[0060] Among them, the facial expression feature represents the expression state feature corresponding to the facial movements of a human face, such as blinking, opening the mouth, shaking the head, etc. Expression recognition refers to separating a specific expression state from a given static image or dynamic video sequence. Expression feature extraction refers to locating and extracting the organ features, texture regions, and predefined feature points of a human face.

[0061] In one embodiment, the terminal can obtain each frame of static second human face image, recognize the expression state feature of each frame, and then determine the overall facial expression feature of the second user through the expression state features of all frames, thereby being able to effectively extract the facial expression feature from the human face video.

[0062] In one embodiment, the terminal can directly analyze the human face images in a continuous dynamic video sequence, and recognize and extract the overall facial expression feature corresponding to the second user from the dynamic video frame sequence.

[0063] S208, respectively fuse the extracted facial expression features into the three-dimensional projection image for expression reconstruction to obtain a synthesized human face video.

[0064] Among them, expression reconstruction refers to obtaining new expression features by adjusting the expression features in a human face to obtain a human face image with a specific expression. By performing expression reconstruction on the basis of expression classification and quantization, human face expression images corresponding to various expressions can be obtained.

[0065] A synthesized human face refers to a human face generated by performing human face fusion processing on two or more human face images, and the generated human face simultaneously has the appearance features of two human faces. The human face fusion technology is to identify the image features, facial features, and expression features of the user-uploaded photo through a face recognition algorithm, and fuse these identified features onto a template image.

[0066] Specifically, the obtained facial expression features of each human face can be a continuous sequence of expression features. The three-dimensional projection image also includes the initial expression feature corresponding to the first user's human face, where the initial expression feature can be the first human face expression feature obtained after performing expression normalization processing on the first human face image.

[0067] After the terminal extracts the facial expression features of the second user from the second human face video, it respectively fuses the facial expression features into the three-dimensional projection image for expression reconstruction. Input the facial expression features of each human face into the three-dimensional projection image of the first user, perform expression feature conversion in the corresponding expression area, convert the feature points in the three-dimensional projection image into facial expression features, and obtain a three-dimensional human face image with expression features. The obtained three-dimensional human face image is a type of synthesized human face image.

[0068] For example, the subspace deformation transmission method can be adopted to transmit the facial expression features of the second user to the three-dimensional projection image corresponding to the first user in real time without affecting the target facial features. Assuming that the facial features of the first user and the second user are relatively fixed, after extracting the facial expression features from the facial video of the first user, the facial expression features, the three-dimensional projection image, and the initial expression features therein are used as inputs for conversion, and the target three-dimensional facial image with expression is directly output in the reduced subspace of the parameter prior. Thus, the expression features in the three-dimensional projection image can be accurately and effectively reconstructed.

[0069] Since the extracted dynamic expression features are of consecutive frames, after each obtained facial expression feature is respectively fused into the three-dimensional projection image for expression reconstruction, a three-dimensional facial image including facial expressions of consecutive frames can be obtained. The terminal then generates a synthetic facial video based on these consecutive-frame three-dimensional facial images, and thus a facial video containing expression responses can be effectively generated.

[0070] S210, project the synthetic facial video onto a three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition.

[0071] Among them, the three-dimensional solid model is a stereoscopic model corresponding to an actually existing tangible and physical specific object, and is used to display the effect of a solid face after projecting the three-dimensional facial image onto the three-dimensional solid model. For example, the three-dimensional solid model can be a 3D face mold, a 3D silicone face mask, etc.

[0072] After the terminal generates a synthetic facial video with expression responses, the synthetic facial video is projected onto a three-dimensional solid model for playback, so that a solid face model with expression responses can be displayed on the three-dimensional solid model after projection. Specifically, the terminal can project the synthetic facial video onto the three-dimensional solid model for playback through a projection device in the terminal, or can also project the synthetic facial video onto the three-dimensional solid model for playback through a projection device connected to the terminal. For example, a micro-projector with a high resolution (such as 1080*1920) can be used for projection.

[0073] Specifically, the projected three-dimensional solid model is used for face recognition. For example, when performing face recognition, it is necessary to collect images or videos of the user's face for recognition. In a face recognition system based on tag response, the user needs to make an expression response according to the instructions, and collect the corresponding response video for face recognition. When collecting the response video during the face recognition process, the video of the projected three-dimensional solid model in this embodiment can be collected through an authentication device, so as to obtain a face video with an expression response, and then perform identity authentication based on face recognition on the collected face video. By performing face recognition on the face video synthesized from the face of the first user and the expression of the second user, the effect of the face recognition system can be effectively verified, which is beneficial to further maintain and update the face recognition system to improve the security of face recognition.

[0074] Reference Figure 3 , Figure 3 shows a schematic flowchart of expression reconstruction for the first user's face image and the second user's image in one embodiment. That is, the face features are extracted from the first face image of the first user, the face expression features are extracted from each second face image of the second user, and the obtained face expression features are converted into the three-dimensional projection image corresponding to the face features of the first user for expression reconstruction, so as to obtain a synthesized face image with an expression response.

[0075] In one embodiment, before the terminal projects the synthesized face video onto the three-dimensional solid model for playback, it can first project the three-dimensional projection image corresponding to the obtained first face image onto the three-dimensional solid model. By continuously adjusting the distance between the projection device and the three-dimensional solid model and keeping them facing each other, it is ensured that the three-dimensional projection image projected onto the three-dimensional solid model is a relatively accurate frontal face image. Then, by adjusting the focal length of the projection device, the face texture projected onto the three-dimensional solid model is made as clear as possible. The terminal can directly project the synthesized face video onto the three-dimensional solid model after obtaining the synthesized face video containing the expression response during the face recognition process, so that the authentication device can collect the face image presented by the three-dimensional solid model for face recognition authentication. Reference Figure 4 , Figure 4 shows a schematic diagram of projecting the synthesized face video onto the three-dimensional solid model in one embodiment. For example, the three-dimensional solid model can use a common silicone face mold, which has a high simulation of human face skin and low cost, and can be reused by replacing the projected face picture. Therefore, it can effectively save the cost of verifying the face recognition system.

[0076] In the above-mentioned face image processing method, after obtaining the first face image of the first user, a corresponding three-dimensional projection image is generated based on the face features of the first face image, thereby being able to effectively generate a face image for projection onto a three-dimensional solid model and presenting the effect of a real face. By obtaining the second face video of the second user, after extracting the face expression features from each second face image in the second face video, the extracted face expression features are respectively fused into the three-dimensional projection image for expression reconstruction. Thereby, the expression features in the three-dimensional projection image can be accurately and effectively reconstructed, and a synthetic face video with high authenticity and including expression responses can be effectively generated. The synthetic face video is projected onto the three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition. By performing face recognition on the projection of the synthetic face video that fuses the face of the first user and the expression of the second user on the three-dimensional solid model, the effect of the face recognition system can be effectively verified, which is beneficial for further improving the security of face recognition. By generating a synthetic face video with high accuracy and authenticity and projecting it onto a reusable three-dimensional solid model for face recognition, the cost of verification can be effectively saved, and the efficiency of the face recognition system can be effectively improved.

[0077] In one embodiment, generating a three-dimensional projection image based on the face features of the first face image includes: obtaining the two-dimensional coordinate information of the feature points in the face features; determining the three-dimensional face features of the first face image based on the two-dimensional coordinate information; performing three-dimensional mapping processing on the three-dimensional face features to generate a three-dimensional projection image corresponding to the first face image.

[0078] Among them, the two-dimensional coordinate refers to a coordinate system composed of two mutually perpendicular number axes with a common origin in the same plane. The three dimensions refer to a space system formed by adding a direction vector to the plane two-dimensional system. The pose estimation in the field of computer vision is represented by the relative translation amount of the target and the rotation matrix. The pose estimation can be obtained by the transformation relationship matrix between the coordinates of n points of the target in the 3D world coordinate system and the corresponding point set projected onto the 2D image coordinate system.

[0079] The terminal can establish a mapping relationship between 2D feature points and 3D feature points of a human face based on a preset perspective projection model. After the terminal obtains the first human face image of the first user, it extracts the feature points corresponding to each facial feature from the first human face image and establishes the two-dimensional coordinate information corresponding to each feature point. The terminal can respectively match and map the feature points in the first human face image with the feature points in the preset three-dimensional human face base to obtain the three-dimensional coordinate points corresponding to the feature points in the preset three-dimensional human face base, obtain the depth information of the two-dimensional coordinate information in the three-dimensional space according to the three-dimensional coordinate points, and determine the three-dimensional coordinates of the two-dimensional coordinate information mapped in the three-dimensional space according to the depth information, so as to obtain the three-dimensional coordinate information corresponding to the human face features. Furthermore, based on the two-dimensional coordinate information and the three-dimensional coordinate information, the three-dimensional human face features of the first human face image are determined.

[0080] After the terminal obtains the three-dimensional human face features of the first human face image based on the three-dimensional mapping relationship according to the two-dimensional coordinate information of the feature points, it further performs three-dimensional mapping processing on the three-dimensional human face image. For example, it can be based on the principle of the perspective projection model, so as to obtain the three-dimensional projection image corresponding to the first human face image. The obtained three-dimensional projection image can be directly projected onto the three-dimensional solid model, so that the effect of the real human face can be accurately and effectively displayed on the three-dimensional solid model.

[0081] In one embodiment, determining the three-dimensional human face features of the first human face image according to the two-dimensional coordinate information includes: estimating the human face pose according to the feature points; obtaining the mapping coordinate information of the preset three-dimensional human face base mapped on the two-dimensional plane; determining the three-dimensional mapping parameters corresponding to the feature points according to the two-dimensional coordinate information and the mapping coordinate information; updating the feature points and the corresponding three-dimensional mapping parameters according to the human face pose and the two-dimensional coordinate information; generating three-dimensional human face features based on the updated feature points and three-dimensional mapping parameters.

[0082] Among them, the human face pose estimation is mainly to obtain the angular information of the face orientation. Generally, it can be represented by a rotation matrix, a rotation vector, a quaternion or Euler angles (these four quantities can also be converted to each other). The three-dimensional human face base can be a pre-obtained general three-dimensional human face base model, or a three-dimensional human face base model trained based on a large number of human face model samples.

[0083] After the terminal extracts the feature points corresponding to each facial feature from the first human face image, it estimates the human face pose according to the feature points of the human face. Furthermore, based on the preset three-dimensional human face base model, a two-dimensional and three-dimensional mapping model is established, and the human face pose is corrected to obtain the corresponding 3D projection image.

[0084] Specifically, the terminal can obtain the mapping coordinate information of the preset three-dimensional face basis mapped on the two-dimensional plane, and determine the three-dimensional coordinate information corresponding to the feature points and the three-dimensional mapping parameters according to the two-dimensional coordinate information and the mapping coordinate information. The terminal can then estimate the face pose according to the three-dimensional coordinate information corresponding to the feature points and the three-dimensional mapping parameters. The terminal can also establish a two-dimensional and three-dimensional mapping model based on the face feature points, the two-dimensional coordinate information, and the three-dimensional mapping parameters. The terminal then updates the feature points and the corresponding three-dimensional mapping parameters according to the face pose and the two-dimensional coordinate information, and thus can generate the three-dimensional face feature corresponding to the first user according to the updated three-dimensional mapping parameters and the face features. The obtained three-dimensional face feature includes the three-dimensional coordinate information corresponding to the two-dimensional coordinate points of each feature point.

[0085] For example, on the basis of a preset three-dimensional deformation model, a new 2D-3D mapping model can be established, and the 3D deformation model mapping is used to correct the face pose. The formula of the specific mapping model can be as follows:

[0086]

[0087] where F 2Dloca represents the position of the feature points on the 2D plane, represents the average shape of the face, k is the scale factor, M1 is the orthographic projection matrix, R is the 3×3 rotation matrix, T 3D represents the feature point transformation vector, P id represents the facial shape feature of this person, λ id represents the shape weight, P exp represents the facial expression feature of this person, λ exp represents the expression weight.

[0088] Specifically, the terminal can first initialize λ id and λ exp to zero, and use the weak perspective projection model to roughly estimate the face pose according to the extracted facial feature points, then update the facial feature points according to the face orientation, and then calculate each three-dimensional mapping parameter through the above formula, so as to effectively obtain the three-dimensional face feature corresponding to the first user.

[0089] In one embodiment, performing three-dimensional mapping processing based on the three-dimensional face feature to generate a three-dimensional projection image corresponding to the first face image includes: constructing a three-dimensional face mapping matrix according to the three-dimensional face feature; positioning the face contour of the face feature; and adjusting the boundary of the three-dimensional face mapping matrix based on the face contour feature in the face feature to obtain a three-dimensional projection image corresponding to the first face image.

[0090] Among them, a matrix is a set of complex or real numbers arranged in a rectangular array. The three-dimensional face mapping matrix can be a linear mapping matrix, which is a quantitative representation of a linear mapping and is used to represent the mapping relationship of facial feature points between a two-dimensional image and a three-dimensional solid model. The face contour can include the local contours of the facial features and the outer contour of the face, and the face contour is an important feature among the facial features of a face. Facial features can include local feature point information and overall feature information.

[0091] After the face image is normalized for face pose and expression, there is usually a blank area between the background and the face. The terminal can adjust the boundary position of the three-dimensional projection image according to the face contour.

[0092] Specifically, after the terminal obtains the three-dimensional face features corresponding to the first user based on the facial feature points, two-dimensional coordinate information, and three-dimensional mapping parameters, it constructs a three-dimensional face mapping matrix according to the three-dimensional face features and locates the face contour of the facial features. The terminal then adjusts the boundary of the three-dimensional face mapping matrix according to the face contour features in the facial features, thereby obtaining a three-dimensional projection image corresponding to the face of the first user with relatively high accuracy.

[0093] Specifically, the formula for adjusting the boundary of the three-dimensional face mapping matrix can be as follows:

[0094]

[0095] Among them, (x 2_new , y 2_new ) represents the new position of the connection anchor point 2 to be solved, (x 1_con , y 1_con ) represents the predefined facial contour position of the boundary anchor point 1, and (x1, y1) and (x2, y2) represent the coordinates before adjustment.

[0096] In one embodiment, the terminal can automatically adjust the boundary of the three-dimensional face mapping matrix according to a preset program, thereby effectively obtaining a three-dimensional projection image corresponding to the face of the first user.

[0097] In another embodiment, manual anchor point adjustment can also be performed on the terminal. By adjusting the anchor points around the face to make these anchor points coincide with the face contour as much as possible, that is, moving the boundary anchor points to the predefined positions on the 3D face model and trying to keep the spatial distance unchanged. To achieve the adjustment of the boundary of the three-dimensional face mapping matrix, thereby effectively obtaining a three-dimensional projection image corresponding to the face of the first user.

[0098] In one embodiment, the above face image processing method further includes: extracting the illumination parameters and facial parameters corresponding to the first face image; performing face detail filling on the three-dimensional face mapping matrix based on the illumination parameters and facial parameters.

[0099] Among them, illumination is a factor affecting human face imaging and the composition of human face images, and illumination change is a key factor affecting the performance of face recognition. The illumination parameters of a human face represent some parameters of the human face under light reflection, such as including parameters such as light intensity, glossiness, specular highlight, ambient light, directional light, specular reflection parameters, and ambient reflection parameters. Facial parameters refer to some parameters of the human face, such as including illumination reflection parameters of the face under light, as well as facial texture parameters, facial surface color parameters, etc. For example, illumination change can be modeled, based on representing the changes caused by illumination in a suitable subspace, estimating model parameters according to the face features in the first human face image to estimate the illumination parameters of the human face. For example, algorithms such as subspace projection method, quotient function method, illumination cone method, and based on spherical harmonic basis images can be used to achieve this.

[0100] When the side angle of the face in the original human face image is too large, it is necessary to further fill the invisible area. And simulate the reflection situation of the human face under light to obtain the illumination parameters of the human face, and then fill in the face details according to the symmetry of the human face and illumination parameters, etc.

[0101] Specifically, after the terminal obtains the three-dimensional face features corresponding to the first user based on face feature points, two-dimensional coordinate information, and three-dimensional mapping parameters, it extracts the illumination parameters and facial parameters corresponding to the first human face image, and fills in the face details of the three-dimensional face mapping matrix based on the illumination reflection parameters and facial parameters.

[0102] For example, the linear combination based on spherical harmonic reflection basis can be used to simulate the reflection situation of the human face under light to obtain the illumination parameters in the human face image. Among them, the formula for extracting the illumination parameters corresponding to the first human face image can be as follows:

[0103]

[0104] Among them, β is a 9-dimensional illumination parameter, P corresponds to the 3D pixel point, and D represents the spherical harmonic reflection basis.

[0105] The terminal can then supplement the face details by mirroring according to the symmetry of the human face. Specifically. The terminal can adopt an image editing method based on the Poisson equation to seamlessly insert the source object into the image. To obtain the Poisson partial differential equation with boundary conditions, the specific expression can be as follows:

[0106]

[0107] Among them, pic represents the processed image to be solved, △ represents the Laplacian operator, The Laplacian value of the texture to be inserted is denoted as, Ω represents the editing region, αΩ represents the boundary of the editing region, and pic0 represents the input image to be processed.

[0108] In one embodiment, extracting facial expression features based on each second facial image in the second facial video includes: extracting the second facial images of the key frames in the second facial video; extracting the facial parameters in each frame of the second facial images; estimating the normal expression distribution according to the facial parameters in each frame of the second facial images, and obtaining the facial expression features of the second user based on the normal expression distribution.

[0109] Among them, the second facial video can be a facial video that makes an expression response according to a specified instruction, and the facial video includes multiple consecutive frames of images including a face. The facial parameters include parameters such as facial feature points, facial poses, and local expression features. The local expression features are the states of various parts of the face, such as opening eyes, closing eyes, opening the mouth, etc. Among them, the key frames include at least two or more video frames.

[0110] When the terminal extracts facial expression features based on each second facial image in the second facial video, it can extract only the key frame facial images corresponding to the expression response. Specifically, the terminal can identify the expression change actions in the facial video, determine the key frames corresponding to the expression response according to the expression change actions, and then extract the second facial images of the key frames in the second facial video.

[0111] In one of the embodiments, after the terminal identifies the expression change actions in the facial video and determines the video frames corresponding to the expression response, the terminal can further extract the key frames related to the expression change actions from these frames. For example, it can extract a part of the more important frames from these video frames according to the expression change amplitude, and determine these extracted video frames as the key frames. Thus, it can effectively obtain only the important frames related to the expression change actions from the facial video to reduce the computing resources for the terminal to process the facial images.

[0112] The terminal extracts the second facial images corresponding to the key frames from the facial video and generates a key frame set. Then the terminal extracts features from each frame of the second facial images in the key frame set, extracts the facial feature points and expression state features in each frame of the second facial images, and obtains the facial parameters according to the facial feature points and expression state features. Then the terminal estimates the normal expression distribution according to the facial parameters in each frame of the second facial images, and the facial expression features of the second user can be obtained based on the normal expression distribution.

[0113] For example, the terminal can estimate all parameters on k key frames of the input video sequence simultaneously. The parameters to be estimated mainly include face parameters, global flags, inline functions, unknown poses of each frame, and lighting parameters, etc., to capture the user's expression. For example, an iterative reweighted least squares (IRLS) solver can be used to solve the normal equations corresponding to the face parameters of the entire key frame set simultaneously to estimate the normal expression distribution, and then obtain dynamic face expression features according to the normal expression distribution.

[0114] In one embodiment, the extracted face expression features are respectively fused into the three-dimensional projection image for expression reconstruction, and the obtained synthetic face video includes: performing model reconstruction on the three-dimensional projection image based on each second face image to obtain a three-dimensional face image; performing expression conversion on the face features in the three-dimensional face image based on each face expression feature to obtain a three-dimensional face image including each face expression feature; generating a synthetic face video according to the three-dimensional face image including each face expression feature.

[0115] Among them, the three-dimensional face image refers to a projection image used to project onto a three-dimensional solid model to display a three-dimensional stereoscopic face.

[0116] After the terminal extracts the three-dimensional projection image corresponding to the first user's face and extracts each face expression feature corresponding to the second user's face from the face video, without affecting the face features of the first user, the expression of the second user's face is transmitted to the first user's face. The face expression features of the second user's face and the expression of the first user's face are converted to obtain a face image with an expression.

[0117] When performing face modeling on a face, the face model can include the position coordinates of face key feature points (such as the brow bone, eye contour, nose, lips, etc.), the reflectivity of key feature points, and parts such as face expressions. The position coordinates of face key points and the reflectivity of key feature points are mainly used to characterize the shape of the face. When performing expression reconstruction, it is necessary to first establish a neutral face model for the first user's face and the second user's face, that is, a face model without expression. Therefore, taking the face model of the second user and the neutral face model as inputs, a face model with an expression can be output.

[0118] Specifically, after the terminal obtains the second face image of the second user, based on the second face image and the three-dimensional projection image corresponding to the first face image, model reconstruction is performed to obtain a target face model, that is, a neutral face model between the first user's face and the second user's face. Thus, a reconstructed three-dimensional face image is obtained, and the three-dimensional face image is the image obtained by projecting the three-dimensional target face model onto a two-dimensional plane.

[0119] The terminal further performs expression conversion on the facial features in the three-dimensional facial image based on each facial expression feature. Specifically, the terminal inputs each facial expression feature of the second user into the target facial model corresponding to the three-dimensional facial image, and replaces the original expression feature according to the dynamic facial expression features, so as to convert the original expression of the target facial model into the facial expression of the second user. Thus, a three-dimensional facial image including each facial expression feature can be effectively obtained.

[0120] For example, the terminal can adopt subspace deformation transmission technology for expression reconstruction, and re-model the second user's face through the extracted facial expression features of the second user. Keep the position coordinates and reflectivity of the key feature points of the first user's face unchanged, and modify the original expression part of the target facial model. Thus, expression conversion can be efficiently achieved without affecting the facial shape features.

[0121] After performing expression conversion on the facial features in the three-dimensional facial image based on the dynamic facial expression features, a continuous multi-frame three-dimensional facial image including dynamic facial expression features can be effectively obtained. The terminal then synthesizes the continuous-frame three-dimensional facial images into a corresponding video, thereby generating a corresponding synthesized facial video.

[0122] In this embodiment, by performing expression reconstruction and conversion on the three-dimensional facial image corresponding to the first user's face based on each facial expression feature of the second user's face, a three-dimensional facial image corresponding to the first user's face with an expression can be accurately and effectively obtained, and then a synthesized facial video with an expression response can be effectively obtained.

[0123] In one embodiment, as Figure 5 shown, a facial image processing method is provided, including the following steps:

[0124] S502, obtain a first facial image of a first user.

[0125] S504, generate a corresponding three-dimensional projection image based on the facial features of the first facial image.

[0126] S506, obtain a second facial video of a second user, and extract facial expression features based on each second facial image in the second facial video.

[0127] S508, perform model reconstruction on the three-dimensional projection image based on each second facial image to obtain a three-dimensional facial image.

[0128] S510, identify the expression area corresponding to the facial expression feature; perform expression transformation on the feature points in the expression area of the three-dimensional projection image according to the facial expression feature to obtain a three-dimensional facial image including the facial expression feature.

[0129] S512. Generate a synthetic face video based on a three-dimensional face image including various facial expression features.

[0130] S514. Project the synthetic face video onto a three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition.

[0131] Among them, the expression area refers to the area corresponding to the key feature points that change with the expression movement in the face with an expression. The expression area can be a local feature area or an overall face feature area. For example, when the expression is "open mouth" or "smile", the expression area is only the lip area; when "laugh out loud" or "shake the head", the expression area can be the entire face area. Among them, the expression transformation can adopt the method of affine transformation. Affine transformation, also known as affine mapping, refers to a linear transformation in a vector space in geometry followed by a translation to transform into another vector space.

[0132] After the terminal obtains the facial expression features corresponding to the second face image from the second face video, it identifies the expression area corresponding to the expression features, and then performs expression conversion on the expression area part. Specifically, the terminal performs model reconstruction based on the second face image and the three-dimensional projection image corresponding to the first face image. After obtaining the target face model, it inputs the various facial expression features of the second user into the target face model corresponding to the three-dimensional face image, and performs expression transformation on the feature points corresponding to the expression area in the target face model according to the facial expression features. Specifically, it can perform affine transformation on the facial expression features and the feature points corresponding to the expression area, so as to achieve expression conversion. Perform expression transformation on the original expression features according to the dynamic various facial expression features, so as to convert the original expression of the target face model into the facial expression of the second user.

[0133] For example, taking the expression of the second user's face as "open mouth" as an example, after the terminal obtains the facial expression features corresponding to the second face image from the second face video, it can identify that the expression area is "open mouth". Among them, the facial expression features corresponding to the "open mouth" expression area can include the lip part and the oral cavity part. When performing transformation on the facial expression features and the feature points corresponding to the expression area in the target face model, the lip features of the first user can be retained, and only the oral cavity part is modified and transformed to achieve the synthesis of the oral cavity with the mouth retained. Thus, the expression of the second user can be effectively incorporated into the face of the first user.

[0134] In one embodiment, as Figure 6 shown, a face image processing method is provided, including the following steps:

[0135] S602. Obtain the first face image of the first user.

[0136] S604. Generate a corresponding three-dimensional projection image based on the facial features of the first facial image.

[0137] S606. Obtain the second facial video of the second user, and extract facial expression features based on each second facial image in the second facial video.

[0138] S608. Perform model reconstruction on the three-dimensional projection image based on each second facial image to obtain a three-dimensional facial image.

[0139] S610. Identify the expression area corresponding to the facial expression feature; obtain the facial expression image of the first user corresponding to the expression area and the facial expression feature.

[0140] S612. Stitch the target expression feature of the expression area in the facial expression image to the three-dimensional facial image to obtain a three-dimensional facial image including the facial expression feature.

[0141] S614. Generate a synthetic facial video based on the three-dimensional facial image including each facial expression feature.

[0142] S616. Project the synthetic facial video onto a three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for facial recognition.

[0143] Among them, the facial expression image of the first user can be obtained from the expression samples corresponding to the first user. The expression samples can be offline samples of the target user obtained in advance, including a sequence of facial expression pictures of the target user. Among them, the sequence of facial expression pictures can include a complete face or a partial expression area.

[0144] After the terminal performs model reconstruction based on the second facial image and the three-dimensional projection image corresponding to the first facial image to obtain a target facial model, it inputs each facial expression feature of the second user into the target facial model corresponding to the three-dimensional facial image, and performs an expression transformation on the expression area in the target facial model according to the facial expression feature. Specifically, the terminal obtains an expression picture sequence matching the facial expression feature from the expression samples corresponding to the first user according to the facial expression feature, stitches the corresponding expression area in the obtained expression picture sequence to the three-dimensional facial image corresponding to the first user, and fuses the target expression feature in the expression picture sequence into the facial features in the three-dimensional facial image, so as to obtain a three-dimensional facial image of the first user including the facial expression feature. Thus, a sequence of three-dimensional facial images that is consistent with the expression of the second user and has a high degree of fidelity can be effectively obtained.

[0145] For example, taking the expression of the second user's face as "open mouth" as an example, after the terminal obtains the facial expression features corresponding to the second face image from the second face video, it can identify that the expression area is "open mouth". Among them, the facial expression features corresponding to the "open mouth" expression area may include features such as the opening degree of the upper and lower lips, the distance and deflection degree of the left and right corners of the mouth, etc. The terminal then obtains an expression picture sequence that matches the facial expression features from the expression samples corresponding to the first user according to the facial expression features. Specifically, the terminal can calculate the distance between the sample frame and the description matrix of the "lip" expression frame, and the description matrix includes various expression feature parameters. Then the terminal can adopt the K-means algorithm. For example, taking 10 frames as a cluster, extract the frame with the smallest distance in each cluster, that is, the frame closest to the target mouth shape of the "lip" expression as the representative of the cluster, and finally select the frame with the smallest distance among all cluster representatives as the final matching result. By matching the mouth shape with the highest matching degree to the target face from the offline samples according to the feature similarity measure, a realistic oral cavity image can be effectively generated. Thus, an expression picture sequence of the first user with a relatively high matching degree to the second user's expression can be extracted. Selecting frames that can match the specified expression action from various mouth shape pictures of the target person and splicing the extracted expression frames into the three-dimensional face image of the first user, a three-dimensional face image sequence that is consistent with the second user's expression and has a relatively high degree of realism can be effectively obtained.

[0146] In one embodiment, the above-mentioned face image processing further includes: obtaining a first face video of the first user; extracting face features based on each first face image in the first face video; generating a corresponding three-dimensional projection image according to the face features of each first face image.

[0147] Among them, the first user face video refers to a video including the face of the first user, including a sequence of first user face images with consecutive frames.

[0148] After the terminal obtains the first face video of the first user, it extracts the first face images with consecutive frames from the first user's face video. The terminal respectively extracts the face features in each first face image, and then generates a corresponding three-dimensional projection image according to the face features of each first face image. Specifically, the terminal estimates the face pose based on the dynamic feature points corresponding to each first face image, and performs face modeling through three-dimensional model mapping based on the two-dimensional feature points to obtain the three-dimensional coordinate information corresponding to the face features. The terminal corrects the face pose according to the face orientation to obtain the face features after expression normalization, and then generates a three-dimensional projection image based on the face features after three-dimensional deformation mapping and correction. By extracting the dynamic face features corresponding to the first user's face from the face video, a three-dimensional face model with relatively high accuracy can be effectively constructed, and then a corresponding three-dimensional projection image can be effectively generated.

[0149] In one embodiment, asFigure 7 As shown, a face image processing method is provided, which specifically includes the following steps:

[0150] S702, obtain the first face video of the first user.

[0151] S704, extract face features based on each first face image in the first face video.

[0152] S706, generate corresponding three-dimensional projection images according to the face features of each first face image.

[0153] S708, obtain the second face video of the second user, and extract face expression features based on each second face image in the second face video.

[0154] S710, respectively fuse the extracted face expression features into each three-dimensional projection image for expression reconstruction to obtain a synthesized face video.

[0155] S712, project the synthesized face video onto a three-dimensional solid model for playback, and the projected three-dimensional solid model is used for face recognition.

[0156] After the terminal extracts face features based on each first face image in the first face video, corresponding three-dimensional projection images are respectively generated according to the face features of each first face image, thereby obtaining a sequence of three-dimensional projection images of consecutive frames.

[0157] The terminal obtains the second face video of the second user. After extracting face expression features based on each second face image in the second face video, the extracted face expression features are further respectively fused into each three-dimensional projection image for expression reconstruction. Specifically, the terminal can respectively fuse the face expression features corresponding to each frame into the three-dimensional projection image corresponding to each frame of the face image for expression reconstruction according to the image frame sequence, and respectively obtain the three-dimensional face images corresponding to each frame. The terminal then synthesizes a face video based on the three-dimensional face images of each frame with dynamic expression responses. By extracting the face expression features corresponding to each frame in the face video of the second user and respectively synthesizing them into consecutive frames of three-dimensional projection images, a synthesized face video with a relatively high authenticity and dynamic expression response can be effectively generated.

[0158] In a specific embodiment, as Figure 8 shown, a face image processing method is provided, including the following steps:

[0159] S802, obtain the first face image of the first user.

[0160] S804, obtain the two-dimensional coordinate information corresponding to the feature points in the face features, and estimate the face pose according to the feature points.

[0161] S806. Obtain the mapping coordinate information of the preset three-dimensional human face base mapped on the two-dimensional plane; determine the three-dimensional mapping parameters corresponding to the feature points according to the two-dimensional coordinate information and the mapping coordinate information.

[0162] S808. Update the feature points and the corresponding three-dimensional mapping parameters according to the human face pose and the two-dimensional coordinate information; generate three-dimensional human face features based on the updated feature points and three-dimensional mapping parameters.

[0163] S810. Perform three-dimensional mapping processing on the three-dimensional human face features to generate a three-dimensional projection image corresponding to the first human face image.

[0164] S812. Construct a three-dimensional human face mapping matrix according to the three-dimensional human face features; locate the human face contour of the human face features.

[0165] S814. Adjust the boundary of the three-dimensional human face mapping matrix based on the human face contour features in the human face features to obtain a three-dimensional projection image corresponding to the first human face image.

[0166] S816. Extract the second human face image of the key frame in the second human face video; extract the human face parameters in each frame of the second human face image.

[0167] S818. Estimate the normal expression distribution according to the human face parameters in each frame of the second human face image, and obtain the human face expression features of the second user based on the normal expression distribution.

[0168] S820. Perform model reconstruction on the three-dimensional projection image based on each second human face image to obtain a three-dimensional human face image.

[0169] S822. Perform expression conversion on the human face features in the three-dimensional human face image based on each human face expression feature to obtain a three-dimensional human face image including each human face expression feature.

[0170] S824. Generate a synthetic human face video according to the three-dimensional human face image including each human face expression feature.

[0171] S826. Project the synthetic human face video onto a three-dimensional solid model for playing, and the three-dimensional solid model after projection is used for human face recognition.

[0172] In this embodiment, by performing human face recognition on the projection of the synthetic human face video fused with the human face of the first user and the expression of the second user on the three-dimensional solid model, the effect of the human face recognition system can be effectively verified, which is beneficial to further improving the security of human face recognition. By generating a synthetic human face video with high accuracy and authenticity and projecting it onto a reusable three-dimensional solid model for human face recognition, the cost of verification can be effectively saved, and the efficiency of the human face recognition system can be effectively improved.

[0173] The present application also provides an application scenario, which applies the above-mentioned face image processing method to verify the effect of face recognition. Specifically, when the identity verification device needs to perform face recognition on the first user for identity authentication, the terminal acquires the first face image of the first user, and generates a corresponding three-dimensional projection image based on the face features of the first face image. The second user makes an expression response according to the instruction issued by the identity authentication, and the terminal then acquires the second face video of the second user responding to the identity authentication, extracts the face expression features based on each second face image in the second face video, and respectively fuses the extracted face expression features into the three-dimensional projection image for expression reconstruction, generating a synthetic face video including the expression response. The terminal projects the synthetic face video onto a three-dimensional solid model for playback, and the identity verification device acquires the face video corresponding to the projected three-dimensional solid model to perform face recognition, obtaining a recognition result. Thus, the effect of face recognition can be verified according to the recognition result.

[0174] The present application also additionally provides an application scenario, which applies the above-mentioned face image processing method to test a face recognition system. Specifically, when the face recognition system needs to perform face recognition on the first user for identity authentication, the terminal acquires the first face image of the first user, and generates a corresponding three-dimensional projection image based on the face features of the first face image. The second user makes an expression response according to the instruction issued by the identity authentication, and the terminal then acquires the second face video of the second user responding to the identity authentication, extracts the face expression features based on each second face image in the second face video, and respectively fuses the extracted face expression features into the three-dimensional projection image for expression reconstruction, generating a synthetic face video including the expression response. The terminal projects the synthetic face video onto a three-dimensional solid model for playback, and the face video corresponding to the projected three-dimensional solid model is collected through the identity verification device, and face recognition is performed using the collected face video to conduct an attack test on the face recognition system, obtaining a face recognition result. Thus, test result data can be obtained by using the face recognition result and the data during the recognition process. By conducting an attack test on the face recognition system, the security of the face recognition system can be further effectively improved.

[0175] It should be understood that although Figure 2 、 5 each step in the flowchart of -8 is shown in sequence according to the arrow indication, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 、 5At least a part of the steps in -8 may include multiple steps or multiple stages. These steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0176] In one embodiment, as Figure 9 shown, a face image processing apparatus 900 is provided. This apparatus can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the apparatus includes: an image acquisition module 902, an image conversion module 904, an expression extraction module 906, an expression reconstruction module 908, and a video projection module 910, where:

[0177] The image acquisition module 902 is configured to acquire a first face image of a first user;

[0178] The image conversion module 904 is configured to generate a corresponding three - dimensional projection image based on the face features of the first face image;

[0179] The expression extraction module 906 is configured to acquire a second face video of a second user and extract face expression features based on each second face image in the second face video;

[0180] The expression reconstruction module 908 is configured to respectively fuse the extracted face expression features into the three - dimensional projection image for expression reconstruction to obtain a synthesized face video;

[0181] The video projection module 910 is configured to project the synthesized face video onto a three - dimensional solid model for playing, and the three - dimensional solid model after projection is used for face recognition.

[0182] In one embodiment, the image conversion module 904 is further configured to identify occluded feature points in the first face image; calculate the lateral angle based on the feature points in the face features; and correct the occluded feature points based on the lateral angle to obtain the face features of the first face image.

[0183] In one embodiment, the image conversion module 904 is further configured to acquire two - dimensional coordinate information corresponding to the feature points in the face features; determine the three - dimensional face features of the first face image based on the two - dimensional coordinate information; and perform three - dimensional mapping processing on the three - dimensional face features to generate a three - dimensional projection image corresponding to the first face image.

[0184] In one embodiment, the image conversion module 904 is further configured to estimate the face pose according to the feature points; obtain the mapping coordinate information of the preset three-dimensional face basis mapped on the two-dimensional plane; determine the three-dimensional mapping parameters corresponding to the feature points according to the two-dimensional coordinate information and the mapping coordinate information; update the feature points and the corresponding three-dimensional mapping parameters according to the face pose and the two-dimensional coordinate information; and generate three-dimensional face features based on the updated feature points and three-dimensional mapping parameters.

[0185] In one embodiment, the image conversion module 904 is further configured to construct a three-dimensional face mapping matrix according to the three-dimensional face features; locate the face contour of the face features; and adjust the boundary of the three-dimensional face mapping matrix based on the face contour features in the face features to obtain a three-dimensional projection image corresponding to the first face image.

[0186] In one embodiment, the image conversion module 904 is further configured to extract the illumination parameters and facial parameters corresponding to the first face image; and perform face detail filling on the three-dimensional face mapping matrix based on the illumination parameters and the facial parameters.

[0187] In one embodiment, the expression extraction module 906 is further configured to extract the second face image of the key frame in the second face video; extract the face parameters in each frame of the second face image; estimate the normal expression distribution according to the face parameters in each frame of the second face image, and obtain the face expression features of the second user based on the normal expression distribution.

[0188] In one embodiment, the expression reconstruction module 908 is further configured to perform model reconstruction on the three-dimensional projection image based on each second face image to obtain a three-dimensional face image; perform expression conversion on the face features in the three-dimensional face image based on each face expression feature to obtain a three-dimensional face image including each face expression feature; and generate a synthesized face video according to the three-dimensional face image including each face expression feature.

[0189] In one embodiment, the expression reconstruction module 908 is further configured to identify the expression area corresponding to the face expression feature; and perform expression transformation on the feature points in the expression area of the three-dimensional projection image according to the face expression feature to obtain a three-dimensional face image including the face expression feature.

[0190] In one embodiment, the expression reconstruction module 908 is further configured to identify the expression area corresponding to the face expression feature; obtain the face expression image corresponding to the first user, the expression area, and the face expression feature; and splice the target expression feature in the expression area of the face expression image to the three-dimensional face image to obtain a three-dimensional face image including the face expression feature.

[0191] In one embodiment, the image acquisition module 902 is further configured to acquire a first face video of a first user; the image conversion module 904 is further configured to extract face features based on each first face image in the first face video; and generate corresponding three-dimensional projection images according to the face features of each first face image.

[0192] In one embodiment, the expression reconstruction module 908 is further configured to respectively fuse the extracted face expression features into each three-dimensional projection image for expression reconstruction, so as to obtain a synthesized face video.

[0193] For the specific limitations of the face image processing device, reference can be made to the limitations on the face image processing method in the foregoing text, which will not be elaborated here. Each module in the foregoing face image processing device can be implemented in whole or in part by software, hardware, and their combination. The foregoing modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the foregoing modules.

[0194] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program stored in the non-volatile storage medium to run. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a face image processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0195] Those skilled in the art can understand that Figure 10 the structure shown in

[0196] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented.

[0197] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0198] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0199] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0200] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A face image processing method, characterized in that, The method includes: Obtaining a first face image of a first user; Generating a corresponding three-dimensional projection image based on the face features of the first face image; Obtaining a second face video of a second user, and extracting face expression features based on each second face image in the second face video; the second face video is a response video recorded when the second user makes an expression response according to an instruction; Performing model reconstruction based on each second face image and the three-dimensional projection image to obtain a three-dimensional neutral face model; Performing expression conversion on the face features of the neutral face model according to the face expression features to obtain a synthesized face video; Projecting the synthesized face video onto a three-dimensional solid model for playback, and the three-dimensional solid model after projection is used for face recognition.

2. The method according to claim 1, wherein The method further includes: Identifying occluded feature points in the first face image; Calculating a lateral angle according to the feature points in the face features; Correcting the occluded feature points based on the lateral angle to obtain the face features of the first face image.

3. The method according to claim 1, characterized in that The generating a corresponding three-dimensional projection image based on the face features of the first face image includes: Obtaining two-dimensional coordinate information corresponding to the feature points in the face features; Determining three-dimensional face features of the first face image based on the two-dimensional coordinate information; Performing three-dimensional mapping processing based on the three-dimensional face features to generate a three-dimensional projection image corresponding to the first face image.

4. The method according to claim 3, characterized in that, The determining three-dimensional face features of the first face image based on the two-dimensional coordinate information includes: Estimating a face pose according to the feature points; Obtaining mapping coordinate information of a preset three-dimensional face basis mapped on a two-dimensional plane; Determining three-dimensional mapping parameters corresponding to the feature points according to the two-dimensional coordinate information and the mapping coordinate information; Updating the feature points and corresponding three-dimensional mapping parameters according to the face pose and the two-dimensional coordinate information; Generating the three-dimensional face features based on the updated feature points and three-dimensional mapping parameters.

5. The method according to claim 3, wherein The performing three-dimensional mapping processing based on the three-dimensional face features to generate a three-dimensional projection image corresponding to the first face image includes: Constructing a three-dimensional face mapping matrix according to the three-dimensional face features; Locating the face contour of the face features; Adjusting the boundary of the three-dimensional face mapping matrix based on the face contour features in the face features to obtain a three-dimensional projection image corresponding to the first face image.

6. The method according to claim 5, wherein The method further includes: Extracting illumination parameters and facial parameters corresponding to the first face image; Performing face detail filling on the three-dimensional face mapping matrix based on the illumination parameters and the facial parameters.

7. The method according to claim 1, wherein The extracting face expression features based on each second face image in the second face video includes: Extracting second face images of key frames in the second face video; Extracting face parameters in each frame of the second face image; Estimating a normal expression distribution according to the face parameters in each frame of the second face image, and obtaining the face expression features of the second user based on the normal expression distribution.

8. The method according to claim 1, characterized in that, The performing expression conversion on the face features of the neutral face model according to the face expression features to obtain a synthesized face video includes: Obtain the three-dimensional face image obtained by projecting the neutral face model onto a two-dimensional plane; Based on each of the face expression features, perform expression conversion on the face features in the three-dimensional face image to obtain a three-dimensional face image including each of the face expression features; Generate the synthesized face video according to the three-dimensional face image including each of the face expression features.

9. The method according to claim 8, characterized in that, The performing expression conversion on the face features in the three-dimensional face image based on each of the face expression features to obtain a three-dimensional face image including each of the face expression features includes: Identify the expression area corresponding to the face expression feature; Perform expression transformation on the feature points in the expression area of the three-dimensional projection image according to the face expression feature to obtain a three-dimensional face image including the face expression feature.

10. The method according to claim 8, wherein The performing expression conversion on the face features in the three-dimensional face image based on each of the face expression features to obtain a three-dimensional face image including each of the face expression features includes: Identify the expression area corresponding to the face expression feature; Obtain the face expression image corresponding to the first user, the expression area, and the face expression feature; Stitch the target expression feature in the expression area of the face expression image to the three-dimensional face image to obtain a three-dimensional face image including the face expression feature.

11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: Obtain the first face video of the first user; Extract face features based on each first face image in the first face video; Generate corresponding three-dimensional projection images according to the face features of each first face image.

12. The method according to claim 11, wherein The method further includes: Fuse each of the extracted face expression features into each of the three-dimensional projection images for expression reconstruction to obtain a synthesized face video.

13. A face image processing device, characterized in that, The device includes: An image acquisition module, configured to acquire a first face image of a first user; An image conversion module, configured to generate a corresponding three-dimensional projection image based on the face features of the first face image; An expression extraction module, configured to acquire a second face video of a second user, and extract face expression features based on each second face image in the second face video; the second face video is a response video recorded when the second user makes an expression response according to an instruction; A model reconstruction module, configured to perform model reconstruction based on each second face image and the three-dimensional projection image to obtain a three-dimensional neutral face model; An expression conversion module, configured to perform expression conversion on the face features of the neutral face model according to the face expression feature to obtain a synthesized face video; A video projection module, configured to project the synthesized face video onto a three-dimensional solid model for playing, and the three-dimensional solid model after projection is used for face recognition.

14. The face image processing device according to claim 13, wherein The image conversion module is further configured to identify occluded feature points in the first face image; calculate a lateral angle according to the feature points in the face features; correct the occluded feature points based on the lateral angle to obtain the face features of the first face image.

15. The face image processing device according to claim 13, characterized in that, The image conversion module is further configured to obtain two-dimensional coordinate information corresponding to feature points in the facial features; determine three-dimensional facial features of the first facial image based on the two-dimensional coordinate information; perform three-dimensional mapping processing based on the three-dimensional facial features to generate a three-dimensional projection image corresponding to the first facial image.

16. The face image processing device according to claim 15, wherein The image conversion module is further configured to estimate a facial pose according to the feature points; obtain mapping coordinate information of a preset three-dimensional human face basis mapped on a two-dimensional plane; determine three-dimensional mapping parameters corresponding to the feature points according to the two-dimensional coordinate information and the mapping coordinate information; update the feature points and the corresponding three-dimensional mapping parameters according to the facial pose and the two-dimensional coordinate information; generate the three-dimensional facial features based on the updated feature points and three-dimensional mapping parameters.

17. The face image processing device according to claim 15, characterized in that, The image conversion module is further configured to construct a three-dimensional human face mapping matrix according to the three-dimensional facial features; locate the facial contour of the facial features; adjust the boundary of the three-dimensional human face mapping matrix based on the facial contour features in the facial features to obtain a three-dimensional projection image corresponding to the first facial image.

18. The face image processing apparatus according to claim 17, wherein The image conversion module is further configured to extract illumination parameters and facial parameters corresponding to the first facial image; perform facial detail filling on the three-dimensional human face mapping matrix based on the illumination parameters and the facial parameters.

19. The face image processing device according to claim 13, wherein The expression extraction module is further configured to extract a second facial image of a key frame in the second facial video; extract facial parameters in each frame of the second facial image; estimate a normal expression distribution according to the facial parameters in each frame of the second facial image, and obtain a facial expression feature of the second user based on the normal expression distribution.

20. The face image processing device according to claim 13, wherein The expression conversion module is further configured to obtain a three-dimensional facial image obtained by projecting a neutral human face model on a two-dimensional plane; perform expression conversion on the facial features in the three-dimensional facial image based on each of the facial expression features to obtain a three-dimensional facial image including each of the facial expression features; generate the synthesized facial video according to the three-dimensional facial image including each of the facial expression features.

21. The face image processing device according to claim 20, wherein, The expression conversion module is further configured to identify an expression area corresponding to the facial expression feature; perform expression transformation on the feature points in the expression area of the three-dimensional projection image according to the facial expression feature to obtain a three-dimensional facial image including the facial expression feature.

22. The face image processing device according to claim 20, wherein The expression conversion module is further configured to identify an expression area corresponding to the facial expression feature; obtain a facial expression image of the first user corresponding to the expression area and the facial expression feature; splice the target expression feature of the expression area in the facial expression image to the three-dimensional facial image to obtain a three-dimensional facial image including the facial expression feature.

23. The face image processing apparatus according to any one of claims 13 to 22, characterized in that The image acquisition module is further configured to acquire a first facial video of a first user; the image conversion module is further configured to extract facial features based on each first facial image in the first facial video; generate a corresponding three-dimensional projection image according to the facial features of each first facial image.

24. The face image processing device according to claim 23, wherein The expression conversion module is further configured to respectively fuse the extracted facial expression features into the three-dimensional projection images for expression reconstruction, so as to obtain a synthesized face video.

25. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.

26. A computer-readable storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 12 is implemented.

27. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Human face expression synthesis method and device, storage medium and computer equipment

    CN107610209A

  • Face image processing method and device and storage medium

    CN108985220A