A method and device for correcting facial posture
By acquiring the initial three-dimensional face model and the target three-dimensional face model, calculating the image deformation field and fusing it, the problem of face posture deflection in video calls is solved, the timing continuity and stability of face posture are achieved, and the dependence on a large amount of training data is avoided.
Patent Information
- Application Number
- CN202310017675.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-01-06
AI Technical Summary
In video call scenarios, due to the deviation of the position and orientation of the camera and the face, the face posture deflects, affecting the communication experience. The existing deep learning-based methods have problems such as loss of resolution and image quality, and the mismatch between the face and the background after editing, and it is difficult to ensure the timing continuity and stability of the face posture in the timing continuous frames.
By obtaining the initial three-dimensional face model and the target three-dimensional face model of the initial frame, the image deformation field is calculated, and the initial frame and the target face frame are fused based on the target boundary of the initial frame, the face pose correction is achieved, avoiding dependence on a large amount of training data.
The correction of face poses without a large amount of training data is realized, ensuring the timing continuity and stability of face poses in timing continuous frames, and solving the problem of inconsistency in the foreground background.
Smart Images

Figure CN116012250B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to image processing technology, and in particular to a method and device for correcting facial posture. Background Art
[0002] In a video call scenario, due to the deviation between the position and orientation of the camera and the person's face, the head posture of the person presented on the screen is deflected, affecting the interactive experience of both parties.
[0003] Currently, in the field of image processing, generative adversarial networks based on deep learning are generally used to edit facial images. However, using such methods for facial editing has problems such as loss of resolution and image quality, and mismatch between the edited face and the background. At the same time, such methods are heavily dependent on the amount of data and require the collection of a large amount of data for different scenes and character postures. In addition, it is difficult to ensure the temporal continuity and stability of facial postures in temporally continuous frames. Summary of the Invention
[0004] The embodiments of the present application provide a method and device for correcting facial posture, which can correct facial posture without requiring a large amount of training data, and can ensure the temporal continuity and stability of facial posture in temporally continuous frames.
[0005] The present invention provides a method for correcting facial posture, which may include:
[0006] Acquire an initial 3D face model and a target 3D face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame;
[0007] Obtaining an image deformation field according to the initial three-dimensional face model and the target three-dimensional face model;
[0008] deforming the initial frame according to the image deformation field to obtain the target face frame;
[0009] Obtaining the target boundary of the initial frame;
[0010] Based on the target boundary of the initial frame, the initial frame and the target face frame are fused to obtain a target frame.
[0011] In an exemplary embodiment of the present application, obtaining an initial 3D face model and a target 3D face model corresponding to the initial frame may include:
[0012] Obtaining an initial three-dimensional face model corresponding to the initial frame through three-dimensional reconstruction;
[0013] Performing posture correction on the initial three-dimensional face model to obtain the target three-dimensional face model.
[0014] In an exemplary embodiment of the present application, obtaining an image deformation field based on the initial three-dimensional face model and the target three-dimensional face model may include:
[0015] Projecting the initial three-dimensional face model onto an initial two-dimensional plane to obtain an initial face image;
[0016] Projecting the target three-dimensional face model onto a target two-dimensional plane to obtain a target face image;
[0017] The image deformation field is obtained according to the pixel correspondence between the initial face image and the target face image.
[0018] In an exemplary embodiment of the present application, obtaining the target boundary of the initial frame may include:
[0019] Obtaining an initial boundary of the initial frame;
[0020] Based on the initial boundary, determining an optimization region;
[0021] In the optimization region, an object boundary of the initial frame is determined based on a difference between a foreground portion and a background portion of the initial frame.
[0022] In an exemplary embodiment of the present application, obtaining the initial boundary of the initial frame may include:
[0023] Acquire the facial contour boundary of the initial facial image;
[0024] An initial boundary of the initial frame is determined according to a facial contour boundary of the initial facial image.
[0025] In an exemplary embodiment of the present application, the area within the initial boundary of the initial frame may be the optimization area.
[0026] In an exemplary embodiment of the present application, obtaining the initial boundary of the initial frame may include:
[0027] When the initial frame is not the first frame of the frame sequence, the initial boundary of the initial frame is determined according to the target boundary of a frame preceding the initial frame.
[0028] In an exemplary embodiment of the present application, determining the optimization region based on the initial boundary may include:
[0029] Respectively obtaining a frame preceding the initial frame and a head posture corresponding to the initial frame;
[0030] Calculating a posture difference between a frame preceding the initial frame and a head posture corresponding to the initial frame;
[0031] The optimization region is determined at the initial boundary according to the posture difference.
[0032] In an exemplary embodiment of the present application, fusing the initial frame and the target face frame based on the target boundary of the initial frame to obtain the target frame may include:
[0033] According to the target boundary of the initial frame, obtaining the background to be fused of the initial frame and the optimized boundary of the target face frame;
[0034] Acquiring a foreground to be fused of the target face frame according to the optimized boundary of the target face frame;
[0035] The background to be fused and the foreground to be fused are fused to obtain the target frame.
[0036] The present application also provides a facial posture correction device, which may include:
[0037] A face reconstruction module, configured to obtain an initial three-dimensional face model and a target three-dimensional face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame;
[0038] a deformation field calculation module, which obtains an image deformation field according to the initial three-dimensional face model and the target three-dimensional face model;
[0039] An image deformation module, configured to deform the initial frame according to the image deformation field to obtain the target face frame;
[0040] A boundary acquisition module, configured to acquire a target boundary of the initial frame;
[0041] The image fusion module is used to fuse the initial frame and the target face frame based on the target boundary of the initial frame to obtain a target frame.
[0042] In an exemplary embodiment of the present application, the face reconstruction module may be used to:
[0043] Obtaining an initial three-dimensional face model corresponding to the initial frame through three-dimensional reconstruction;
[0044] Performing posture correction on the initial three-dimensional face model to obtain the target three-dimensional face model.
[0045] In an exemplary embodiment of the present application, the deformation field calculation module may be used to:
[0046] Projecting the initial three-dimensional face model onto an initial two-dimensional plane to obtain an initial face image;
[0047] Projecting the target three-dimensional face model onto a target two-dimensional plane to obtain a target face image;
[0048] The image deformation field is obtained according to the pixel correspondence between the initial face image and the target face image.
[0049] In an exemplary embodiment of the present application, the boundary acquisition module may be used to:
[0050] Obtaining an initial boundary of the initial frame;
[0051] Based on the initial boundary, determining an optimization region;
[0052] In the optimization region, an object boundary of the initial frame is determined based on a difference between a foreground portion and a background portion of the initial frame.
[0053] In an exemplary embodiment of the present application, the image fusion module may be used to:
[0054] According to the target boundary of the initial frame, obtaining the background to be fused of the initial frame and the optimized boundary of the target face frame;
[0055] Acquiring a foreground to be fused of the target face frame according to the optimized boundary of the target face frame;
[0056] The background to be fused and the foreground to be fused are fused to obtain the target frame.
[0057] Compared to related technologies, the embodiments of the present application may include: obtaining an initial 3D face model and a target 3D face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame; obtaining an image deformation field based on the initial 3D face model and the target 3D face model; deforming the initial frame based on the image deformation field to obtain the target face frame; obtaining a target boundary of the initial frame; and fusing the initial frame and the target face frame based on the target boundary of the initial frame to obtain the target frame. Through this embodiment, facial posture correction can be achieved without the need for a large amount of training data, and the temporal continuity and stability of facial posture in temporally continuous frames can be guaranteed.
[0058] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. Other advantages of the present application can be realized and obtained through the solutions described in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings are used to provide an understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0060] Figure 1 This is a flowchart of the face posture correction method according to an embodiment of the present application;
[0061] Figure 2 This is a flowchart of a method for obtaining an initial 3D face model and a target 3D face model corresponding to an initial frame according to an embodiment of the present application;
[0062] Figure 3 This is a flowchart of a method for obtaining an image deformation field based on an initial three-dimensional face model and a target three-dimensional face model according to an embodiment of the present application;
[0063] Figure 4 This is a flowchart of a method for obtaining an object boundary of an initial frame according to an embodiment of the present application;
[0064] Figure 5 This is a flowchart of a method for fusing an initial frame and a target face frame based on a target boundary of the initial frame to obtain a target frame according to an embodiment of the present application;
[0065] Figure 6 This is a block diagram of the facial posture correction device according to an embodiment of the present application. DETAILED DESCRIPTION
[0066] This application describes multiple embodiments, but this description is exemplary rather than restrictive, and it will be apparent to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described herein. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.
[0067] This application includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The embodiments, features, and elements disclosed in this application may also be combined with any conventional features or elements to form a unique inventive solution defined by the claims. Any features or elements of any embodiment may also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this application may be implemented individually or in any appropriate combination. Therefore, except for the limitations made according to the appended claims and their equivalents, the embodiments are not subject to other limitations. In addition, various modifications and changes may be made within the scope of protection of the appended claims.
[0068] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can be changed and still remain within the spirit and scope of the embodiments of the present application.
[0069] The embodiment of the present application provides a method for correcting facial posture. Figure 1 As shown, the method may include steps S101-S105:
[0070] S101, obtaining an initial 3D face model and a target 3D face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame;
[0071] In an exemplary embodiment, the frame sequence is a temporally continuous frame sequence of no less than two frames. For example, the frame sequence may be a video frame sequence. The initial frame may be any frame in the frame sequence. For example, the initial frame may be the first frame in the frame sequence. All frames in the frame sequence are temporally continuous and contain facial information.
[0072] In an exemplary embodiment, the frame sequence may include only one frame, which is the initial frame;
[0073] S102, obtaining an image deformation field according to the initial 3D face model and the target 3D face model;
[0074] In an exemplary embodiment of the present application, the image deformation field corresponds to a deformation relationship between the initial three-dimensional face model and the target three-dimensional face model;
[0075] S103, deforming the initial frame according to the image deformation field to obtain a target face frame;
[0076] In an exemplary embodiment of the present application, since the image deformation field corresponds to the deformation relationship between the initial 3D face model and the target 3D face model, that is, the image deformation field corresponds to the deformation relationship between the initial frame and the frame corresponding to the target 3D face model, the initial frame is deformed according to the image deformation field to obtain the frame corresponding to the target 3D face model, that is, the target face frame;
[0077] S104, obtaining the target boundary of the initial frame;
[0078] If the face image in the target face frame is directly pasted back to the initial frame, problems such as the foreground and background being inconsistent will occur. Therefore, in an exemplary embodiment of the present application, a target boundary of the initial frame is obtained, which is a boundary that minimizes the difference between the foreground and background.
[0079] S105, based on the target boundary of the initial frame, fusing the initial frame and the target face frame to obtain a target frame;
[0080] Directly pasting the facial image in the target face frame back to the initial frame according to the target boundary of the initial frame still has problems such as artifacts. Therefore, in an exemplary embodiment of the present application, the initial frame and the target face frame are fused based on the target boundary of the initial frame to obtain the target frame.
[0081] The facial posture correction method provided in the embodiment of the present application can correct facial posture based on a single frame or multiple frames, avoiding dependence on a large amount of training data, and solving the problem of incoordination between foreground and background after facial posture correction; and when correcting facial posture on temporally continuous frames, historical frame information is utilized to achieve continuity of posture changes in time, solving the problem of difficulty in maintaining temporal continuity and stability of facial posture in temporally continuous frames.
[0082] The embodiment of the present application provides a method for obtaining an initial 3D face model and a target 3D face model corresponding to an initial frame, such as Figure 2 As shown, the method may include steps S201-S202:
[0083] S201, obtaining an initial 3D face model corresponding to an initial frame through 3D reconstruction;
[0084] 3D reconstruction can be achieved through methods such as visual geometry or deep learning;
[0085] The 3D reconstruction method based on visual geometry completes the conversion from a 2D image to a 3D model by extracting facial features from the initial frame and performing comparison and splicing. That is, the conversion from the initial frame to the initial 3D face model.
[0086] The 3D reconstruction method based on deep learning reconstructs the initial frame in 3D through the learning and fitting capabilities of deep neural networks;
[0087] In an exemplary embodiment, first, three-dimensional prior information of a face in an initial frame is obtained, and an initial three-dimensional face model is obtained based on the three-dimensional prior information;
[0088] S202, performing posture correction on the initial three-dimensional face model to obtain a target three-dimensional face model;
[0089] In an exemplary embodiment, the parameter values of the 3D facial parameters of the initial 3D facial model may be adjusted so that the head posture of the 3D facial model is a desired posture, thereby obtaining a target 3D facial model.
[0090] The embodiment of the present application provides a method for obtaining an image deformation field based on an initial three-dimensional face model and a target three-dimensional face model, such as Figure 3 As shown, the method may include steps S301-S302:
[0091] S301, projecting the initial three-dimensional face model onto an initial two-dimensional plane to obtain an initial face image;
[0092] In an exemplary embodiment of the present application, the initial three-dimensional face model is projected onto an initial two-dimensional plane perpendicular to the coronal plane of the initial three-dimensional face model, and the initial two-dimensional plane is parallel to the coronal plane of the initial three-dimensional face model;
[0093] S302, projecting the target three-dimensional face model onto a target two-dimensional plane to obtain a target face image;
[0094] In an exemplary embodiment of the present application, the target three-dimensional face model is projected onto a target two-dimensional plane perpendicular to the coronal plane of the target three-dimensional face model, and the target two-dimensional plane is parallel to the coronal plane of the target three-dimensional face model;
[0095] S303, obtaining an image deformation field according to a pixel correspondence relationship between the initial face image and the target face image;
[0096] Since the target three-dimensional face model is obtained by performing posture correction on the initial three-dimensional face model, the initial three-dimensional face model and the target three-dimensional face model have the same topological structure, and an initial face image and a target face image obtained by projecting the initial three-dimensional face model and the target three-dimensional face model have a pixel correspondence relationship; the pixel correspondence relationship between the initial face image and the target face image is obtained, and an image deformation field for transforming from the initial face image to the target face image can be obtained based on the pixel correspondence relationship;
[0097] In an exemplary embodiment, the image deformation field may be obtained by a three-dimensional model rendering technology based on the pixel correspondence between the initial face image and the target face image.
[0098] The embodiment of the present application provides a method for obtaining the target boundary of the initial frame, such as Figure 4 As shown, the method may include steps S401-S403:
[0099] S401, obtaining the initial boundary of the initial frame;
[0100] In an exemplary embodiment, step S401 may include:
[0101] Obtain the facial contour boundary of the initial face image;
[0102] In this exemplary embodiment, the outer contour boundary of the initial face image can be used as its face contour boundary.
[0103] Obtaining an initial boundary of an initial frame according to a facial contour boundary of an initial face image;
[0104] In this exemplary embodiment, since the initial facial image is generated by projecting the initial three-dimensional facial model corresponding to the initial frame, the facial contour boundary of the initial facial image can be mapped back onto the initial three-dimensional facial model to obtain the three-dimensional facial contour boundary on the initial three-dimensional facial model; the three-dimensional facial contour boundary on the initial three-dimensional facial model is then projected back onto the initial frame to obtain the initial boundary of the initial frame;
[0105] In an exemplary embodiment, step S401 may include:
[0106] When the initial frame is not the first frame of the frame sequence, determining the initial boundary of the initial frame according to the target boundary of the previous frame of the initial frame;
[0107] In this exemplary embodiment, the frame sequence is a temporally continuous frame sequence of no less than two frames. For example, the frame sequence may be a video frame sequence. The initial frame may be any frame other than the first frame in the frame sequence. All frames in the frame sequence are temporally continuous and contain facial information.
[0108] In this exemplary embodiment, a target boundary of a frame preceding an initial frame is first obtained, in a manner to be described later. The target boundary of the frame preceding the initial frame is mapped back onto a 3D face model corresponding to the frame preceding the initial frame to obtain a 3D target boundary on the 3D face model. A 3D initial boundary having the same position as the 3D target boundary is determined on the initial 3D face model corresponding to the initial frame. The 3D initial boundary is projected back onto the initial frame to obtain an initial boundary of the initial frame.
[0109] S402, determining an optimization region based on the initial boundary;
[0110] In an exemplary embodiment, the face region within the initial boundary is the optimized region;
[0111] In an exemplary embodiment, step S402 may include:
[0112] Get the head posture corresponding to the previous frame and the initial frame of the initial frame respectively;
[0113] Calculate the head posture difference between the previous frame and the initial frame;
[0114] According to the attitude difference, the optimization area is determined at the initial boundary;
[0115] In this exemplary embodiment, the frame sequence is a temporally continuous frame sequence of no less than two frames. For example, the frame sequence may be a video frame sequence. The initial frame may be any frame other than the first frame in the frame sequence. All frames in the frame sequence are temporally continuous and contain facial information.
[0116] In this exemplary embodiment, the head posture corresponding to the frame preceding the initial frame is subtracted from the head posture corresponding to the initial frame to obtain a posture difference value. Based on the posture difference value, an area of a certain width at the initial boundary can be set as an optimization area. The specific position can be preset. For example, the initial boundary can be the outer boundary or the inner boundary of the optimization area. The width of the optimization area is obtained by subtracting the image mask within the initial boundary after performing dilation and erosion operations respectively. The larger the posture difference value, the larger the kernel used for the dilation and erosion operations, and the larger the width of the optimization area obtained by subtraction.
[0117] S403, determining the target boundary of the initial frame based on the difference between the foreground and background parts of the initial frame within the optimization area;
[0118] In an exemplary embodiment of the present application, a graph cut algorithm may be used to determine a boundary within the optimization region that minimizes the difference between the foreground and background parts of the initial frame. This boundary is the target boundary of the initial frame.
[0119] The embodiment of the present application provides a method for fusing the initial frame and the target face frame based on the target boundary of the initial frame to obtain the target frame, such as Figure 5 As shown, the method may include steps S501-S503:
[0120] S501, determining the optimized boundary of the to-be-fused background of the initial frame and the target face frame according to the target boundary of the initial frame;
[0121] In an exemplary embodiment of the present application, the area within the target boundary of the initial frame is used as the mask portion required for image fusion, and the area outside the target boundary of the initial frame is used as the background to be fused required for image fusion;
[0122] S502, obtaining a fused foreground of the target face frame according to the optimized boundary of the target face frame;
[0123] In an exemplary embodiment, the target boundary of the initial frame can be mapped back to the initial 3D face model to obtain an initial 3D boundary on the initial 3D face model; a 3D face boundary at the same position as the initial 3D boundary is determined on the target 3D face model corresponding to the target face frame; the 3D face boundary is projected back to the target face frame to obtain an optimized boundary of the target face frame;
[0124] In an exemplary embodiment of the present application, the portion within the optimized boundary of the target face frame is used as the foreground to be fused;
[0125] S503, fusing the background to be fused and the foreground to be fused to obtain a target frame;
[0126] The background to be fused and the foreground to be fused can be fused by using image fusion methods such as Laplace fusion method or Poisson image fusion method;
[0127] In an exemplary embodiment, the background to be fused and the foreground to be fused are fused using a Laplacian fusion method. First, a Gaussian pyramid of the initial frame, the target face frame, and the mask portion is constructed, and a Laplacian residual pyramid of the target face frame and the initial frame is constructed from this. The corresponding layers of the two Laplacian residual pyramids are fused using the masked Gaussian pyramid. Finally, the image is reconstructed by upsampling and fusion of the Laplacian residual pyramid to obtain a fused target frame.
[0128] The facial posture correction method provided in the embodiment of the present application can correct facial posture based on a single frame or multiple frames, avoiding dependence on a large amount of training data, and solving the problem of incoordination between foreground and background after facial posture correction; and when correcting facial posture on temporally continuous frames, historical frame information is utilized to achieve continuity of posture changes in time, solving the problem of difficulty in maintaining temporal continuity and stability of facial posture in temporally continuous frames.
[0129] The embodiment of the present application provides a face posture correction device 10, such as Figure 6 As shown, the device may include:
[0130] A face reconstruction module 101 is configured to obtain an initial 3D face model and a target 3D face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame;
[0131] In an exemplary embodiment, the frame sequence is a temporally continuous frame sequence of not less than two frames, for example, the frame sequence may be a video frame sequence; the initial frame may be any frame in the frame sequence, for example, the initial frame may be the first frame in the frame sequence; all frames in the frame sequence are temporally continuous and contain facial information;
[0132] In an exemplary embodiment, the frame sequence may include only one frame, which is the initial frame;
[0133] A deformation field calculation module 102 is configured to obtain an image deformation field based on the initial 3D face model and the target 3D face model;
[0134] In an exemplary embodiment of the present application, the image deformation field corresponds to a deformation relationship between the initial three-dimensional face model and the target three-dimensional face model;
[0135] An image deformation module 103 is used to deform the initial frame according to the image deformation field to obtain a target face frame;
[0136] In the exemplary embodiment of the present application, since the image deformation field corresponds to the deformation relationship between the initial 3D face model and the target 3D face model, that is, the image deformation field corresponds to the deformation relationship between the initial frame and the frame corresponding to the target 3D face model; the image deformation module 103 deforms the initial frame according to the image deformation field to obtain a frame corresponding to the target 3D face model, that is, the target face frame;
[0137] Boundary acquisition module 104, used to obtain the target boundary of the initial frame;
[0138] If the face image in the target face frame is directly pasted back to the initial frame, problems such as inconsistency between the foreground and the background will occur. Therefore, in an exemplary embodiment of the present application, the boundary acquisition module 104 acquires the target boundary of the initial frame, which is the boundary that minimizes the difference between the foreground and the background.
[0139] An image fusion module 105 is configured to fuse the initial frame and the target face frame based on the target boundary of the initial frame to obtain a target frame;
[0140] Directly pasting the facial image in the target face frame back to the initial frame according to the target boundary of the initial frame still causes problems such as artifacts. Therefore, in an exemplary embodiment of the present application, the image fusion module 105 fuses the initial frame and the target face frame based on the target boundary of the initial frame to obtain the target frame.
[0141] The facial posture correction device provided in the embodiment of the present application can correct facial posture based on a single frame or multiple frames, avoiding dependence on a large amount of training data, and solving the problem of incoordination between foreground and background after facial posture correction; and when correcting facial posture on temporally continuous frames, historical frame information is utilized to achieve continuity of posture changes in time, solving the problem of difficulty in maintaining temporal continuity and stability of facial posture in temporally continuous frames.
[0142] In an exemplary embodiment of the present application, the face reconstruction module 101 may be used to:
[0143] Obtain an initial 3D face model corresponding to the initial frame through 3D reconstruction;
[0144] Perform posture correction on the initial 3D face model to obtain the target 3D face model;
[0145] In this exemplary embodiment, the face reconstruction module 101 can implement three-dimensional reconstruction based on methods such as visual geometry or deep learning;
[0146] The 3D reconstruction method based on visual geometry completes the conversion from a 2D image to a 3D model by extracting facial features from the initial frame and performing comparison and splicing. That is, the conversion from the initial frame to the initial 3D face model.
[0147] The 3D reconstruction method based on deep learning reconstructs the initial frame in 3D through the learning and fitting capabilities of deep neural networks;
[0148] In this exemplary embodiment, the face reconstruction module 101 may first obtain three-dimensional prior information of the face in the initial frame, and obtain an initial three-dimensional face model based on the three-dimensional prior information;
[0149] In this exemplary embodiment, the parameter values of the 3D facial parameters of the initial 3D facial model may be adjusted so that the head posture of the 3D facial model is the required frontal posture, thereby obtaining a target 3D facial model.
[0150] In an exemplary embodiment of the present application, the deformation field calculation module 102 may be used to:
[0151] Projecting the initial three-dimensional face model onto the initial two-dimensional plane to obtain an initial face image;
[0152] Project the target three-dimensional face model onto the target two-dimensional plane to obtain the target face image;
[0153] According to the pixel correspondence between the initial face image and the target face image, an image deformation field is obtained;
[0154] In this exemplary embodiment, the initial three-dimensional face model is projected onto an initial two-dimensional plane perpendicular to the coronal plane of the initial three-dimensional face model, and the initial two-dimensional plane is parallel to the coronal plane of the initial three-dimensional face model;
[0155] In this exemplary embodiment, the target three-dimensional face model is projected onto a target two-dimensional plane perpendicular to the coronal plane of the target three-dimensional face model, and the target two-dimensional plane is parallel to the coronal plane of the target three-dimensional face model;
[0156] Since the target three-dimensional face model is obtained by performing posture correction on the initial three-dimensional face model, the initial three-dimensional face model and the target three-dimensional face model have the same topological structure, and there is a pixel correspondence between the initial face image and the target face image obtained by projecting the initial three-dimensional face model and the target three-dimensional face model; the pixel correspondence between the initial face image and the target face image is obtained, and based on the pixel correspondence, an image deformation field transformed from the initial face image to the target face image can be obtained.
[0157] In an exemplary embodiment of the present application, the boundary acquisition module 104 may be configured to:
[0158] Get the initial boundaries of the initial frame;
[0159] Based on the initial boundary, determine the optimization area;
[0160] In the optimization area, the target boundary of the initial frame is determined based on the difference between the foreground and background parts of the initial frame;
[0161] In an exemplary embodiment, obtaining the initial boundary of the initial frame may include:
[0162] Obtain the facial contour boundary of the initial face image;
[0163] Obtaining an initial boundary of an initial frame according to a facial contour boundary of an initial face image;
[0164] In this exemplary embodiment, the outer contour boundary of the initial face image may be used as its face contour boundary;
[0165] In this exemplary embodiment, since the initial facial image is generated by projecting the initial three-dimensional facial model corresponding to the initial frame, the facial contour boundary of the initial facial image can be mapped back onto the initial three-dimensional facial model to obtain the three-dimensional facial contour boundary on the initial three-dimensional facial model; the three-dimensional facial contour boundary on the initial three-dimensional facial model is then projected back onto the initial frame to obtain the initial boundary of the initial frame;
[0166] In an exemplary embodiment, obtaining the initial boundary of the initial frame may include:
[0167] When the initial frame is not the first frame of the frame sequence, determining the initial boundary of the initial frame according to the target boundary of the previous frame of the initial frame;
[0168] In this exemplary embodiment, the frame sequence is a temporally continuous frame sequence of no less than two frames. For example, the frame sequence may be a video frame sequence. The initial frame may be any frame other than the first frame in the frame sequence. All frames in the frame sequence are temporally continuous and contain facial information.
[0169] In this exemplary embodiment, a target boundary of a frame preceding an initial frame is first obtained, in a manner to be described later. The target boundary of the frame preceding the initial frame is mapped back onto a 3D face model corresponding to the frame preceding the initial frame to obtain a 3D target boundary on the 3D face model. A 3D initial boundary having the same position as the 3D target boundary is determined on the initial 3D face model corresponding to the initial frame. The 3D initial boundary is projected back onto the initial frame to obtain an initial boundary of the initial frame.
[0170] In an exemplary embodiment, the face region within the initial boundary is the optimized region;
[0171] In an exemplary embodiment, based on the initial boundary, determining the optimization region may include:
[0172] Get the head posture corresponding to the previous frame and the initial frame of the initial frame respectively;
[0173] Calculate the head posture difference between the previous frame and the initial frame;
[0174] According to the attitude difference, the optimization area is determined at the initial boundary;
[0175] In this exemplary embodiment, the frame sequence is a temporally continuous frame sequence of no less than two frames. For example, the frame sequence may be a video frame sequence. The initial frame may be any frame other than the first frame in the frame sequence. All frames in the frame sequence are temporally continuous and contain facial information.
[0176] In this exemplary embodiment, the head posture corresponding to the frame preceding the initial frame is subtracted from the head posture corresponding to the initial frame to obtain a posture difference value. Based on the posture difference value, an area of a certain width at the initial boundary can be set as an optimization area. The specific position can be preset. For example, the initial boundary can be the outer boundary or the inner boundary of the optimization area. The width of the optimization area is obtained by subtracting the image mask within the initial boundary after performing dilation and erosion operations respectively. The larger the posture difference value, the larger the kernel used for the dilation and erosion operations, and the larger the width of the optimization area obtained by subtraction.
[0177] In this exemplary embodiment, a boundary that minimizes the difference between the foreground and background parts of the initial frame may be determined by a graph cut algorithm, and this boundary is the target boundary of the initial frame.
[0178] In an exemplary embodiment of the present application, the image fusion module 105 may be used to:
[0179] According to the target boundary of the initial frame, the optimized boundary of the background to be fused of the initial frame and the target face frame is determined;
[0180] According to the optimized boundary of the target face frame, the foreground to be fused of the target face frame is obtained;
[0181] Fuse the background to be fused and the foreground to be fused to obtain the target frame;
[0182] In an exemplary embodiment of the present application, the area within the target boundary of the initial frame is used as the mask portion required for image fusion, and the area outside the target boundary of the initial frame is used as the background to be fused required for image fusion;
[0183] In this exemplary embodiment, the target boundary of the initial frame can be mapped back to the initial three-dimensional face model to obtain an initial three-dimensional boundary on the initial three-dimensional face model; a three-dimensional face boundary at the same position as the initial three-dimensional boundary is determined on the target three-dimensional face model corresponding to the target face frame; and the three-dimensional face boundary is projected back to the target face frame to obtain an optimized boundary of the target face frame.
[0184] In this exemplary embodiment, the portion within the optimized boundary of the target face frame is used as the foreground to be fused;
[0185] In this exemplary embodiment, the background to be fused and the foreground to be fused may be fused by an image fusion method such as a Laplace fusion method or a Poisson image fusion method;
[0186] The Laplacian fusion method is used to achieve the fusion of the background and foreground to be fused. First, it is necessary to construct a Gaussian pyramid of the initial frame, the target face frame, and the mask part, and then construct a Laplacian residual pyramid of the target face frame and the initial frame; the corresponding layers of the two Laplacian residual pyramids are fused using the masked Gaussian pyramid; finally, the image is reconstructed by upsampling and fusion of the Laplacian residual pyramid to obtain the fused target frame.
[0187] The facial posture correction device provided in the embodiment of the present application can correct facial posture based on a single frame or multiple frames, avoiding dependence on a large amount of training data, and solving the problem of incoordination between foreground and background after facial posture correction; and when correcting facial posture on temporally continuous frames, historical frame information is utilized to achieve continuity of posture changes in time, solving the problem of difficulty in maintaining temporal continuity and stability of facial posture in temporally continuous frames.
[0188] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
Claims
1. A face posture correction method, characterized in that: The method comprises: Acquire an initial 3D face model and a target 3D face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame; Obtaining an image deformation field according to the initial three-dimensional face model and the target three-dimensional face model; deforming the initial frame according to the image deformation field to obtain a target face frame; Obtaining the target boundary of the initial frame; Based on the target boundary of the initial frame, the initial frame and the target face frame are fused to obtain a target frame.
2. The face posture correction method according to claim 1, wherein: The step of obtaining an initial three-dimensional face model and a target three-dimensional face model corresponding to the initial frame includes: Obtaining an initial three-dimensional face model corresponding to the initial frame through three-dimensional reconstruction; Performing posture correction on the initial three-dimensional face model to obtain the target three-dimensional face model.
3. The face posture correction method according to claim 1, wherein: The step of obtaining an image deformation field according to the initial three-dimensional face model and the target three-dimensional face model includes: Projecting the initial three-dimensional face model onto an initial two-dimensional plane to obtain an initial face image; Projecting the target three-dimensional face model onto a target two-dimensional plane to obtain a target face image; The image deformation field is obtained according to the pixel correspondence between the initial face image and the target face image.
4. The face posture correction method according to claim 3, wherein: The obtaining of the target boundary of the initial frame includes: Obtaining an initial boundary of the initial frame; Based on the initial boundary, determining an optimization region; In the optimization region, an object boundary of the initial frame is determined based on a difference between a foreground portion and a background portion of the initial frame.
5. The face posture correction method according to claim 4, characterized in that: The obtaining of the initial boundary of the initial frame includes: Acquire the facial contour boundary of the initial facial image; An initial boundary of the initial frame is determined according to a facial contour boundary of the initial facial image.
6. The face posture correction method according to claim 5, characterized in that: The area within the initial boundary of the initial frame is the optimized area.
7. The face posture correction method according to claim 4, characterized in that: The obtaining of the initial boundary of the initial frame includes: When the initial frame is not the first frame of the frame sequence, the initial boundary of the initial frame is determined according to the target boundary of a frame preceding the initial frame.
8. The face posture correction method according to claim 7, characterized in that: The step of determining an optimization region based on the initial boundary includes: Respectively obtaining a frame preceding the initial frame and a head posture corresponding to the initial frame; Calculating a posture difference between a frame preceding the initial frame and a head posture corresponding to the initial frame; The optimization region is determined at the initial boundary according to the posture difference.
9. The face posture correction method according to any one of claims 1 to 8, characterized in that: The step of fusing the initial frame and the target face frame based on the target boundary of the initial frame to obtain the target frame includes: According to the target boundary of the initial frame, obtaining the background to be fused of the initial frame and the optimized boundary of the target face frame; Acquiring a foreground to be fused of the target face frame according to the optimized boundary of the target face frame; The background to be fused and the foreground to be fused are fused to obtain the target frame.
10. A facial posture correction device, characterized in that: The device comprises: A face reconstruction module, configured to obtain an initial three-dimensional face model and a target three-dimensional face model corresponding to an initial frame, wherein the initial frame is any frame in a frame sequence, and the frame sequence includes at least one frame; a deformation field calculation module, which obtains an image deformation field according to the initial three-dimensional face model and the target three-dimensional face model; An image deformation module, configured to deform the initial frame according to the image deformation field to obtain a target face frame; A boundary acquisition module, configured to acquire a target boundary of the initial frame; The image fusion module is used to fuse the initial frame and the target face frame based on the target boundary of the initial frame to obtain a target frame.
11. The face posture correction device according to claim 10, characterized in that: The face reconstruction module is used to: Obtaining an initial three-dimensional face model corresponding to the initial frame through three-dimensional reconstruction; Performing posture correction on the initial three-dimensional face model to obtain the target three-dimensional face model.
12. The face posture correction device according to claim 10, characterized in that: The deformation field calculation module is used to: Projecting the initial three-dimensional face model onto an initial two-dimensional plane to obtain an initial face image; Projecting the target three-dimensional face model onto a target two-dimensional plane to obtain a target face image; The image deformation field is obtained according to the pixel correspondence between the initial face image and the target face image.
13. The face posture correction device according to claim 12, characterized in that: The boundary acquisition module is used to: Obtaining an initial boundary of the initial frame; Based on the initial boundary, determining an optimization region; In the optimization region, an object boundary of the initial frame is determined based on a difference between a foreground portion and a background portion of the initial frame.
14. The face posture correction device according to any one of claims 10 to 13, characterized in that: The image fusion module is used to: According to the target boundary of the initial frame, obtaining the background to be fused of the initial frame and the optimized boundary of the target face frame; Acquiring a foreground to be fused of the target face frame according to the optimized boundary of the target face frame; The background to be fused and the foreground to be fused are fused to obtain the target frame.
Citation Information
Patent Citations
Method, device and system of three-dimensional face reconstruction and computer storage medium
CN108876893A
Face image deformation method and device, electronic equipment and storage medium
CN113986105A