Face reconstruction method, apparatus, device, medium and product
By performing two-dimensional facial key point detection and three-dimensional facial parameter prediction on the target image, combined with fine-tuning and difference representation data update, the problem of inaccurate facial reconstruction in existing technologies is solved, and more efficient three-dimensional facial model construction and audio adjustment tasks in videos are achieved.
Patent Information
- Application Number
- PCT/CN2024/139769
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2024-12-16
- Publication Date
- 2025-10-16
AI Technical Summary
Existing technologies have difficulty in accurately reconstructing faces in scenarios where audio adjustment is required in videos or other scenarios where face reconstruction is required, resulting in poor results in building three-dimensional face models and performing audio adjustment tasks in videos.
By acquiring the target image, two-dimensional facial key point detection and three-dimensional facial parameter prediction are performed, combined with fine-tuning processing, a three-dimensional face model is constructed, and the reconstruction result is updated through difference characterization data to ensure accuracy and stability in two-dimensional and three-dimensional space.
The accuracy and stability of facial reconstruction have been improved, enabling better results in tasks such as building three-dimensional facial models and adjusting audio in videos.
Smart Images

Figure CN2024139769_16102025_PF_FP_ABST
Abstract
Description
A face reconstruction method, device, equipment, medium and product
[0001] The present application claims priority to the Chinese patent application No. 202410418018.7, filed on April 8, 2024, and entitled "A face reconstruction method, device, equipment, medium and product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of data processing, and in particular to a face reconstruction method, device, equipment, medium and product. BACKGROUND
[0003] For some application scenarios, such as audio adjustment in video scenarios or other scenarios with face reconstruction requirements, these application scenarios may have the following requirements: face reconstruction is performed on a single image or a video, so that subsequent tasks such as three-dimensional face model construction, two-dimensional video conversion to three-dimensional video, or audio adjustment in video can be completed using the face reconstruction result. SUMMARY
[0004] The present application provides a face reconstruction method, device, equipment, medium and product, which is beneficial to improve the face reconstruction effect.
[0005] In order to achieve the above-mentioned purpose, the technical scheme provided by the present application is as follows:
[0006] The present application provides a face reconstruction method, which comprises:
[0007] Obtaining a target image;
[0008] Performing two-dimensional face key point detection processing on the target image to obtain a two-dimensional face key point detection result of the target image, and performing three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image;
[0009] According to the two-dimensional face key point detection result, performing fine-tuning processing on the three-dimensional face parameter prediction result to obtain a three-dimensional face reconstruction result corresponding to the target image.
[0010] In a possible implementation, the fine-tuning processing comprises:
[0011] According to the three-dimensional face parameter prediction result, performing initialization processing on the three-dimensional face reconstruction result;
[0012] According to the three-dimensional face reconstruction result, constructing a three-dimensional face model corresponding to the target image;
[0013] mapping the three-dimensional face model to a two-dimensional image space to obtain a two-dimensional face key point mapping result corresponding to the target image;
[0014] updating the three-dimensional face reconstruction result according to difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result.
[0015] In a possible implementation, the updating the three-dimensional face reconstruction result according to difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result includes:
[0016] updating the three-dimensional face reconstruction result according to difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result, and difference representation data between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result.
[0017] In a possible implementation, a constraint intensity of the difference representation data between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result on the updating is weaker than a constraint intensity of the difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result on the updating.
[0018] In a possible implementation, the target image is a frame image in a reference video, and the reference video includes a previous frame image of the target image.
[0019] The three-dimensional face reconstruction result is updated according to a time sequence loss corresponding to the three-dimensional face reconstruction result and / or a time sequence loss corresponding to the two-dimensional face key point mapping result.
[0020] The time sequence loss corresponding to the three-dimensional face reconstruction result is determined according to motion state representation data between a three-dimensional face parameter prediction result of the target image and three-dimensional face parameter information of the previous frame image, and motion state representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter information of the previous frame image; the three-dimensional face parameter information of the previous frame image is determined according to a three-dimensional face parameter prediction result of the previous frame image and / or a three-dimensional face reconstruction result corresponding to the previous frame image.
[0021] The time sequence loss corresponding to the two-dimensional face key point mapping result is determined according to motion state representation data between a two-dimensional face key point detection result of the target image and a two-dimensional face key point detection result of the previous frame image, and motion state representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the previous frame image.
[0022] In a possible implementation, the updating the three-dimensional face reconstruction result comprises:
[0023] updating other parameters in the three-dimensional face reconstruction result except the face identity parameter.
[0024] In a possible implementation, the target image refers to any frame image in the reference video.
[0025] The face identity parameter in the three-dimensional face reconstruction result corresponding to the target image is determined according to an average value of face identity parameters in three-dimensional face parameter prediction results of at least two frame images in the reference video, and the at least two frame images include the target image.
[0026] In a possible implementation, the target image refers to any frame image in the reference video.
[0027] The method further comprises:
[0028] generating a video corresponding to an audio sequence according to the three-dimensional face reconstruction result corresponding to each frame image in the reference video, wherein a subject presented in the video corresponding to the audio sequence is consistent with a subject presented in the reference video, and the video corresponding to the audio sequence is used to describe face state changes of the subject under the audio sequence.
[0029] The present application provides a face reconstruction device, comprising:
[0030] an acquisition unit configured to acquire a target image;
[0031] a processing unit configured to perform two-dimensional face key point detection processing on the target image to obtain a two-dimensional face key point detection result of the target image, and perform three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image;
[0032] a fine-tuning unit configured to perform fine-tuning processing on the three-dimensional face parameter prediction result according to the two-dimensional face key point detection result to obtain a three-dimensional face reconstruction result corresponding to the target image.
[0033] The present application provides an electronic device, comprising a processor and a memory.
[0034] The memory is configured to store instructions or computer programs.
[0035] The processor is configured to execute the instructions or computer programs in the memory, so that the electronic device performs the face reconstruction method provided by the present application.
[0036] The application provides a computer readable medium, wherein instructions or a computer program are stored in the computer readable medium, and the instructions or the computer program enable a device to perform a face reconstruction method provided by the application when the instructions or the computer program are executed on the device.
[0037] The application provides a computer program product, which comprises a computer program carried on a non-transitory computer readable medium, and the computer program comprises program codes for performing a face reconstruction method provided by the application.
[0038] In the technical solution provided by the application, for a target image, such as a single image or an i-th image in a video, after the target image is obtained, first, two-dimensional face key point detection processing is performed on the target image to obtain a two-dimensional face key point detection result of the target image, so that the two-dimensional face key point detection result can represent a face state of an object in the target image in a two-dimensional space, and three-dimensional face parameter prediction processing is performed on the target image to obtain a three-dimensional face parameter prediction result of the target image, so that the three-dimensional face parameter prediction result can more accurately represent a face state of the object in a three-dimensional space; then, the three-dimensional face parameter prediction result is fine-tuned according to the two-dimensional face key point detection result to obtain a three-dimensional face reconstruction result corresponding to the target image. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0040] FIG. 1 is a flowchart of a face reconstruction method provided by an embodiment of the application;
[0041] FIG. 2 is a schematic diagram of an implementation flow of an audio adjustment task in a video provided by an embodiment of the application;
[0042] FIG. 3 is a schematic diagram of a face reconstruction flow provided by an embodiment of the application;
[0043] FIG. 4 is a structural schematic diagram of a face reconstruction device provided by an embodiment of the application;
[0044] FIG. 5 is a structural schematic diagram of an electronic device provided by an embodiment of the application. DETAILED DESCRIPTION
[0045] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0046] In order to better understand the technical solutions provided in the present application, the face reconstruction method provided in the present application will be described in conjunction with some drawings. As shown in FIG. 1, the face reconstruction method provided in the embodiments of the present application includes the following S1-S4. Wherein, FIG. 1 is a flow chart of a face reconstruction method provided in the embodiments of the present application.
[0047] S1: Obtain a target image.
[0048] Wherein, the target image refers to an image that needs to be processed by face reconstruction, and the implementation of the target image is not limited in the present application. In order to facilitate understanding, two cases will be described in the following.
[0049] Case 1, in some application scenarios, such as single image processing scenarios, the above-mentioned target image can refer to an image involved in a single image processing task, such as an image specified by a user or an image provided by other tasks, etc. It should be noted that the present application does not limit the single image processing task, for example, the single image processing task at least involves face reconstruction processing for a single image.
[0050] Case 2, in some application scenarios, such as video processing scenarios similar to audio adjustment in video, the above-mentioned target image can refer to any one frame image in a reference video, such as the i-th frame image, i is a positive integer, i≤I, I is a positive integer, I represents the total number of frames of images in the reference video. Wherein, the reference video refers to a video involved in a certain video processing task, which needs to be processed by face reconstruction, such as the reference video shown in FIG. 2 or FIG. 3; and the implementation of the reference video is not limited in the present application, for example, the reference video can be implemented by using a single-person broadcast video. In addition, the implementation of the video processing task is not limited in the present application, for example, the video processing task at least involves face reconstruction processing for part or all of the images in the video. For example, the video processing task can be implemented by using tasks similar to video translation, video error correction, and partial audio replacement in video.
[0051] It can be seen that in a possible implementation, the above-mentioned target image can refer to the i-th frame image in the reference video, i is a positive integer, i≤I, I is a positive integer, I represents the total number of frames of images in the reference video.
[0052] In addition, the present application does not limit the manner of obtaining the target image.
[0053] S2: performing a two-dimensional face key point detection process on the target image to obtain a two-dimensional face key point detection result of the target image.
[0054] The two-dimensional face key point detection result of the target image is used to describe the face state, such as the expression state, of the object in the target image in the two-dimensional space. The present application does not limit the implementation of the two-dimensional face key point detection result. For example, the two-dimensional face key point detection result can be implemented by using any existing or future two-dimensional face key point, such as two-dimensional landmarks. It should be noted that the present application does not limit the implementation of the object. For example, the object can be implemented by using an animal or a virtual image.
[0055] In addition, the present application does not limit the implementation of the two-dimensional face key point detection process in S2. For example, the two-dimensional face key point detection process can be implemented by using any existing or future method capable of performing a two-dimensional face key point detection process on an image, such as by using a pre-constructed machine learning model having a two-dimensional face key point detection function.
[0056] S3: performing a three-dimensional face parameter prediction process on the target image to obtain a three-dimensional face parameter prediction result of the target image.
[0057] The three-dimensional face parameter prediction result of the target image is used to describe the face state, such as the expression state, of the object in the target image in the three-dimensional space.
[0058] In addition, the present application does not limit the implementation of the three-dimensional face parameter prediction result of the target image, for example, it can include face identity (Identity document, ID) parameters, face expression parameters, and face pose parameters. Among them, the face identity parameters are used to describe the facial features of the object in the target image, such as facial contour, facial feature distribution, and the like, so that the three-dimensional face model constructed based on the face identity parameters can represent the face state of the object in the case of no expression and no pose. The face expression parameters are used to describe the expression state of the object in the target image, so that the three-dimensional face model constructed based on the face expression parameters can represent the face state of the object in the case of no ID and no pose; and the present application does not limit the implementation of the face expression parameters, for example, the face expression parameters can be implemented by using blendshape coefficients. The face pose parameters are used to describe the face pose of the object in the target image, such as front face, side face, and the like, so that the three-dimensional face model constructed based on the face pose parameters can represent the face state of the object in the case of no ID and no expression; and the present application does not limit the implementation of the face pose parameters, for example, the face pose parameters can include rotation, translation, scaling, and the like. It can be seen that in a possible implementation, the three-dimensional face parameter prediction result of the target image can be implemented by using a three-dimensional face morphable model (3D Morphable Model, 3DMM).
[0059] In addition, the present application does not limit the implementation of the S3, for example, in some application scenarios, in order to better improve the face reconstruction effect, the S3 can be specifically: using a pre-constructed three-dimensional face parameter prediction model to perform three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image. Among them, the three-dimensional face parameter prediction model is used to perform three-dimensional face parameter prediction processing on the input data of the three-dimensional face parameter prediction model; and the present application does not limit the implementation of the three-dimensional face parameter prediction model, for example, it can be implemented by using any existing or future three-dimensional face parameter prediction model, such as a machine learning model.
[0060] In addition, the present application does not limit the implementation of the S3, for example, in some application scenarios, in order to better improve the face reconstruction effect, the S3 can be specifically: using a pre-constructed three-dimensional face parameter prediction model to perform three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image. Among them, the three-dimensional face parameter prediction model is used to perform three-dimensional face parameter prediction processing on the input data of the three-dimensional face parameter prediction model; and the present application does not limit the implementation of the three-dimensional face parameter prediction model, for example, it can be implemented by using any existing or future three-dimensional face parameter prediction model, such as a machine learning model.
[0061] Based on the related content of S3 above, in some application scenarios, such as video processing scenarios, for the i-th frame image in the reference video, after obtaining the i-th frame image, the i-th frame image can be input into the pre-constructed three-dimensional face parameter prediction model, so that the three-dimensional face parameter prediction model can perform three-dimensional face parameter prediction processing on the i-th frame image, and obtain and output the three-dimensional face parameter prediction result of the i-th frame image. Wherein, because the three-dimensional face parameter prediction model has good three-dimensional face parameter prediction performance, so that the three-dimensional face parameter prediction result obtained by predicting the i-th frame image by using the three-dimensional face parameter prediction model can more accurately represent the face state of the object in the i-th frame image in the three-dimensional space, so that the three-dimensional face reconstruction result for the i-th frame image can be better determined by taking the three-dimensional face parameter prediction result as the initial value. i is a positive integer, i≤I, I is a positive integer, and I represents the total number of frames of images in the reference video.
[0062] S4: According to the two-dimensional face key point detection result of the target image, the three-dimensional face parameter prediction result of the target image is fine-tuned to obtain the three-dimensional face reconstruction result corresponding to the target image.
[0063] Wherein, the three-dimensional face reconstruction result corresponding to the target image refers to the fine-tuning result of the three-dimensional face parameter prediction result of the target image, so that the three-dimensional face reconstruction result corresponding to the target image can more accurately represent the face state of the object in the three-dimensional space in the target image.
[0064] In addition, the present application does not limit the implementation of the three-dimensional face reconstruction result corresponding to the target image, such as the implementation of the three-dimensional face reconstruction result corresponding to the target image is similar to the implementation of the three-dimensional face parameter prediction result of the target image. It can be seen that in a possible implementation, the three-dimensional face reconstruction result corresponding to the target image can include face identification parameters, face expression parameters and face posture parameters.
[0065] In addition, the present application does not limit the implementation of S4 above, such as in some application scenarios, S4 can be specifically: input the two-dimensional face key point detection result of the target image and the three-dimensional face parameter prediction result of the target image into the pre-constructed parameter fine-tuning model, so that the parameter fine-tuning model can fine-tune the three-dimensional face parameter prediction result of the target image according to the two-dimensional face key point detection result of the target image, and obtain and output the three-dimensional face reconstruction result corresponding to the target image. Wherein, the parameter fine-tuning model refers to a pre-constructed model with three-dimensional face parameter fine-tuning function, such as a certain machine learning model, etc.; and the present application does not limit the implementation of the parameter fine-tuning model.
[0066] In addition, in order to better improve the reconstruction effect, the present application further provides a possible implementation manner of the above S4, in which the S4 can specifically include the following steps 11-14.
[0067] Step 11: initializing the three-dimensional face reconstruction result corresponding to the target image according to the three-dimensional face parameter prediction result of the target image.
[0068] It should be noted that the present application does not limit the implementation manner of the above step 11, for example, it can specifically be that the three-dimensional face parameter prediction result of the target image is determined as the initial value of the three-dimensional face reconstruction result corresponding to the target image.
[0069] It can be seen that in a possible implementation manner, when the three-dimensional face parameter prediction result of the target image includes the face identity parameter, the face expression parameter and the face pose parameter, the above step 11 can specifically be that the face identity parameter in the three-dimensional face reconstruction result corresponding to the target image is initialized by using the face identity parameter in the three-dimensional face parameter prediction result, so that the initial value of the face identity parameter in the three-dimensional face reconstruction result is consistent with the face identity parameter in the three-dimensional face parameter prediction result; the face expression parameter in the three-dimensional face reconstruction result is initialized by using the face expression parameter in the three-dimensional face parameter prediction result, so that the initial value of the face expression parameter in the three-dimensional face reconstruction result is consistent with the face expression parameter in the three-dimensional face parameter prediction result; and the face pose parameter in the three-dimensional face reconstruction result is initialized by using the face pose parameter in the three-dimensional face parameter prediction result, so that the initial value of the face pose parameter in the three-dimensional face reconstruction result is consistent with the face pose parameter in the three-dimensional face parameter prediction result.
[0070] Step 12: constructing the three-dimensional face model corresponding to the target image according to the three-dimensional face reconstruction result corresponding to the target image.
[0071] The three-dimensional face model corresponding to the target image refers to the three-dimensional face model constructed according to the three-dimensional face reconstruction result corresponding to the target image, so that the model can present the face state of the object in the target image in the three-dimensional space.
[0072] In addition, the present application does not limit the implementation manner of the above step 12, for example, when the three-dimensional face reconstruction result corresponding to the target image is implemented by using the 3DMM, the step 12 can be implemented by using any method capable of constructing the three-dimensional face model based on the 3DMM.
[0073] Step 13: mapping the three-dimensional face model corresponding to the target image to the two-dimensional image space to obtain a two-dimensional face key point mapping result corresponding to the target image.
[0074] The two-dimensional face key point mapping result corresponding to the target image is obtained by mapping the three-dimensional face model corresponding to the target image to the two-dimensional image space, so that the two-dimensional face key point mapping result can represent the state of the three-dimensional face model in the two-dimensional space, thereby enabling the two-dimensional face key point mapping result to represent the face state of the object in the target image in the two-dimensional space to a certain extent.
[0075] In addition, the present application does not limit the implementation of step 13 above, for example, it can be specifically: first map the three-dimensional face model corresponding to the target image into a two-dimensional image; then perform two-dimensional face key point detection processing on the two-dimensional image to obtain the two-dimensional face key point mapping result corresponding to the target image. It should be noted that the present application does not limit the acquisition method of the two-dimensional image, for example, it can be implemented by using any method that can map the three-dimensional face model back to the two-dimensional image.
[0076] Step 14: updating the three-dimensional face reconstruction result corresponding to the target image according to the difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, and returning to continue executing step 12 and subsequent steps above until a preset stopping condition is reached.
[0077] The difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image is used to represent the difference between the two-dimensional face key point mapping result and the two-dimensional face key point detection result, so that the difference representation data can represent the error of the three-dimensional face reconstruction result corresponding to the target image in the current round under the re-projection of the two-dimensional key points, thereby enabling the difference representation data to represent the accuracy of the three-dimensional face reconstruction result corresponding to the target image in the current round to a certain extent, and the present application does not limit the determination process of the difference representation data, for example, it can be implemented by using a loss function that can measure the difference between different two-dimensional face key points.
[0078] In addition, the present application does not limit the implementation of step 14 above, for example, it can be specifically: according to the difference between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, update the three-dimensional face reconstruction result corresponding to the target image, so that the error presented by the updated three-dimensional face reconstruction result in the re-projection of the two-dimensional key points is smaller than the error presented by the three-dimensional face reconstruction result before the update in the re-projection of the two-dimensional key points, and continue to execute step 12 above and the subsequent steps based on the updated three-dimensional face reconstruction result, to start the next round of process, and so on iteratively until the preset stopping condition is reached. Wherein, the preset stopping condition refers to the condition required to be reached at the end of the iteration cycle, such as the error presented by the three-dimensional face reconstruction result of the current round in the re-projection of the two-dimensional key points is lower than the preset error threshold and the like. It can be seen that in one possible implementation, the three-dimensional face reconstruction result can be iteratively updated by continuously reducing the error presented by the three-dimensional face reconstruction result corresponding to the target image in the re-projection of the two-dimensional key points.
[0079] It has been found through research that if only the error presented by the three-dimensional face reconstruction result in the re-projection of the two-dimensional key points is considered in the updating process, the following two defects may occur: ① When the two-dimensional face key point detection result of the target image is inaccurate, it will lead to inaccurate guiding information for updating, thereby causing the three-dimensional face reconstruction result corresponding to the target image to be updated in the wrong guiding direction, and further affecting the three-dimensional face reconstruction effect. ② When the error presented in the re-projection of the two-dimensional key points is blindly reduced, it may lead to poor face stability and continuity in three-dimensional space.
[0080] It has also been found through research that for the three-dimensional face parameter prediction result of the target image above, the three-dimensional face parameter prediction result can more accurately represent the face state of the object in the target image in the three-dimensional space, so the three-dimensional face reconstruction result corresponding to the target image can be determined in a change range not too far from the three-dimensional face parameter prediction result, which is conducive to improving the accuracy and efficiency.
[0081] Based on the findings shown in the above two paragraphs, in order to better improve the reconstruction effect, the present application also provides one possible implementation of step 14 above, in which the step 14 can be specifically: according to the difference between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, and the difference between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image, update the three-dimensional face reconstruction result corresponding to the target image, and return to continue to execute step 12 above and the subsequent steps until the preset stopping condition is reached.
[0082] In addition, for the difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image, the difference representation data is used to represent the difference between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result, so that the difference representation data can represent the relative distance between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result in the current round, so that the difference representation data can represent the change degree of the three-dimensional face reconstruction result relative to the three-dimensional face parameter prediction result in the current round to some extent, and then the subsequent can constrain the change of the three-dimensional face reconstruction result relative to the three-dimensional face parameter prediction result not to be too drastic by means of the difference representation data when updating, so as to improve the reconstruction accuracy.
[0083] In addition, in order to better improve the reconstruction effect, the above two difference representation data can satisfy the following constraint: the constraint strength of the difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image on updating is weaker than the constraint strength of the difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image on updating, so that the subsequent updating process based on the constraint can achieve better optimization effect. It should be noted that the application does not limit the implementation manner of the constraint, for example, it can control the constraint strength of the two difference representation data on updating through a hyperparameter.
[0084] Based on the above three paragraphs, in some application scenarios, for the three-dimensional face reconstruction result corresponding to the target image in the current round, the updating process of the three-dimensional face reconstruction result not only refers to the above-mentioned "difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image" influencing factor, but also refers to the above-mentioned "difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image" influencing factor, so that the gap between the updated three-dimensional face reconstruction result and the three-dimensional face parameter prediction result of the target image is smaller, and the error of the updated three-dimensional face reconstruction result in the re-projection of the two-dimensional key points is smaller than that of the three-dimensional face reconstruction result before updating. This can ensure that the optimization is carried out along the guidance direction of the two-dimensional face key point detection result without deviating too far from the three-dimensional face parameter prediction result, thereby facilitating the balance between accuracy and stability.
[0085] It should be noted that the present application does not limit the implementation of the preset stop condition corresponding to the updating process in the above paragraph, for example, the preset stop condition can specifically include that the loss of the three-dimensional face reconstruction result of the current round is lower than a preset loss threshold. For another example, the preset stop condition can specifically include that the change rate of the loss of the three-dimensional face reconstruction result of the current round is lower than a preset change rate threshold. For another example, the preset stop condition can specifically include that the number of updates of the three-dimensional face reconstruction result reaches a preset number threshold. Wherein, the loss is used to represent the performance of the three-dimensional face reconstruction result, such as accuracy + stability, etc.; and the loss is determined according to the difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, and the difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image.
[0086] In fact, in some application scenarios, such as video processing scenarios, in order to better improve the reconstruction effect, the present application also provides a possible implementation of the updating process of the three-dimensional face reconstruction result corresponding to the target image as described above. In this implementation, when the target image refers to a frame image in a reference video, and the reference video includes a previous frame image of the target image, the three-dimensional face reconstruction result corresponding to the target image can be updated according to the time sequence loss corresponding to the three-dimensional face reconstruction result and / or the time sequence loss corresponding to the two-dimensional face key point mapping result described above, so that the updating process of the three-dimensional face reconstruction result satisfies the time sequence constraint in the reference video.
[0087] For the target image and the previous frame image shown in the above paragraph, the following constraint is satisfied between them: the arrangement position of the target image in the reference video is adjacent to the arrangement position of the previous frame image in the reference video, and the arrangement position of the target image in the reference video is later than the arrangement position of the previous frame image in the reference video.
[0088] In addition, in order to avoid interference caused by action switching in the video, the reference video described above can be divided into multiple segments, so that each video segment is used to describe the change of an action, and different video segments are used to describe different actions, so that subsequent time sequence constraints can be performed on each video segment. Based on this, in a possible implementation, the following constraint can be satisfied between the target image and the previous frame image: the target image and the previous frame image come from the same video segment, the arrangement position of the target image in the video segment is adjacent to the arrangement position of the previous frame image in the video segment, and the arrangement position of the target image in the video segment is later than the arrangement position of the previous frame image in the video segment.
[0089] For the three-dimensional face reconstruction result corresponding to the target image pair, the temporal loss corresponding to the three-dimensional face reconstruction result is used to represent the state presented by the three-dimensional face reconstruction result in the temporal constraint; and the temporal loss corresponding to the three-dimensional face reconstruction result is determined according to the motion state representation data between the three-dimensional face parameter prediction result of the target image and the three-dimensional face parameter information of the previous frame image, and the motion state representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter information of the previous frame image. Wherein, the three-dimensional face parameter information of the previous frame image is used to describe the face state of the object in the three-dimensional space in the previous frame image; and the three-dimensional face parameter information of the previous frame image is determined according to the three-dimensional face parameter prediction result of the previous frame image and / or the three-dimensional face reconstruction result corresponding to the previous frame image, so that the three-dimensional face parameter information of the previous frame image can include the three-dimensional face parameter prediction result of the previous frame image, and / or the three-dimensional face reconstruction result corresponding to the previous frame image. Wherein, the three-dimensional face parameter prediction result of the previous frame image is obtained by performing three-dimensional face parameter prediction processing on the previous frame image; and the implementation of the three-dimensional face parameter prediction result of the previous frame image is similar to the implementation of the three-dimensional face parameter prediction result of the target image. The three-dimensional face reconstruction result corresponding to the previous frame image is obtained by fine-tuning the three-dimensional face parameter prediction result of the previous frame image according to the two-dimensional face key point detection result of the previous frame image; and the implementation of the three-dimensional face reconstruction result corresponding to the previous frame image is similar to the implementation of the three-dimensional face reconstruction result corresponding to the target image. Wherein, the two-dimensional face key point detection result of the previous frame image is obtained by performing two-dimensional face key point detection processing on the previous frame image; and the implementation of the two-dimensional face key point detection result of the previous frame image is similar to the implementation of the two-dimensional face key point detection result of the target image.
[0090] In addition, the motion state representation data between the three-dimensional face parameter prediction result of the target image and the three-dimensional face parameter information of the previous frame image is used to represent the motion state of the object represented by the three-dimensional face parameter prediction result relative to the object in the three-dimensional space in the previous frame image, such as speed, acceleration, and the like, so that the motion state representation data can represent the motion state of the object in the target image relative to the object in the previous frame image in the three-dimensional space to some extent, so that the motion state representation data can represent the time sequence constraint required to be met by the three-dimensional face reconstruction result corresponding to the target image to some extent; and the application does not limit the implementation of the motion state representation data, for example, it can include speed and / or acceleration, and the like. In addition, the application also does not limit the acquisition method of the motion state representation data, for example, it can be implemented by using any method that can measure the motion state between two frames of data.
[0091] In addition, the motion state representation data between the three-dimensional face parameter prediction result of the target image and the three-dimensional face parameter information of the previous frame image is used to represent the motion state of the object represented by the three-dimensional face parameter prediction result relative to the object in the three-dimensional space in the previous frame image, such as speed, acceleration, and the like, so that the motion state representation data can represent the motion state of the object in the target image relative to the object in the previous frame image in the three-dimensional space to some extent, so that the motion state representation data can represent the time sequence constraint required to be met by the three-dimensional face reconstruction result corresponding to the target image to some extent; and the application does not limit the implementation of the motion state representation data, for example, it can include speed and / or acceleration, and the like. In addition, the application also does not limit the acquisition method of the motion state representation data, for example, it can be implemented by using any method that can measure the motion state between two frames of data.
[0092] In addition, the application does not limit the determination method of the time sequence loss corresponding to the three-dimensional face reconstruction result, for example, it can be specifically: first, calculate the difference representation data between the motion state representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter information of the previous frame image and the motion state representation data between the three-dimensional face parameter prediction result of the target image and the three-dimensional face parameter information of the previous frame image, so that the difference representation data can represent the state of the three-dimensional face reconstruction result in terms of time sequence constraint, such as constraint satisfaction degree or whether the constraint is met, and the like; and then determine the time sequence loss corresponding to the three-dimensional face reconstruction result according to the difference representation data.
[0093] For the two-dimensional face key point mapping result corresponding to the target image, the time sequence loss corresponding to the two-dimensional face key point mapping result is used to represent the state of the two-dimensional face key point mapping result in the time sequence constraint; and the time sequence loss corresponding to the two-dimensional face key point mapping result is determined according to the motion state representation data between the two-dimensional face key point detection result of the target image and the two-dimensional face key point detection result of the previous frame image, and the motion state representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the previous frame image. Wherein, the two-dimensional face key point detection result of the previous frame image is used to describe the face state of the object in the two-dimensional space in the previous frame image.
[0094] In addition, for the motion state representation data between the two-dimensional face key point detection result of the target image and the two-dimensional face key point detection result of the previous frame image, the motion state representation data is used to represent the motion state of the object represented by the two-dimensional face key point detection result relative to the object in the two-dimensional space in the previous frame image, such as speed, acceleration and the like, so that the motion state representation data can represent the motion state of the object in the target image relative to the object in the previous frame image in the two-dimensional space to a certain extent, so that the motion state representation data can represent the time sequence constraint required to be met by the two-dimensional face key point mapping result corresponding to the target image to a certain extent; and the application does not limit the implementation of the motion state representation data, for example, it can include speed and / or acceleration and the like. In addition, the application does not limit the acquisition method of the motion state representation data, for example, it can be implemented by using any method that can measure the motion state between two frames of data existing at present or appearing in the future.
[0095] In addition, for the motion state representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the previous frame image, the motion state representation data is used to represent the motion state of the object represented by the two-dimensional face key point mapping result relative to the object in the two-dimensional space in the previous frame image, such as speed, acceleration and the like; and the application does not limit the implementation of the motion state representation data, for example, it can include speed and / or acceleration and the like. In addition, the application does not limit the acquisition method of the motion state representation data, for example, it can be implemented by using any method that can measure the motion state between two frames of data existing at present or appearing in the future.
[0096] Also, the application does not limit the determination method of the time sequence loss corresponding to the two-dimensional face key point mapping result above, for example, it can be specifically: first, calculate the difference representation data between the motion state representation data between the two-dimensional face key point mapping result of the target image and the two-dimensional face key point detection result of the previous frame image and the motion state representation data between the two-dimensional face key point detection result of the target image and the two-dimensional face key point detection result of the previous frame image, so that the difference representation data can represent the state of the two-dimensional face key point mapping result in the time sequence constraint, such as the constraint satisfaction degree or whether the constraint is satisfied; then, according to the difference representation data, determine the time sequence loss corresponding to the two-dimensional face key point mapping result.
[0097] Based on the related content of the time sequence constraint above, in one possible implementation, step 14 above can be: according to the difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, the difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image, the time sequence loss corresponding to the three-dimensional face reconstruction result, and the time sequence loss corresponding to the two-dimensional face key point mapping result, update the three-dimensional face reconstruction result corresponding to the target image, so that the error of the updated three-dimensional face reconstruction result in the re-projection of the two-dimensional key point is smaller than the error of the three-dimensional face reconstruction result before updating in the re-projection of the two-dimensional key point, and the updated three-dimensional face reconstruction result and the three-dimensional face parameter prediction result of the target image are smaller, and the updated three-dimensional face reconstruction result presents smaller loss in the time sequence constraint, so that the updated three-dimensional face reconstruction result can better represent the face state of the object in the target image, so that the subsequent can continue to execute step 12 and the subsequent steps based on the updated three-dimensional face reconstruction result. Iterative loop until the preset stopping condition is reached.
[0098] It should be noted that the present application does not limit the implementation of the preset stop condition in the above paragraph, for example, the preset stop condition can specifically include that the loss of the three-dimensional face reconstruction result of the current round is lower than a preset loss threshold. For another example, the preset stop condition can specifically include that the change rate of the loss of the three-dimensional face reconstruction result of the current round is lower than a preset change rate threshold. For another example, the preset stop condition can specifically include that the number of updates of the three-dimensional face reconstruction result reaches a preset number threshold. Wherein, the loss can be determined according to the difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, the difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image, the time sequence loss corresponding to the three-dimensional face reconstruction result, and the time sequence loss corresponding to the two-dimensional face key point mapping result.
[0099] In fact, in some application scenarios, such as the face identification parameter, the face expression parameter and the face posture parameter are decoupled from each other, in order to better improve the reconstruction effect, the present application also provides an implementation of the above step of "updating the three-dimensional face reconstruction result corresponding to the target image", in which implementation, the step can be specifically updating the parameters other than the face identification parameter in the three-dimensional face reconstruction result, such as the face expression parameter and the face posture parameter. It can be seen that for the three-dimensional face reconstruction result corresponding to the target image in the current round, the update process for the three-dimensional face reconstruction result can be completed by fixing the face identification parameter in the three-dimensional face reconstruction result and updating the face expression parameter and the face posture parameter in the three-dimensional face reconstruction result, which can effectively avoid the interference of the update process on the face identification parameter, thereby facilitating the improvement of the reconstruction effect.
[0100] Based on the related content of steps 11 to 14 above, in some application scenarios, for the three-dimensional face reconstruction result corresponding to the target image in the current round, after obtaining the three-dimensional face reconstruction result, a three-dimensional face model can be first constructed according to the three-dimensional face reconstruction result; then the three-dimensional face model is mapped into the two-dimensional face key point mapping result corresponding to the target image; then, according to the difference representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the target image, the difference representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter prediction result of the target image, the time sequence loss corresponding to the three-dimensional face reconstruction result, and the time sequence loss corresponding to the two-dimensional face key point mapping result, the loss of the three-dimensional face reconstruction result is determined, so that the loss can represent the performance, such as accuracy + stability, of the three-dimensional face reconstruction result corresponding to the target image in the current round; finally, under the premise of fixing the face identity parameter in the three-dimensional face reconstruction result, the face expression parameter and the face pose parameter in the three-dimensional face reconstruction result are updated according to the loss of the three-dimensional face reconstruction result, to obtain an updated three-dimensional face reconstruction result, so that the face identity parameter in the updated three-dimensional face reconstruction result remains consistent with the face identity parameter in the three-dimensional face reconstruction result before updating, and the updated three-dimensional face reconstruction result includes the updated face expression parameter and the updated face pose parameter, so that the updated three-dimensional face reconstruction result has better performance, so that subsequent steps 12 and subsequent steps can be continued based on the updated three-dimensional face reconstruction result, and the iteration is repeated until the preset stopping condition is reached, so that the final obtained three-dimensional face reconstruction result can more accurately represent the face state of the object in the target image in the three-dimensional space, thereby improving the face reconstruction effect. Wherein, the initial value of the three-dimensional face reconstruction result is determined according to the three-dimensional face parameter prediction result of the target image, so that the initial value of the three-dimensional face reconstruction result can more accurately represent the face state of the object in the target image, thereby enabling subsequent rapid and stable convergence based on the initial value, thereby improving the reconstruction effect.
[0101] In fact, in some application scenarios, such as the face identity parameter, the face expression parameter, and the face pose parameter are decoupled from each other, in order to better improve the reconstruction effect, the application further provides an acquisition manner of the face identity parameter in the three-dimensional face reconstruction result corresponding to the target image. In this manner, when the target image refers to any one frame image in the reference video, the face identity parameter in the three-dimensional face reconstruction result corresponding to the target image can be determined according to the average value of the face identity parameters in the three-dimensional face parameter prediction results of at least two frame images in the reference video. Among them, the at least two frame images include the target image; and the application does not limit the implementation of the at least two frame images, for example, the at least two frame images can refer to all images in the reference video. For another example, after the reference video is divided into a plurality of video segments, the at least two frame images can refer to all images in the video segment including the target image, so that the at least two frame images can better describe the face features of the object in the target image.
[0102] It can be seen that in a possible implementation, when the above target image refers to any one frame image in the reference video, such as the i-th frame image, the acquisition process of the face identity parameter in the three-dimensional face reconstruction result corresponding to the target image can include: first, acquiring the three-dimensional face parameter prediction result of each frame image in the reference video; and then calculating the average value of the face identity parameters in the three-dimensional face parameter prediction results of all images in the reference video as the face identity parameter in the three-dimensional face reconstruction result corresponding to the target image, so that the three-dimensional face reconstruction result corresponding to the target image can more accurately represent the face features of the object in the target image, such as the face contour, the distribution of the facial features, and the like. This is conducive to improving the face reconstruction effect. Among them, i is a positive integer, i≤I, I is a positive integer, and I represents the total number of frames of images in the reference video.
[0103] Based on the related content of S1 to S4 above, it can be known that, for the face reconstruction method provided by the embodiments of the present application, after the target image is obtained, firstly, the target image is subjected to two-dimensional face key point detection processing to obtain a two-dimensional face key point detection result of the target image, so that the two-dimensional face key point detection result can represent the face state of the object in the target image in a two-dimensional space, and the target image is subjected to three-dimensional face parameter prediction processing to obtain a three-dimensional face parameter prediction result of the target image, so that the three-dimensional face parameter prediction result can more accurately represent the face state of the object in a three-dimensional space; then, the three-dimensional face parameter prediction result is subjected to fine-tuning processing according to the two-dimensional face key point detection result to obtain a three-dimensional face reconstruction result corresponding to the target image, so that the three-dimensional face reconstruction result can more accurately represent the face state of the object in the three-dimensional space, which is beneficial to improve the face reconstruction effect. Among them, because the three-dimensional face parameter prediction result can more accurately represent the face state of the object in the three-dimensional space, so that the fine-tuning processing based on the three-dimensional face parameter prediction result can quickly and stably reach convergence, so that the three-dimensional face reconstruction result obtained by fine-tuning processing can more accurately represent the face state of the object in the three-dimensional space, which is beneficial to improve the face reconstruction effect, such as accuracy, stability and efficiency.
[0104] In addition, the present application does not limit the execution subject of the face reconstruction method provided by the embodiments of the present application. For example, the face reconstruction method provided by the embodiments of the present application can be applied to a terminal device or a server. For another example, the face reconstruction method provided by the embodiments of the present application can also be implemented by means of data interaction process between the terminal device and the server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server or a cloud server.
[0105] In addition, the present application does not limit the application scenario of the face reconstruction method provided by the embodiments of the present application. In order to facilitate understanding, the following will be described in conjunction with examples.
[0106] As an example, when the face reconstruction method provided by the present application is applied to a certain video processing task, such as a video audio adjustment task, the face reconstruction method can include the following steps 21-25.
[0107] Step 21: obtaining a target image, the target image being any one frame image in a reference video.
[0108] In the present application, in some application scenarios, such as a certain video processing scenario, after a reference video is obtained, the i-th image in the reference video can be regarded as a target image, so that the three-dimensional face reconstruction result corresponding to the i-th image can be determined by means of the related processing procedure of the target image. Wherein, i is a positive integer, i≤I, I is a positive integer, I represents the total number of frames of images in the reference video.
[0109] In addition, for the related content of step 21, please refer to the related content of S1 above.
[0110] Step 22: performing two-dimensional face key point detection processing on the target image to obtain a two-dimensional face key point detection result of the target image.
[0111] It should be noted that the related content of step 22 can be referred to the related content of S2 above.
[0112] Step 23: performing three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image.
[0113] It should be noted that the related content of step 23 can be referred to the related content of S3 above.
[0114] Step 24: performing fine-tuning processing on the three-dimensional face parameter prediction result of the target image according to the two-dimensional face key point detection result of the target image to obtain a three-dimensional face reconstruction result corresponding to the target image.
[0115] It should be noted that the related content of step 24 can be referred to the related content of S4 above.
[0116] Step 25: generating a video corresponding to an audio sequence according to the three-dimensional face reconstruction results corresponding to each frame of image in the reference video, the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence is used to describe the face state change of the object under the audio sequence.
[0117] Wherein, the audio sequence refers to the audio required for reference when processing the reference video, such as mouth shape adjustment processing; and the present application does not limit the implementation of the audio sequence, for example, in some application scenarios, the audio sequence satisfies the following constraint: the total number of frames in the audio sequence is consistent with the total number of frames in the reference video.
[0118] In addition, for the above audio sequence, the video corresponding to the audio sequence refers to the video generated for the audio sequence, so that the video satisfies the following constraint: the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence is used to describe the face state change of the object under the audio sequence.
[0119] Further, the present application does not limit the implementation of the above step 25, for example, it can be implemented by means of any one of existing or future methods capable of generating a video based on some three-dimensional face reconstruction results and an audio sequence. As another example, in some application scenarios, in order to better improve the video generation effect, the present application also provides a possible implementation of the step 25, in which the step 25 can specifically include: generating, by a pre-constructed video generation model, a video corresponding to the audio sequence according to the three-dimensional face reconstruction results corresponding to all images in the reference video, so that the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence is used to describe the face state change of the object under the audio sequence. Wherein, the video generation model refers to a pre-constructed model with video generation function, such as a machine learning model, etc., and the present application does not limit the implementation of the video generation model.
[0120] Based on the above steps 21 to 25, in some application scenarios, such as video processing scenarios, after obtaining a reference video, first, the two-dimensional face key point detection processing is performed on the i-th frame image in the reference video to obtain the two-dimensional face key point detection result of the target image, and the three-dimensional face parameter prediction processing is performed on the i-th frame image to obtain the three-dimensional face parameter prediction result of the target image, i is a positive integer, i≤I, I is a positive integer, and I represents the total number of frames of images in the reference video; then, the three-dimensional face parameter prediction result of the i-th frame image is fine-tuned according to the two-dimensional face key point detection result of the i-th frame image to obtain the three-dimensional face reconstruction result corresponding to the i-th frame image, i is a positive integer, i≤I, I is a positive integer, and I represents the total number of frames of images in the reference video; finally, a video corresponding to an audio sequence is generated according to the three-dimensional face reconstruction results corresponding to all images in the reference video, so that the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence is used to describe the face state change of the object under the audio sequence, so as to realize the adjustment of the face state change of the object in the reference video based on the audio sequence. Wherein, since these three-dimensional face reconstruction results can accurately represent the face state of the object in the reference video, such as face features, expression features similar to large expression amplitude, etc. presented when speaking, so that the video generated based on these three-dimensional face reconstruction results can better represent the face state change of the object under the audio sequence, so as to improve the video generation effect.
[0121] Based on the face reconstruction method provided in the embodiments of the present application, the embodiments of the present application further provide a face reconstruction device, which is explained and described below in combination with FIG. 4. FIG. 4 is a structural schematic diagram of a face reconstruction device provided in the embodiments of the present application. It should be noted that the technical details of the face reconstruction device provided in the embodiments of the present application refer to the related content of the face reconstruction method described above.
[0122] As shown in FIG. 4, the face reconstruction device 400 provided in the embodiments of the present application includes:
[0123] The acquisition unit 401 is configured to acquire a target image.
[0124] The processing unit 402 is configured to perform two-dimensional face key point detection processing on the target image to obtain a two-dimensional face key point detection result of the target image, and perform three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image.
[0125] The fine-tuning unit 403 is configured to perform fine-tuning processing on the three-dimensional face parameter prediction result according to the two-dimensional face key point detection result to obtain a three-dimensional face reconstruction result corresponding to the target image.
[0126] In a possible implementation, the fine-tuning unit 403 is specifically configured to: perform initialization processing on the three-dimensional face reconstruction result according to the three-dimensional face parameter prediction result; construct a three-dimensional face model corresponding to the target image according to the three-dimensional face reconstruction result; map the three-dimensional face model to a two-dimensional image space to obtain a two-dimensional face key point mapping result corresponding to the target image; and update the three-dimensional face reconstruction result according to difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result.
[0127] In a possible implementation, the fine-tuning unit 403 is specifically configured to: update the three-dimensional face reconstruction result according to difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result, and difference representation data between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result.
[0128] In a possible implementation, the constraint intensity of the difference representation data between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result on the update is weaker than the constraint intensity of the difference representation data between the two-dimensional face key point mapping result and the two-dimensional face key point detection result on the update.
[0129] In a possible implementation, the target image refers to a frame image in a reference video, and the reference video includes a previous frame image of the target image.
[0130] The three-dimensional face reconstruction result is updated according to the time sequence loss corresponding to the three-dimensional face reconstruction result and / or the time sequence loss corresponding to the two-dimensional face key point mapping result.
[0131] The time sequence loss corresponding to the three-dimensional face reconstruction result is determined according to motion state representation data between the three-dimensional face parameter prediction result of the target image and the three-dimensional face parameter information of the previous frame image, and motion state representation data between the three-dimensional face reconstruction result corresponding to the target image and the three-dimensional face parameter information of the previous frame image; the three-dimensional face parameter information of the previous frame image is determined according to the three-dimensional face parameter prediction result of the previous frame image and / or the three-dimensional face reconstruction result corresponding to the previous frame image.
[0132] The time sequence loss corresponding to the two-dimensional face key point mapping result is determined according to motion state representation data between the two-dimensional face key point detection result of the target image and the two-dimensional face key point detection result of the previous frame image, and motion state representation data between the two-dimensional face key point mapping result corresponding to the target image and the two-dimensional face key point detection result of the previous frame image.
[0133] In a possible implementation, the fine-tuning unit 403 is specifically configured to update parameters other than the face identification parameter in the three-dimensional face reconstruction result.
[0134] In a possible implementation, the target image refers to any one frame image in a reference video; the face identification parameter in the three-dimensional face reconstruction result corresponding to the target image is determined according to an average value of face identification parameters in three-dimensional face parameter prediction results of at least two frame images in the reference video, and the at least two frame images include the target image.
[0135] In a possible implementation, the target image refers to any one frame image in a reference video;
[0136] The face reconstruction apparatus 400 further includes:
[0137] The generating unit is configured to generate a video corresponding to an audio sequence according to the three-dimensional face reconstruction result corresponding to each frame image in the reference video, the object presented in the video corresponding to the audio sequence being consistent with the object presented in the reference video, and the video corresponding to the audio sequence being used to describe face state changes of the object under the audio sequence.
[0138] Based on the related content of the face reconstruction apparatus 400, the working principle of the face reconstruction apparatus 400 provided in the application can include: after obtaining a target image, first, performing two-dimensional face key point detection processing on the target image to obtain a two-dimensional face key point detection result of the target image, so that the two-dimensional face key point detection result can represent the face state of the object in the target image in a two-dimensional space, and performing three-dimensional face parameter prediction processing on the target image to obtain a three-dimensional face parameter prediction result of the target image, so that the three-dimensional face parameter prediction result can more accurately represent the face state of the object in a three-dimensional space; then, according to the two-dimensional face key point detection result, fine-tuning processing is performed on the three-dimensional face parameter prediction result to obtain a three-dimensional face reconstruction result corresponding to the target image, so that the three-dimensional face reconstruction result can more accurately represent the face state of the object in a three-dimensional space, which is beneficial to improve the face reconstruction effect. Among them, because the three-dimensional face parameter prediction result can more accurately represent the face state of the object in a three-dimensional space, so that the fine-tuning processing based on the three-dimensional face parameter prediction result can quickly and stably converge, so that the three-dimensional face reconstruction result obtained by fine-tuning processing can more accurately represent the face state of the object in a three-dimensional space, which is beneficial to improve the face reconstruction effect, such as accuracy, stability and efficiency.
[0139] In addition, the embodiment of the application further provides an electronic device, the device includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any embodiment of the face reconstruction method provided by the embodiment of the application.
[0140] Referring to FIG. 5, a structural schematic diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in FIG. 5 is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0141] As shown in FIG. 5, the electronic device 500 can include a processing apparatus (e.g., a central processing unit, a graphics processing unit, etc.) 501 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage apparatus 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing apparatus 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0142] Generally, the following apparatuses can be connected to the I / O interface 505: input apparatuses 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output apparatuses 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage apparatuses 508 including, for example, a magnetic tape, a hard disk, etc.; and communication apparatuses 509. The communication apparatuses 509 can allow the electronic device 500 to perform wireless or wired communication with other devices to exchange data. Although FIG. 5 shows the electronic device 500 with various apparatuses, it should be understood that all of the illustrated apparatuses are not required to be implemented or possessed. More or fewer apparatuses can be alternatively implemented or possessed.
[0143] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication apparatuses 509, or installed from the storage apparatuses 508, or installed from the ROM 502. When the computer program is executed by the processing apparatus 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0144] The electronic device provided by the embodiments of the present disclosure and the method provided by the above-mentioned embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiments can be referred to the above-mentioned embodiments, and the present embodiments have the same beneficial effects as the above-mentioned embodiments.
[0145] The embodiments of the present application also provide a computer readable medium, wherein instructions or computer programs are stored in the computer readable medium, and when the instructions or computer programs are run on a device, the device is caused to perform any of the embodiments of the face reconstruction method provided by the embodiments of the present application.
[0146] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as a carrier wave in a propagated data signal, which bears computer-readable program code. Such a propagated data signal can take many forms, including but not limited to electro-magnetic, optical or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer a program for use by or in connection with an instruction execution system, apparatus or device. Program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0147] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0148] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.
[0149] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device can execute the method described above.
[0150] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0151] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0152] The units involved in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit / module does not constitute a limitation on the unit itself.
[0153] The functions described in the above description can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0154] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0155] It should be noted that the various embodiments described in the specification are progressive and each embodiment focuses on the differences from other embodiments. The same and similar parts between embodiments can be mutually referred to. For the system or device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0156] It should be understood that in this application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0157] It is also to be noted that, as used in the specification and the appended claims, the singular forms "a," "an" and "the" include plural referents unless otherwise indicated. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or "contains" are used in either the detailed description and the claims, such terms are intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or other elements.
[0158] The embodiments disclosed herein can each be implemented as a method, apparatus, or article of manufacture using programming instructions. The embodiments disclosed herein can be implemented using software, firmware, hardware, or a combination thereof. The various elements of the disclosed embodiments, as well as the procedural aspects of the disclosed embodiments, can be implemented using a variety of programming instructions, software, firmware, or other programming instructions. In one embodiment, programming instructions are distributed via a computer medium, such as a compact disc, diskette, tape, file, or other computer medium. In another embodiment, programming instructions are downloaded into a computer from a network connection, such as the Internet, a local area network, a wide area network, or other network connection.
[0159] The above description of disclosed embodiments provides enough information to enable those with ordinary skill in the art to make and use the application. Various modifications to these embodiments will be readily apparent to those with ordinary skill in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Accordingly, the application is not to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A facial reconstruction method, wherein: The method comprises: Acquire the target image; Performing a two-dimensional facial key point detection process on the target image to obtain a two-dimensional facial key point detection result of the target image, and performing a three-dimensional facial parameter prediction process on the target image to obtain a three-dimensional facial parameter prediction result of the target image; The three-dimensional facial parameter prediction result is fine-tuned based on the two-dimensional facial key point detection result to obtain a three-dimensional facial reconstruction result corresponding to the target image.
2. The method according to claim 1, wherein The fine-tuning process includes: Initializing the three-dimensional face reconstruction result according to the three-dimensional face parameter prediction result; constructing a three-dimensional face model corresponding to the target image based on the three-dimensional face reconstruction result; Mapping the three-dimensional facial model to a two-dimensional image space to obtain a two-dimensional facial key point mapping result corresponding to the target image; The three-dimensional face reconstruction result is updated according to the difference representation data between the two-dimensional facial key point mapping result and the two-dimensional facial key point detection result.
3. The method according to claim 2, wherein: The updating of the three-dimensional face reconstruction result based on the difference representation data between the two-dimensional facial key point mapping result and the two-dimensional facial key point detection result includes: The three-dimensional face reconstruction result is updated according to the difference representation data between the two-dimensional facial landmark mapping result and the two-dimensional facial landmark detection result, and the difference representation data between the three-dimensional face reconstruction result and the three-dimensional face parameter prediction result.
4. The method according to claim 3, wherein: The constraint strength of the update by the difference representation data between the three-dimensional facial reconstruction result and the three-dimensional facial parameter prediction result is weaker than the constraint strength of the update by the difference representation data between the two-dimensional facial key point mapping result and the two-dimensional facial key point detection result.
5. The method according to claim 2, wherein: The target image refers to a frame image in a reference video, and the reference video includes a frame image previous to the target image; The three-dimensional face reconstruction result is updated according to the temporal loss corresponding to the three-dimensional face reconstruction result and / or the temporal loss corresponding to the two-dimensional facial key point mapping result; The temporal loss corresponding to the 3D facial reconstruction result is determined based on motion state representation data between the 3D facial parameter prediction result of the target image and the 3D facial parameter information of the previous frame image, and motion state representation data between the 3D facial reconstruction result corresponding to the target image and the 3D facial parameter information of the previous frame image; the 3D facial parameter information of the previous frame image is determined based on the 3D facial parameter prediction result of the previous frame image and / or the 3D facial reconstruction result corresponding to the previous frame image; The temporal loss corresponding to the two-dimensional facial key point mapping result is determined based on the motion state representation data between the two-dimensional facial key point detection result of the target image and the two-dimensional facial key point detection result of the previous frame image, and the motion state representation data between the two-dimensional facial key point mapping result corresponding to the target image and the two-dimensional facial key point detection result of the previous frame image.
6. The method according to any one of claims 2 to 5, wherein: The updating of the three-dimensional face reconstruction result includes: Update other parameters in the three-dimensional face reconstruction result except the face identification parameter.
7. The method according to claim 1, wherein The target image refers to any frame image in the reference video; The facial identification parameters in the three-dimensional facial reconstruction result corresponding to the target image are determined based on an average value of the facial identification parameters in the three-dimensional facial parameter prediction results of at least two frames of images in the reference video, where the at least two frames of images include the target image.
8. The method according to claim 1, wherein The target image refers to any frame image in the reference video; The method further comprises: Based on the three-dimensional facial reconstruction results corresponding to each frame image in the reference video, a video corresponding to the audio sequence is generated, where the object presented in the video corresponding to the audio sequence is consistent with the object presented in the reference video, and the video corresponding to the audio sequence is used to describe the changes in the facial state of the object under the audio sequence.
9. A facial reconstruction device, wherein: include: an acquisition unit, configured to acquire a target image; a processing unit, configured to perform two-dimensional facial key point detection processing on the target image to obtain a two-dimensional facial key point detection result of the target image, and perform three-dimensional facial parameter prediction processing on the target image to obtain a three-dimensional facial parameter prediction result of the target image; A fine-tuning unit is used to fine-tune the three-dimensional facial parameter prediction result based on the two-dimensional facial key point detection result to obtain a three-dimensional facial reconstruction result corresponding to the target image.
10. An electronic device, wherein: The device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 8.
11. A computer-readable medium, wherein: The computer-readable medium stores instructions or a computer program, and when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 8.
12. A computer program product, wherein: The method comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Face key point detection method and device, computer equipment and storage medium
CN109657583A
Three-dimensional face image correction method and device, electronic device and storage medium
CN110533777A
Three-dimensional Face Reconstruction Method
CN111223175A
Three-dimensional face reconstruction method, electronic equipment and computer readable storage medium
CN114067059A
Three-dimensional face fitting method and device, electronic equipment and storage medium
CN117593493A