A method, apparatus, device and medium for improving expression base expression capability
By automatically adjusting the semantic parameters and residual data of the facial expression parameterization model, the problems of large memory consumption and insufficient expression accuracy in traditional facial animation are solved, and efficient expression base expression effect is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NETEASE (HANGZHOU) NETWORK CO LTD
- Filing Date
- 2022-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
In traditional facial animation, the method of using multi-view 3D reconstruction to directly drive facial expressions consumes a lot of memory and is not easy to edit. The general Blendshape model is not accurate enough to express different faces. Existing technologies require additional manual operation to improve accuracy, which is inefficient.
By extracting semantic parameters from 3D reconstructed data based on the initial expression parameterization model, the initial parameterization result is generated. Then, the expression residual data is generated through error analysis to adjust the model and automatically optimize the expression basis to improve the expression accuracy.
Without requiring external human input data, it improves the expressive effect of facial expressions, reduces labor costs, increases processing efficiency, and achieves higher expressive accuracy.
Smart Images

Figure CN115994963B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer animation technology, and more specifically, to a method, apparatus, device, and medium for enhancing the expressive power of facial expressions. Background Technology
[0002] Traditional facial animation uses multi-view 3D reconstructed facial expressions to directly drive the facial model. This method consumes a lot of memory and is inconvenient for animators to edit. With the rapid development of computer technology, the reconstructed facial expression source data is transformed into a parameterized expression sequence to reduce memory usage.
[0003] Currently, the general-purpose Blendshape model is widely used for parameterizing reconstructed facial expressions. However, using this model to obtain parameters for reconstructed facial expressions can lead to insufficient accuracy in representing different faces. To improve the accuracy of the Blendshape model, existing techniques primarily involve introducing externally acquired specific facial expressions. These acquisition methods mainly include filming specific actors or manually depicting them. All of these methods require additional manual intervention, resulting in low efficiency. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method, apparatus, device and medium for improving the expressive power of facial expressions, which can automatically improve the expressive effect of facial expressions without the need for manual external input data, thereby reducing labor costs and improving processing efficiency.
[0005] In a first aspect, embodiments of this application provide a method for enhancing the expressive power of facial expressions, the method comprising:
[0006] The initial semantic parameters of each frame of 3D reconstruction data in a series of consecutive frames of 3D reconstruction data are extracted based on the initial expression parameterization model, and the initial parameterization result is generated by driving the preset facial model through the first expression base corresponding to the initial semantic parameters.
[0007] The initial parameterization results under each first expression basis are compared with the 3D reconstruction data to determine whether the initial parameterization results under the first expression basis and the 3D reconstruction data match.
[0008] If there is a mismatch, then based on the initial parameterization results under the first expression basis and the error information of the 3D reconstruction data, expression residual data is generated;
[0009] The initial expression parameterization model is adjusted using the expression residual data, and a second semantic parameter vector is extracted from each frame of the three-dimensional reconstruction data in the continuous multi-frame three-dimensional reconstruction data based on the adjusted second parameterization model. The second expression basis corresponding to the second semantic parameter vector drives the facial model to generate a second parameterization result that matches the three-dimensional reconstruction data.
[0010] Secondly, embodiments of this application provide an apparatus for enhancing the expressive power of facial expressions, the apparatus comprising:
[0011] The extraction module is used to extract the initial semantic parameters of each frame of 3D reconstruction data in a series of consecutive frames of 3D reconstruction data based on the initial expression parameterization model, and to drive the preset facial model to generate the initial parameterization result through the first expression base corresponding to the initial semantic parameters.
[0012] The error analysis module is used to compare the initial parameterization results and the 3D reconstruction data under each first expression basis to determine whether the initial parameterization results and the 3D reconstruction data under the first expression basis match.
[0013] The generation module is used to generate expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis if there is a mismatch.
[0014] The adjustment module is used to adjust the initial expression parameterization model using the expression residual data, and extract the second semantic parameter vector of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data based on the adjusted second parameterization model, and drive the facial model to generate a second parameterization result that matches the 3D reconstruction data through the second expression basis corresponding to the second semantic parameter vector.
[0015] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above for improving the expressive power of facial expressions.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method for improving the expressive power of facial expressions described above.
[0017] The technical solutions provided by the embodiments of this application may include the following beneficial effects:
[0018] The method of this application includes: extracting initial semantic parameters of each frame of 3D reconstruction data in a series of consecutive frames of 3D reconstruction data based on an initial expression parameterization model, and driving a preset facial model to generate an initial parameterization result through a first expression basis corresponding to the initial semantic parameters; comparing the initial parameterization result under each first expression basis with the 3D reconstruction data to determine whether the initial parameterization result under the first expression basis and the 3D reconstruction data match; if they do not match, generating expression residual data based on the error information between the initial parameterization result under the first expression basis and the 3D reconstruction data; adjusting the initial expression parameterization model using the expression residual data, and extracting a second semantic parameter vector of each frame of 3D reconstruction data in the series of consecutive frames of 3D reconstruction data based on the adjusted second parameterization model, and driving the facial model to generate a second parameterization result that matches the 3D reconstruction data through a second expression basis corresponding to the second semantic parameter vector. This application can analyze the first expression basis obtained through the initial expression parameterization model, determine the error information of the initial parameterization result and the 3D reconstruction data under the first expression basis, and then generate expression residual data; use the expression residual data to adjust the initial expression parameterization model to obtain the second parameterization model; the second expression basis with better expression effect can be obtained through the second parameterization model; no manual input of external data is required, thus improving efficiency.
[0019] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating a method for enhancing the expressive power of facial features according to an embodiment of this application is shown.
[0022] Figure 2 This illustration shows a schematic diagram of a first key point provided in an embodiment of this application;
[0023] Figure 3 A schematic diagram of an apparatus for enhancing the expressive power of facial features, provided in an embodiment of this application, is shown.
[0024] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0026] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0027] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0028] Traditional facial animation uses multi-view 3D reconstructed facial expressions to directly drive the facial model. This method consumes a lot of memory and is inconvenient for animators to edit. With the rapid development of computer technology, the reconstructed facial expression source data is transformed into a parameterized expression sequence to reduce memory usage.
[0029] Currently, the general-purpose Blendshape model is widely used for parametric transformation of reconstructed facial expressions. However, using this model to obtain parameters for reconstructed facial expressions can lead to insufficient accuracy across different faces. To improve the accuracy of the Blendshape model, existing techniques primarily involve introducing externally acquired specific facial expressions. These acquisition methods mainly include filming with specific actors or manually depicting them. While existing techniques can improve the expressiveness of Blendshape through artists manually adding correction terms and actors providing specific expressions, these methods require a significant amount of manual work. In practical applications, when the input consists of a large number of high-precision 3D facial expression sequences, ensuring the accuracy of the parametric results requires substantial manual correction. Furthermore, optimizing Blendshape with specific expressions cannot guarantee that the accuracy improvement from the acquired expressions will cover all input frames.
[0030] Based on this, embodiments of this application provide a method, apparatus, device, and medium for enhancing the expressive power of facial expressions, which are described below through embodiments.
[0031] Figure 1 The diagram illustrates a flowchart of a method for enhancing the expressive power of facial features according to an embodiment of this application, wherein the method includes steps S101-S104; specifically:
[0032] S101. Extract the initial semantic parameters of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data based on the initial expression parameterization model, and drive the preset facial model to generate the initial parameterization result through the first expression base corresponding to the initial semantic parameters.
[0033] S102. Compare the initial parameterization results and 3D reconstruction data under each first expression basis to determine whether the initial parameterization results and 3D reconstruction data under the first expression basis match.
[0034] S103. If there is no match, then based on the error information of the initial parameterization result and the three-dimensional reconstruction data under the first expression basis, expression residual data is generated.
[0035] S104. The initial expression parameterization model is adjusted using the expression residual data, and the second semantic parameter vector of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data is extracted based on the adjusted second parameterization model. The second expression basis corresponding to the second semantic parameter vector drives the facial model to generate a second parameterization result that matches the 3D reconstruction data.
[0036] This application can analyze the first expression basis obtained through the initial expression parameterization model, determine the error information of the initial parameterization result and the 3D reconstruction data under the first expression basis, and then generate expression residual data; use the expression residual data to adjust the initial expression parameterization model to obtain the second parameterization model; the second expression basis with better expression effect can be obtained through the second parameterization model; no manual input of external data is required, thus improving efficiency.
[0037] It should be noted that the methods for enhancing facial expression capabilities provided in the embodiments of this application can all be implemented based on artificial intelligence. Artificial intelligence (AI) is a comprehensive discipline that utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Basic AI technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0038] In the embodiments of this application, the main artificial intelligence technologies involved include computer vision (image) and other fields. For example, it may involve video processing, video semantic understanding (VSU), and face recognition in computer vision.
[0039] Video semantic understanding includes target recognition, target detection and localization, etc.; face recognition includes face 3D reconstruction, face detection, face tracking, etc.
[0040] The method for enhancing the expressive power of facial expressions provided in this application can be applied to a processing device, which can be a terminal device or a server. The processing device can be a terminal device, such as a smart terminal, computer, personal digital assistant (PDA), tablet computer, etc. The processing device can also be a server, such as a standalone server or a cluster server.
[0041] The method for enhancing the expressive power of facial expressions provided in this application can be applied to various scenarios suitable for virtual avatars, such as customer service reminders, game commentary, and game characters with arbitrary facial features. In these scenarios, the method provided in this application can determine a more vivid facial expression base, and then drive the virtual avatar based on this facial expression base and a facial model.
[0042] The following describes some embodiments of this application in detail. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0043] S101. Based on the initial expression parameterization model, extract the initial semantic parameters of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data, and drive the preset facial model to generate the initial parameterization result through the first expression base corresponding to the initial semantic parameters.
[0044] This application is mainly based on Blendshape technology, which is a semantic-based facial expression parameterization technology. This technology uses a set of semantic facial models (parametric models) to parameterize facial expression information into vector form. This vector contains multiple semantic parameters, through which a corresponding expression base can be determined. The expression base is used to drive the facial model of the 3D virtual object to make various expressions.
[0045] To obtain facial expression animations of virtual objects, a series of 3D reconstruction data needs to be generated. This 3D reconstruction data can be directly used to obtain facial expression animations of virtual objects. However, this method requires a large amount of memory. Therefore, this application uses the initial expression parameterization model in Blendshape technology to parametrically fit these reference expressions, obtaining a first semantic vector corresponding to each reference expression. The first semantic vector contains multiple initial semantic parameters, each of which determines a first expression basis. Using these first expression bases to jointly drive a preset facial model, the virtual object can display rich expressions (i.e., generate the initial parameterization result of the virtual object). This method not only obtains facial expression animations of virtual objects but also significantly reduces memory usage. The facial model in this application can be a person, an animal, or even a doll with a face. For ease of description, the following embodiments use a human face as an example.
[0046] The virtual object expressions obtained through the aforementioned general initial expression parameterization model suffer from inaccuracies. Specifically, the first expression base obtained by driving a preset facial model results in an unnatural first displayed expression. In existing technologies, to improve the accuracy of the first expression base, the required expression is either manually drawn or performed by professional actors before filming. This method requires significant manpower and is inefficient. Furthermore, when a large number of expressions are needed, the above methods are no longer sufficient to address the low accuracy of the initial expression parameterization model. Therefore, this application provides a method to enhance expression expression. After the first expression base drives a preset facial model to generate initial parameterization results, by judging the accuracy of the initial parameterization results, expression residual data can be automatically generated to adjust the initial expression parameterization model, thereby improving the expressive effect of the expression base.
[0047] S102. Compare the initial parameterization results and 3D reconstruction data under each first expression basis to determine whether the initial parameterization results and 3D reconstruction data under the first expression basis match.
[0048] This application obtains multiple initial semantic parameters from the initial expression parameterization model after parametrically fitting the 3D reconstructed data. The first language parameter determines the first expression basis for representing the reference expression. All the first expression bases together drive the preset facial model to express the first displayed expression. Ideally (the result desired by the designer), the initial parameterization result should match the 3D reconstructed data, indicating that the accuracy of the first expression basis obtained by the initial expression parameterization model meets the requirements. However, in practice, the first expression basis obtained by the initial expression parameterization model is often not accurate enough. Therefore, after obtaining the initial parameterization result, it is necessary to judge whether the initial parameterization result is the expression image with the required accuracy.
[0049] When determining whether the initial parameterization result under the first expression basis matches the 3D reconstruction data, this application requires comparing the initial parameterization result with the 3D reconstruction data. When a face displays an expression, it exhibits many details, such as upturned corners of the mouth, squinting eyes, and arched eyebrows. To determine whether the initial parameterization result and the 3D reconstruction data match, these details in both data need to be compared to confirm their compatibility.
[0050] Comparing every detail in the initial facial expression graphic and the 3D reconstructed data would consume a significant amount of processing time. To improve comparison efficiency, this application sets a preset number of first keypoints in the 3D reconstructed data. These first keypoints include points around the mouth, points around the eyes, etc. Figure 2 As shown, the black dots in the image represent the first key points. These first key points allow for the differentiation of different facial expressions; for example, the displacement changes of key points around the mouth during a smile differ from those during crying. This application improves comparison efficiency by using less data for comparison.
[0051] When comparing key points, the first key point in the 3D reconstructed data is already determined. Therefore, it is also necessary to determine a second key point representing the same location as the first key point in the initial parametric results. This can be done by referring to the parametric fitting process of the first key point in the initial facial expression parametric model. After determining the second key point in the initial parametric results, the first and second key points need to be compared. In this comparison, the first and second key points can be placed in the same coordinate system, and the error information between them, i.e., the distance between the second and first key points, can be calculated using their coordinates.
[0052] In practical implementation, multiple first keypoints are typically selected, and when there are multiple first keypoints, the number of second keypoints is also corresponding. When determining the error information between the initial parameterization result and the 3D reconstruction data under the first expression basis, this application first determines the error between each first keypoint and its corresponding second keypoint. Then, the average of the errors between all first keypoints and their corresponding second keypoints in the initial parameterization result is used as the error information between the initial parameterization result and the 3D reconstruction data. For example, the distance between each first keypoint and its corresponding second keypoint is calculated, and the average of these distances is used as the error information between the initial parameterization result and the 3D reconstruction data under the first expression basis.
[0053] After obtaining the error information between the initial parameterization result and the 3D reconstruction data, this application needs to compare the error information with a preset first error threshold to determine whether the initial parameterization result matches the 3D reconstruction data. When setting the first error threshold, based on practical work habits, it is generally set to the maximum error value that meets the work requirements. That is, if the error information is greater than or equal to the first error threshold, the displayed expression image under the first expression base and the 3D reconstruction data do not match; if the error information is less than the first error threshold, the displayed expression image under the first expression base and the 3D reconstruction data match. In other words, if the distance between the second key point in the generated initial parameterization result and the first key point in the 3D reconstruction data is large, this application considers the initial parameterization result to have a large error compared to the 3D reconstruction data (the initial parameterization result fails to vividly display the expression in the 3D reconstruction data). If the distance between the second key point in the generated initial parameterization result and the first key point in the 3D reconstruction data is small, this application considers the initial parameterization result to vividly display the expression in the 3D reconstruction data and can be used in subsequent work.
[0054] S103. If there is no match, then based on the error information of the initial parameterization result and the three-dimensional reconstruction data under the first expression basis, expression residual data is generated.
[0055] When it is determined that the displayed facial expression image under the first expression basis does not match the 3D reconstruction data, the initial facial expression parameterization model needs to be adjusted to improve the accuracy of the first language parameters output by the adjusted initial facial expression parameterization model. In the prior art, after using the initial parameterization result and the reference expression, the facial expression image that needs to be improved by the initial facial expression parameterization model is obtained manually, thereby realizing the adjustment of the initial facial expression parameterization model.
[0056] The purpose of this application is to automate the adjustment of the initial expression parameterization model. Specifically, this application only requires inputting 3D reconstruction data, and through a series of adjustments, can ultimately output an expression base that meets accuracy requirements. This allows the virtual object to display an expression that is identical to the reference expression or whose error is within acceptable limits. The automated adjustment process for the initial expression parameterization model is as follows: First, the error information between the initial parameterization result and the 3D reconstruction data under the first expression base is determined. Then, expression residual data for adjusting the initial expression parameterization model is generated based on the error information. Specifically, the expression residual data is generated by superimposing the error information of the initial parameterization result and the 3D reconstruction data under the first expression base with a preset neutral expression model. Here, the neutral expression model represents an image without any expression. In other words, this application determines the error information between the initial parameterization result and the 3D reconstruction data by comparing them, and then combines this error information with the neutral expression model to compensate for the differences between the initial parameterization result and the 3D reconstruction data, making the supplemented initial parameterization result more closely match the 3D reconstruction data.
[0057] When adjusting the initial expression parameterization model using expression residual data, this application considers that adjusting the initial expression parameterization model with the first expression residual data will affect subsequent expression residual data (since the two generated expression residual data are quite similar, their adjustment effects are similar). Therefore, this application does not generate expression residual data for every first expression basis, but instead selects a target expression basis from the first expression basis. Then, expression residual data is generated based on the initial parameterization results under the selected target expression basis, the error information of the 3D reconstruction data, and the neutral expression model. Thus, the initial expression parameterization model is adjusted only using the expression residual data generated based on the target expression basis, improving adjustment efficiency.
[0058] In this embodiment, as an optional implementation, a second error threshold is set when selecting target expression bases from the first expression bases. First expression bases with error information greater than or equal to the second error threshold are selected as target expression bases. Expression residual data is generated based on the initial parameterization results and error information of the 3D reconstruction data under the target expression base, and a preset neutral expression model. In this embodiment, all first expression bases are considered comprehensively, and the first expression base with larger error information is selected as the target expression base. An equivalent embodiment is as follows: all first expression bases are sorted according to the magnitude of their error information. A first number of first expression bases are selected as target expression bases in descending order of error information. Then, expression residual data is generated based on the initial parameterization results and error information of the 3D reconstruction data under the target expression base, and a preset neutral expression model. For example, the error information of the first expression base A (i.e., the average error distance between the first key point and the second key point) is 0.50 mm, the error information of the first expression base B is 0.60 mm, the error information of the first expression base C is 0.90 mm, the error information of the first expression base D is 0.23 mm, the error information of the first expression base E is 0.45 mm, the error information of the first expression base F is 0.36 mm, and the error information of the first expression base G is 0.78 mm. This application sets a second error threshold of 0.70 mm, thus selecting the first expression bases C and G as the target expression bases.
[0059] In the above embodiments, in order to further improve efficiency, when selecting a target expression base from the first expression base, this application selects only one target expression base from the first expression base. Here, the target expression base is the first expression base with the largest error information in the first expression base.
[0060] In this embodiment, as an optional implementation, considering the impact of the first expression residual data on the adjustment of the initial expression parameterization model on subsequent expression residual data, and also considering that such impact is relatively small among the expression residual data generated from the first expression containing different expressions, this application, when selecting target expression bases from the first expression bases, first classifies all the first expression bases according to the expressions they contain, and then selects target expression bases from different types of first expression bases. For example, this application believes that the correlation between expression residual data generated from the first expression base containing a smile and expression residual data generated from the first expression base containing a cry is relatively small. When selecting target expression bases from the first expression bases, if the error information of the first expression base containing a smile is large, while the error information of the first expression base containing a cry is small, then the target expression bases selected through the above embodiment, after adjusting the initial expression parameterization model, will greatly improve the parameterization fitting effect for images containing smiles, while the improvement effect for images containing cries is relatively small. In order to improve the overall performance of the initial expression parameterization model with a single adjustment, this application selects target expression bases from the first expression bases under each type. The specific selection process is the same as that in the previous embodiment, and will not be repeated here.
[0061] For example, for expression bases containing a smile: the error information of first expression base A is 0.50mm, the error information of first expression base B is 0.89mm, the error information of first expression base C is 0.90mm, the error information of first expression base D is 0.23mm, and the error information of first expression base E is 0.45mm. For expression bases containing a crying expression: the error information of first expression base F is 0.36mm, and the error information of first expression base G is 0.83mm. Selecting one target expression base from each of the different types of first expression bases yields target expression bases with error information of 0.90mm for first expression base C and 0.83mm for first expression base G.
[0062] S104. The initial expression parameterization model is adjusted using the expression residual data, and the second semantic parameter vector of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data is extracted based on the adjusted second parameterization model. The second expression basis corresponding to the second semantic parameter vector drives the facial model to generate a second parameterization result that matches the 3D reconstruction data.
[0063] This application obtains facial expression residual data without requiring external manual input, using the aforementioned method. To improve the accuracy of the output of the initial facial expression parameterization model, this application uses the facial expression residual data to adjust the initial facial expression parameterization model. When parametrically fitting a reference facial expression, the initial facial expression parameterization model transforms the 3D reconstructed data through different facial expression dimensions, thereby obtaining the initial semantic parameters of the facial expression dimensions corresponding to the 3D reconstructed data. This application believes that the initial facial expression parameterization model initially contains a first number of facial expression dimensions. After inputting the 3D reconstructed data into the initial facial expression parameterization model, a first number of initial semantic parameters can be obtained. Based on the initial semantic parameters under each facial expression dimension, the first facial expression basis corresponding to the initial semantic parameters is determined, thereby driving the facial model to exhibit the initial parameterization result.
[0064] The facial expression residual data was determined using the above method. This residual data reflects the error information between the initial parameterization result and the 3D reconstruction data, specifically in the facial expression dimension of the initial parameterization result. Due to the lack of facial expression dimensions in the initial facial expression parameterization model, some semantic parameters are missing from the initial semantic parameters output from it. The purpose of determining the facial expression residual data in this application is to supplement this missing facial expression dimension (i.e., the additional correction dimension of the initial facial expression parameterization model). After supplementing the missing facial expression dimension, a new initial facial expression parameterization model is obtained. This new model outputs more facial expression dimension initial semantic parameters than the original model. These more facial expression dimension initial semantic parameters allow for the determination of a more comprehensive first facial expression basis, thereby driving the virtual object to exhibit a more realistic first display expression.
[0065] For example, in the 3D reconstructed data, the left corner of the mouth is upturned by 0.3mm, while in the generated initial parametric result, the left corner of the mouth is upturned by 0.2mm. In this case, the initial parametric result is missing 0.1mm in the left corner of the mouth upturn dimension compared to the 3D reconstructed data. This application first calculates the 0.1mm error between the initial parametric result and the 3D reconstructed data, and then uses the 0.1mm left corner of the mouth upturn as the expression residual data in the neutral expression model. The expression residual data is added to the initial expression parametric model to obtain a new initial expression parametric model. After inputting the 3D reconstructed data with a 0.3mm left corner of the mouth upturn into the new initial expression parametric model, the initial parametric result with a 0.3mm left corner of the mouth upturn output by the new initial expression parametric model can be obtained.
[0066] After obtaining a new initial expression parameterization model, this application inputs the 3D reconstruction data back into the new initial expression parameterization model to obtain new initial semantic parameters output by the new first parameter model. The new initial semantic parameters, corresponding to a new first expression basis, drive a preset facial model to generate new initial parameterization results. Then, the initial parameterization results under each new first expression basis are compared with the 3D reconstruction data to determine if they match. If they still do not match, the above adjustment process is repeated until the initial parameterization results under the first expression basis match the 3D reconstruction data. When the initial parameterization results under each first expression basis match the 3D reconstruction data, this initial expression parameterization model is used as the second parameterization model. This second parameterization model is the required parameterization model.
[0067] Figure 3 This illustration shows a structural schematic diagram of a device for enhancing the expressive power of facial features according to an embodiment of this application. The device includes:
[0068] The extraction module is used to extract the initial semantic parameters of each frame of 3D reconstruction data in a series of consecutive frames of 3D reconstruction data based on the initial expression parameterization model, and to drive the preset facial model to generate the initial parameterization result through the first expression base corresponding to the initial semantic parameters.
[0069] The error analysis module is used to compare the initial parameterization results and the 3D reconstruction data under each first expression basis to determine whether the initial parameterization results and the 3D reconstruction data under the first expression basis match.
[0070] The generation module is used to generate expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis if there is a mismatch.
[0071] The adjustment module is used to adjust the initial expression parameterization model using the expression residual data, and extract the second semantic parameter vector of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data based on the adjusted second parameterization model, and drive the facial model to generate a second parameterization result that matches the 3D reconstruction data through the second expression basis corresponding to the second semantic parameter vector.
[0072] The 3D reconstruction data contains a preset number of first key points; the comparison of the initial parameterization results under each first expression basis with the 3D reconstruction data includes:
[0073] Based on the first key point contained in the three-dimensional reconstruction data under the first expression basis, a second key point representing the same position as the first key point is determined in the initial parameterization result of the first expression basis.
[0074] Compare the position of the first key point with the position of the corresponding second key point.
[0075] The error analysis module is used to obtain error information of the initial parameterization result and the three-dimensional reconstruction data under the first expression basis by comparing the position of the first key point with the position of the corresponding second key point;
[0076] The method determines whether the displayed facial expression image and the 3D reconstruction data under the first facial expression base match in the following way:
[0077] If the error information is greater than or equal to the first error threshold, the displayed facial expression image and the 3D reconstruction data under the first expression base do not match.
[0078] If the error information is less than the first error threshold, the displayed facial expression image and the 3D reconstruction data under the first expression base are matched.
[0079] The step of generating expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis includes:
[0080] Based on the initial parameterization results and error information of the three-dimensional reconstruction data under the first expression basis, the target expression basis is selected from the first expression basis;
[0081] Based on the initial parameterization results under the target expression basis, the error information of the 3D reconstruction data, and the preset neutral expression model, expression residual data is generated.
[0082] The step of generating expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis includes:
[0083] Based on the expressions contained in the first expression base, the first expression base is classified to obtain different types of the first expression base;
[0084] Based on the initial parameterization results and error information of the three-dimensional reconstruction data under the first expression basis, target expression bases are selected from the first expression bases of each type.
[0085] Based on the initial parameterization results under the target expression basis, the error information of the 3D reconstruction data, and the preset neutral expression model, expression residual data is generated.
[0086] The initial facial expression parameterization model contains a first number of facial expression dimensions; adjusting the initial facial expression parameterization model using the facial expression residual data to obtain the adjusted second parameterization model includes:
[0087] Based on the facial expression residual data, generate additional correction dimensions for the initial facial expression parameterization model;
[0088] The additional correction dimension is added to the initial expression parameterization model to obtain a second parameterization model containing a second number of expression dimensions.
[0089] The second parameterized model is obtained in the following way:
[0090] The initial expression parameterization model is adjusted using the expression residual data to obtain a new initial expression parameterization model, and then expression residual data is generated again to adjust the new initial expression parameterization model.
[0091] When the initial parameterization result under each first expression basis matches the 3D reconstruction data, the initial expression parameterization model is used as the second parameterization model.
[0092] like Figure 4 As shown, this application provides an electronic device for executing the method for improving the expressive power of facial expressions in this application. The device includes a memory, a processor, a bus, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for improving the expressive power of facial expressions.
[0093] Specifically, the aforementioned memory and processor can be general-purpose memory and processor, without any specific limitations. When the processor runs the computer program stored in the memory, it can execute the aforementioned method for improving the expressive power of the expression base.
[0094] Corresponding to the method for improving the expressive power of facial expressions in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the method for improving the expressive power of facial expressions described above.
[0095] Specifically, the storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the storage medium is run, it can execute the aforementioned method for improving the expressive power of the expression base.
[0096] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0099] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0100] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0101] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for enhancing the expressive power of facial expressions, characterized in that, The method includes: The initial semantic parameters of each frame of 3D reconstruction data in a series of consecutive frames of 3D reconstruction data are extracted based on the initial expression parameterization model, and the initial parameterization result is generated by driving the preset facial model through the first expression base corresponding to the initial semantic parameters. The initial parameterization results under each first expression basis are compared with the 3D reconstruction data to determine whether the initial parameterization results under the first expression basis and the 3D reconstruction data match. If there is a mismatch, then based on the initial parameterization results under the first expression basis and the error information of the 3D reconstruction data, expression residual data is generated; The initial expression parameterization model is adjusted using the expression residual data, and the second semantic parameter vector of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data is extracted based on the adjusted second parameterization model. The second expression basis corresponding to the second semantic parameter vector drives the facial model to generate a second parameterization result that matches the 3D reconstruction data. The initial facial expression parameterization model contains a first number of facial expression dimensions; adjusting the initial facial expression parameterization model using the facial expression residual data to obtain the adjusted second parameterization model includes: Based on the facial expression residual data, generate additional correction dimensions for the initial facial expression parameterization model; The additional correction dimension is added to the initial expression parameterization model to obtain a second parameterization model containing a second number of expression dimensions.
2. The method according to claim 1, characterized in that, The 3D reconstruction data contains a preset number of first key points; The comparison of the initial parameterization results and 3D reconstruction data under each first expression basis includes: Based on the first key point contained in the three-dimensional reconstruction data under the first expression basis, a second key point representing the same position as the first key point is determined in the initial parameterization result of the first expression basis. Compare the position of the first key point with the position of the corresponding second key point.
3. The method according to claim 2, characterized in that, The method further includes: By comparing the position of the first key point with the position of the corresponding second key point, error information of the initial parameterization result and the three-dimensional reconstruction data under the first expression basis is obtained; The method determines whether the displayed facial expression image and the 3D reconstruction data under the first facial expression base match in the following way: If the error information is greater than or equal to the first error threshold, the displayed facial expression image and the 3D reconstruction data under the first expression base do not match. If the error information is less than the first error threshold, the displayed facial expression image and the 3D reconstruction data under the first expression base are matched.
4. The method according to claim 1, characterized in that, The step of generating expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis includes: Based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis, the target expression basis is selected from the first expression basis; Based on the initial parameterization results under the target expression basis, the error information of the 3D reconstruction data, and the preset neutral expression model, expression residual data is generated.
5. The method according to claim 1, characterized in that, The step of generating expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis includes: Based on the expressions contained in the first expression base, the first expression base is classified to obtain different types of the first expression base; Based on the initial parameterization results and error information of the three-dimensional reconstruction data under the first expression basis, target expression bases are selected from the first expression bases of each type. Based on the initial parameterization results under the target expression basis, the error information of the 3D reconstruction data, and the preset neutral expression model, expression residual data is generated.
6. The method according to claim 1, characterized in that, The second parameterized model is obtained in the following way: The initial expression parameterization model is adjusted using the expression residual data to obtain a new initial expression parameterization model, and then new expression residual data is generated again to adjust the new initial expression parameterization model. When the initial parameterization result under each first expression basis matches the 3D reconstruction data, the initial expression parameterization model is used as the second parameterization model.
7. A device for enhancing the expressive power of facial expressions, characterized in that, The device includes: The extraction module is used to extract the initial semantic parameters of each frame of 3D reconstruction data in a series of consecutive frames of 3D reconstruction data based on the initial expression parameterization model, and to drive the preset facial model to generate the initial parameterization result through the first expression base corresponding to the initial semantic parameters. The error analysis module is used to compare the initial parameterization results and the 3D reconstruction data under each first expression basis to determine whether the initial parameterization results and the 3D reconstruction data under the first expression basis match. The generation module is used to generate expression residual data based on the initial parameterization results and error information of the 3D reconstruction data under the first expression basis if there is a mismatch. The adjustment module is used to adjust the initial expression parameterization model using the expression residual data, and extract the second semantic parameter vector of each frame of 3D reconstruction data in the continuous multi-frame 3D reconstruction data based on the adjusted second parameterization model, and drive the facial model to generate a second parameterization result that matches the 3D reconstruction data through the second expression basis corresponding to the second semantic parameter vector; The initial expression parameterization model contains a first number of expression dimensions; the adjustment module uses the expression residual data to adjust the initial expression parameterization model to obtain an adjusted second parameterization model, including: Based on the facial expression residual data, generate additional correction dimensions for the initial facial expression parameterization model; The additional correction dimension is added to the initial expression parameterization model to obtain a second parameterization model containing a second number of expression dimensions.
8. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the method for enhancing the expressive power of facial expressions as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method for enhancing the expressive power of an emoji as described in any one of claims 1 to 6.