Digital human animation generation method and device, electronic equipment and storage medium
By acquiring the control parameters of the first digital human and using the expression parameter prediction model and weighted data to drive the second digital human to display expressions, the problem of low efficiency in digital human facial expression animation generation is solved, and automated and efficient expression transfer is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AVATAR WORKS INC
- Filing Date
- 2023-03-14
- Publication Date
- 2026-08-04
AI Technical Summary
The generation efficiency of digital human facial expression animation is low, relying on manual operation, which is time-consuming and costly, and cannot be automated.
By acquiring the control parameters of the first digital human under a specified facial expression animation, the control parameters are predicted using an expression parameter prediction model, and weighted based on preset weight data of the facial region, ultimately driving the second digital human to display the expression sequence in the specified facial expression animation.
It improves the efficiency of digital human facial expression animation generation, realizes automated expression transfer and display, and reduces labor costs.
Smart Images

Figure CN116363264B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, electronic device, and storage medium for generating digital human animation. Background Technology
[0002] Digital humans are virtual simulations of the human body at different levels of form and function, using information science methods.
[0003] Currently, the production of digital human animation relies entirely on manual operation, which is time-consuming, inefficient, and costly in terms of manpower. After creating expressions for one digital human according to multiple specified expressions, if it is necessary to create expressions for a second different digital human according to the same specified expressions, the animator needs to repeat the operations performed on the first digital human on the second digital human. Since there is no alternative automated process, the efficiency of generating digital human facial expression animation is greatly reduced.
[0004] Therefore, how to improve the generation efficiency of digital human facial expression animation is a technical problem that needs to be solved.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] This application provides a method, apparatus, electronic device, and storage medium for generating digital human animations, in order to improve the generation efficiency of digital human facial expression animations.
[0007] In a first aspect, a method for generating digital human animation is provided. The method includes: obtaining control parameters of a first digital human under a specified facial expression animation, the control parameters being composed of first parameters of multiple different facial regions; inputting the control parameters into an expression parameter prediction model, obtaining predicted control parameters based on the output of the expression parameter prediction model, the predicted control parameters being composed of multiple second parameters corresponding to the first parameters; weighting each of the second parameters based on multiple preset weight data corresponding to each of the facial regions to obtain target control parameters corresponding to the predicted control parameters; driving the second digital human based on the target control parameters to obtain a target facial expression animation, in which the second digital human displays expressions according to the expression sequence in the specified facial expression animation; wherein the expression parameter prediction model is trained based on the correlation between multiple control parameters of the first digital human and the second digital human under the same expression.
[0008] Secondly, a digital human animation generation apparatus is provided, the apparatus comprising: an acquisition module for acquiring control parameters of a first digital human under a specified facial expression animation, the control parameters being composed of first parameters of multiple different facial regions; a prediction module for inputting the control parameters into an expression parameter prediction model and obtaining predicted control parameters based on the output of the expression parameter prediction model, the predicted control parameters being composed of multiple second parameters corresponding to the first parameters; a weighting module for weighting each of the second parameters based on multiple preset weight data corresponding to each of the facial regions to obtain target control parameters corresponding to the predicted control parameters; and a driving module for driving the second digital human based on the target control parameters to obtain a target facial expression animation, in which the second digital human displays expressions according to the expression sequence in the specified facial expression animation; wherein the expression parameter prediction model is trained based on the correlation between multiple control parameters of the first digital human and the second digital human under the same expression.
[0009] Thirdly, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the digital human animation generation method of the first aspect by executing the executable instructions.
[0010] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method for generating digital human animation as described in the first aspect.
[0011] By applying the above technical solution, control parameters of a first digital human under a specified facial expression animation are obtained. These control parameters consist of first parameters from multiple different facial regions. The control parameters are input into an expression parameter prediction model, and predicted control parameters are obtained based on the output of the model. These predicted control parameters consist of multiple second parameters corresponding to the first parameters. Each second parameter is weighted based on multiple preset weight data corresponding to each facial region to obtain a target control parameter corresponding to the predicted control parameter. The second digital human is driven based on the target control parameter to obtain a target facial expression animation, in which the second digital human displays expressions according to the expression sequence in the specified facial expression animation. The expression parameter prediction model is trained based on the correlation between multiple control parameters of the first and second digital humans under the same expression. By training an expression parameter prediction model based on the correlation between control parameters of different digital humans, the expression control parameters of one digital human can drive the expression of another digital human, improving the generation efficiency of digital human facial expression animations. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a method for generating digital human animation according to an embodiment of the present invention is shown.
[0014] Figure 2 A flowchart illustrating the detection of abnormal segments in a target facial expression animation is shown in an embodiment of the present invention;
[0015] Figure 3 This invention illustrates a flowchart of the process for determining preset weight data in an embodiment of the invention.
[0016] Figure 4 A flowchart illustrating the process of obtaining the facial expression parameter prediction model in an embodiment of the present invention is shown;
[0017] Figure 5 A schematic diagram of the structure of a digital human animation generation device according to an embodiment of the present invention is shown;
[0018] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] It should be noted that other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated in the claims section.
[0021] It should be understood that this application is not limited to the precise structure described below and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
[0022] This application can be used in a wide variety of general-purpose or special-purpose computing environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, distributed computing environments including any of the above devices, etc.
[0023] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0024] The following is combined Figures 1-4 This application describes a method for generating digital human animations according to exemplary embodiments thereof. It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application can be applied to any applicable scenario.
[0025] This application provides a method for generating digital human animation, such as... Figure 1 As shown, the method includes the following steps:
[0026] Step S101: Obtain the control parameters of the first digital human under the specified facial expression animation. The control parameters are composed of first parameters of multiple different facial regions.
[0027] In this embodiment, a specified facial expression animation corresponding to the first digital human is pre-generated, and the first digital human displays facial expressions according to the corresponding facial expression sequence. In order to efficiently generate a target facial expression animation corresponding to the second digital human, the control parameters of the first digital human under the specified facial expression animation are transferred to the second digital human, so that the second digital human displays facial expressions according to the facial expression sequence in the specified facial expression animation.
[0028] The digital face is pre-divided into multiple different facial regions (such as forehead, eyes, cheeks, nose, mouth, etc.), and each facial region is equipped with a corresponding parameter controller. These parameter controllers, along with the corresponding parameter controllers of the first digital human, form the first parameters. The control parameters are obtained by acquiring the first parameters of the first digital human under a specified facial expression animation. Specifically, each animation frame of the specified facial expression animation can be acquired firstly, and each first parameter can be extracted sequentially from each animation frame. After extraction, the control parameters are obtained. To improve extraction efficiency, keyframes can also be selected from each animation frame according to a preset expression type (such as laughing, crying, smiling, etc.), and each first parameter can be extracted sequentially from each keyframe. After extraction, the control parameters are obtained.
[0029] Step S102: Input the control parameters into the facial expression parameter prediction model, and obtain the predicted control parameters based on the output of the facial expression parameter prediction model. The predicted control parameters are composed of multiple second parameters corresponding to the first parameters.
[0030] In this embodiment, in order to apply the expressions on the first digital human to the second digital human, an expression parameter prediction model is pre-trained. This expression parameter prediction model is trained based on the correlation between multiple control parameters of the first and second digital humans under the same expression. After obtaining the control parameters, the control parameters are input into the expression parameter prediction model, and the expression parameter prediction model outputs multiple second parameters corresponding to the first parameters. These second parameters constitute the predicted control parameters.
[0031] Step S103: Weight each of the second parameters based on multiple preset weight data corresponding to each of the facial regions to obtain the target control parameters corresponding to the predictive control parameters.
[0032] Different facial regions have varying degrees of influence on overall facial expressions. For example, the mouth and eye areas have a greater impact on expressions, while the nose has a smaller impact. Therefore, multiple preset weight data are set in advance based on the degree of influence of different facial regions on overall facial expressions. After obtaining the predictive control parameters, each secondary parameter is weighted based on each preset weight data to obtain the target control parameters corresponding to the predictive control parameters. This allows for more precise facial expression control based on the target control parameters.
[0033] Optionally, the preset weight data can be manually set based on experience, or the convolutional neural network model can be trained in advance according to the correlation between each facial region and different weight parameters, and the preset weight data can be predicted based on the trained convolutional neural network model.
[0034] Step S104: Drive the second digital human based on the target control parameters to obtain the target facial expression animation.
[0035] After obtaining the target control parameters, the second digital human is driven based on the target control parameters, so that the second digital human displays expressions according to the expression sequence in the specified expression animation, thus obtaining the target expression animation.
[0036] Among them, the target control parameters can be applied to the second digital human in rendering engines such as UE4, Unity, and Maya to drive the second digital human and obtain the target facial expression animation.
[0037] By applying the above technical solution, control parameters of a first digital human under a specified facial expression animation are obtained. These control parameters consist of first parameters from multiple different facial regions. The control parameters are input into an expression parameter prediction model, and predicted control parameters are obtained based on the output of the model. These predicted control parameters consist of multiple second parameters corresponding to the first parameters. Each second parameter is weighted based on multiple preset weight data corresponding to each facial region to obtain a target control parameter corresponding to the predicted control parameter. The second digital human is driven based on the target control parameter to obtain a target facial expression animation, in which the second digital human displays expressions according to the expression sequence in the specified facial expression animation. The expression parameter prediction model is trained based on the correlation between multiple control parameters of the first and second digital humans under the same expression. By training an expression parameter prediction model based on the correlation between control parameters of different digital humans, the expression control parameters of one digital human can drive the expression of another digital human, improving the generation efficiency of digital human facial expression animations.
[0038] In some embodiments of this application, after driving the second digital human based on the target control parameters to obtain the target facial expression animation, such as... Figure 2 As shown, the method also includes the following steps:
[0039] Step S21: Continuously obtain the similarity between every two adjacent frames in the target facial expression animation.
[0040] Because the target control parameters may contain inaccurate descriptions, the second digital human's facial expressions in the target facial animation may exhibit abnormal phenomena such as unsmoothness or stillness. Therefore, after obtaining the target facial animation, corresponding anomaly detection is performed to obtain abnormal segments, and these abnormal segments are corrected.
[0041] Anomaly detection can be performed by assessing the similarity between every two adjacent frames. Specifically, the similarity between every two adjacent frames in a target facial expression animation can be continuously obtained. This involves first acquiring the facial feature vectors of every two adjacent frames, which represent the positional information of key facial points. The similarity is then determined based on the distance between the two corresponding facial feature vectors. Any algorithm, including cosine similarity, dot product, Euclidean distance, Pearson correlation coefficient, and Jaccard coefficient, can be used to calculate the distance between the two corresponding facial feature vectors.
[0042] Step S22: If the similarity between two adjacent frames is less than a preset threshold, the two adjacent frames are taken as two target frames, and abnormal segments including the two target frames are identified.
[0043] Understandably, similarity represents the similarity of facial expression changes between two adjacent frames. Higher similarity indicates a smoother and more natural facial expression change between the two adjacent frames, while lower similarity indicates a less smooth facial expression change between the two adjacent frames. Therefore, a preset threshold can be set, and the similarity can be compared with this threshold. If the similarity is less than the preset threshold, an anomaly of unevenness can be confirmed between the current two adjacent frames. These two adjacent frames can then be marked as an abnormal segment. If multiple pairs of consecutive adjacent frames are abnormal segments, they can be merged into one abnormal segment.
[0044] Step S23: Correct the abnormal segment according to the number of frames of the abnormal segment.
[0045] After identifying the abnormal segments, the number of frames in the abnormal segments is determined, and the abnormal segments are corrected based on the number of frames. By performing anomaly detection on the target facial animation and correcting the detected abnormal segments, the smoothness of the target facial animation is further improved.
[0046] In some embodiments of this application, correcting the abnormal segment based on the number of frames in the abnormal segment includes:
[0047] If the number of frames is not less than a preset number, a first target segment matching the abnormal segment is obtained from a preset material library, and the abnormal segment is replaced based on the first target segment;
[0048] If the number of frames is less than the preset number, the two frames adjacent to the abnormal segment are used as reference frames. The two reference frames are interpolated based on a preset interpolation algorithm to obtain a second target segment. The abnormal segment is then replaced by the second target segment.
[0049] If the number of frames is not less than the preset number, it indicates that the abnormal segment contains a large number of frames. The abnormal segment can be corrected by replacing the entire abnormal segment. The material to replace the abnormal segment is retrieved from a preset material library. This library stores multiple expression frames corresponding to different expressions, with each expression frame being a separate material. First, the vector similarity between the abnormal segment and each material set in the preset material library is determined. Then, the material sets are sorted according to the vector similarity, and the material set with the highest similarity is selected to form the first target segment. The abnormal segment is then replaced based on this first target segment. It should be noted that when calculating vector similarity, the corresponding vector similarity can be calculated frame by frame, then accumulated to obtain the overall vector similarity. The material set with the highest similarity is then determined based on the overall vector similarity, resulting in the first target segment.
[0050] If the number of frames is less than a preset number, it indicates that the number of frames in the abnormal segment is insufficient. First, the two frames immediately before and after the abnormal segment are used as reference frames. Then, interpolation processing is performed on the two reference frames based on a preset interpolation algorithm to obtain a second target segment for replacement. The abnormal segment is then replaced based on the second target segment. The preset interpolation algorithm can be a linear interpolation algorithm or a spherical interpolation algorithm, whichever is appropriate for those skilled in the art.
[0051] By varying the number of frames in the abnormal segments, appropriate methods are used to correct them, making the expressions in the target facial animation smoother and more natural.
[0052] In some embodiments of this application, before weighting each of the second parameters based on multiple preset weight data to obtain the target control parameter corresponding to the predicted control parameter, such as Figure 3 As shown, the method further includes the following steps:
[0053] Step S31: Obtain multiple facial images from facial video data of multiple real faces.
[0054] Multiple real faces are captured in advance using camera equipment to obtain facial video data, and multiple facial images can be obtained from the facial video data of multiple real faces.
[0055] In some embodiments of this application, obtaining multiple facial images from facial video data of multiple real faces includes:
[0056] Image stream data is obtained based on the facial video data;
[0057] The image stream data is processed using a preset image enhancement algorithm to obtain the facial image.
[0058] In this embodiment, image stream data can be obtained by sampling facial video data. However, the expressiveness of some subtle facial expressions in the image stream data may not be sufficient. Therefore, a preset image enhancement algorithm is used to perform image enhancement processing on the image stream data to obtain facial images, thereby making the facial images more accurately represent the subtle facial expressions of a real human face.
[0059] In some embodiments of this application, image enhancement processing is performed on the image stream data based on a preset image enhancement algorithm to obtain the facial image, including: centering the face in the image stream data based on a face pose correction algorithm; and / or adjusting the brightness parameters of the image stream data using the OpenCV algorithm library; and / or performing contrast-limited adaptive histogram equalization using the OpenCV algorithm library to enhance the contrast of facial expression textures in the image stream data; and / or using the OpenCV algorithm library and the Numpy data processing library to project Lambert light onto the image stream data.
[0060] In this embodiment, face detection technology can first be used to perform face recognition on the obtained image stream data to obtain the location region of the face. Then, noise reduction and cropping processing can be performed on the environment around the face to remove as much noise scene as possible in the image stream data. Then, a face pose correction algorithm can be used to center-align the face. Alternatively, the HSV detection mode in OpenCV can be used to convert the original RGB three-channel model of the image stream data into an HSV color model, and the brightness parameter of the image stream data can be adjusted by controlling the brightness. Alternatively, the adaptive histogram equalization createCLAHE method in OpenCV with contrast limitation can be used to obtain a more suitable full-image pigment distribution. Furthermore, the Lambert light source model can be introduced through OpenCV algorithm library and NumPy matrix calculation to form diffuse light source noise and enhance the data.
[0061] Image enhancement is performed using various methods to more accurately represent the subtle facial expressions of a real human face.
[0062] Step S32: Divide the facial image into regions according to each facial region, and set an expression control function corresponding to each facial region. The expression control function is equipped with control weight data.
[0063] After obtaining multiple facial images, the facial images are divided into different facial regions. Then, expression control functions are set that correspond one-to-one with each facial region. Each expression control function controls each facial region. The expression control function contains control weight data, and it can be understood that each control weight data also corresponds one-to-one with each facial region.
[0064] Step S33: Based on the first preset neural network model, perform correlation training on each of the facial image data and each of the control weight data, and obtain the target neural network model after training.
[0065] A first preset neural network model is pre-constructed. Facial image data and control weight data are used as training data to perform correlation training on the first preset neural network model. After training, the target neural network model is obtained. The first preset neural network model can be a convolutional neural network model.
[0066] Step S34: Determine the preset weight data based on the output data of the target neural network model.
[0067] After obtaining the trained target neural network model, the facial image of the second digital human is input into the target neural network model, and the output data of the target neural network model is used as the preset weight data, so that each weight data is more consistent with the expression of a real human face, thereby making the expression of the second digital human more natural.
[0068] In some embodiments of this application, after determining the preset weight data based on the output data of the target neural network model, the method further includes:
[0069] The preset weight data is filtered based on a preset filtering algorithm, thereby further improving the smoothness of the preset weight data. Optionally, the preset filtering algorithm is... A filter is added because the weight data output by the target neural network model is not time-dependent, i.e., there is no smoothness between frames. The filter can smooth the output data, making the effect of subsequent real-time driving of digital faces more consistent with real human faces.
[0070] In some embodiments of this application, before inputting the control parameters into the facial expression parameter prediction model, such as Figure 4 As shown, the method also includes the following steps:
[0071] Step S41: Establish a second preset neural network model according to the preset network structure.
[0072] In this embodiment, a preset network structure is specified, and a second preset neural network model is established according to the preset network structure. The second preset neural network model can be a convolutional neural network, and those skilled in the art can flexibly adopt different network structures as the preset network structure.
[0073] Step S42: The first expression control parameters of the first digital human under multiple specified expressions are used as input, and the second expression control parameters of the second digital human under multiple specified expressions are used as labels to perform supervised learning training on the second preset neural network model.
[0074] In this embodiment, a supervised learning training method is used to train the second preset neural network model. Specifically, the first expression control parameters of the first digital human under multiple specified expressions are used as input, and the second expression control parameters of the second digital human under multiple specified expressions are used as labels to establish a regression task for training.
[0075] Step S43: When the preset training stopping condition is met, the expression parameter prediction model is obtained.
[0076] In this embodiment, the preset training stopping condition can be that the loss function is lower than a preset threshold or that a preset number of iterations is reached. When the preset training stopping condition is met, the facial expression parameter prediction model is obtained, thereby improving the reliability and accuracy of the facial expression parameter prediction model.
[0077] In some embodiments of this application, after establishing a second preset neural network model according to a preset network structure, the method further includes:
[0078] Acquire a first animation of the first digital human under multiple specified expressions, and a second animation of the second digital human under multiple specified expressions;
[0079] Based on preset expression types, extract multiple first keyframes from the first animation and multiple second keyframes from the second animation;
[0080] Each first expression control parameter is obtained from each first keyframe, and each second expression control parameter is obtained from each second keyframe.
[0081] In this embodiment, after establishing a second preset neural network model according to a preset network structure, the first and second animations of the first and second digital humans with the same expression are first obtained. The first and second animations include multiple animation frames. The expressiveness of the expressions in some animation frames is relatively weak, so only the keyframes with strong expressiveness are selected, thereby improving efficiency. Specifically, based on preset expression types (such as laughing, crying, smiling, etc.), multiple first keyframes are extracted from the first animation, and multiple second keyframes are extracted from the second animation. Then, each first expression control parameter is obtained from each first keyframe, and each second expression control parameter is obtained from each second keyframe, thereby obtaining the first expression control parameters and the second expression control parameters more efficiently.
[0082] This application also proposes a device for generating digital human animation, such as... Figure 5 As shown, the device includes:
[0083] The acquisition module 501 is used to acquire control parameters of the first digital human under a specified facial expression animation, the control parameters being composed of first parameters of multiple different facial regions; the prediction module 502 is used to input the control parameters into an expression parameter prediction model, and obtain predicted control parameters based on the output of the expression parameter prediction model, the predicted control parameters being composed of multiple second parameters corresponding to the first parameters; the weighting module 503 is used to weight each of the second parameters based on multiple preset weight data corresponding to each of the facial regions, to obtain target control parameters corresponding to the predicted control parameters; the driving module 504 is used to drive the second digital human based on the target control parameters to obtain a target facial expression animation, in which the second digital human displays expressions according to the expression sequence in the specified facial expression animation; wherein, the expression parameter prediction model is trained based on the correlation between multiple control parameters of the first digital human and the second digital human under the same expression.
[0084] In specific application scenarios, the device further includes an anomaly detection module, used to: continuously acquire the similarity between every two adjacent frames in the target facial expression animation; if the similarity between the current two adjacent frames is less than a preset threshold, take the current two adjacent frames as two target frames, and determine an abnormal segment including the two target frames; and correct the abnormal segment according to the number of frames in the abnormal segment.
[0085] In specific application scenarios, the anomaly detection module is specifically used for: if the number of frames is not less than a preset number, obtaining a first target segment that matches the abnormal segment from a preset material library, and replacing the abnormal segment based on the first target segment; if the number of frames is less than the preset number, taking two frames adjacent to the abnormal segment as reference frames, interpolating the two reference frames based on a preset interpolation algorithm to obtain a second target segment, and replacing the abnormal segment based on the second target segment.
[0086] In specific application scenarios, the device further includes a weighting module, used for: acquiring multiple facial images from facial video data of multiple real faces; dividing the facial images according to each facial region, and setting an expression control function corresponding to each facial region, wherein the expression control function is configured with control weight data; performing correlation training on each facial image data and each control weight data based on a first preset neural network model, and obtaining a target neural network model after training; and determining the preset weight data according to the output data of the target neural network model.
[0087] In specific application scenarios, the weighting module is specifically used to: obtain image stream data based on the facial video data; and perform image enhancement processing on the image stream data based on a preset image enhancement algorithm to obtain the facial image.
[0088] In specific application scenarios, the device further includes a training module, used to: establish a second preset neural network model according to a preset network structure; take the first expression control parameters of the first digital person under multiple specified expressions as input, and the second expression control parameters of the second digital person under multiple specified expressions as labels, and perform supervised learning training on the second preset neural network model; and obtain the expression parameter prediction model when the preset training stopping condition is met.
[0089] In specific application scenarios, the training module is further configured to: acquire a first animation of the first digital person under multiple specified expressions, and a second animation of the second digital person under multiple specified expressions; extract multiple first keyframes from the first animation based on preset expression types, and extract multiple second keyframes from the second animation; obtain each first expression control parameter from each first keyframe, and obtain each second expression control parameter from each second keyframe.
[0090] By applying the above technical solutions, the digital human animation generation device includes: an acquisition module for acquiring control parameters of a first digital human under a specified facial expression animation, wherein the control parameters are composed of first parameters of multiple different facial regions; a prediction module for inputting the control parameters into an expression parameter prediction model and obtaining predicted control parameters based on the output of the expression parameter prediction model, wherein the predicted control parameters are composed of multiple second parameters corresponding to the first parameters; a weighting module for weighting each of the second parameters based on multiple preset weight data corresponding to each of the facial regions to obtain target control parameters corresponding to the predicted control parameters; and a driving module for driving the second digital human based on the target control parameters to obtain a target facial expression animation, wherein the second digital human displays expressions according to the expression sequence in the specified facial expression animation; wherein the expression parameter prediction model is trained based on the correlation between multiple control parameters of the first and second digital humans under the same expression, and the expression parameter prediction model is trained based on the correlation between the control parameters of different digital humans, thereby enabling the expression control parameters of one digital human to drive the expression of another digital human, improving the generation efficiency of digital human facial expression animation.
[0091] This invention also provides an electronic device, such as... Figure 6As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.
[0092] Memory 603 is used to store the processor's executable instructions;
[0093] Processor 601 is configured to execute the following via executing the executable instructions:
[0094] The process involves obtaining control parameters for a first digital human under a specified facial expression animation, the control parameters being composed of first parameters for multiple different facial regions; inputting the control parameters into an expression parameter prediction model, and obtaining predicted control parameters based on the output of the expression parameter prediction model, the predicted control parameters being composed of multiple second parameters corresponding to the first parameters; weighting each second parameter based on multiple preset weight data corresponding to each facial region to obtain target control parameters corresponding to the predicted control parameters; driving the second digital human based on the target control parameters to obtain a target facial expression animation, in which the second digital human displays expressions according to the expression sequence in the specified facial expression animation; wherein, the expression parameter prediction model is trained based on the correlation between multiple control parameters of the first and second digital humans under the same expression.
[0095] The aforementioned communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0096] The communication interface is used for communication between the aforementioned terminal and other devices.
[0097] The memory may include RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0098] The processors mentioned above can be general-purpose processors, including CPUs (Central Processing Units), NPs (Network Processors), etc.; they can also be DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0099] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the method for generating digital human animation as described above.
[0100] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the method for generating digital human animation as described above.
[0101] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0102] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0103] The various embodiments in this specification are described in a related manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0104] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method of generating digital human animation, the method comprising: The method includes: Obtain control parameters for a first digital human under a specified facial expression animation, wherein the control parameters consist of first parameters for multiple different facial regions; The control parameters are input into the facial expression parameter prediction model, and the predicted control parameters are obtained based on the output of the facial expression parameter prediction model. The predicted control parameters are composed of multiple second parameters corresponding to the first parameters. The second parameter is weighted based on multiple preset weight data corresponding to each of the facial regions to obtain the target control parameter corresponding to the predictive control parameter. The second digital human is driven based on the target control parameters to obtain a target facial expression animation, in which the second digital human displays facial expressions according to the facial expression sequence in the specified facial expression animation; The facial expression parameter prediction model is trained based on the correlation between multiple control parameters of the first digital person and the second digital person under the same facial expression. After driving the second digital human based on the target control parameters to obtain the target facial expression animation, the method further includes: Continuously acquire the similarity between every two adjacent frames in the target facial expression animation; If the similarity between two adjacent frames is less than a preset threshold, the two adjacent frames are taken as two target frames, and abnormal segments including the two target frames are identified. Correcting the abnormal segment based on the number of frames in the abnormal segment includes: if the number of frames is not less than a preset number, obtaining a first target segment matching the abnormal segment from a preset material library, and replacing the abnormal segment based on the first target segment; if the number of frames is less than the preset number, taking two frames adjacent to the abnormal segment as reference frames, interpolating the two reference frames based on a preset interpolation algorithm to obtain a second target segment, and replacing the abnormal segment based on the second target segment.
2. The method of claim 1, wherein, Before weighting each of the second parameters based on multiple preset weight data to obtain the target control parameter corresponding to the predicted control parameter, the method further includes: Multiple facial images are obtained from facial video data of multiple real faces; The facial image is divided into regions according to each of the facial regions, and an expression control function corresponding to each of the facial regions is set. The expression control function is set with control weight data. Based on the first preset neural network model, the facial image data and the control weight data are correlated and trained to obtain the target neural network model after training. The preset weight data is determined based on the output data of the target neural network model.
3. The method of claim 2, wherein, The process of acquiring multiple facial images from facial video data of multiple real faces includes: Image stream data is obtained based on the facial video data; The image stream data is processed using a preset image enhancement algorithm to obtain the facial image.
4. The method of claim 1, wherein, Before inputting the control parameters into the facial expression parameter prediction model, the method further includes: Establish a second preset neural network model according to the preset network structure; The first expression control parameters of the first digital human under multiple specified expressions are used as input, and the second expression control parameters of the second digital human under multiple specified expressions are used as labels to perform supervised learning training on the second preset neural network model. When the preset training stopping condition is met, the expression parameter prediction model is obtained.
5. The method of claim 4, wherein, After establishing a second preset neural network model according to a preset network structure, the method further includes: Acquire a first animation of the first digital human under multiple specified expressions, and a second animation of the second digital human under multiple specified expressions; Based on preset expression types, extract multiple first keyframes from the first animation and multiple second keyframes from the second animation; Each first expression control parameter is obtained from each first keyframe, and each second expression control parameter is obtained from each second keyframe.
6. A device for generating digital human animation, characterized in that, The device includes: The acquisition module is used to acquire the control parameters of the first digital human under a specified facial expression animation, wherein the control parameters are composed of first parameters of multiple different facial regions; The prediction module is used to input the control parameters into the facial expression parameter prediction model and obtain the predicted control parameters based on the output of the facial expression parameter prediction model. The predicted control parameters are composed of multiple second parameters corresponding to the first parameters. The weighting module is used to weight each of the second parameters based on multiple preset weight data corresponding to each of the facial regions, so as to obtain the target control parameters corresponding to the predictive control parameters. The driving module is used to drive the second digital human based on the target control parameters to obtain a target facial expression animation, in which the second digital human displays facial expressions according to the facial expression sequence in the specified facial expression animation; The facial expression parameter prediction model is trained based on the correlation between multiple control parameters of the first digital person and the second digital person under the same facial expression. After driving the second digital human based on the target control parameters to obtain the target facial expression animation, the driving module is also used for: Continuously acquire the similarity between every two adjacent frames in the target facial expression animation; If the similarity between two adjacent frames is less than a preset threshold, the two adjacent frames are taken as two target frames, and abnormal segments including the two target frames are identified. Correcting the abnormal segment based on the number of frames in the abnormal segment includes: if the number of frames is not less than a preset number, obtaining a first target segment matching the abnormal segment from a preset material library, and replacing the abnormal segment based on the first target segment; if the number of frames is less than the preset number, taking two frames adjacent to the abnormal segment as reference frames, interpolating the two reference frames based on a preset interpolation algorithm to obtain a second target segment, and replacing the abnormal segment based on the second target segment.
7. An electronic device, comprising: include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method for generating digital human animation according to any one of claims 1 to 5 by executing the executable instructions.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for generating digital human animation as described in any one of claims 1 to 5.