Face and head driving method and device of virtual character image, equipment and medium

By obtaining user voice characteristics and driving prediction models, generating facial expressions and head motion control parameters, the problem of insufficient mimicry in the existing technology is solved, and high simulation and naturalness of virtual characters are achieved.

CN120339554APending Publication Date: 2025-07-18IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209423.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing virtual character facial driving technology is based on traditional computer vision algorithms or deep learning technology. It has insufficient mimicry and the generated virtual character images lack flexibility and nature, making it difficult to achieve high simulation of facial expressions and head movements.

Method used

By obtaining the target voice input by the user, performing feature extraction and inputting a pre-trained prediction model, obtaining facial expression control parameters and head motion control parameters. Based on these parameters, the facial and head models of virtual characters are driven, and the prediction model is obtained based on the speech characteristics of sample speech and controller parameter label training.

Benefits of technology

It realizes high simulation of facial expressions and head movements of virtual characters, improves mimicry, flexibility and nature, and meets the needs of efficient and real-time voice driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339554A_ABST
    Figure CN120339554A_ABST
Patent Text Reader

Abstract

The invention provides a face and head driving method and device for a virtual character image, equipment and a medium, and the method comprises the steps: obtaining a target voice inputted by a user, carrying out the feature extraction of the target voice, and obtaining the voice features of the target voice; the voice features are input into a pre-trained prediction model, target controller parameters corresponding to target voice output by the prediction model are obtained, and the target controller parameters comprise facial expression control parameters and head motion control parameters; based on the facial expression control parameters, performing facial expression driving on a pre-constructed head model of the virtual character image, and based on the head motion control parameters, performing head motion driving on the head model; wherein the prediction model is obtained by training based on the voice features of the sample voice and the controller parameter tag of the sample voice. According to the method, the simulation degree, flexibility and naturalness of the facial expression and the head movement of the virtual character can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment, and medium for driving the face and head of a virtual character image. Background Art

[0002] With the rapid development of virtual reality and augmented reality technologies, people have put forward higher requirements for the realistic expressiveness of virtual characters. Especially in the fields of entertainment, film and television, social networking, and education, there is a strong demand for realistic and natural facial expressions and head movements of characters. Existing facial driving technologies mainly rely on traditional computer vision algorithms or use deep learning technologies to drive the facial expressions of virtual characters, with insufficient realism, making it difficult to achieve a high degree of simulation of facial expressions and head movements, and the generated virtual character images lack flexibility and naturalness. Summary of the Invention

[0003] The present invention provides a method, device, equipment, and medium for driving the face and head of a virtual character image to solve the defect that in the prior art, based on traditional computer vision algorithms or using deep learning technologies to drive the face of a virtual character, the realism is insufficient, and the generated virtual character images lack flexibility and naturalness.

[0004] The present invention provides a method for driving the face and head of a virtual character image, including: Obtaining a target voice input by a user, extracting features from the target voice to obtain voice features of the target voice; Inputting the voice features into a pre-trained prediction model to obtain target controller parameters corresponding to the target voice output by the prediction model, where the target controller parameters include facial expression control parameters and head movement control parameters; Based on the facial expression control parameters, driving the facial expressions of a pre-constructed head model of a virtual character image, and based on the head movement control parameters, driving the head movement of the head model; Wherein, the prediction model is trained based on the voice features of sample voices and the controller parameter labels of the sample voices.

[0005] In some embodiments, the driving the head movement of the head model based on the head movement control parameters includes: Extracting features from the facial expression control parameters to obtain facial expression implicit features, and extracting features from the head movement control parameters to obtain head movement implicit features; Predict the head vertex positions of the virtual character image based on the implicit facial expression features, and predict the head rotation matrix and head translation vector of the virtual character image based on the implicit head movement features; Concatenate the implicit facial expression features and the implicit head movement features to obtain the concatenated features, and predict the head movement area of the virtual character image based on the concatenated features; Drive the head movement of the head model of the virtual character image based on the head vertex positions, head rotation matrix, head translation vector, and head movement area.

[0006] In some embodiments, the method further includes: Obtain the video of the real person corresponding to the target voice; Process the video to obtain the eye region image of the real person, and extract features from the eye region image to obtain the line-of-sight features; Render the virtual character image based on the head vertex positions, implicit facial expression features, implicit head movement features, and line-of-sight features to obtain the video of the virtual character image.

[0007] In some embodiments, the method further includes: Obtain the operation information of the user; Rotate and / or scale the virtual character image based on the operation information and the preset camera pose transformation parameters.

[0008] In some embodiments, the construction process of the head model of the virtual character image includes: Obtain the 4D face video dataset of the real person; Construct the head model of the virtual character image based on the 4D face video dataset, and the head model includes a controller; Bind the head model to the corresponding head bone system, and the head bone system and the controller are used to drive the head model to generate facial expressions and head movements; Determine the weights of each bone node of the head bone system.

[0009] In some embodiments, the training process of the prediction model includes: Obtain the sample voice, extract features from the sample voice to obtain the sample voice features of the sample voice; Determine the controller parameter tags of the sample voice; Input the sample voice features into the pre-constructed initial prediction model to obtain the predicted values of the controller parameters corresponding to the sample voice output by the initial prediction model; Calculate a loss function value based on the predicted value of the controller parameters corresponding to the sample voice and the controller parameter label of the sample voice; Iteratively optimize the parameters of the initial prediction model based on the loss function value to obtain the prediction model.

[0010] In some embodiments, the controller parameter label includes a facial expression control parameter label and a head movement control parameter label. Determining the controller parameter label of the sample voice includes: Obtain the sample video corresponding to the sample voice, and determine the sample expression and sample head movement corresponding to the sample voice; Obtain initial facial expression control parameters based on the sample expression, and obtain initial head movement control parameters based on the sample head movement; Drive the facial expression of the head model based on the initial facial expression control parameters to generate a predicted facial expression, and drive the head movement of the head model based on the initial head movement control parameters to generate a predicted head movement; Iteratively optimize the weights and positions of the first bone nodes corresponding to the head model, and the initial facial expression control parameters based on the predicted facial expression and the sample expression to obtain the facial expression control parameter label; Iteratively optimize the weights and positions of the second bone nodes corresponding to the head model, and the initial head movement control parameters based on the predicted head movement and the sample head movement to obtain the head movement control parameter label.

[0011] The present invention also provides a facial and head driving device for a virtual character image, including: A first acquisition unit for acquiring a target voice input by a user, extracting features of the target voice, and obtaining a voice feature of the target voice; A prediction unit for inputting the voice feature into a pre-trained prediction model to obtain target controller parameters corresponding to the target voice output by the prediction model, where the target controller parameters include facial expression control parameters and head movement control parameters; A control unit for driving the facial expression of a pre-constructed head model of a virtual character image based on the facial expression control parameters, and driving the head movement of the head model based on the head movement control parameters; Wherein, the prediction model is trained based on the voice feature of the sample voice and the controller parameter label of the sample voice.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for driving the face and head of the virtual character image as described in any one of the above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for driving the face and head of the virtual character image as described in any one of the above is implemented.

[0014] The method, device, equipment, and medium for driving the face and head of the virtual character image provided by the present invention obtain the target voice input by the user, extract the features of the target voice to obtain the voice features of the target voice, input the voice features into a pre-trained prediction model to obtain the target controller parameters corresponding to the target voice. The target controller parameters include facial expression control parameters and head movement control parameters. Based on the facial expression control parameters, facial expression driving is performed on the head model of the pre-constructed virtual character image. Based on the head movement control parameters, head movement driving is performed on the head model, which can achieve a high degree of simulation of the facial expressions and head movements of the virtual character image, and improve the authenticity, flexibility, and naturalness of the virtual character image. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic flowchart of the method for driving the face and head of the virtual character image provided by the embodiment of the present invention.

[0017] Figure 2 It is a schematic flowchart of the training process of the prediction model provided by the embodiment of the present invention.

[0018] Figure 3 It is a schematic structural diagram of the device for driving the face and head of the virtual character image provided by the embodiment of the present invention.

[0019] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in combination with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the scope of protection of the present invention.

[0021] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the present invention means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0022] Currently, the voice-driven technologies for virtual human expressions and head movements mainly include the following several solutions: rule-based voice-driven methods, mapping methods based on audio features and facial movements, driving methods based on key point tracking combined with voice emotions, deep learning-based voice-driven methods, and driving methods based on the fusion of voice and facial control parameters.

[0023] Although the above technical solutions have achieved certain results in voice-driven virtual human expressions and head movements, there are still the following defects in terms of fineness, emotional expression, real-time performance, etc.: (1) Lack of detailed performance: The virtual human expressions generated by the existing technologies mainly focus on rough lip shapes and basic expressions, and cannot well display micro-expressions and complex emotional changes.

[0024] (2) High difficulty in synchronizing voice and facial movements: Due to the complex synchronization and coordination between voice signals and facial controller parameters, the generated expressions and head movements often lack naturalness.

[0025] (3) High requirements for real-time performance and computing: The use of deep learning algorithms and high-dimensional control parameters requires a large amount of computing resources, which affects the practical application of voice driving in low-latency scenarios and is difficult to meet the requirements of efficient and real-time voice driving.

[0026] To this end, an embodiment of the present invention provides a method for driving the face and head of a virtual character image. By obtaining the target speech input by the user, feature extraction is performed on the target speech to obtain the speech features of the target speech; the speech features are input into a pre-trained prediction model to obtain the target controller parameters corresponding to the target speech, and the target controller parameters include facial expression control parameters and head movement control parameters; based on the facial expression control parameters, facial expression driving is performed on the head model of the pre-constructed virtual character image, and based on the head movement control parameters, head movement driving is performed on the head model. The present invention can achieve a high degree of simulation of the subtle facial expressions and head movements of the virtual character image, and improve the realism, flexibility and naturalness of the virtual character image.

[0027] Figure 1 It is a schematic flowchart of the method for driving the face and head of a virtual character image provided by an embodiment of the present invention. As Figure 1 shown, a method for driving the face and head of a virtual character image is provided, including the following steps: Step 110, Step 120, Step 130. The process steps of this method are only a possible implementation manner of the present invention.

[0028] Step 110: Obtain the target speech input by the user, perform feature extraction on the target speech, and obtain the speech features of the target speech.

[0029] Among them, the virtual character image refers to a virtual character with a visual image created by digital technology, and the virtual character image is widely used in fields such as movies, animations, and games.

[0030] Among them, the speech features include features such as intonation, tone, speech rate, and volume.

[0031] Optionally, use the conformer network to perform feature extraction on the target speech to obtain the speech features of the target speech; the conformer network is a hybrid network structure that combines a convolutional neural network (CNN) and a Transformer, aiming to retain local features and global representations through parallel processing.

[0032] Optionally, the target speech input by the user is obtained in real time, and feature extraction is performed on the target speech.

[0033] Step 120: Input the speech features into a pre-trained prediction model to obtain the target controller parameters corresponding to the target speech output by the prediction model, and the target controller parameters include facial expression control parameters and head movement control parameters.

[0034] Among them, the prediction model is trained based on the speech features of the sample speech and the controller parameter labels of the sample speech.

[0035] Among them, the facial expression control parameters are used to control the facial expressions of the virtual character, and the facial expression control parameters include the motion control parameters of key parts such as eyes, eyebrows, and mouth; the head motion control parameters are used to control the head movements of the virtual character, and the head motion control parameters include the motion control parameters of key parts such as the mandible and cervical vertebra.

[0036] Optionally, the prediction model includes a multi-layer Long Short-Term Memory (LSTM) network; LSTM can effectively capture long-term dependencies in the sequence, thereby improving the performance of the prediction model.

[0037] Optionally, based on the prediction model, combined with the first mapping table of the facial expression control parameters and facial expressions of the pre-constructed virtual character image, and the second mapping table of the head motion control parameters and head movements, predict the target controller parameters corresponding to the target speech.

[0038] It can be understood that predicting the target controller parameters corresponding to the target speech through the pre-trained prediction model has high accuracy and efficiency, strong real-time performance, and can meet the high-efficiency and real-time speech-driven requirements.

[0039] Step 130, based on the facial expression control parameters, perform facial expression driving on the head model of the pre-constructed virtual character image, and based on the head motion control parameters, perform head motion driving on the head model.

[0040] Among them, the head model is a highly complex and delicate three-dimensional model, and the head model includes key head bones and facial key points.

[0041] Optionally, filter the predicted target controller parameter sequence to obtain a relatively smooth target controller parameter sequence signal, so as to achieve stable control of the facial expressions and head movements of the virtual character image.

[0042] In some embodiments, the construction process of the head model of the virtual character image includes: Obtain a 4D face video dataset of a real person; Based on the 4D face video dataset, construct a head model of the virtual character image, and the head model includes a controller; Bind the head model to the corresponding head bone system, and the head bone system and the controller are used to drive the head model to generate facial expressions and head movements; Determine the weights of each bone node of the head bone system.

[0043] Optionally, use a camera array to collect 4D face video data of a real person to construct a 4D face video dataset with rich expression details.

[0044] Optionally, process the 4D face video dataset to obtain multiple frames of face images, and based on the multi-view reconstruction algorithm, construct a non-parametric model corresponding to each frame of face image; the multi-view reconstruction algorithm is a technology that uses image data from multiple viewpoints to construct a three-dimensional model.

[0045] Optionally, construct a parametric model corresponding to each frame of face image through non-rigid registration; non-rigid registration is a method used in the fields of computer vision and image processing to match and align non-rigid objects at different viewpoints or time points.

[0046] Among them, the head bone system is constructed based on real people. The head bone system includes multiple bone nodes, such as key bones like the skull, mandible, and cervical vertebra, and also includes key points on the face, such as the eyes, nose, mouth, and ears; the key points on the face are attached to the corresponding bone nodes; each bone node is associated with a controller.

[0047] Optionally, bind the head model to the corresponding head bone system to obtain the initial position and orientation of multiple (such as 870) control points and the skinning weights of each control point, and finally fit frame by frame to obtain the multi-dimensional controller parameters (870 * 6, that is, the 6-dimensional controller parameters of 870 control points) under the corresponding facial expressions and head movements, and establish a mapping table between facial expressions and head movements and the multi-dimensional controller parameters to achieve more refined head super-humanoid speech driving and better display micro-expressions and complex emotional changes; the multi-dimensional controller parameters include facial expression control parameters (829 * 6) and head movement control parameters (41 * 6).

[0048] It should be noted that during the fitting process of the controller parameters, the controller parameters will affect the rotation and scaling of the bones, and the transformation of the bones is transmitted to the facial mesh through Linear Blend Skinning (LBS) to achieve natural expression and pose transitions. After using LBS to complete the initial fitting, repeated optimization is still required, including fine-tuning the weight distribution, bone positions, and controller parameters to ensure smooth and natural facial deformation.

[0049] Optionally, calculate the weight of each bone node according to the distance between each bone node and the facial vertex and the topological structure of the head model.

[0050] In an embodiment of the present invention, by obtaining a target voice input by a user, extracting features from the target voice to obtain voice features of the target voice, inputting the voice features into a pre-trained prediction model, and obtaining target controller parameters corresponding to the target voice, where the target controller parameters include facial expression control parameters and head movement control parameters, driving facial expressions of a head model of a pre-constructed virtual character image based on the facial expression control parameters, and driving head movement of the head model based on the head movement control parameters, highly realistic facial expressions and head movement of the virtual character image can be achieved, and the realism, flexibility, and naturalness of the virtual character image are improved.

[0051] In some embodiments, in step 130, driving head movement of the head model based on the head movement control parameters includes: Step 131: Extracting features from the facial expression control parameters to obtain implicit facial expression features, and extracting features from the head movement control parameters to obtain implicit head movement features; Step 132: Predicting the head vertex positions of the virtual character image based on the implicit facial expression features, and predicting the head rotation matrix and head translation vector of the virtual character image based on the implicit head movement features; Step 133: Concatenating the implicit facial expression features and the implicit head movement features to obtain concatenated features, and predicting the head movement area of the virtual character image based on the concatenated features; Step 134: Driving head movement of the head model of the virtual character image based on the head vertex positions, head rotation matrix, head translation vector, and head movement area.

[0052] Optionally, encoding the facial expression control parameters to obtain 256-dimensional implicit facial expression features, and encoding the head movement control parameters to obtain 256-dimensional implicit head movement features.

[0053] Optionally, predicting a head vertex position without head movement based on the implicit facial expression features.

[0054] Optionally, the head movement area includes the area where the head is connected to the body and the facial area above the neck.

[0055] Optionally, within the head movement area, rotating the head vertices based on the head rotation matrix and translating the head vertices based on the head translation vector.

[0056] It can be understood that by decoupling head movement and facial expressions, independent control of head movement and facial expressions can be achieved, improving the fineness, coordination, naturalness, and expressiveness of the head movement and facial expressions of the virtual character.

[0057] In some embodiments, the above method further includes: Obtaining a video of the real person corresponding to the target voice; Processing the video to obtain an eye region image of the real person, and extracting features from the eye region image to obtain line-of-sight features; Rendering the virtual character image based on the head vertex position, implicit facial expression features, implicit head movement features, and line-of-sight features to obtain a video of the virtual character image.

[0058] Optionally, detecting the eye region image, calculating the distance ratio between the pupil center and the upper and lower eyelids, and the distance ratio between the pupil center and the left and right eyelids to obtain line-of-sight features.

[0059] Optionally, the virtual character image is rendered using the Model View Projection (MVP) algorithm.

[0060] Among them, the MVP algorithm is used to present the object in the three-dimensional space on the two-dimensional screen, and this process is completed through three consecutive matrix transformations, namely, model transformation, view transformation, and projection transformation.

[0061] Optionally, through operations such as translation, rotation, and scaling, the model transformation of the head model of the virtual character image is performed to convert it from the local coordinate system to the world coordinate system; the view transformation of the head model is performed to use the position of the observer as the new origin to reposition the head model in the scene; the projection transformation of the head model is performed to convert the three-dimensional coordinates into two-dimensional screen coordinates.

[0062] In some embodiments, the method further includes: Obtaining the operation information of the user; Rotating and / or scaling the virtual character image based on the operation information and the pre-set camera pose transformation parameters.

[0063] Among them, the operation information of the user refers to the mouse movement information.

[0064] It should be noted that the user can achieve the rotation of the virtual character at different angles through the relative movement of the mouse between the front and back frames, and achieve the scale change of the virtual character through the zoom of the mouse wheel.

[0065] Optionally, determining the correspondence between the mouse movement (x, y, scale) and the camera pose (x_rad, y_rad, camera_dist_scale) to obtain the camera pose transformation parameters.

[0066] It can be understood that by obtaining the operation information of the user, based on the operation information and the pre-set camera pose transformation parameters, the virtual character image is rotated and / or scaled, which is convenient for optimizing the virtual character, has strong interactivity, and improves the user experience.

[0067] In some embodiments, an improved MVP renderer is used to optimize the edges of the virtual character image to make it smoother.

[0068] Optionally, traverse the voxels of 1024*8*8*8 of the edge of the virtual character image, traverse to obtain the small voxel blocks that the light will pass through within the image mask range, so as to obtain a 1024*8*8*8 template mask, and use Gaussian kernels with different sizes and different variances to blur the template mask to optimize the edge of the virtual character image.

[0069] Figure 2 It is a schematic flow chart of the training process of the prediction model provided by the embodiment of the present invention. As Figure 2 shown, in some embodiments, the training process of the prediction model includes: Step 210: Obtain a sample voice, extract features from the sample voice to obtain the sample voice features of the sample voice; Step 220: Determine the controller parameter label of the sample voice; Step 230: Input the sample voice features into the pre-constructed initial prediction model to obtain the predicted value of the controller parameters corresponding to the sample voice output by the initial prediction model; Step 240: Calculate the loss function value based on the predicted value of the controller parameters corresponding to the sample voice and the controller parameter label of the sample voice; Step 250: Iteratively optimize the parameters of the initial prediction model based on the loss function value to obtain the prediction model.

[0070] Among them, the sample voice features include features such as intonation, tone, speech rate, and volume.

[0071] Optionally, use the conformer network to extract features from the sample voice to obtain the sample voice features of the sample voice.

[0072] Optionally, determine the sample expression and sample head movement corresponding to the sample voice, and based on the sample expression and sample head movement, determine the controller parameter label of the corresponding head model, which is the controller parameter label of the sample voice.

[0073] In some embodiments, the controller parameter label includes a facial expression control parameter label and a head movement control parameter label. Determining the controller parameter label of the sample voice includes: Obtain a sample video corresponding to the sample voice, and determine the sample expression and sample head movement corresponding to the sample voice; Based on the sample expression, obtain the initial facial expression control parameters, and based on the sample head movement, obtain the initial head movement control parameters; Based on the initial facial expression control parameters, drive the facial expression of the head model to generate a predicted facial expression. Based on the initial head movement control parameters, drive the head movement of the head model to generate a predicted head movement; Based on the predicted facial expression and the sample expression, iteratively optimize the weights and positions of the first bone nodes corresponding to the head model, as well as the initial facial expression control parameters, to obtain the facial expression control parameter labels; Based on the predicted head movement and the sample head movement, iteratively optimize the weights and positions of the second bone nodes corresponding to the head model, as well as the initial head movement control parameters, to obtain the head movement control parameter labels.

[0074] Among them, the first bone node refers to the bone node associated with the facial expression of the head model; the second bone node refers to the bone node associated with the head movement of the head model.

[0075] It can be understood that by determining the sample expression and sample head movement corresponding to the sample voice, obtaining the initial facial expression control parameters based on the sample expression, obtaining the initial head movement control parameters based on the sample head movement, and iteratively optimizing the initial facial expression control parameters and the initial head movement control parameters, the accuracy of the controller parameter labels is improved, thereby improving the performance of the prediction model.

[0076] Next, the facial and head driving device for a virtual character image provided by an embodiment of the present invention will be described. The facial and head driving device for a virtual character image described below can be mutually corresponding and referred to the facial and head driving method for a virtual character image described above.

[0077] Figure 3 It is a schematic structural diagram of a facial and head driving device for a virtual character image provided by an embodiment of the present invention. As Figure 3 shown, the facial and head driving device 300 for the virtual character image includes: A first acquisition unit 310, configured to acquire a target voice input by a user, perform feature extraction on the target voice, and obtain the voice feature of the target voice; A prediction unit 320, configured to input the voice feature into a pre-trained prediction model, and obtain target controller parameters corresponding to the target voice output by the prediction model, where the target controller parameters include facial expression control parameters and head movement control parameters; A control unit 330, configured to perform facial expression driving on a head model of a pre - constructed virtual character image based on facial expression control parameters, and perform head movement driving on the head model based on head movement control parameters; Wherein, the prediction model is trained based on the speech features of the sample speech and the controller parameter tags of the sample speech.

[0078] Optionally, performing head movement driving on the head model based on the head movement control parameters and the facial expression control parameters includes: Performing feature extraction on the facial expression control parameters to obtain implicit facial expression features, and performing feature extraction on the head movement control parameters to obtain implicit head movement features; Predicting the head vertex positions of the virtual character image based on the implicit facial expression features, and predicting the head rotation matrix and the head translation vector of the virtual character image based on the implicit head movement features; Concatenating the implicit facial expression features and the implicit head movement features to obtain the concatenated features, and predicting the head movement area of the virtual character image based on the concatenated features; Performing head movement driving on the head model of the virtual character image based on the head vertex positions, the head rotation matrix, the head translation vector, and the head movement area.

[0079] Optionally, the facial and head driving device 300 of the virtual character image further includes: A second acquisition unit, configured to acquire a video of a real person corresponding to the target speech; A data processing unit, configured to process the video to obtain an eye region image of the real person, and perform feature extraction on the eye region image to obtain a line - of - sight feature; A rendering unit, configured to render the virtual character image based on the head vertex positions, the implicit facial expression features, the implicit head movement features, and the line - of - sight feature to obtain a video of the virtual character image.

[0080] Optionally, the facial and head driving device 300 of the virtual character image further includes: A third acquisition unit, configured to acquire operation information of the user; An adjustment unit, configured to rotate and / or scale the virtual character image based on the operation information and pre - set camera pose transformation parameters.

[0081] Optionally, the construction process of the head model of the virtual character image includes: Acquiring a 4D face video data set of a real person; Based on the 4D face video data set, constructing a head model of the virtual character image, where the head model includes a controller; Bind the head model to the corresponding head bone system, and the head bone system and the controller are used to drive the head model to generate facial expressions and head movements; Determine the weights of each bone node of the head bone system.

[0082] Optionally, the training process of the prediction model includes: Obtain a sample voice, extract features from the sample voice to obtain the sample voice features of the sample voice; Determine the controller parameter labels of the sample voice; Input the sample voice features into a pre-constructed initial prediction model to obtain the predicted values of the controller parameters corresponding to the sample voice output by the initial prediction model; Based on the predicted values of the controller parameters corresponding to the sample voice and the controller parameter labels of the sample voice, calculate the loss function value; Based on the loss function value, iteratively optimize the parameters of the initial prediction model to obtain the prediction model.

[0083] Optionally, the controller parameter labels include facial expression control parameter labels and head movement control parameter labels. Determining the controller parameter labels of the sample voice includes: Obtain the sample video corresponding to the sample voice, and determine the sample expression and sample head movement corresponding to the sample voice; Obtain the initial facial expression control parameters based on the sample expression, and obtain the initial head movement control parameters based on the sample head movement; Based on the initial facial expression control parameters, drive the facial expression of the head model to generate a predicted facial expression. Based on the initial head movement control parameters, drive the head movement of the head model to generate a predicted head movement; Based on the predicted facial expression and the sample expression, iteratively optimize the weights and positions of the first bone nodes corresponding to the head model, as well as the initial facial expression control parameters, to obtain the facial expression control parameter labels; Based on the predicted head movement and the sample head movement, iteratively optimize the weights and positions of the second bone nodes corresponding to the head model, as well as the initial head movement control parameters, to obtain the head movement control parameter labels.

[0084] It should be noted here that the facial and head driving device of the virtual character image provided by the embodiments of the present invention can implement all the method steps implemented by the embodiments of the above-mentioned facial and head driving method of the virtual character image, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.

[0085] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present invention, asFigure 4 As shown in Figure 4 , the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call logic instructions in the memory 430 to execute a method for driving the face and head of a virtual character image. The method includes: obtaining a target voice input by a user, extracting features of the target voice to obtain voice features of the target voice; inputting the voice features into a pre-trained prediction model to obtain target controller parameters corresponding to the target voice output by the prediction model. The target controller parameters include facial expression control parameters and head movement control parameters; based on the facial expression control parameters, performing facial expression driving on a head model of a pre-constructed virtual character image, and based on the head movement control parameters and / or the facial expression control parameters, performing head movement driving on the head model; wherein, the prediction model is trained based on the voice features of sample voices and the controller parameter labels of the sample voices.

[0086] In addition, when the logic instructions in the above-mentioned memory 430 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0087] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for driving the face and head of a virtual character image provided by the above-mentioned various methods. The method includes: obtaining a target voice input by a user, extracting features from the target voice to obtain the voice features of the target voice; inputting the voice features into a pre-trained prediction model to obtain the target controller parameters corresponding to the target voice output by the prediction model, where the target controller parameters include facial expression control parameters and head movement control parameters; based on the facial expression control parameters, performing facial expression driving on the head model of the pre-constructed virtual character image, and based on the head movement control parameters and / or facial expression control parameters, performing head movement driving on the head model; wherein, the prediction model is trained based on the voice features of the sample voice and the controller parameter labels of the sample voice.

[0088] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for driving the face and head of a virtual character image, characterized in that, Including: Obtain the target speech input by the user, extract features from the target speech to obtain the speech features of the target speech; Input the speech features into a pre-trained prediction model to obtain the target controller parameters corresponding to the target speech output by the prediction model, where the target controller parameters include facial expression control parameters and head movement control parameters; Based on the facial expression control parameters, perform facial expression driving on the head model of the pre-constructed virtual character image, and based on the head movement control parameters, perform head movement driving on the head model; Among them, the prediction model is trained based on the speech features of the sample speech and the controller parameter labels of the sample speech.

2. The method for driving the face and head of a virtual character image according to claim 1, wherein The performing head movement driving on the head model based on the head movement control parameters includes: Extract features from the facial expression control parameters to obtain facial expression implicit features, and extract features from the head movement control parameters to obtain head movement implicit features; Based on the facial expression implicit features, predict the head vertex positions of the virtual character image, and based on the head movement implicit features, predict the head rotation matrix and head translation vector of the virtual character image; Concatenate the facial expression implicit features and the head movement implicit features to obtain the concatenated features, and based on the concatenated features, predict the head movement area of the virtual character image; Based on the head vertex positions, head rotation matrix, head translation vector and head movement area, perform head movement driving on the head model of the virtual character image.

3. The method for driving the face and head of a virtual character image according to claim 2, wherein The method further includes: Obtain the video of the real person corresponding to the target speech; Process the video to obtain the eye region image of the real person, and extract features from the eye region image to obtain the line-of-sight features; Based on the head vertex positions, facial expression implicit features, head movement implicit features and line-of-sight features, render the virtual character image to obtain the video of the virtual character image.

4. The method for driving the face and head of a virtual character image according to claim 1, characterized in that, The method further includes: Obtain the operation information of the user; Based on the operation information and the pre-set camera pose transformation parameters, perform rotation and / or scale scaling on the virtual character image.

5. The method for driving the face and head of a virtual character image according to any one of claims 2-4, characterized in that, The construction process of the head model of the virtual character image includes: Obtain the 4D face video dataset of the real person; Based on the 4D face video dataset, construct the head model of the virtual character image, and the head model includes a controller; Bind the head model to the corresponding head bone system, and the head bone system and the controller are used to drive the head model to generate facial expressions and head movements; Determine the weights of each bone node of the head bone system.

6. The method for driving the face and head of a virtual character image according to claim 1, wherein, The training process of the prediction model includes: Obtain the sample speech, extract features from the sample speech to obtain the sample speech features of the sample speech; Determine the controller parameter labels of the sample speech; Input the sample speech features into the pre-constructed initial prediction model to obtain the predicted values of the controller parameters corresponding to the sample speech output by the initial prediction model; Calculate a loss function value based on the predicted value of the controller parameters corresponding to the sample speech and the controller parameter labels of the sample speech; Iteratively optimize the parameters of the initial prediction model based on the loss function value to obtain the prediction model.

7. The facial and head driving method for the virtual character image according to claim 6, wherein The controller parameter labels include facial expression control parameter labels and head movement control parameter labels. Determining the controller parameter labels of the sample speech includes: Obtain the sample video corresponding to the sample speech, and determine the sample expression and sample head movement corresponding to the sample speech; Obtain initial facial expression control parameters based on the sample expression, and obtain initial head movement control parameters based on the sample head movement; Drive the facial expression of the head model based on the initial facial expression control parameters to generate a predicted facial expression, and drive the head movement of the head model based on the initial head movement control parameters to generate a predicted head movement; Iteratively optimize the weights and positions of the first bone nodes corresponding to the head model, and the initial facial expression control parameters based on the predicted facial expression and the sample expression to obtain the facial expression control parameter labels; Iteratively optimize the weights and positions of the second bone nodes corresponding to the head model, and the initial head movement control parameters based on the predicted head movement and the sample head movement to obtain the head movement control parameter labels.

8. A facial and head driving device for a virtual character image, characterized in that, Includes: A first acquisition unit for acquiring a target speech input by a user, extracting features from the target speech to obtain speech features of the target speech; A prediction unit for inputting the speech features into a pre-trained prediction model to obtain target controller parameters corresponding to the target speech output by the prediction model, where the target controller parameters include facial expression control parameters and head movement control parameters; A control unit for driving the facial expression of the head model of a pre-constructed virtual character image based on the facial expression control parameters, and driving the head movement of the head model based on the head movement control parameters; Wherein, the prediction model is trained based on the speech features of the sample speech and the controller parameter labels of the sample speech.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the facial and head driving method of the virtual character image according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the facial and head driving method of the virtual character image according to any one of claims 1 to 7.