A method and device for driving the expression of a virtual character
By obtaining face images for key points extraction and using expression coefficient model to drive virtual characters' expressions, the problems of large network consumption and unreal expressions in the existing technology are solved, and real and natural virtual characters' expressions are driven, with small network consumption and fast speed.
Patent Information
- Application Number
- CN202210518926.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-13
AI Technical Summary
The existing virtual character expression driver methods have high requirements for network stability and bandwidth, resulting in large traffic consumption, poor timeliness and fluency, and the expression is not realistic and natural enough, affecting the audience's experience.
By acquiring face images for key points extraction, the pre-trained expression coefficient model is used to learn the correspondence between face key points and expression coefficients, and drive the expressions of virtual characters, including step-by-step learning and deep learning optimization of the blink and non-blink expression coefficient models.
It realizes real, natural and accurate virtual character expression drivers, with small network consumption, fast driving speed, strong generalization, and improved user experience.
Smart Images

Figure CN114821734B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for driving the expressions of virtual characters. Background Art
[0002] With the rapid development of artificial intelligence technology, virtual characters have become very popular products on the market. Using virtual avatars for live social interactions, performances, live broadcasts, etc. has become a popular trend on the Internet. Using a customized cartoonized virtual avatar as one's own endorsement can not only ensure that everyone can have a satisfactory appearance, but also greatly protect personal privacy, and can also capture people's facial expressions and movements in real time, so it is sought after by people. The presentation forms of virtual characters are diverse, mainly depending on different driving methods, such as voice driving, text driving, and image driving, etc.
[0003] Currently, the mainstream virtual avatar expression driving method is to use facial expression capture technology to drive virtual characters. However, this virtual avatar expression driving method has high requirements for network stability and bandwidth, and consumes a large amount of traffic. If the network stability is low or the bandwidth is small, it will affect the timeliness and fluency of driving; moreover, the expressions of virtual characters are not real, natural, and accurate enough, affecting the viewing experience of the audience. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and device for driving the expressions of virtual characters, which can quickly drive virtual characters in a relatively real, natural, and accurate manner, with extremely low network consumption, fast driving speed, and strong generalization ability.
[0005] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for driving the expressions of virtual characters is provided, including:
[0006] Obtain a face image, and perform key point extraction on the face image to obtain face key points;
[0007] Input the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image;
[0008] Perform expression driving on the virtual character according to the expression coefficient value.
[0009] Optionally, performing key point extraction on the face image includes: inputting the face image into a pre-trained face key point recognition model for key point extraction, and the face key point recognition model is obtained by training face images with key point annotations.
[0010] Optionally, before inputting the face key points into a pre-trained expression coefficient model, the method further includes: restricting the face image to a picture of a specified size, and normalizing the face key points.
[0011] Optionally, normalizing the face key points includes: obtaining the abscissa and ordinate of the center point of the face key points, as well as the height and width of the face image; for each face key point, calculating the normalized abscissa of the face key point according to the abscissa of the face key point, the abscissa of the center point, and the height and width of the face image, and calculating the normalized ordinate of the face key point according to the ordinate of the face key point, the ordinate of the center point, and the height and width of the face image; translating the center point of the face key points to the upper left corner of the face image of the specified size, so as to translate the normalized abscissa and the normalized ordinate, and complete the normalization process of the face key points.
[0012] Optionally, the expression coefficient model is trained in the following manner: obtaining a set of face images; extracting key points from each face image to obtain the face key points corresponding to each face image; for each face image, obtaining the blend shape values of the face image based on a face model reconstruction method, and using the blend shape values to label the face key points corresponding to the face image; learning the face key points corresponding to the labeled face images to obtain the expression coefficient model.
[0013] Optionally, the pre-trained expression coefficient model includes a blinking expression coefficient model and a non-blinking expression coefficient model; inputting the face key points into the pre-trained expression coefficient model to obtain the expression coefficient value corresponding to the face image includes: extracting eye key points from the face key points, and inputting the eye key points into the blinking expression coefficient model to obtain the blinking expression coefficient value corresponding to the face image; inputting the face key points into the non-blinking expression coefficient model to obtain the non-blinking expression coefficient value corresponding to the face image; obtaining the expression coefficient value corresponding to the face image according to the blinking expression coefficient value and the non-blinking expression coefficient value.
[0014] Optionally, the blinking expression coefficient model is trained in the following manner: obtaining a set of face images; extracting key points from each face image to obtain the face key points corresponding to each face image; for each face image, obtaining the blend shape values of the face image based on a face model reconstruction method, and using the blend shape values to label the face key points corresponding to the face image; for the face key points corresponding to the labeled face images, extracting eye key points from the face key points, and learning the eye key points through a deep learning optimization algorithm to obtain the blinking expression coefficient model.
[0015] Optionally, the non-blink expression coefficient model is trained as follows: obtaining a set of face images; extracting key points for each face image to obtain the face key points corresponding to each face image; for each face image, obtaining the blend shape value of the face image based on a face model reconstruction method, and using the blend shape value to label the face key points corresponding to the face image; performing a coordinate format transformation on the face key points corresponding to the labeled face image to obtain a square matrix, and learning the square matrix through a deep learning optimization algorithm to obtain the non-blink expression coefficient model.
[0016] Optionally, when the non-blink expression coefficient model is trained, the loss function includes a first loss function corresponding to the full-face expression other than the blink expression and a second loss function corresponding to the open-mouth expression.
[0017] Optionally, before performing a coordinate format transformation on the face key points corresponding to the labeled face image to obtain a square matrix, it further includes: in the case where the coordinates of the face key points corresponding to the labeled face image cannot be directly transformed into a square matrix, supplementing with 0 values until the set of key point coordinates composed of the coordinates of the face key points corresponding to the labeled face image and the 0 values can be subjected to a coordinate format transformation to obtain a square matrix.
[0018] According to another aspect of the embodiments of the present invention, there is provided an apparatus for driving the expression of a virtual character, including:
[0019] A key point acquisition module, configured to acquire a face image and extract key points from the face image to obtain face key points;
[0020] A coefficient value calculation module, configured to input the face key points into a pre-trained expression coefficient model to obtain the expression coefficient value corresponding to the face image;
[0021] An expression driving module, configured to drive the expression of the virtual character according to the expression coefficient value.
[0022] According to yet another aspect of the embodiments of the present invention, there is provided an electronic device for driving the expression of a virtual character, including: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the one or more processors to implement the method for driving the expression of a virtual character provided by the embodiments of the present invention.
[0023] According to still another aspect of the embodiments of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method for driving the expression of a virtual character provided by the embodiments of the present invention is implemented.
[0024] One embodiment of the above invention has the following advantages or beneficial effects: By obtaining a face image and extracting key points from the face image to obtain face key points; inputting the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image; according to the expression coefficient value, a technical solution for driving the expression of a virtual character, by learning the correspondence between the position coordinates of the face key points and the expression coefficient model to drive the expression of the virtual character, enabling the virtual character model to show the same expression as a real person, can drive the virtual character quickly, realistically, naturally and accurately, with extremely low network consumption, fast driving speed, and strong generalization ability.
[0025] The further effects of the above non-conventional optional methods will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0027] Figure 1 is a schematic diagram of the main steps of the method for driving the expression of a virtual character according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of the recognition result of face key points according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of the normalized face key point result obtained by normalizing the face key points according to an embodiment of the present invention;
[0030] Figure 4 is a schematic diagram of the network structure of the blink expression coefficient model according to an embodiment of the present invention;
[0031] Figure 5 is a schematic diagram of the network structure of the non-blink expression coefficient model according to an embodiment of the present invention;
[0032] Figure 6 is a schematic diagram of the main modules of the device for driving the expression of a virtual character according to an embodiment of the present invention;
[0033] Figure 7 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0034] Figure 8 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The exemplary embodiments of the present invention will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0036] In the technical solution of the present invention, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0037] To solve the technical problems existing in the prior art, the present invention provides a method and device for driving the expressions of virtual characters. By learning the correspondence between the position coordinates of facial key points and an expression coefficient model (in the embodiments of the present invention, for example, a deformator), the expressions of virtual characters are driven. Among them, the deformator, that is, blendshape, refers to the parameters that can drive a 3D character model, and it can be assigned different weights, and different weights determine the degree of opening and closing of the action. The present invention mainly learns the corresponding expression coefficient model based on the input image information to drive the model expression, so that the model shows the same expression as a real person. The present invention can drive virtual characters quickly, relatively realistically, naturally and accurately, with extremely low network consumption, fast driving speed, and strong generalization.
[0038] A method for driving the expressions of virtual characters provided by the present invention first obtains a facial image through a camera, outputs an expression coefficient value as an expression control parameter through an expression coefficient model, and then drives a 3D model of a virtual character to perform an expression display.
[0039] Figure 1 It is a schematic diagram of the main steps of the method for driving the expressions of virtual characters according to the embodiments of the present invention. As Figure 1 shown, the method for driving the expressions of virtual characters in the embodiments of the present invention mainly includes the following steps S101 to step S103.
[0040] Step S101: Obtain a facial image, and extract key points from the facial image to obtain facial key points. According to the technical solution of the present invention, a person's photo can be obtained through a camera and processed to obtain a facial image. For example, object detection and matting are performed on the person's photo to obtain a facial image. Then, key points are extracted from the facial image to obtain facial key points. Among them, when performing key point extraction, it can be based on an existing facial key point recognition model for extraction. After obtaining the facial key points, the coordinate values corresponding to each facial key point can be obtained, including the abscissa x and the ordinate y. Subsequently, when performing processing operations on the facial key points, they are all carried out for the coordinate values of the facial key points.
[0041] In an embodiment of the present invention, when performing key point extraction on the face image, the implementation principle is, for example: input the face image into a pre-trained face key point recognition model for key point extraction, and the face key point recognition model is obtained by training face images with key point annotations. The face key point recognition model is a deep learning model. To meet the generalization of the model, face pictures of multiple people can be collected, a training data set is formed by annotating the key points of each face picture, and then the face key point recognition model is trained through deep learning. Figure 2 It is a schematic diagram of the face key point recognition result of the embodiment of the present invention. In the embodiment of the present invention, the face key points recognized by the face key point recognition model are as Figure 2 shown. In this embodiment, a 300-point face key point recognition model can be selected to perform face key point extraction, and a total of 300 face key points are extracted. Combining its abscissa x and ordinate y, there are a total of 600 coordinate values.
[0042] Step S102: Input the face key points into a pre-trained expression coefficient model to obtain the expression coefficient value corresponding to the face image. The expression coefficient model is used to obtain the face expression according to the face key points, and then generate the expression coefficient value for driving the expression of the virtual character.
[0043] According to an embodiment of the present invention, before inputting the face key points into a pre-trained expression coefficient model, it further includes: restricting the face image to a picture of a specified size, and normalizing the face key points. By normalizing the face key points, it is convenient for subsequent data processing, simplifies the model complexity, and improves the model operation efficiency.
[0044] According to the technical solution of the embodiment of the present invention, the normalization process of the face key points may specifically include: obtaining the abscissa and ordinate of the center point of the face key points, as well as the height and width of the face image; for each face key point, calculating the normalized abscissa of the face key point according to the abscissa of the face key point, the abscissa of the center point, and the height and width of the face image, and calculating the normalized ordinate of the face key point according to the ordinate of the face key point, the ordinate of the center point, and the height and width of the face image; translating the center point of the face key points to the upper left corner of the face image of the specified size to translate the normalized abscissa and the normalized ordinate, and complete the normalization process of the face key points.
[0045] In the embodiment of the present invention, if the x coordinates of all face key points are denoted as x_points and the y coordinates of all face key points are denoted as y_points, then the entire normalization process is as follows:
[0046] Solve for the minimum abscissa value of all face key points:
[0047] min_x_val = min(x_points) (1)
[0048] Solve for the maximum abscissa value of all face key points:
[0049] max_x_val = max(x_points) (2)
[0050] Solve for the minimum ordinate value of all face key points:
[0051] min_y_val = min(y_points) (3)
[0052] Solve for the maximum ordinate value of all face key points:
[0053] max_y_val = max(y_points) (4)
[0054] Obtain the width of the face image based on the maximum and minimum abscissa values:
[0055] width = max_x_val - min_x_val (5)
[0056] Obtain the height of the face image based on the maximum and minimum ordinate values:
[0057] height = max_y_val - min_y_val (6)
[0058] Solve for the abscissa of the center point of all face key points:
[0059] mean_x = (max_x_val + min_x_val) / 2 (7)
[0060] Solve for the ordinate of the center point of all face key points:
[0061] mean_y = (max_y_val + min_y_val) / 2 (8)
[0062] Subtract the abscissa of the center point from the abscissas of all face key points to update the abscissas of the face key points:
[0063] x_points = x_points - mean_x (9)
[0064] Subtract the ordinate of the center point from the ordinates of all face key points to update the ordinates of the face key points:
[0065] y_points = y_points - mean_y (10)
[0066] Through the above formulas (9) and (10), the center point of the facial key points can be moved to the center of the face;
[0067] Calculate the normalized abscissa of each facial key point:
[0068]
[0069] Calculate the normalized ordinate of each facial key point:
[0070]
[0071] According to the above formulas (11) and (12), the abscissa and ordinate of the facial key points are normalized by dividing them by the maximum value of the height and width of the facial image; then scaled to the pre-specified image size; in order to avoid the phenomenon of being close to the edge, multiply by a coefficient of 0.9, thus obtaining the normalized abscissa and normalized ordinate of each facial key point;
[0072] Translate the normalized abscissa of each facial key point:
[0073] x_points = x_points + (m_width / 2) (13)
[0074] Translate the normalized ordinate of each facial key point:
[0075] y_points = y_points + (m_height / 2) (14)
[0076] Through the above formulas (13) and (14), the abscissa and ordinate of the center point are translated to the upper left corner of the set picture again. At this time, the abscissa value of the facial key points is between (0, m_width), and the ordinate value is between (0, m_height).
[0077] Figure 3 It is a schematic diagram of the normalized facial key point result obtained by normalizing the facial key points in the embodiment of the present invention, which shows the normalized facial key points obtained by normalizing the facial key points in the embodiment of the present invention. The position of the white pixels in the figure is the coordinate value of the extracted facial key points. In the embodiment of the present invention, m_width = 96 and m_height = 96 can be set, that is, the values of images with different resolutions input into the subsequent model will be normalized to between 0 and 96. After dividing by 96, they are normalized to between 0 and 1 and input into the expression coefficient model for training, which can reduce the complexity of data processing and improve the model training efficiency.
[0078] According to an embodiment of the present invention, the expression coefficient model is trained, for example, in the following manner: obtaining a set of face images; extracting key points from each face image to obtain the face key points corresponding to each face image; for each face image, obtaining the hybrid deformation value of the face image based on the face model reconstruction method, and using the hybrid deformation value to label the face key points corresponding to the face image; and learning the face key points corresponding to the labeled face images to obtain the expression coefficient model.
[0079] In the embodiment of the present invention, in order to meet the generalization of the model, it is necessary to capture various expressions of multiple people. Only by comprehensively capturing the expressions of various people can the network learn more expression features. Specifically, it is required that the recording personnel make various exaggerated expressions such as opening the mouth, pursing the lips, smiling, blinking, etc. At the same time, pictures in the normal speech state are also recorded to form a set of face images. For each face image in the set of face images, key points need to be extracted to obtain the face key points corresponding to each face image. When extracting the key points, they can be extracted according to the method described in step S101, and finally the face key points corresponding to each face image are obtained. Similarly, after the face key points are extracted and before the face key points are labeled, the face key points corresponding to each extracted face image can also be normalized, and the normalization of the face key points can be implemented according to the aforementioned normalization method.
[0080] Before learning the correspondence between facial keypoints and expression coefficient values, it is necessary to label the facial keypoints corresponding to the facial image using the blendshape value (blendshape value) of the facial image. When determining the blendshape value of the facial image, training data is first made. The input of the network is a facial picture, and what is required to be obtained is the blendshape value corresponding to each picture. Blendshape refers to a technique in which a single mesh realizes the combination between many predefined shapes and any quantity, and is called a deformation target in some animation production software such as Maya and 3dsMax. Mixing multiple shapes can express expressions such as smiling, frowning, and closing eyes. Currently, in the field of 3DMM (3D morphable model, 3D deformation model), the three-dimensional image of the corresponding person can be reconstructed from a single picture. The facial reconstruction of a single picture is based on some three-dimensional models in special formats. Such models contain changeable vertex coordinates to make the model achieve the deformation effect, and at the same time contain blendshape values that can control facial expressions. In the embodiment of the present invention, according to the facial images in the obtained facial image set mentioned above, the corresponding blendshape value of each facial image can be obtained through the facial model reconstruction method. And the blendshape value obtained by the model is used as the label of the facial keypoints corresponding to the facial image of the present invention to label the facial keypoints corresponding to the facial image. After that, the facial keypoints corresponding to the labeled facial image can be learned to obtain an expression coefficient model.
[0081] According to another embodiment of the present invention, the pre-trained expression coefficient model includes a blinking expression coefficient model and a non-blinking expression coefficient model. Among them, the blinking expression coefficient model is only used to calculate the expression coefficient value corresponding to the blinking expression, and the non-blinking expression coefficient model is used to calculate the expression coefficient values corresponding to other facial expressions except the blinking expression. In the embodiment of the present invention, a step-by-step network is used to learn the relationships of different facial parts. When driving the expression change of a virtual character, opening the mouth will affect the change of the eyes. Therefore, the present invention proposes to learn the mapping relationship between facial keypoints and facial expressions by parts. In addition, since a normal person will slightly close the other eye when blinking one eye, but this effect is not desired to occur on the 3D model. Therefore, the present invention proposes that the left and right eyes learn the corresponding expression coefficient values respectively to obtain the blinking expression coefficient. That is to say, there are two blinking expression coefficient models in the present invention, which are respectively used to calculate the expression coefficient values of the left eye and the right eye.
[0082] According to one of the embodiments of the present invention, when inputting the facial keypoints into the pre-trained expression coefficient model to obtain the expression coefficient value corresponding to the facial image, it may specifically include:
[0083] Extract eye key points from the face key points, and input the eye key points into the blink expression coefficient model to obtain the blink expression coefficient value corresponding to the face image;
[0084] Input the face key points into the non-blink expression coefficient model to obtain the non-blink expression coefficient value corresponding to the face image;
[0085] Obtain the expression coefficient value corresponding to the face image according to the blink expression coefficient value and the non-blink expression coefficient value.
[0086] Specifically, the blink expression coefficient model can be trained in the following way: Obtain a set of face images; Extract key points for each face image to obtain the face key points corresponding to each face image; For each face image, obtain the blend shape value of the face image based on the face model reconstruction method, and use the blend shape value to label the face key points corresponding to the face image; For the face key points corresponding to the labeled face image, extract eye key points from the face key points, and use a deep learning optimization algorithm to learn the eye key points to obtain the blink expression coefficient model. Among them, when training the blink expression coefficient model, only the coordinate information of dozens of eye key points related to the eyes in the labeled face key points needs to be learned.
[0087] Figure 4 It is a schematic diagram of the network structure of the blink expression coefficient model of the embodiment of the present invention. As Figure 4 shown, the network of the blink expression coefficient model of the embodiment of the present invention mainly applies a three-layer fully connected network, and an activation (RELU) function is added in the middle of the fully connected network. Since the input parameters of the network are few, only the coordinate information of dozens of key points of the eyes, there is no need for too many complex network layers. When training the blink expression coefficient model, the loss function is set as the mean square loss function MSE loss function, and the model output is a 1*1 parameter weight value for controlling blinking. Therefore, the loss is as follows formula (15), where y is the label value corresponding to the face key points, and y' is the predicted expression coefficient value,
[0088] loss = MSE(y, y') (15).
[0089] The training method of the blink expression coefficient model can select the deep learning optimization algorithm Adam training method, where opt params is the network parameter, and weight decay is the regularization parameter. Then the blink expression coefficient model opt is as follows formula (16):
[0090] opt = Adam(opt params , weight decay= 0.0001) (16).
[0091] According to an embodiment of the present invention, the non-blink expression coefficient model is trained as follows: obtaining a set of face images; extracting key points for each face image to obtain the face key points corresponding to each face image; for each face image, obtaining the mixed deformation value of the face image based on the face model reconstruction method, and using the mixed deformation value to label the face key points corresponding to the face image; transforming the coordinate format of the face key points corresponding to the labeled face image to obtain a square matrix, and using a deep learning optimization algorithm to learn the square matrix to obtain the non-blink expression coefficient model. The non-blink expression coefficient model mainly learns the mapping relationship of other expressions of the face except blinking, so all face key points can be input for the training of the network model.
[0092] Among them, when transforming the coordinate format of the face key points corresponding to the labeled face image to obtain a square matrix, if the coordinates of the face key points corresponding to the labeled face image cannot be directly transformed into a square matrix, 0 values are used for supplementation until the key point coordinate set composed of the coordinates of the face key points corresponding to the labeled face image and the 0 values can be transformed in coordinate format to obtain a square matrix. For example, in combination with the foregoing embodiment, the number of extracted face key points is 300, and there are 600 coordinate values in total in combination with its abscissa x and ordinate y, which cannot be directly converted into a square matrix, so some 0 values need to be supplemented. Specifically, in the embodiment of the present invention, 300 non-face key points (such as Figure 3 the black pixel points therein) can be added, and the pixel values of the black pixel points are 0, to form a key point coordinate set, so that 900 values in the key point coordinate set can be formatted and converted to obtain a 30*30 square matrix for input to the network for network training. Specifically, it can be implemented through a function in python. Its principle is actually to arrange 900 values in the key point coordinate set into a 30-row and 30-column square matrix, that is, to form a grayscale image and send it into the network for training.
[0093] Figure 5 is a schematic diagram of the network structure of the non-blink expression coefficient model according to an embodiment of the present invention. As Figure 5As shown, the input of the non-blinking expression coefficient model in the embodiment of the present invention is the position coordinates of the key points of the entire face. To meet the network input format, the key point coordinate information format is converted and reshaped into a 30*30 picture. The network structure of the non-blinking expression coefficient model includes a convolutional layer (the first 5 network layers) and a fully connected layer (the last 3 network layers). For example, 25*8*8 represents a convolutional layer with a convolutional kernel size of 8*8 and 25 in number, and 1*1800 represents that the parameters of the fully connected layer are 1800. The output of the network is a group of 1*51 vectors, and this vector is the expression coefficient value of the non-blinking expression coefficient model. The 51 expression coefficient values respectively drive 51 parts of the face to complete actions such as opening the mouth, grinning, pouting, and closing the mouth. When training the non-blinking expression coefficient model, by using MSE as the loss function and using the Adam optimization method for parameter update, among them, the training method of the non-blinking expression coefficient model is the same as that of the blinking expression coefficient model. The generated non-blinking expression coefficient model opt can be seen in formula (16).
[0094] In addition, when training the non-blinking expression coefficient model, the present invention takes into account that the expression coefficient value corresponding to the open-mouth expression is particularly important. Therefore, when training the non-blinking expression coefficient model, the loss function includes a first loss function corresponding to the entire face expression except the blinking expression and a second loss function corresponding to the open-mouth expression. The loss function is shown in the following formula (17):
[0095] Loss = MSE(y, y') + λ * MSE(mouth_open_label, mouth_open_predict) (17).
[0096] In the above formula (17), the first term MSE(y, y') is the first loss function corresponding to the entire face expression except the blinking expression, where y is the label value corresponding to the face key points, and y' is the predicted expression coefficient value; the second term is the second loss function corresponding to the open-mouth expression, where λ is the weight coefficient, mouth_open_label is the label value of the open-mouth expression coefficient value, and mouth_open_predict is the predicted value of the open-mouth expression coefficient value. In the embodiment of the present invention, λ can be set to 0.3, which can ensure that the learned open-mouth value is more correct and occupies a larger weight.
[0097] Step S103: Perform expression driving on the virtual character according to the expression coefficient value.
[0098] According to the above steps S101 to S103, by obtaining a face image and performing key point extraction on the face image to obtain face key points; inputting the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image; according to the expression coefficient value, a technical solution for driving the expression of a virtual character, by learning the corresponding relationship between the position coordinates of the face key points and the expression coefficient model to drive the expression of the virtual character, enabling the virtual character model to show the same expression as a real person, can drive the virtual character quickly, realistically, naturally, and accurately, with extremely low network consumption, fast driving speed, and strong generalization ability.
[0099] Figure 6 is a schematic diagram of the main modules of the device for driving the expression of a virtual character according to an embodiment of the present invention. As Figure 6 shown, the device 600 for driving the expression of a virtual character according to an embodiment of the present invention mainly includes a key point acquisition module 601, a coefficient value calculation module 602, and an expression driving module 603.
[0100] The key point acquisition module 601 is used to obtain a face image and perform key point extraction on the face image to obtain face key points;
[0101] The coefficient value calculation module 602 is used to input the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image;
[0102] The expression driving module 603 is used to drive the expression of the virtual character according to the expression coefficient value.
[0103] According to an embodiment of the present invention, the key point acquisition module 601 can also be used to: perform key point extraction by inputting the face image into a pre-trained face key point recognition model, and the face key point recognition model is obtained by training face images with key point annotations.
[0104] According to another embodiment of the present invention, the device 600 for driving the expression of a virtual character may further include a key point normalization processing module (not shown in the figure), which is used to: before inputting the face key points into a pre-trained expression coefficient model, limit the face image to a picture of a specified size and perform normalization processing on the face key points.
[0105] According to another embodiment of the present invention, the key point normalization processing module (not shown in the figure) can also be used to: obtain the abscissa and ordinate of the center point of the face key points, as well as the height and width of the face image; for each face key point, calculate the normalized abscissa of the face key point according to the abscissa of the face key point, the abscissa of the center point, and the height and width of the face image, and calculate the normalized ordinate of the face key point according to the ordinate of the face key point, the ordinate of the center point, and the height and width of the face image; translate the center point of the face key point to the upper left corner of the face image with a specified size, so as to translate the normalized abscissa and the normalized ordinate, and complete the normalization processing of the face key points.
[0106] According to another embodiment of the present invention, the expression coefficient model is trained in the following manner: obtain a set of face images; perform key point extraction on each face image to obtain the face key points corresponding to each face image; for each face image, obtain the hybrid deformation value of the face image based on the face model reconstruction method, and use the hybrid deformation value to label the face key points corresponding to the face image; perform learning on the face key points corresponding to the labeled face images to obtain the expression coefficient model.
[0107] According to another embodiment of the present invention, the pre-trained expression coefficient model includes a blinking expression coefficient model and a non-blinking expression coefficient model;
[0108] The coefficient value calculation module 602 can also be used to: extract eye key points from the face key points, and input the eye key points into the blinking expression coefficient model to obtain the blinking expression coefficient value corresponding to the face image; input the face key points into the non-blinking expression coefficient model to obtain the non-blinking expression coefficient value corresponding to the face image; obtain the expression coefficient value corresponding to the face image according to the blinking expression coefficient value and the non-blinking expression coefficient value.
[0109] According to another embodiment of the present invention, the blinking expression coefficient model is trained in the following manner: obtain a set of face images; perform key point extraction on each face image to obtain the face key points corresponding to each face image; for each face image, obtain the hybrid deformation value of the face image based on the face model reconstruction method, and use the hybrid deformation value to label the face key points corresponding to the face image; for the face key points corresponding to the labeled face images, extract eye key points from the face key points, and perform learning on the eye key points through a deep learning optimization algorithm to obtain the blinking expression coefficient model.
[0110] According to another embodiment of the present invention, the non-blink expression coefficient model is trained as follows: obtaining a set of face images; extracting key points for each face image to obtain the face key points corresponding to each face image; for each face image, obtaining the hybrid deformation value of the face image based on the face model reconstruction method, and using the hybrid deformation value to label the face key points corresponding to the face image; performing coordinate format transformation on the face key points corresponding to the labeled face image to obtain a square matrix, and learning the square matrix through a deep learning optimization algorithm to obtain the non-blink expression coefficient model.
[0111] According to another embodiment of the present invention, when the non-blink expression coefficient model is trained, the loss function includes a first loss function corresponding to the full-face expression other than the blink expression and a second loss function corresponding to the open-mouth expression.
[0112] According to another embodiment of the present invention, the device 600 for driving the virtual character expression may further include a coordinate format transformation module (not shown in the figure), configured to: before performing coordinate format transformation on the face key points corresponding to the labeled face image to obtain a square matrix, in the case where the coordinates of the face key points corresponding to the labeled face image cannot be directly transformed into a square matrix, use 0 values for supplementation until the key point coordinate set composed of the coordinates of the face key points corresponding to the labeled face image and the 0 values can be subjected to coordinate format transformation to obtain a square matrix.
[0113] According to the technical solution of the embodiment of the present invention, by obtaining a face image, extracting key points from the face image to obtain face key points; inputting the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image; and driving the virtual character expression according to the expression coefficient value, the technical solution drives the virtual character expression by learning the corresponding relationship between the position coordinates of the face key points and the expression coefficient model, enables the virtual character model to show the same expression as a real person, can drive the virtual character quickly, realistically, naturally and accurately, has extremely low network consumption, fast driving speed, and strong generalization ability.
[0114] Figure 7 An exemplary system architecture 700 is shown that can apply the method for driving a virtual character expression or the device for driving a virtual character expression according to the embodiments of the present invention.
[0115] As Figure 7 shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705. The network 704 is used to provide a medium for a communication link between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0116] Users can use terminal devices 701, 702, and 703 to interact with server 705 through network 704 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 701, 702, and 703, such as image processing applications, image key point extraction applications, image feature extraction applications, camera software, etc. (only for examples).
[0117] Terminal devices 701, 702, and 703 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0118] Server 705 can be a server that provides various services, such as a background management server that supports virtual character expression driving requests sent by users using terminal devices 701, 702, and 703 (only for examples). The background management server can obtain a face image for the received virtual character expression driving request and other data, and perform key point extraction on the face image to obtain face key points; input the face key points into a pre-trained expression coefficient model to obtain the expression coefficient value corresponding to the face image; perform expression driving and other processing on the virtual character according to the expression coefficient value, and feedback the processing result (such as the virtual character expression display result - only for examples) to the terminal device.
[0119] It should be noted that the method for driving virtual character expressions provided in the embodiments of the present invention is generally executed by server 705. Correspondingly, the device for driving virtual character expressions is generally set in server 705.
[0120] It should be understood that Figure 7 the numbers of the terminal devices, network, and server in
[0121] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, and server. Figure 8 Figure 8 The terminal device or server shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0122] Figure 8 As shown in Figure 8As shown, computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0123] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0124] Specifically, according to an embodiment disclosed by the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment disclosed by the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by a central processing unit (CPU) 801, the above functions defined in the system of the present invention are executed.
[0125] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0127] The units or modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a key point acquisition module, a coefficient value calculation module, and an expression driving module. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases. For example, the key point acquisition module can also be described as "a module for acquiring a face image and performing key point extraction on the face image to obtain face key points".
[0128] On the other hand, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: acquiring a face image, and performing key point extraction on the face image to obtain face key points; inputting the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image; and driving the expression of a virtual character according to the expression coefficient value.
[0129] According to the technical solution of the embodiments of the present invention, by acquiring a face image, performing key point extraction on the face image to obtain face key points, inputting the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image, and driving the expression of a virtual character according to the expression coefficient value, the technical solution drives the expression of the virtual character by learning the corresponding relationship between the position coordinates of the face key points and the expression coefficient model, enables the virtual character model to show the same expression as a real person, can drive the virtual character quickly, accurately, and naturally in a relatively real manner, has extremely low network consumption, fast driving speed, and strong generalization ability.
[0130] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for driving the expression of a virtual character, characterized in that, Including: Obtain a face image, and perform key point extraction on the face image to obtain face key points; Input the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image; the pre-trained expression coefficient model includes a blinking expression coefficient model and a non-blinking expression coefficient model; Perform expression driving on the virtual character according to the expression coefficient value; Among them, inputting the face key points into a pre-trained expression coefficient model to obtain an expression coefficient value corresponding to the face image includes: extracting eye key points from the face key points, and inputting the eye key points into the blinking expression coefficient model to obtain a blinking expression coefficient value corresponding to the face image; inputting the face key points into the non-blinking expression coefficient model to obtain a non-blinking expression coefficient value corresponding to the face image; obtaining an expression coefficient value corresponding to the face image according to the blinking expression coefficient value and the non-blinking expression coefficient value.
2. The method according to claim 1, wherein Performing key point extraction on the face image includes: Performing key point extraction by inputting the face image into a pre-trained face key point recognition model, and the face key point recognition model is obtained by training face images with key point annotations.
3. The method according to claim 1, wherein Before inputting the face key points into a pre-trained expression coefficient model, it further includes: Restricting the face image to a picture of a specified size, and performing normalization processing on the face key points.
4. The method according to claim 3, wherein Performing normalization processing on the face key points includes: Obtaining the abscissa and ordinate of the center point of the face key points, as well as the height and width of the face image; For each face key point, calculating the normalized abscissa of the face key point according to the abscissa of the face key point, the abscissa of the center point, and the height and width of the face image, and calculating the normalized ordinate of the face key point according to the ordinate of the face key point, the ordinate of the center point, and the height and width of the face image; Translating the center point of the face key points to the upper left corner of the face image of the specified size to perform translation on the normalized abscissa and the normalized ordinate, and completing the normalization processing of the face key points.
5. The method according to claim 1, wherein The expression coefficient model is obtained by training in the following manner: Obtain a set of face images; Perform key point extraction on each face image to obtain face key points corresponding to each face image; For each face image, obtain the blend shape value of the face image based on the face model reconstruction method, and use the blend shape value to label the face key points corresponding to the face image; Learn the face key points corresponding to the labeled face images to obtain the expression coefficient model.
6. The method according to claim 1, characterized in that The blinking expression coefficient model is obtained by training in the following manner: Obtain a set of face images; Perform key point extraction on each face image to obtain face key points corresponding to each face image; For each face image, obtain the blend shape value of the face image based on the face model reconstruction method, and use the blend shape value to label the face key points corresponding to the face image; For the facial key points corresponding to the face image after marking, extract the eye key points from the facial key points, and use a deep learning optimization algorithm to learn the eye key points to obtain the blinking expression coefficient model.
7. The method according to claim 1, characterized in that, The non-blinking expression coefficient model is trained in the following manner: Obtain a set of face images; Extract key points for each face image to obtain the facial key points corresponding to each face image; For each face image, obtain the hybrid deformation value of the face image based on the face model reconstruction method, and use the hybrid deformation value to mark the facial key points corresponding to the face image; Perform coordinate format transformation on the facial key points corresponding to the marked face image to obtain a square matrix, and use a deep learning optimization algorithm to learn the square matrix to obtain the non-blinking expression coefficient model.
8. The method according to claim 7, wherein When the non-blinking expression coefficient model is trained, the loss function includes a first loss function corresponding to the full-face expression other than the blinking expression and a second loss function corresponding to the open-mouth expression.
9. The method according to claim 7, characterized in that Before performing coordinate format transformation on the facial key points corresponding to the marked face image to obtain a square matrix, it further includes: In the case where the coordinates of the facial key points corresponding to the marked face image cannot be directly transformed into a square matrix, use 0 values for supplementation until the key point coordinate set composed of the coordinates of the facial key points corresponding to the marked face image and the 0 values can be subjected to coordinate format transformation to obtain a square matrix.
10. A device for driving the expressions of a virtual character, characterized in that, It includes: A key point acquisition module, configured to acquire a face image and extract key points from the face image to obtain facial key points; A coefficient value calculation module, configured to input the facial key points into a pre-trained expression coefficient model to obtain the expression coefficient value corresponding to the face image; the pre-trained expression coefficient model includes a blinking expression coefficient model and a non-blinking expression coefficient model; The coefficient value calculation module is further configured to: extract eye key points from the facial key points, and input the eye key points into the blinking expression coefficient model to obtain the blinking expression coefficient value corresponding to the face image; Input the facial key points into the non-blinking expression coefficient model to obtain the non-blinking expression coefficient value corresponding to the face image; Obtain the expression coefficient value corresponding to the face image according to the blinking expression coefficient value and the non-blinking expression coefficient value; An expression driving module, configured to drive the expression of the virtual character according to the expression coefficient value.
11. An electronic device for driving the expressions of a virtual character, characterized in that, It includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Lightweight human face key point detection method based on Gaussian heat map regression
CN112580515A
Three-dimensional facial expression rendering method and device, storage medium and electronic device
CN114360018A