Data processing methods and apparatuses, electronic devices and media for virtual avatars

By acquiring facial key points from the target image and adjusting them using expression base and expression coefficients to generate a second facial key point image, the problem of high computational resource consumption in virtual human portrait video generation is solved, and the generation effect and stability are improved.

CN119648877BActive Publication Date: 2025-10-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411767448.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-10-28
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing technologies consume excessive computing resources and produce unstable results when generating realistic and expressive virtual human videos, making it difficult to achieve efficient facial expression synchronization.

Method used

By acquiring facial key points from the target image, adjusting them using the expression base and expression coefficients, a second facial key point image is generated. The expression coefficients are then directly manipulated to predict the generation effect and adjust computational resources, reducing subsequent computational consumption.

Benefits of technology

It achieves computational resource conservation when generating virtual human portrait videos and improves the stability and generation effect of facial expression synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648877B_ABST
    Figure CN119648877B_ABST
Patent Text Reader

Abstract

This disclosure provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for virtual avatars, relating to the field of artificial intelligence, and particularly to the fields of image processing, digital humans, and deep learning. The implementation scheme is as follows: acquiring a target image including the face of a target object; extracting facial key points from the target image to obtain a first facial key point image; obtaining a first set of expression coefficients based on the first facial key point image and a preset set of expression bases; adjusting the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients; obtaining a second facial key point image based on the second set of expression coefficients and a set of expression bases; and obtaining a first image corresponding to the target image after transforming the facial expression based on the second facial key point image, the expression image, and the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and more particularly to the fields of deep learning, image processing, and digital human technology. Specifically, it relates to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for virtual avatars. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] With the development of artificial intelligence technology, virtual digital humans have been widely used in live streaming, news broadcasting, voice prompts, and other fields. Typically, the virtual digital human needs to perform actions and expressions synchronized with the audio to be broadcast, resulting in audio-driven video. Generating realistic and expressive portrait videos from a single facial image through audio-driven technology has broad application prospects, covering multiple fields from digital media to games and film production. Summary of the Invention

[0004] This disclosure provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for virtual avatars.

[0005] According to one aspect of this disclosure, a data processing method for virtual avatars is provided, comprising: acquiring a target image including the face of a target object; extracting facial key points based on the target image to obtain a first facial key point image; obtaining a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases; adjusting the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the target image after changing the facial expression; obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases; and obtaining a first image corresponding to the target image after changing the facial expression based on the second facial key point image and the target image.

[0006] According to another aspect of this disclosure, a model training method is provided, comprising: acquiring a first target image including a face of a target object and a first label image; extracting facial key points based on the first target image to obtain a first facial key point image; obtaining a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases; adjusting the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the first target image after changing facial expression; obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases; obtaining a first image corresponding to the target image after changing facial expression through a video generation model based on the second facial key point image and the first target image; determining a first loss value based on the first image and the first label image through a preset loss function; and adjusting the parameter values ​​of the video generation model based on the first loss value.

[0007] According to another aspect of this disclosure, a data processing apparatus for virtual avatars is provided, comprising: a first acquisition unit configured to acquire a target image including the face of a target object; a first key point extraction unit configured to extract facial key points based on the target image to obtain a first facial key point image; a first calculation unit configured to obtain a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases; a second calculation unit configured to adjust the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the target image after changing the facial expression; a second acquisition unit configured to obtain a second facial key point image based on the second set of expression coefficients and the preset set of expression bases; and a first generation unit configured to obtain a first image corresponding to the target image after changing the facial expression based on the second facial key point image and the target image.

[0008] According to another aspect of this disclosure, a model training apparatus is provided, comprising: a third acquisition unit configured to acquire a first target image including a face of a target object and a first label image; a second key point extraction unit configured to extract facial key points based on the first target image to obtain a first facial key point image; a third calculation unit configured to obtain a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases; and a fourth calculation unit configured to adjust the corresponding expression coefficients in the first set of expression coefficients. The system comprises: a first image and a second image; a second generation unit; a third calculation unit; and an adjustment unit. The fourth acquisition unit is configured to obtain a second set of expression coefficients corresponding to the facial expression transformation of the first target image; a fifth calculation unit is configured to determine a first loss value based on the first image and the first label image using a preset loss function; and a sixth calculation unit is configured to adjust the parameter values ​​of the video generation model based on the first loss value.

[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.

[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.

[0011] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.

[0012] According to one or more embodiments of this disclosure, facial key points are directly manipulated based on expression base and expression coefficients to obtain corresponding second facial key point images. Since the computation time and resources consumed in generating the image corresponding to the target image after facial expression transformation are much greater than the generation of the second facial key point image, the generation effect of subsequent images can be judged in advance and adjusted in time through the second facial key point image, thereby saving computational resources.

[0013] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Attached Figure Description

[0014] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;

[0016] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;

[0017] Figure 3 A schematic diagram of a video generation model for generating video according to an embodiment of the present disclosure is shown;

[0018] Figure 4 A flowchart of a model training method according to an embodiment of the present disclosure is shown;

[0019] Figure 5 A structural block diagram of a data processing apparatus according to an embodiment of the present disclosure is shown;

[0020] Figure 6 A structural block diagram of a model training apparatus according to an embodiment of the present disclosure is shown; and

[0021] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to define the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0024] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0025] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0026] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0027] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of data processing methods.

[0028] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.

[0029] exist Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0030] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to input target images, edit facial expression coefficients, or display images / videos. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0031] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0032] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0033] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0034] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0035] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0036] In some embodiments, server 120 may be a distributed system server or a server integrated with blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and virtual private servers (VPS) services.

[0037] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as images, emoji base files, and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located remotely to server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0038] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0039] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0040] In early studies, researchers achieved face reconstruction by constructing parametric models of faces (such as 3DMM). 3DMM can model features such as shape, expression, texture, and angle. However, the face model rendering algorithm based on 3DMM has poor performance and cannot generate high-precision textures, teeth, and other detailed areas.

[0041] Recently, deep learning-based methods have been widely studied due to their excellent video generation performance. The two most representative approaches are GAN-based methods (such as the StyleGAN series) and diffusion model-based methods (such as Hallo, Follow-Your-Emoji, EchoMimic, and Aniportrait). GAN-based methods generate more realistic portraits, but their diversity is significantly affected by data distribution, and the training process is unstable, prone to pattern collapse. Diffusion model-based methods, on the other hand, can generate high-quality, high-resolution portrait videos with better diversity, but require more computational resources.

[0042] Therefore, according to embodiments of this disclosure, a data processing method for virtual avatars is provided to generate corresponding portrait videos based on audio drivers. Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown, such as... Figure 2As shown, method 200 includes: acquiring a target image including the face of a target object (step 210); extracting facial key points based on the target image to obtain a first facial key point image (step 220); obtaining a first set of expression coefficients corresponding to facial expressions in the target image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases (step 230); adjusting the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the facial expression transformation of the target image (step 240); obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases (step 250); and obtaining an image corresponding to the facial expression transformation of the target image based on the second facial key point image and the target image (step 260).

[0043] According to embodiments of this disclosure, facial key points are directly manipulated based on expression base and expression coefficients to obtain corresponding second facial key point images. Since the computation time and resources consumed in generating the image corresponding to the target image after facial expression transformation are much greater than the generation of the second facial key point image, the generation effect of subsequent images can be judged in advance and adjusted in time through the second facial key point image, thereby saving computational resources.

[0044] In step 210, a target image including the face of the target object is obtained.

[0045] In this disclosure, the target object is not limited to people; it can also be an animal, or anthropomorphic animals, objects, etc., without limitation. For example, taking a real person as the target object, the generated image can be an image obtained by transforming the facial expression of the person in the target image. In some examples, the target object is used to generate a virtual digital human. This virtual digital human can be a two-dimensional virtual digital human generated based on the target object.

[0046] In some embodiments, the generated image may include not only the face of the target object, but also the background area in the target image other than the face of the target object.

[0047] In step 220, facial key points are extracted based on the target image to obtain a first facial key point image.

[0048] Specifically, in some examples, facial landmark detection refers to locating key regions in a facial image using algorithms, such as eyebrows, eyes, nose, mouth, and facial contours. During the detection process, the system returns the coordinates of these key points, thereby enabling precise recognition and analysis of the target object's face.

[0049] In some examples, facial landmark detection can employ various suitable landmark annotation methods, such as 68 points, 96 / 98 points, and 106 / 186 points, without limitation. For instance, when annotating a face with 68 points, the facial landmarks are divided into internal landmarks and contour landmarks. The internal landmarks include 51 landmarks (eyebrows, eyes, nose, mouth), while the contour landmarks include 17 landmarks. This yields a facial landmark image. This image can include the coordinate or positional information of each landmark.

[0050] According to some embodiments, the first facial key point image and the second facial key point image include coordinate information of pupil key points.

[0051] In some examples, pupil keypoints can be further included in facial landmark detection. For eye-related applications such as face recognition, expression changing, and eye movement tracking, accurate pupil location is crucial. Using two keypoints to represent the left and right pupils provides more accurate positional information, aiding in subsequent, more refined expression change analysis and processing.

[0052] In some examples, 3D face reconstruction techniques can be used to reconstruct the face of the target object from the first image, thereby obtaining the facial key point image. For example, open-source plugins such as Media Pipe and FaceNet can be used to extract the 3D coordinates of the facial key points. Alternatively, it can be understood that 2D coordinates of the facial key points can also be extracted using OpenCV, etc., without limitation.

[0053] In step 230, based on the first facial key point image and a preset set of expression bases, a first set of expression coefficients corresponding to the facial expressions in the target image is obtained, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases.

[0054] In this disclosure, a pre-defined set of expression bases consists of deformable networks representing different expressions. Each deformable network is formed by changing a 3D average face model of an animated character under different expressions, and can also be referred to as other shapes of the basic shape, or deformation targets. For example, the basic shape can be a default shape, such as a face without expression. Other shapes of the basic shape are used for blending / deformation, which are different expressions (such as blinking the left eye, blinking the left eye, chin turning to the left when pouting). These are collectively referred to as blending shapes or deformation targets.

[0055] In this disclosure, expression coefficients are used to represent the amplitude (or degree) of expression change corresponding to other shapes of a basic shape (i.e., expression bases). For example, if an expression base represents a mouth deformation based on a basic shape (such as an expressionless face) to represent a jaw open, then the expression coefficients corresponding to that expression base can be used to represent the degree of mouth deformation (i.e., the size of the mouth opening). For example, the range corresponding to its expression coefficients is [0,1], where an expression coefficient of 0 indicates a closed mouth and an expression coefficient of 1 indicates a fully open mouth.

[0056] For example, the driving of a virtual face model can be implemented based on formula (1):

[0057]

[0058] Where B0 represents the average face model of the virtual face model, i.e., the basic shape mentioned above, and the neutral face model is a face model without expression. i Let α be the i-th expression basis among the n expression bases used to form a virtual face model. i Indicates the base of the expression B i The expression coefficient.

[0059] In some examples, based on a first facial keypoint image and a preset set of expression bases, the first set of expression coefficients corresponding to the facial expressions in the target image can be determined by formula (1). It is understood that any other suitable set of expression bases and driving formulas can also be used to determine the first set of expression coefficients, and no restriction is placed here.

[0060] In some examples, the preset set of expression bases represents a complete set of expression bases, meaning that various forms of facial expressions can be obtained through mixing / deformation using this set of expression bases. For example, this set of expression bases may include 52 expressions from ARKit, where each expression includes a corresponding expression name and an expression coefficient representing its current expression amplitude. Taking the ARKit expression base representing a wide-open left eye (eyeWideLeft) as an example, its expression coefficient gradually increases from 0 (i.e., eyeWideLeft = 0) to 1 (i.e., eyeWideLeft = 1), which in its corresponding facial model represents the left eye gradually changing from an unlabeled state to a wide-open state.

[0061] In some examples, any set of expression coefficients can correspond one-to-one with a pre-defined set of complete expression bases. When a facial expression does not involve a certain expression change in a certain part, such as when the eyes have no expression, the expression coefficient in that set of expression coefficients corresponding to that expression change in that part can be 0.

[0062] In step 240, the corresponding expression coefficients in the first set of expression coefficients are adjusted to obtain the second set of expression coefficients corresponding to the facial expression transformation of the target image; and in step 250, a second facial key point image is obtained based on the second set of expression coefficients and the preset set of expression bases.

[0063] As described above, the expression coefficient is used to represent the amplitude (or degree) of expression change corresponding to other shapes of the basic shape (i.e., the expression base). By changing the corresponding expression coefficient, the degree of expression change of the corresponding parts of the face can be changed, thereby achieving the effect of expression transformation.

[0064] According to some embodiments, the target image and the second facial key point image have the same dimensions.

[0065] In this embodiment, the target image and the second facial key point image have the same dimension, that is, the data between the images used to generate the final expression transformation image is aligned, which helps to improve the generation effect of the expression transformation image.

[0066] According to some embodiments, obtaining a first image corresponding to the target image after changing the facial expression based on the second facial key point image and the target image includes: generating an expression image based on the second facial key point image, wherein the expression image is generated based on the connection of facial key points related to the expression in the second facial key point image; and obtaining a first image corresponding to the target image after changing the facial expression based on the expression image and the target image.

[0067] It is understandable that when the second facial landmark image used to generate the expression image has the same dimension as the target image, the expression image generated based on the second facial landmark image will also have the same dimension as the target image. For example, both the target image and the second facial landmark image may be two-dimensional images, or both may be three-dimensional images. In this case, data alignment between the images used to finally generate the expression transformation image will further improve the generation effect of the expression transformation image.

[0068] Typically, the input target image is a two-dimensional image. Therefore, according to some embodiments, both the target image and the second facial key point image are two-dimensional key point images, and each expression base in the preset set of expression bases and the first facial key point image are three-dimensional key point images.

[0069] Understandably, 3D face data contains one more dimension of depth information than 2D face data. Therefore, 3D face recognition has advantages over 2D face recognition in both recognition accuracy and liveness detection accuracy. In the above embodiment, the expression base formed by the 3D keypoint image can achieve more refined expression changes.

[0070] Therefore, further, in order to improve the generation effect of the expression transformation image, according to some embodiments, obtaining the second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: obtaining a third facial key point image after the expression transformation of the target object's face based on the second set of expression coefficients and the preset set of expression bases, wherein the third facial key point image includes the coordinate information of three-dimensional key points; and obtaining the second facial key point image based on the coordinate information of the three-dimensional key points and a preset first transformation matrix, wherein the second facial key point image includes the coordinate information of two-dimensional key points.

[0071] In some examples, there is a mapping relationship between the 3D coordinates of the facial model and the corresponding 2D coordinates of the image. This relationship can be represented by a 3x4 affine transformation matrix. It's important to note that both the 3D coordinates of the facial model and the corresponding 2D coordinates of the image are normalized, each having one more dimension than before; that is, the 2D coordinates X of the image are... 2D =(X i ,Y i ,1) T The three-dimensional coordinates of the face model X 3D =(X i ,Y i Z i ,1) T The two-dimensional to three-dimensional coordinate transformation is achieved according to formula (2):

[0072] X 2D =P×X 3D Formula (2)

[0073] Here, P is the affine transformation matrix that needs to be obtained, which, when applied to a three-dimensional coordinate point, can yield the coordinates of the corresponding two-dimensional point.

[0074] For example, in face recognition, by detecting facial key points (including eyes, nose, mouth, etc.), the 3D key points of the face model and the corresponding 2D key points of the image are obtained. Then, the affine transformation matrix (i.e., the first transformation matrix) is calculated based on the coordinates of the key points.

[0075] In some examples, the corresponding affine transformation matrix can be obtained using any suitable algorithm, including but not limited to the Gold Standard Algorithm. The obtained affine transformation matrix, when applied to a 3D coordinate point, yields the coordinates of the corresponding 2D point, thus producing the aligned face image.

[0076] In some examples, the affine transformation matrix between the 3D and 2D key point coordinates of the corresponding facial model can be pre-calculated. Subsequently, the second facial key point image obtained after adjusting the expression coefficients can be dimensionally transformed based on this affine transformation matrix.

[0077] According to some embodiments, obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: obtaining a third facial key point image after performing expression transformation on the target object's face based on the second set of expression coefficients and the preset set of expression bases, wherein the third facial key point image includes coordinate information of three-dimensional key points; obtaining the second facial key point image based on the coordinate information of the three-dimensional key points and a preset second transformation matrix, wherein the preset second transformation matrix is ​​used to perform operations on the target object's face in three-dimensional space, wherein the operations include at least one of the following: rotation and translation.

[0078] In this embodiment, rotation and translation operations can be used to control the rotation and movement of the face, such as achieving fine control effects like tilting the head up 5 degrees or moving it 5 centimeters to the left.

[0079] In some examples, to estimate the correspondence between two-dimensional and three-dimensional coordinate points, the scaling factor, rotation matrix, and translation matrix can be further derived from the affine transformation relationship described above. Using the corresponding rotation and translation matrices, rotation and translation operations on the target object's face in three-dimensional space can then be performed.

[0080] According to some embodiments, the second facial key point image includes the coordinate information of the pupil key points. By controlling the pupil position, refined facial expression changes can be achieved.

[0081] As mentioned above, facial landmark detection can locate key regions in a facial image, including eyebrows, eyes, nose, mouth, and facial contours. Typically, the key points in the eye region are those within the eye socket area. During the detection process, the system returns the coordinate information of these key points, such as the coordinates of key points arranged around the eye socket.

[0082] According to some embodiments, the coordinate information of the pupil key point is determined based on the coordinate information of the orbital region key point in the second facial key point image.

[0083] In some examples, after obtaining a second facial keypoint image, the coordinates of the pupil keypoint can be located and determined based on the coordinates of the orbital region keypoints in the second facial keypoint image. Here, the pupil keypoint location can be determined based on any suitable position and number of orbital keypoints. For example, the pupil keypoint coordinates can be determined based on the coordinates of keypoints at the four vertices of the orbit (up, down, left, and right).

[0084] It is understandable that the pupil key point position is set to the middle of the eye, but it can also be set to a suitable other position of the eye based on the changed expression, and there are no restrictions here.

[0085] According to some embodiments, adjusting the corresponding expression coefficients in the first set of expression coefficients to obtain the second set of expression coefficients corresponding to the transformed facial expression of the target image includes: adjusting the corresponding expression coefficients in the first set of expression coefficients multiple times to obtain multiple second sets of expression coefficients, wherein each of the multiple second sets of expression coefficients corresponds to a transformed facial expression.

[0086] Specifically, the first set of expression coefficients can be adjusted multiple times, with each adjustment corresponding to a different facial expression of the target object. Through these multiple adjustments, various different facial expressions of the target object can be achieved.

[0087] According to some embodiments, adjusting the corresponding expression coefficients in the first set of expression coefficients multiple times to obtain multiple second sets of expression coefficients includes: acquiring video data including the face of a first object; dividing the video data into frames to obtain multiple second images including the face of the first object; extracting facial key points from each of the multiple second images to obtain a fourth facial key point image corresponding to each of the multiple second images; obtaining a third set of expression coefficients corresponding to each of the multiple second images based on the fourth facial key point image and the preset set of expression bases; for each of the multiple second images, determining the difference between the corresponding expression coefficient in the third set of expression coefficients corresponding to the second image and the corresponding expression coefficient in the third set of expression coefficients corresponding to the third image to obtain a set of expression coefficient differences; for each second image, adjusting the corresponding expression coefficient in the first set of expression coefficients based on the set of expression coefficient differences corresponding to the second image to obtain the multiple second sets of expression coefficients.

[0088] In this embodiment, a video-driven method is used to adjust the facial expression of the target object based on the facial expression of the first object in the video, so as to achieve high-precision facial video generation based on the target image.

[0089] In this disclosure, the acquired video data includes the face of a first object. The video data contains multiple video frame images of the face of the first object. The face in the video frame images can be in any pose and with any expression, as long as the face is clearly visible, so as to obtain reliable facial features.

[0090] In the above embodiments, the difference between the third set of expression coefficients corresponding to the third image and the first set of expression coefficients corresponding to the target image is minimized. This "difference" can be determined based on the sum of the differences between corresponding sets of expression coefficients or a weighted sum thereof. For example, assuming the third set of expression coefficients is {0.2, 0.3, 0.7, 0.1} and the first set of expression coefficients is {0.5, 0.4, 0.6, 0.1}, then the difference between the third set of expression coefficients and the first set of expression coefficients is {0.3, 0.1, 0.1, 0}, and the difference can be (0.3 + 0.1 + 0.1 + 0) = 0.5. Alternatively, the difference can be determined based on the weighted sum of the differences between the set of expression coefficients.

[0091] In some examples, the weight value corresponding to the corresponding expression coefficient can be determined based on the degree of influence of its corresponding expression base on the expression change. For example, the weight corresponding to the expression base of the mouth is greater than the weight corresponding to the expression base of the eyes.

[0092] In some embodiments, for each second image, adjusting the corresponding expression coefficient in the first set of expression coefficients based on the set of expression coefficient differences corresponding to the second image to obtain the plurality of second sets of expression coefficients includes: for each second image, adjusting the corresponding expression coefficient in the first set of expression coefficients based on the set of expression coefficient differences corresponding to the second image to obtain the plurality of fourth sets of expression coefficients; determining the difference between the corresponding expression coefficient in the third set of expression coefficients corresponding to the third image and the corresponding expression coefficient in the first set of expression coefficients corresponding to the target image; and obtaining the plurality of second sets of expression coefficients based on the plurality of fourth sets of expression coefficients and the expression coefficient differences calculated based on the third image and the target image.

[0093] Specifically, assuming the above difference is A, the expression coefficient corresponding to the target image is a (assuming only one table is included), and the expression coefficients corresponding to multiple image frames of the video data are a sequence: [b,c,d,e,f,g] (assuming each image frame also includes only the same expression), and assuming, for example, the difference between e and a is the smallest, i.e., A = ea. Then, the obtained multiple second set of expression coefficients are a sequence: [a+(be)+A,a+(ce)+A,a+(de)+A,a+(ee)+A,a+(fe)+A,a+(ge)+A].

[0094] Therefore, according to some embodiments, obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: obtaining a second facial key point image sequence based on the plurality of second sets of expression coefficients and the preset set of expression bases, wherein the images in the second facial key point image sequence correspond one-to-one with the plurality of second sets of expression coefficients.

[0095] In step 260, based on the second facial key point image and the target image, a first image corresponding to the target image after changing the facial expression is obtained.

[0096] According to some embodiments, obtaining a first image corresponding to the target image after changing the facial expression based on the second facial key point image and the target image includes: generating an expression image based on the second facial key point image, wherein the expression image is generated based on the connection of facial key points related to the expression in the second facial key point image; and obtaining a first image corresponding to the target image after changing the facial expression based on the expression image and the target image.

[0097] In some examples, the facial expression image may include the eyes, mouth, eyebrows, and lines connecting key points of the facial contour below the eyebrows or eyes. Typically, the nose shows little or no change in facial expressions, so it can be ignored in the facial expression image, and no lines are drawn connecting key points of the nose.

[0098] In the facial landmark detection example above, which includes pupil landmarks, the expression image can be further enhanced with pupil landmark information. This provides more accurate facial location information, aiding in subsequent, more refined analysis and processing.

[0099] In the case where a second facial key point image sequence has been generated as described above, according to some embodiments, obtaining a first image corresponding to the target image after changing the facial expression of the target image based on the second facial key point image and the target image includes: generating an expression image sequence based on the second facial key point image sequence; and obtaining multiple first images corresponding to the target image after changing the facial expression multiple times based on the expression image sequence and the target image, wherein the multiple images correspond one-to-one with the expression images in the expression image sequence.

[0100] In some examples, multiple first images can also be feature maps.

[0101] According to some embodiments, the corresponding first image after changing the facial expression of the target image is obtained by the following operations: extracting image features from the target image to obtain first image features; extracting image features from the expression image to obtain second image features; inputting the first image features and the second image features into a preset diffusion model to obtain third image features; and obtaining the corresponding first image after changing the facial expression of the target image based on the third image features.

[0102] In this disclosure, image feature extraction is the process of extracting useful information from an image. This information is typically represented in numerical, vector, or symbolic form and is not directly represented by the image itself. These features help computers "understand" the content of an image, thereby enabling image recognition and classification. Image features typically include geometric features, shape features, amplitude features, histogram features, and color features, etc.

[0103] In some examples, an image encoder can be used to extract image features from a target image to obtain initial image features. An image encoder is a component used to process visual information; it transforms image data into a format that can be further analyzed by a model. This typically involves feature extraction, i.e., extracting useful information from an image, such as color, texture, shape, and object location.

[0104] According to some embodiments, extracting image features from the target image to obtain first image features includes: inputting the target image into a variational autoencoder to obtain the first image features.

[0105] In some examples, image feature extraction can be performed using deep learning-based neural networks, such as convolutional neural networks (CNNs). CNNs automatically learn features through multiple layers of networks, eliminating the need for manually setting feature extraction rules. VGG and ResNet are two well-known CNN architectures that extract image features through deep network structures.

[0106] It is understood that image feature extraction can be achieved by any suitable method in the embodiments of this disclosure, and no limitation is made herein.

[0107] According to some embodiments, the method further includes: performing video synthesis based on the plurality of first images to obtain a video generated based on the target image.

[0108] According to some embodiments, obtaining a video generated based on the target image based on the plurality of first images includes: inputting the plurality of first images into a variational autodecoder to obtain a video generated based on the target image.

[0109] According to some embodiments, video synthesis based on the plurality of first images to obtain a video generated based on the target image includes: extracting image features from the target image to obtain first image features; extracting image features from the expression image sequence to obtain a second image feature sequence; inputting the first image features and the second image feature sequence into an image generation module to obtain a fourth image feature sequence, wherein the image features in the fourth image features are image features generated based on the target image that correspond to the corresponding image features in the second image feature sequence; inputting the fourth image feature sequence into a video synthesis module to obtain the plurality of first images, wherein the first image among the plurality of first images is a feature map, wherein the video synthesis model is used to achieve the smoothness of the video generated based on the plurality of first images; and obtaining a video generated based on the target image based on the plurality of first images.

[0110] In some examples, each image feature in the fourth image feature sequence is an image feature generated based on the target image that corresponds to a corresponding image feature in the second image feature sequence. Each of the plurality of first images is a feature map.

[0111] According to some embodiments, the corresponding second image features are obtained through a preset linear attention network.

[0112] Figure 3 A schematic diagram of a video generation model for generating video according to an embodiment of the present disclosure is shown. Figure 3 As shown, the target image is input to an image encoder to obtain first image features; the expression image sequence is input to a keypoint encoder to obtain a second image feature sequence; the first image features and the second image feature sequence are sequentially input to the image generation module and video synthesis module in the diffusion model to obtain multiple first images (i.e., feature maps). These multiple first images are then processed by an image decoder to obtain a video generated based on the target image. In some examples, the main body of the diffusion model can adopt the Unet framework, with the image generation module responsible for generating the target object and supplementing the background of a single image, and the video synthesis module responsible for the smooth and stable generation of the entire video segment.

[0113] According to embodiments of this disclosure, such as Figure 4As shown, a model training method 400 is also provided, including: acquiring a first target image including the face of a target object and a first label image (step 410); extracting facial key points based on the first target image to obtain a first facial key point image (step 420); obtaining a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases (step 430); adjusting the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the facial expression transformation of the first target image (step 440); obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases (step 450); obtaining a first image corresponding to the facial expression transformation of the target image through a video generation model based on the second facial key point image and the first target image (step 460); determining a first loss value based on the first image and the first label image through a preset loss function (step 470); and adjusting the parameter values ​​of the video generation model based on the first loss value (step 480).

[0114] According to some embodiments, the target image and the second facial key point image have the same dimensions.

[0115] In this embodiment, the target image and the second facial key point image have the same dimension, that is, the data between the images used to generate the final expression transformation image is aligned, which helps to improve the generation effect of the expression transformation image.

[0116] According to some embodiments, obtaining a first image corresponding to the transformation of facial expression on the target image by means of a video generation model based on the second facial key point image and the first target image includes: generating a first expression image based on the second facial key point image, wherein the first expression image is generated based on the connection of facial key points related to expression in the second facial key point image; and inputting the first expression image and the target image into the video generation model to obtain the first image corresponding to the transformation of facial expression on the target image.

[0117] It is understandable that when the second facial landmark image used to generate the expression image has the same dimension as the target image, the expression image generated based on the second facial landmark image will also have the same dimension as the target image. For example, both the target image and the second facial landmark image may be two-dimensional images, or both may be three-dimensional images. In this case, data alignment between the images used to finally generate the expression transformation image is more conducive to improving the model's learning and training performance.

[0118] Typically, the input target image is a two-dimensional image. Therefore, according to some embodiments, both the first target image and the second facial key point image are two-dimensional key point images, and each expression base in the preset set of expression bases and the first facial key point image are three-dimensional key point images.

[0119] Understandably, 3D face data contains one more dimension of depth information than 2D face data. Therefore, 3D face recognition has advantages over 2D face recognition in both recognition accuracy and liveness detection accuracy. In the above embodiment, the expression base formed by the 3D keypoint image can achieve more refined expression changes.

[0120] Therefore, further, in order to improve the generation effect of the expression transformation image, according to some embodiments, obtaining the second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: obtaining a third facial key point image after the expression transformation of the target object's face based on the second set of expression coefficients and the preset set of expression bases, wherein the third facial key point image includes the coordinate information of three-dimensional key points; and obtaining the second facial key point image based on the coordinate information of the three-dimensional key points and a preset first transformation matrix, wherein the second facial key point image includes the coordinate information of two-dimensional key points.

[0121] According to some embodiments, obtaining a second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: obtaining a third facial key point image after performing expression transformation on the target object's face based on the second set of expression coefficients and the preset set of expression bases, wherein the third facial key point image includes coordinate information of three-dimensional key points; obtaining the second facial key point image based on the coordinate information of the three-dimensional key points and a preset second transformation matrix, wherein the preset second transformation matrix is ​​used to perform operations on the target object's face in three-dimensional space, wherein the operations include at least one of the following: rotation and translation.

[0122] According to some embodiments, the second facial key point image includes the coordinate information of the pupil key points. By controlling the pupil position, refined facial expression changes can be achieved.

[0123] According to some embodiments, the coordinate information of the pupil key point is determined based on the coordinate information of the orbital region key point in the second facial key point image.

[0124] According to some embodiments, the video generation model includes a first image encoder, a second image encoder, a diffusion model, and an image decoder. Inputting the first facial expression image and the first target image into the video generation model to obtain a second image corresponding to the target image after facial expression transformation includes: inputting the first target image into the first image encoder to obtain first image features; inputting the first facial expression image into the second image encoder to obtain second image features; inputting the first image features and the second image features into the diffusion model to obtain third image features; and based on the third image features, obtaining a corresponding first image corresponding to the target image after facial expression transformation.

[0125] In some examples, image feature extraction can be performed using deep learning-based neural networks, where at least one of the first and second image encoders can be a deep learning-based neural network, such as a convolutional neural network (CNN). CNNs automatically learn features through multi-layered networks, eliminating the need for manually defined feature extraction rules. VGG and ResNet are two well-known CNN architectures that extract image features through deep network structures.

[0126] It is understood that image feature extraction can be achieved by any suitable method in the embodiments of this disclosure, and no limitation is made herein.

[0127] According to some embodiments, the first image encoder may include a variational autoencoder to obtain first image features. According to some embodiments, a video generated based on a target image can be obtained based on multiple first images. For example, multiple first images are input into a variational autoencoder to obtain a video generated based on the images.

[0128] Therefore, according to some embodiments, adjusting the parameter values ​​of the video generation model based on the first loss value includes: adjusting the parameter values ​​of the second image encoder and the diffusion model based on the first loss value.

[0129] According to some embodiments, the diffusion model includes an image generation module and a video synthesis module, wherein inputting the first image feature and the second image feature into the diffusion model to obtain a third image feature includes: inputting the first image feature and the second image feature into the image generation module to obtain a fourth image feature, wherein the fourth image feature is an image feature generated based on the target image that corresponds to the second image feature; and inputting the fourth image feature into the video synthesis module to obtain the third image feature, wherein the video synthesis model is used to achieve the smoothness of the video when generating a video based on multiple second image features.

[0130] According to some embodiments, adjusting the parameter values ​​of the video generation model based on the first loss value includes: adjusting the parameter values ​​of the second image encoder and the image generation module based on the first loss value.

[0131] According to some embodiments, the model training method of this disclosure further includes: acquiring a plurality of second expression images, a second target image including the face of a target object, and a plurality of second label images corresponding one-to-one with the plurality of second expression images, wherein each of the plurality of second expression images is generated based on the connection of facial key points related to expression in the corresponding facial key point image; inputting the second target image into the first image encoder to obtain a fifth image feature; inputting the plurality of second expression images into the second image encoder to obtain a plurality of sixth image features; inputting the fifth image features and the plurality of sixth image features into the image generation module to obtain a plurality of seventh image features, wherein the plurality of seventh image features correspond one-to-one with the plurality of sixth image features; inputting the plurality of seventh image features into the video synthesis module to obtain a plurality of eighth image features; inputting the plurality of eighth image features into the image decoder to obtain a plurality of second images; determining a second loss value based on the plurality of second images and the plurality of second label images using a preset loss function; and adjusting the parameter values ​​of the video synthesis module based on the second loss value.

[0132] Specifically, in some examples, the training of the video generation model can be divided into two stages. That is, in the first stage, the second image encoder and the image generation module are trained based on a single frame of facial expression images; in the second stage, the video synthesis module is trained based on multiple frames of facial expression images.

[0133] According to some embodiments, the second image encoder includes a linear attention network.

[0134] According to some embodiments, the preset loss function loss is determined based on the following formula:

[0135]

[0136] in, C represents the first tag image or the corresponding second tag image among the plurality of second tag images, D represents the mouth mask image, and a and b are preset hyperparameters.

[0137] In some examples, when When C represents the corresponding second label image among the plurality of second label images, the loss value corresponding to each image among the plurality of second label images / second images can be calculated using the above formula. By adding the loss values ​​corresponding to each of the plurality of second label images / second images, a second loss value is obtained, and the parameter values ​​of the video synthesis module are adjusted based on the second loss value.

[0138] In this embodiment, to improve the clarity and stability of tooth generation, mouth masking loss and overall image loss are used to simultaneously supervise model training. Furthermore, in some examples, by setting appropriate values ​​for 'a' and 'b', the weights of the mouth masking loss can be increased to further improve the clarity and stability of tooth generation.

[0139] In this disclosure, the model trained by the model training method described in any of the above embodiments can be used to implement the data processing method described in any of the embodiments of this disclosure.

[0140] Here, the embodiments for implementing the model training method and the embodiments for implementing the data processing method are described, and the corresponding operations are similar, so they will not be described again here.

[0141] According to embodiments of this disclosure, such as Figure 5 As shown, a data processing apparatus 500 for virtual avatars is also provided, comprising: a first acquisition unit 510 configured to acquire a target image including the face of a target object; a first key point extraction unit 520 configured to extract facial key points based on the target image to obtain a first facial key point image; a first calculation unit 530 configured to obtain a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases; a second calculation unit 540 configured to adjust the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the target image after changing the facial expression; a second acquisition unit 550 configured to obtain a second facial key point image based on the second set of expression coefficients and the preset set of expression bases; and a first generation unit 560 configured to obtain a first image corresponding to the target image after changing the facial expression based on the second facial key point image and the target image.

[0142] Here, the operation of each of the above-mentioned units 510 to 560 of the data processing device 500 is similar to the operation of steps 210 to 260 described above, and will not be repeated here.

[0143] According to embodiments of this disclosure, such as Figure 6As shown, a model training device 600 is also provided, including: a third acquisition unit 610 configured to acquire a first target image including the face of a target object and a first label image; a second key point extraction unit 620 configured to extract facial key points based on the first target image to obtain a first facial key point image; a third calculation unit 630 configured to obtain a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases; and a fourth calculation unit 640 configured to adjust the corresponding expression coefficients in the first set of expression coefficients to achieve the desired effect. The system includes: a second set of expression coefficients corresponding to the transformation of facial expressions on the first target image; a fourth acquisition unit 650 configured to obtain a second facial key point image based on the second set of expression coefficients and a preset set of expression bases; a second generation unit 660 configured to obtain a first image corresponding to the transformation of facial expressions on the target image based on the second facial key point image and the first target image through a video generation model; a fifth calculation unit 670 configured to determine a first loss value based on the first image and the first label image through a preset loss function; and an adjustment unit 680 configured to adjust the parameter values ​​of the video generation model based on the first loss value.

[0144] Here, the operation of each of the above units 610 to 680 of the model training device 600 is similar to the operation of steps 410 to 480 described above, and will not be repeated here.

[0145] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0146] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0147] refer to Figure 7 The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0148] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0149] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0150] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as method 200 or 400. For example, in some embodiments, method 200 or 400 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of method 200 or 400 described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute method 200 or 400 by any other suitable means (e.g., by means of firmware).

[0151] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0152] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0153] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0155] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0156] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0157] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0158] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A data processing method for virtual avatars, comprising: Obtain the target image, including the face of the target object; Facial key points are extracted based on the target image to obtain a first facial key point image; Based on the first facial key point image and a preset set of expression bases, a first set of expression coefficients corresponding to the first facial key point image is obtained, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases. The corresponding expression coefficients in the first set of expression coefficients are adjusted to obtain the second set of expression coefficients corresponding to the facial expression transformation of the target image, including: For each of the multiple second images, the difference between the corresponding expression coefficient in the third set of expression coefficients corresponding to the second image and the corresponding expression coefficient in the third set of expression coefficients corresponding to the third image is determined to obtain a set of expression coefficient differences, wherein the difference between the third set of expression coefficients corresponding to the third image and the first set of expression coefficients corresponding to the target image is the smallest, wherein the multiple second images all include the face of the first object; For each second image, the corresponding expression coefficient in the first set of expression coefficients is adjusted based on the difference of the set of expression coefficients corresponding to the second image to obtain multiple sets of expression coefficients, wherein each set of expression coefficients corresponds to a transformed facial expression. Based on the plurality of second-set expression coefficients and the preset set of expression bases, a second facial key point image sequence is obtained, wherein the images in the second facial key point image sequence correspond one-to-one with the plurality of second-set expression coefficients; and Based on the second facial key point image sequence and the target image, a first image corresponding to the target image after changing the facial expression is obtained.

2. The method as described in claim 1, wherein, The target image and the second facial key point image have the same dimensions.

3. The method of claim 2, wherein, Both the target image and the second facial key point image are two-dimensional key point images, and each expression base in the preset set of expression bases and the first facial key point image are three-dimensional key point images. The process of obtaining the second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: Based on the second set of expression coefficients and the preset set of expression bases, a third facial key point image is obtained after performing expression transformation on the target object's face, wherein the third facial key point image includes the coordinate information of three-dimensional key points; and Based on the coordinate information of the three-dimensional key points and the preset first transformation matrix, the second facial key point image is obtained, wherein the second facial key point image includes the coordinate information of two-dimensional key points.

4. The method according to any one of claims 1-3, wherein, Based on the second set of expression coefficients and the preset set of expression bases, the second facial key point image is obtained as follows: Based on the second set of expression coefficients and the preset set of expression bases, a third facial key point image is obtained after the expression transformation of the target object's face, wherein the third facial key point image includes the coordinate information of three-dimensional key points. Based on the coordinate information of the three-dimensional key points and the preset second transformation matrix, the second facial key point image is obtained. The preset second transformation matrix is ​​used to perform operations on the face of the target object in three-dimensional space, wherein the operations include at least one of the following: rotation and translation.

5. The method of claim 1, wherein, The second facial key point image includes the coordinate information of the pupil key points.

6. The method of claim 5, wherein, The coordinate information of the pupil key point is determined based on the coordinate information of the orbital region key point in the second facial key point image.

7. The method of claim 1, wherein, The plurality of second images and the corresponding third set of expression coefficients for each of the plurality of second images are obtained through the following operations: Acquire video data including the face of the first object; The video data is divided into frames to obtain the plurality of second images including the face of the first object; Facial key points are extracted from each of the plurality of second images to obtain a fourth facial key point image corresponding to each of the plurality of second images; Based on the fourth facial key point image and the preset set of expression bases, the third set of expression coefficients corresponding to each of the plurality of second images are obtained.

8. The method of claim 1, wherein, Based on the second facial key point image and the target image, obtaining the first image corresponding to the target image after changing the facial expression includes: An expression image sequence is generated based on the second facial key point image sequence, wherein the expression images in the expression image sequence are generated based on the connection of facial key points related to expressions in the corresponding second facial key point images; and Based on the expression image sequence and the target image, multiple first images are obtained corresponding to the multiple facial expressions of the target image after changing the facial expressions multiple times, wherein the multiple first images correspond one-to-one with the expression images in the expression image sequence.

9. The method of claim 8, wherein, The following operations are used to obtain the corresponding first image after transforming the facial expression of the target image: Image features are extracted from the target image to obtain the first image features; Image features are extracted from the facial expression image to obtain second image features; The first image feature and the second image feature are input into a preset diffusion model to obtain the third image feature; as well as Based on the third image features, a corresponding first image is obtained after changing the facial expression of the target image.

10. The method of claim 8, further comprising: Video synthesis is performed based on the plurality of first images to obtain a video generated based on the target image.

11. The method of claim 10, wherein, Video synthesis based on the plurality of first images to obtain a video generated based on the target image includes: Image features are extracted from the target image to obtain the first image features; Image feature extraction is performed on the expression image sequence to obtain a second image feature sequence; The first image feature and the second image feature sequence are input into the image generation module to obtain a fourth image feature sequence, wherein the image features in the fourth image feature are image features generated based on the target image that correspond to the corresponding image features in the second image feature sequence; The fourth image feature sequence is input into the video synthesis module to obtain the plurality of first images, wherein the first image among the plurality of first images is a feature map, and the video synthesis model is used to achieve smoothness in the video generated based on the plurality of first images; and A video generated based on the target image is obtained based on the plurality of first images.

12. The method of claim 11, wherein, Extracting image features from the target image to obtain first image features includes: inputting the target image into a variational autoencoder to obtain the first image features; and Obtaining a video generated based on the target image from the plurality of first images includes: inputting the plurality of first images into a variational autodecoder to obtain a video generated based on the target image.

13. The method as described in claim 9 or 11, wherein the corresponding second image features are obtained through a preset linear attention network.

14. A model training method, comprising: Obtain a first target image and a first label image, including the face of the target object; Facial key points are extracted based on the first target image to obtain a first facial key point image; Based on the first facial key point image and a preset set of expression bases, a first set of expression coefficients corresponding to the first facial key point image is obtained, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases. The corresponding expression coefficients in the first set of expression coefficients are adjusted to obtain the second set of expression coefficients corresponding to the facial expression transformation of the first target image; Based on the second set of expression coefficients and the preset set of expression bases, a second facial key point image is obtained; Based on the second facial key point image and the first target image, a first image corresponding to the target image after changing the facial expression is obtained through a video generation model; Based on the first image and the first label image, a first loss value is determined by a preset loss function; as well as The parameter values ​​of the video generation model are adjusted based on the first loss value. The trained video generation model is used in the method of any one of claims 1-13.

15. The method of claim 14, wherein, The target image and the second facial key point image have the same dimensions.

16. The method of claim 14, wherein, Both the first target image and the second facial key point image are two-dimensional key point images, and each expression base in the preset set of expression bases and the first facial key point image are three-dimensional key point images. The process of obtaining the second facial key point image based on the second set of expression coefficients and the preset set of expression bases includes: Based on the second set of expression coefficients and the preset set of expression bases, a third facial key point image is obtained after performing expression transformation on the target object's face, wherein the third facial key point image includes the coordinate information of three-dimensional key points; and Based on the coordinate information of the three-dimensional key points and the preset first transformation matrix, the second facial key point image is obtained, wherein the second facial key point image includes the coordinate information of two-dimensional key points.

17. The method according to any one of claims 14-16, wherein, Based on the second set of expression coefficients and the preset set of expression bases, the second facial key point image is obtained as follows: Based on the second set of expression coefficients and the preset set of expression bases, a third facial key point image is obtained after the expression transformation of the target object's face, wherein the third facial key point image includes the coordinate information of three-dimensional key points. Based on the coordinate information of the three-dimensional key points and the preset second transformation matrix, the second facial key point image is obtained. The preset second transformation matrix is ​​used to perform operations on the face of the target object in three-dimensional space, wherein the operations include at least one of the following: rotation and translation.

18. The method of claim 14, wherein, The second facial key point image includes the coordinate information of the pupil key points.

19. The method of claim 18, wherein, The coordinate information of the pupil key point is determined based on the coordinate information of the orbital region key point in the second facial key point image.

20. The method of claim 14, wherein, Based on the second facial key point image and the first target image, the first image corresponding to the target image after facial expression transformation is obtained through a video generation model includes: A first expression image is generated based on the second facial key point image, wherein the first expression image is generated based on the lines connecting facial key points related to the expression in the second facial key point image; and The first facial expression image and the target image are input into a video generation model to obtain a first image corresponding to the target image after the facial expression is changed.

21. The method of claim 20, wherein, The video generation model includes a first image encoder, a second image encoder, a diffusion model, and an image decoder, wherein inputting the first expression image and the first target image into the video generation model to obtain the second image corresponding to the target image after transforming the facial expression includes: The first target image is input into the first image encoder to obtain the first image features; The first facial expression image is input into the second image encoder to obtain the second image features; The first image features and the second image features are input into the diffusion model to obtain a third image feature; and Based on the third image features, a corresponding first image is obtained after changing the facial expression of the target image.

22. The method of claim 21, wherein, Adjusting the parameter values ​​of the video generation model based on the first loss value includes: adjusting the parameter values ​​of the second image encoder and the diffusion model based on the first loss value.

23. The method of claim 21 or 22, wherein, The diffusion model includes an image generation module and a video synthesis module, wherein inputting the first image features and the second image features into the diffusion model to obtain the third image features includes: The first image feature and the second image feature are input into the image generation module to obtain a fourth image feature, wherein the fourth image feature is an image feature generated based on the target image that corresponds to the second image feature; and The fourth image feature is input into the video synthesis module to obtain the third image feature, wherein the video synthesis module is used to achieve the smoothness of the video when generating a video based on multiple second image features.

24. The method of claim 23, wherein, Adjusting the parameter values ​​of the video generation model based on the first loss value includes: adjusting the parameter values ​​of the second image encoder and the image generation module based on the first loss value.

25. The method of claim 23, further comprising: A plurality of second expression images are acquired, including a second target image of the target object's face, and a plurality of second label images corresponding one-to-one with the plurality of second expression images, wherein each of the plurality of second expression images is generated based on the connection of facial key points related to the expression in the corresponding facial key point image; The second target image is input into the first image encoder to obtain the fifth image feature; The plurality of second facial expression images are input into the second image encoder to obtain a plurality of sixth image features; The fifth image feature and the plurality of sixth image features are input into the image generation module to obtain a plurality of seventh image features, wherein the plurality of seventh image features correspond one-to-one with the plurality of sixth image features; and The plurality of seventh image features are input into the video synthesis module to obtain a plurality of eighth image features; The plurality of eighth image features are input into the image decoder to obtain a plurality of second images; Based on the plurality of second images and the plurality of second label images, a second loss value is determined using a preset loss function; and The parameter values ​​of the video synthesis module are adjusted based on the second loss value.

26. The method according to any one of claims 21-22 and 25, wherein, The second image encoder includes a linear attention network.

27. The method of claim 14 or 25, wherein, The preset loss function Determined based on the following formula: in, This refers to the corresponding second label image among the first label image or the plurality of second label images. This refers to the first image or the corresponding second image among the plurality of second images. This represents a masked image of the mouth. and All of these are preset hyperparameters.

28. A data processing apparatus for virtual avatars, comprising: The first acquisition unit is configured to acquire a target image including the face of the target object; The first key point extraction unit is configured to extract facial key points based on the target image to obtain a first facial key point image; The first calculation unit is configured to obtain a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases. The second calculation unit is configured to adjust the corresponding expression coefficients in the first set of expression coefficients to obtain a second set of expression coefficients corresponding to the facial expression transformation of the target image, including: For each of the multiple second images, the difference between the corresponding expression coefficient in the third set of expression coefficients corresponding to the second image and the corresponding expression coefficient in the third set of expression coefficients corresponding to the third image is determined to obtain a set of expression coefficient differences, wherein the difference between the third set of expression coefficients corresponding to the third image and the first set of expression coefficients corresponding to the target image is the smallest, wherein the multiple second images all include the face of the first object; For each second image, the corresponding expression coefficient in the first set of expression coefficients is adjusted based on the difference of the set of expression coefficients corresponding to the second image to obtain multiple sets of expression coefficients, wherein each set of expression coefficients corresponds to a transformed facial expression. The second acquisition unit is configured to obtain a second facial key point image sequence based on the plurality of second sets of expression coefficients and the preset set of expression bases, wherein the images in the second facial key point image sequence correspond one-to-one with the plurality of second sets of expression coefficients; and The first generation unit is configured to obtain a first image corresponding to the target image after changing the facial expression of the target image, based on the second facial key point image sequence and the target image.

29. A model training device, comprising: The third acquisition unit is configured to acquire a first target image and a first label image, including the face of the target object; The second key point extraction unit is configured to extract facial key points based on the first target image to obtain a first facial key point image. The third calculation unit is configured to obtain a first set of expression coefficients corresponding to the first facial key point image based on the first facial key point image and a preset set of expression bases, wherein the first set of expression coefficients corresponds one-to-one with the set of expression bases. The fourth calculation unit is configured to adjust the corresponding expression coefficients in the first set of expression coefficients to obtain the second set of expression coefficients corresponding to the facial expression transformation of the first target image. The fourth acquisition unit is configured to obtain a second facial key point image based on the second set of expression coefficients and the preset set of expression bases; The second generation unit is configured to obtain a first image corresponding to the target image after changing the facial expression based on the second facial key point image and the first target image through a video generation model; The fifth calculation unit is configured to determine a first loss value based on the first image and the first label image using a preset loss function; as well as The adjustment unit is configured to adjust the parameter values ​​of the video generation model based on the first loss value. The trained video generation model is used in the method of any one of claims 1-13.

30. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-27.

31. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-27.

32. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-27.

Citation Information

Patent Citations

  • Expression generation method and device, equipment, medium and computer program product

    CN113870401A

  • Three-dimensional face reconstruction method, electronic equipment and computer readable storage medium

    CN114067059A