Cartoon digital human real-time expression driving method, electronic equipment and storage medium

By building a cartoon character expression-driven network model based on a fully connected neural network and convolutional structure, the real-time and accuracy problems of cartoon digital human expression-driven on mobile terminal devices are solved, and efficient cartoon digital human expression-driven is realized, improving the social experience.

CN120472510AActive Publication Date: 2025-08-12HONOR DEVICE CO LTD

Patent Information

Application Number
CN202411255295.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-08-12
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

The existing cartoon digital human expression driving method cannot meet the real-time requirements of mobile terminal devices, and the accuracy is low, making it difficult to achieve ideal driving effects.

Method used

A cartoon character expression-driven network model based on a fully connected neural network and convolutional structure is constructed. FaceMesh network is used to reconstruct face 3D, and the cartoon character expression basic coefficient and rotation feature representation are used to render high-precision expressions, and real-time expression-driven is used to combine the BlendShape parameter model.

Benefits of technology

It realizes the high real-time and high-precision expression driver of cartoon digital people, reduces the performance requirements for mobile terminal device processors, and provides a more novel and interesting social experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472510A_ABST
    Figure CN120472510A_ABST
Patent Text Reader

Abstract

The invention provides a cartoon digital human real-time expression driving method, electronic equipment and a storage medium, and relates to the technical field of terminals. The method comprises the following steps that: terminal equipment firstly obtains target images containing face information of a target user according to a preset frequency through a camera, and then inputs each target image into a FaceMesh network frame by frame to carry out 3D reconstruction processing on the face image of the target user so as to obtain face 3D key point information of the target user. Inputting the face 3D key point information of the target user into the cartoon role expression driving network model, and predicting to obtain a cartoon role expression base coefficient and rotation feature representation; the cartoon role expression base coefficient and the rotation feature representation are utilized, a preset BlendShape parameter model is combined, rendering is carried out according to related knowledge of computer graphics, a cartoon role expression animation image corresponding to the target image is obtained, and the cartoon role expression animation image is continuously output. And the cartoon digital human real-time facial expression driving animation corresponding to the target user is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a method for driving real-time expression of a cartoon digital human, an electronic device, and a storage medium. Background Art

[0002] With the continuous development of technology, mobile terminal devices such as mobile phones and tablet computers are increasingly used in people's lives and work, bringing great convenience to people. For example, people can use these terminal devices to take photos, record videos, and communicate socially.

[0003] Currently, when using mobile devices for social communication, users often customize cartoon digital human avatars to enhance social interaction. However, existing methods for driving cartoon digital human expressions often fail to meet the real-time requirements of mobile devices and have low accuracy, making it difficult to achieve ideal driving effects. Summary of the Invention

[0004] In order to solve the above problems, the present application provides a real-time expression driving method for cartoon digital humans, an electronic device and a storage medium. The purpose is to achieve high real-time and high-precision expression driving of cartoon digital humans when users use mobile terminal devices for social communication, thereby providing users with a more novel and interesting social experience.

[0005] In a first aspect, the present application provides a method for driving real-time facial expressions of a cartoon digital human. The method comprises: upon receiving a cartoon digital human generation instruction triggered by a user on a terminal device's shooting interface, the terminal device first responds to the instruction by capturing a target image containing a target user's facial information via a camera at a preset frequency (e.g., 30 fps). Each target image is then input into a FaceMesh network frame by frame for 3D reconstruction of the target user's facial image, and 468 3D facial landmarks are estimated in real time as 3D facial key point information of the target user's face. Subsequently, the 3D facial key point information of the target user is input into a cartoon character expression driving network model to predict the cartoon character's expression base coefficients and rotation feature representation. The cartoon character expression base coefficients and rotation feature representation can then be used in combination with a preset BlendShape parameter model (for a set of 52 predefined facial shapes) to render according to relevant computer graphics knowledge to obtain a cartoon character expression animation image corresponding to the target image. Finally, all the obtained cartoon character expression animation images are concatenated in chronological order, and all the concatenated cartoon character expression animation images are continuously output to form a real-time facial expression driven animation of the cartoon digital human corresponding to the target user.

[0006] It can be seen that in the above-mentioned cartoon digital human real-time expression driving method, this embodiment pre-constructs a cartoon character expression driving network model that is friendly to mobile terminal devices based on a fully connected neural network and convolutional structure. This not only improves the driving accuracy and speed, but also facilitates the quantitative deployment of various deployment tools on mobile terminal devices, thereby reducing the performance requirements for the mobile terminal device processor without reducing the network computing performance. This allows users to achieve high real-time and high-precision expression driving of cartoon digital humans when using mobile terminal devices such as mobile phones for social communication, thereby providing users with a more novel and interesting social experience.

[0007] In one possible implementation, a target image is input into a FaceMesh network for 3D reconstruction of the target user's facial image to obtain 3D facial key point information of the target user, including: cropping the target user's facial region from the target image to obtain a facial region image of the target user; performing rotation and resizing alignment processing on the facial region image of the target user to obtain an aligned facial region image of the target user; and calculating a rotation inverse transformation matrix representation corresponding to the facial region image of the target user; inputting the aligned facial region image of the target user into a FaceMesh network for 3D reconstruction of the target user's facial image to obtain 468 facial key point information of the target user; and using the rotation inverse transformation matrix representation corresponding to the facial region image of the target user and the 468 facial key point information of the target user to calculate the 3D facial key point information of the target user. In this way, the geometric shape of the 3D facial surface can be inferred using machine learning through input from only a single camera, without the need for a dedicated depth sensor, thereby providing real-time performance and meeting the high real-time requirements of mobile terminal devices.

[0008] In one possible implementation, a cartoon character expression driving network model includes a key information encoder, a cartoon character expression base coefficient prediction module, and a rotation representation prediction module. The target user's facial 3D key point information is input into a pre-built cartoon character expression driving network model to predict the cartoon character expression base coefficients and rotation feature representations. This includes: inputting the target user's facial 3D key point information into the pre-built cartoon character expression driving network model's key information encoder for feature extraction to obtain a latent space feature representation; inputting the latent space feature representation into the cartoon character expression base coefficient prediction module of the cartoon character expression driving network model for mean dimensionality reduction to obtain the cartoon character expression base coefficients; and inputting the latent space feature representation into the rotation representation prediction module of the cartoon character expression driving network model for resizing and convolution to obtain the rotation feature representation. This improves driving accuracy and speed, facilitates the quantitative deployment of various deployment tools on mobile terminal devices, reduces the performance requirements of the mobile terminal device processor, and does not reduce network computing performance.

[0009] In one possible implementation, the key information encoder is an encoder built based on a fully connected network structure; the cartoon character expression base coefficient prediction module is a decoder built based on a fully connected network structure; and the rotation representation prediction module is a decoder built based on a convolutional structure.

[0010] In one possible implementation, a cartoon character expression base coefficient and a rotation feature representation are used in combination with a preset blend deformation BlendShape parameter model to render a cartoon character expression animation image corresponding to a target image, including: using the cartoon character expression base coefficient to perform weighted summation on the preset blend deformation BlendShape parameter model to predict a cartoon character expression model corresponding to the target image; using the rotation feature representation, performing a rotation transformation on the predicted cartoon character expression model corresponding to the target image to obtain a cartoon character expression model mapped to a 2D plane, and rendering the model to obtain a cartoon character expression animation; according to the target user's facial area image, the cartoon character expression animation is aligned with the target user's face to obtain a cartoon character expression animation image corresponding to the target image, so as to improve the driving effect of the cartoon character expression animation image.

[0011] In one possible implementation, continuously outputting animated cartoon character expressions to form a real-time facial expression-driven animation of a cartoon digital human corresponding to a target user involves sequentially connecting the animated cartoon character expressions and continuously outputting all connected animated cartoon character expressions to form a real-time facial expression-driven animation of a cartoon digital human corresponding to the target user. This can enhance the target user's social experience.

[0012] In one possible implementation, the objective loss function includes a first objective loss function, a second objective loss function, and a third objective loss function. The cartoon character expression-driven network model is constructed by obtaining a sample image containing sample user facial information; using preset coefficient priors and rotation priors as constraints, the sample image is used to construct input and output data pairs consisting of the sample user's facial 3D key point information, the sample cartoon character expression base coefficients, and the sample rotation feature representation, as a training image dataset; and the initial cartoon character expression-driven network model is trained using the training image dataset, the first objective loss function, the second objective loss function, and the third objective loss function to generate a cartoon character expression-driven network model. In this way, based on an incompletely supervised training strategy, a massive training dataset is generated using a small amount of data, effectively improving the model's training accuracy.

[0013] In one possible implementation, the first objective loss function is used to constrain the cartoon character expression base coefficients output by the cartoon character expression driven network model so that they are as close as possible to the sample cartoon character expression base coefficients in the training image dataset; the second objective loss function is used to constrain the rotation feature representation output by the cartoon character expression driven network model so that it is as close as possible to the sample rotation feature representation in the training image dataset; the third objective loss function is used to simultaneously constrain the cartoon character expression base coefficients and rotation feature representations output by the cartoon character expression driven network model so that the two are as close as possible to the sample cartoon character expression base coefficients and sample rotation feature representations in the training image dataset, respectively.

[0014] In a possible implementation, the rotation feature is represented as a 6D rotation transformation representation.

[0015] In a second aspect, the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor is used to call and execute the computer program to implement the real-time expression driving method of a cartoon digital human as described in any one of the first aspects above.

[0016] In a third aspect, the present application provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor of an electronic device, it is used to implement the real-time expression driving method of a cartoon digital human as described in any one of the first aspects above.

[0017] In a fourth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the cartoon digital human real-time expression driving method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is one of the scenario diagrams provided in the embodiment of the present application;

[0019] Figure 2 A schematic diagram of a mobile terminal device provided in an embodiment of the present application;

[0020] Figure 3 A software structure diagram of a mobile terminal device provided in an embodiment of the present application;

[0021] Figure 4 A flowchart of the real-time expression driving method for a cartoon digital human provided in an embodiment of the present application;

[0022] Figure 5 A schematic diagram of the process of determining 3D key point information of a target user's face using the FaceMesh network provided in an embodiment of the present application;

[0023] Figure 6A schematic diagram of the process of extracting latent space feature representation using a key information encoder provided in an embodiment of the present application;

[0024] Figure 7 The mixer provided in the embodiment of this application m,n (.) Network structure diagram;

[0025] Figure 8 A schematic diagram of the network structure of a cartoon character expression base coefficient prediction module provided in an embodiment of the present application;

[0026] Figure 9 A schematic diagram of the network structure of the rotation representation prediction module provided in an embodiment of the present application;

[0027] Figure 10 This is a schematic diagram of the overall implementation process of real-time expression driving of a cartoon digital human provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of this application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise.

[0029] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0030] The "multiple" involved in the embodiments of the present application means greater than or equal to two. It should be noted that in the description of the embodiments of the present application, the words "first" and "second" are only used for the purpose of distinguishing the description and cannot be understood as indicating or implying relative importance or order.

[0031] In order to enable people skilled in the art to more clearly understand the solution of the present application, the application scenario of the technical solution of the present application will be first explained.

[0032] See also Figure 1 , which shows a scenario schematic diagram provided by an embodiment of the present application.

[0033] In this example scenario, when a user uses their mobile phone for social communication, such as video chatting with others through social media apps installed on their phone, or recording or live streaming videos on short video apps, they typically capture their own facial image to create a cartoon digital human avatar to enhance the social interaction. Furthermore, the user can select or design different avatar elements based on their personal preferences, such as face shape, eyes, hairstyle, and clothing, and then combine these elements to create a unique cartoon digital human avatar for social communication.

[0034] Traditional methods for driving expressions in cartoon digital humans typically begin by acquiring multi-view RGBD images using an image acquisition device. RGB represents the color modes of the three primary colors (red, green, and blue), which can be used to create different color images. D represents a depth map, which provides distance information. Based on the multi-view RGBD images, a facial point cloud is generated, which is then further transformed into a three-dimensional (3D) facial mesh, achieving a two-dimensional (2D) to 3D conversion and completing the expression driving of the cartoon digital human. However, this method is relatively cumbersome in terms of image acquisition and places high demands on the image acquisition device, requiring it to be capable of capturing depth.

[0035] To address the aforementioned issues with traditional driving methods, currently, a method based on an artificial intelligence (AI) neural network model is commonly used to drive cartoon digital human expressions. However, this method still has high production costs and technical barriers, and is also limited by the current IQ level and computing resources of AI, making it difficult to achieve ideal prediction accuracy.

[0036] Moreover, with the increasing use of mobile terminal devices in people's lives and work, such as people can use mobile phones to live broadcast videos or video chat at any time, the implementation of cartoon digital human expression drive on mobile terminal devices also faces the challenge of real-time requirements. This is because the two existing methods of driving cartoon digital human expressions mentioned above cannot meet the real-time requirements of mobile terminal devices. Whether from the perspective of device computing power, performance power consumption, or simultaneous multi-tasking of the device, it will bring a certain burden to the data calculation and processing of the mobile terminal device, thereby affecting the implementation of real-time expression drive of cartoon digital humans; moreover, even if the expression drive technology relies on cloud computing power, the upload and download of data will also increase the input and output delay to a certain extent, which will also affect the real-time performance of the cartoon digital human expression drive.

[0037] Therefore, how to achieve near real-time facial video frame input and cartoon digital human expression-driven video output to achieve a more realistic driving effect is a technical problem that needs to be solved urgently.

[0038] To overcome the above technical problems, this application provides a method for driving cartoon digital human expressions in real time. A mobile-friendly cartoon character expression driving network model is pre-built based on a fully connected neural network and convolutional architecture. This not only improves driving accuracy and speed but also facilitates the quantitative deployment of various deployment tools on mobile devices, thereby reducing the performance requirements for mobile device processors without compromising network computing performance. This allows users to achieve high-precision, real-time expression driving of cartoon digital humans when using mobile devices such as mobile phones for social communication, thereby providing users with a more novel and engaging social experience.

[0039] It should be noted that the real-time expression driving method for cartoon digital humans provided in the embodiments of the present application can be applied to mobile terminal devices (hereinafter referred to as terminal devices or electronic devices) such as mobile phones, tablet computers, personal digital assistants (PDAs), desktop computers, laptop computers, notebook computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, and wearable devices.

[0040] In order to enable people skilled in the art to more clearly understand the cartoon digital human real-time expression driving method provided by this application, the hardware architecture and software system architecture of the electronic device that implements the cartoon digital human real-time expression driving method are first introduced in detail below.

[0041] See also Figure 2 , which shows a schematic diagram of a mobile terminal device provided in an embodiment of the present application.

[0042] like Figure 2 As shown, the mobile terminal device 200 may include a processor 210, a mobile communication module 220, a wireless communication module 230, a sensor module 240, a display screen 250, an internal memory 260, a camera 270, an audio module 280, a speaker 280A, a receiver 280B, a microphone 280C, an earphone interface 280D, an antenna group 1 and an antenna group 2.

[0043] It is understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the mobile terminal device 200. In other embodiments of the present application, the mobile terminal device 200 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0044] The processor 210 may include one or more processing units, for example, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. The controller may generate an operation control signal based on an instruction opcode and a timing signal to control instruction fetching and execution. For example, the camera 270 pre-installed on the mobile terminal device 200 may be used to capture image frames containing a human face at a fixed frame rate. The data of each image frame may then be processed to obtain character expression animation frames. By connecting and outputting all the obtained character expression animation frames, real-time facial expression driving of a cartoon digital human can be achieved.

[0045] Processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 210 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 210. If processor 210 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 210 latency, and thus improves system efficiency.

[0046] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0047] The sensor module 240 can be used to obtain data signals related to various aspects of the mobile terminal device 200 as a basis for implementing corresponding functions. In some embodiments, the sensor module 240 may include but is not limited to an image sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a temperature sensor, a pressure sensor, and the like.

[0048] The display screen 250 is used to display images, videos, and the like, such as a user's customized cartoon digital avatar corresponding to a selfie taken by the mobile terminal device 200. The display screen 250 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, the electronic device 200 may include one or P display screens 250, where P is a positive integer greater than one.

[0049] The internal memory 260 can be used to store computer executable program code, which includes instructions. The internal memory 260 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound acquisition function, an image shooting function, etc.), etc. The data storage area can store data created during the use of the mobile terminal device 200 (such as audio data, image data, etc.), etc. In addition, the internal memory 260 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the mobile terminal device 200 by running instructions stored in the internal memory 260 and / or instructions stored in a memory provided in the processor.

[0050] In some embodiments, the internal memory 260 stores instructions for executing a method for driving real-time expressions of a cartoon digital human. By executing the instructions stored in the internal memory 260, the processor 210 can implement the following functions: pre-training a cartoon character expression driving network model using incomplete supervision using training images and a target loss function based on a fully connected neural network and convolutional structure, and deploying this model to the mobile terminal device 200. Then, upon receiving a cartoon digital human generation instruction triggered by a user on the camera interface of the mobile terminal device 200, the processor 210 can respond to this instruction by capturing an image (hereinafter referred to as a target image) containing facial information of a user (hereinafter referred to as a target user) at a preset frequency (e.g., 30 fps) via the camera 270. Each target image is then input frame by frame into the FaceMesh network for 3D reconstruction of the target user's facial image, resulting in real-time estimation of 468 3D facial landmarks as 3D key point information of the target user's face. Next, the 3D key point information of the target user's face is input into the cartoon character expression driving network model to predict the cartoon character expression base coefficients and rotation feature representation; then the cartoon character expression base coefficients and rotation feature representation can be used, combined with a preset blend deformation (BlendShape) parameter model (referring to a predefined facial shape set), and rendered according to relevant knowledge of computer graphics to obtain the cartoon character expression animation image corresponding to the target image; and all the obtained cartoon character expression animation images are continuously output in sequence to form a high-real-time, high-precision real-time facial expression driven animation of a cartoon digital human corresponding to the target user.

[0051] The camera 270 is used to capture static images or videos. For example, when a user holds the mobile terminal device 200, the camera 270 installed on the mobile terminal device 200 can be used to capture images containing the user's facial information at a fixed frequency (e.g., 30 fps). In some embodiments, the mobile terminal device 200 may include 1 or K cameras 270, where K is a positive integer greater than 1.

[0052] Mobile terminal device 200 implements display functions through a GPU, display screen 250, and an application processor. The GPU is a microprocessor for image processing that connects display screen 250 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or modify display information.

[0053] The mobile terminal device 200 can implement audio functions such as music playback, recording, and voice input and output through the audio module 280, speaker 280A, receiver 280B, microphone 280C, headphone jack 280D, and application processor.

[0054] The audio module 280 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 280 can also be used to encode and decode audio signals. In some embodiments, the audio module 280 can be provided in the processor 210, or some functional modules of the audio module 280 can be provided in the processor 210.

[0055] The speaker 280A, also called a "horn," is used to convert audio electrical signals into sound signals. The mobile terminal device 200 can listen to music or make hands-free calls through the speaker 280A.

[0056] The receiver 280B, also called a "handset", is used to convert audio electrical signals into sound signals. When the mobile terminal device 200 receives a call or voice message, the user can hear the voice by placing the receiver 280B close to the ear.

[0057] The microphone 280C, also known as a "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 280C to input the sound signal into the microphone 280C. The mobile terminal device 200 can be provided with at least one microphone 280C. In other embodiments, the mobile terminal device 200 can be provided with two microphones 280C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 200 can also be provided with three, four or more microphones 280C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.

[0058] The headphone jack 280D is used to connect a wired headphone, and the standard attributes of the interface are not limited.

[0059] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present application is only a schematic illustration and does not constitute a structural limitation on the mobile terminal device 200.

[0060] The wireless communication function of the mobile terminal device 200 can be implemented through the antenna 1, the antenna 2, the mobile communication module 220, the wireless communication module 230, the modem processor and the baseband processor.

[0061] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile terminal device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0062] The mobile communication module 220 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the mobile terminal device 200. The mobile communication module 220 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 220 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 220 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 220 can be set in the processor 210. In some embodiments, at least some of the functional modules of the mobile communication module 220 can be set in the same device as at least some of the modules of the processor 210.

[0063] The wireless communication module 230 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the mobile terminal device 200. The wireless communication module 230 can be one or more devices that integrate at least one communication processing module. The wireless communication module 230 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 230 can also receive the signal to be sent from the processor 210, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0064] In addition, on top of the above components, the mobile terminal device 200 runs an operating system, such as an iOS operating system, an Android operating system, a Windows operating system, etc. Applications can be installed and run on the operating system.

[0065] See also Figure 3 , which shows a schematic diagram of the software structure of the mobile terminal device provided in an embodiment of the present application.

[0066] The software system of the mobile terminal device 200 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the mobile terminal device 200.

[0067] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers: application program layer (APK), application framework layer (Framework), hardware abstraction layer (HAL), driver layer, and hardware layer, from top to bottom.

[0068] The application layer can include a series of application packages (APP). Figure 3As shown, the application package may include applications such as camera, call, navigation, WLAN, Bluetooth, gallery, etc. When the user holds the mobile terminal device 200 to take a photo, the camera application can communicate with the camera-related device in the camera access interface of the framework layer to request camera functions and obtain image data.

[0069] The application framework layer (framework layer) provides application programming interface (API) and programming framework for the application layer. The application framework layer includes some predefined functions. Figure 3 As shown, the application framework layer may include a window manager, a notification manager, a resource manager, and camera access interfaces (including but not limited to camera management and camera devices). The application framework layer enables interaction between camera services and camera APIs. In other words, it provides a unified interface that enables different camera hardware to interact with different camera applications. The framework layer also handles many common aspects of camera functionality, such as autofocus.

[0070] The window manager is used to manage window programs. It can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0071] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0072] In this embodiment, the hardware abstraction layer (HAL) provides a set of standard interfaces, which enables the framework layer to communicate with camera hardware from various manufacturers without having to understand the underlying hardware details. The hardware abstraction layer (HAL) stores the hardware abstraction layer and the camera algorithm library. It should be noted that the real-time expression driving algorithm for cartoon digital humans provided in this application can be stored in the camera algorithm library of the HAL layer, such as Figure 3 shown.

[0073] The real-time expression-driven algorithm for a cartoon digital human first interacts with the user on a shooting interface (e.g., a user taking a selfie in a short video app) to receive a cartoon digital human generation instruction from the user. In response to this instruction, the algorithm then uses a camera to capture image frames containing the human face at a fixed frame rate. The algorithm then processes and processes the data from each image frame to generate character expression animation frames. After all these generated character expression animation frames are connected and output, real-time facial expression driving for the cartoon digital human is achieved. Specifically, upon receiving a cartoon digital human generation instruction triggered by the user on the shooting interface of the mobile terminal device 200, the algorithm can respond to this instruction by capturing a target image containing the user's facial information via the camera 270 at a fixed frequency (e.g., 30 fps). Each target image is then input frame by frame into the FaceMesh network for 3D reconstruction of the user's facial image, resulting in the real-time estimation of 468 3D facial landmarks as 3D key point information for the user's face. Next, the 3D key point information of the user's face is input into the pre-built cartoon character expression-driven network model to predict the cartoon character's expression base coefficients and rotation feature representations; then, the cartoon character's expression base coefficients and rotation feature representations can be used, combined with a predefined set of facial shapes (such as 52 BlendShape parameter models), and rendered according to relevant knowledge of computer graphics to obtain the cartoon character expression animation image corresponding to the target image; and all the obtained cartoon character expression animation images are continuously output in sequence to form a high-real-time, high-precision real-time facial expression-driven animation of a cartoon digital human corresponding to the target user.

[0074] The hardware layer may include the hardware components of the electronic device mentioned above. Figure 3 The camera, image signal processor, digital signal processor and graphics processor were displayed.

[0075] Among them, the camera is used to perform focus and exposure processing for image shooting when acquiring images containing facial information at a fixed frequency (such as 30fps).

[0076] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the mobile terminal device 200 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0077] In this way, the interoperability of different camera applications and hardware devices on the Android platform can be achieved through the application layer, framework layer, HAL layer, driver layer and hardware layer, and the high real-time and high-precision expression driving of cartoon digital people can be realized, thereby providing users with a more novel and interesting social experience.

[0078] The technical solutions involved in the following embodiments can all be implemented in electronic devices having the above-mentioned hardware architecture and software architecture.

[0079] Next, the specific implementation process of the real-time expression driving method for cartoon digital humans provided by this application will be introduced in detail:

[0080] like Figure 4 As shown, the specific implementation process of the cartoon digital human real-time expression driving method may include the following steps S401-S404:

[0081] S401: The terminal device obtains a target image containing facial information of a target user through a camera at a preset frequency.

[0082] In this embodiment, when the terminal device (such as a mobile phone) and the user are in a shooting interface (such as an interface for taking a selfie in a short video application installed on the user's mobile phone, such as Figure 1 As shown in the figure, when the user interacts with the user and receives the cartoon digital human generation instruction issued by the user, the user can first obtain an image containing the user's facial information at a preset frequency through a pre-configured camera (such as the front camera of the mobile phone) on the terminal device (such as a mobile phone) to execute the subsequent step S402.

[0083] It should be noted that, in order to facilitate the subsequent explanation of the real-time expression driving method for cartoon digital humans proposed in this application, the image containing the user's facial information obtained by the terminal device through the camera at a preset frequency in step S401 is now defined as the target image, and the user contained therein is defined as the target user. In this way, by executing the subsequent steps S402-S404, the obtained target image frames are processed and processed to generate a highly real-time and high-precision facial expression driven animation of the cartoon digital human corresponding to the target user.

[0084] Furthermore, it should be noted that this application does not limit the preset frequency of the terminal device acquiring the target image, and it can be set according to actual conditions and empirical values. For example, the preset frequency can be determined to be 30fps, that is, 30 frames of the target image can be captured per second. In addition, this embodiment does not limit the type of target image. For example, the target image can be any color image composed of the three primary colors of red (R), green (G), and blue (B), and it is defined as I. Furthermore, this embodiment does not limit the size of the target image, that is, I∈R h*w*3 , where h and w represent the height and width of the target image I, respectively, and their specific values are not limited; 3 means that the target image I contains pixel information of three channels: red (R), green (G), and blue (B).

[0085] S402: Input the target image into the facial mesh FaceMesh network to perform 3D reconstruction of the target user's facial image to obtain 3D key point information of the target user's face.

[0086] In this embodiment, after the terminal device obtains each frame of target image containing the target user's facial information through the camera at a preset frequency (such as 30fps) in step S401, in order to improve the real-time processing of the target image, the target image can be input into the facial mesh (FaceMesh) network frame by frame for 3D reconstruction of the target user's facial image to obtain the target user's facial 3D key point information for executing the subsequent step S403.

[0087] It should be noted that the reason why this application uses the FaceMesh network to extract the 3D key point information of the target user's face is that, compared to the traditional expression-driven method that requires the use of multi-view RGBD images collected, the FaceMesh network can use machine learning (ML) to infer the geometry of the 3D facial surface with only a single camera input, without the need for a dedicated depth sensor. It uses a lightweight model architecture and GPU acceleration throughout the pipeline to provide real-time performance. Moreover, on this basis, it can also be combined with the FaceTransform module to bridge the gap between facial landmark estimation and practical real-time augmented reality (AR) applications. It also establishes a measurable 3D space and uses the screen position of facial landmarks to estimate facial transformations within the space, thereby meeting the high real-time requirements of mobile terminal devices.

[0088] Specifically, an optional implementation method is as follows: Figure 5 As shown, the specific implementation process of this step S402 may include: first, cropping the target user's facial region from the target image to obtain the target user's facial region image, and this application does not limit the image cropping method, which can be set according to actual conditions and experience. Then, the target user's facial region image is rotated and resized to obtain the aligned target user's facial region image, which is then defined as I i , and make I i The size of meets the size standard of image data that can be processed by the FaceMesh network, that is, I i ∈R 256*256*3 At the same time, the inverse rotation matrix corresponding to the target user's facial area image can also be calculated, which is defined as M rotation , and M rotation ∈R 3*3 .

[0089] Then, the aligned target user face area image I iInput the FaceMesh network to perform 3D reconstruction of the target user's facial image and obtain the target user's 468 facial key point information. The specific processing formula is as follows:

[0090] landmarks=F θ (I i ), (1)

[0091] Among them, F θ (·) represents the image processing function of the FaceMesh network; landmarks represents the target user’s facial area image I after the FaceMesh network is aligned i Detect and track 468 facial landmarks, i.e., landmarks∈R 468*3 These landmarks include various characteristic points on the target user's face, such as the corners of the eyes, the tip of the nose, the corners of the mouth, and more detailed points related to facial expressions. Each landmark consists of x, y, and z coordinates, where the x and y coordinates are normalized to the range [0.0, 1.0] of the image width and height, and the z coordinate represents the depth of the landmark point, with the depth of the center of the human head as the origin. Smaller values indicate that the landmark point is closer to the camera.

[0092] Then, the 468 facial landmarks of the target user and the previously calculated inverse rotation matrix corresponding to the target user's facial region image are used to represent M rotation , you can calculate the target user's facial 3D key point information and define it as landmarks ori , the specific processing formula is as follows:

[0093] landmarks ori =landmarks×M rotation (2)

[0094] Among them, landmarks ori ∈R 468*3 .

[0095] S403: Input the target user's facial 3D key point information into a pre-built cartoon character expression driven network model to predict the cartoon character expression base coefficients and rotation feature representation; wherein, the cartoon character expression driven network model is based on a fully connected neural network and a convolutional structure, and is obtained by incompletely supervised training using training images and a target loss function.

[0096] In this embodiment, the terminal device obtains the target user's facial 3D key point information landmarks in step S402. oriFinally, in order to achieve the ideal driving effect, the target user’s facial 3D key point information landmarks can be further ori The pre-built cartoon character expression driving network model is input to perform prediction processing on the cartoon character, and the expression base coefficient and rotation feature representation of the cartoon character are predicted. The two are defined as w and r respectively, which are used to execute the subsequent step S404 to achieve high-real-time and high-precision expression driving of the cartoon digital human corresponding to the target user.

[0097] It should be noted that in order to improve the real-time expression driving effect of cartoon digital people, this application pre-uses a small amount of sample images containing sample user facial information and a large amount of unlabeled image data generated based on these small amounts of sample images. The training image data set is composed of a large amount of unlabeled image data, and the target loss function is combined to perform incomplete supervision training on the initial cartoon character expression driving network model composed of a fully connected neural network and a convolutional structure, thereby constructing a cartoon character expression driving network model with better prediction effect. In addition, this application does not limit the specific network composition structure of the cartoon character expression driving network model, and it can be selected and set according to actual conditions. An optional implementation method is that the pre-constructed cartoon character expression driving network model may include but is not limited to a key information encoder, a cartoon character expression base coefficient prediction module, and a rotation representation prediction module. For the specific training process of the cartoon character expression driving network model, please refer to the detailed description of the subsequent embodiments, which will not be repeated here.

[0098] Specifically, an optional implementation method is that the specific implementation process of "inputting the target user's facial 3D key point information into the pre-built cartoon character expression driven network model to predict the cartoon character expression basis coefficients and rotation feature representation" in this step S403 may include the following steps S4031-S4033:

[0099] S4031: Input the target user's facial 3D key point information into the key information encoder of the pre-built cartoon character expression-driven network model for feature extraction to obtain a latent space feature representation.

[0100] In this implementation, the terminal device obtains the target user's facial 3D key point information landmarks in step S402. ori After that, the target user’s facial 3D key point information landmarks can be further ori Input the key information encoder of the cartoon character expression-driven network model to perform feature extraction, obtain the latent space feature representation, and define it as feats to perform the subsequent step S4032.

[0101] Among them, this embodiment does not limit the specific composition structure of the key information encoder. In order to reduce the performance requirements of the terminal device processor and meet the real-time requirements of the terminal device as much as possible, the key information encoder used in the pre-construction of the cartoon character expression drive network model in this application is an encoder based on a fully connected network structure. The role of this encoder is to analyze and calculate the input target user's facial 3D key point information landmarks ori , and extract features from it, output a feature tensor in the latent space as the latent space feature representation feats, which provides prior information for the subsequent steps S4032 and S4033 to calculate the cartoon character expression base coefficient and rotation feature representation.

[0102] Specifically, in the feature extraction process, such as Figure 6 As shown in the figure, the key information encoder draws on the structure of the multi-layer perceptron mixer (MLP-Mixer) network, and now inputs the target user's facial 3D key point information landmarks ori (and landmarks ori ∈R 468*3 ) is split into 468 one-dimensional small image block area (patch) tensors with a length of 3 based on a single key point and defined as landmarks ori ={l1, l2, ..., l 468}, where l i ∈R 3 , i = 1, 2, ..., 468, and then normalize all key point information using a single patch tensor as a unit to obtain l′ i =LayerNorm(l i ), then for all l′ i Superposition and transposition are performed in the form of channels to obtain l′∈R 3*468 , then split l' into l'={l'1, l'2, l'3}, l' based on the channel j ∈R 468 , j = 1, 2, 3, l′ corresponding to each channel j Input a fully connected network mixer respectively 1,j (.), all mixers 1,j (l′ j ) will output the results l″ respectively j ∈R 108, where 108 represents a custom dimension example value, which can also be set to other dimensions lower than 468 based on actual conditions and experience. On the one hand, it can extract key features while achieving dimensionality reduction. On the other hand, it also facilitates the subsequent steps S4032 and S4033 to more quickly and accurately calculate the cartoon character expression base coefficients and rotation feature representations, ensuring the proportional consistency of the feature matrix dimensions.

[0103] Next, l″ j ∈R 108 Superposition and transposition are performed in the form of channels to obtain l″∈R 108*3 , then split l″ into l″={l″1, l″2, ..., l″ based on channel 108}, l″ k ∈R 3 , k=1, 2, ..., 108, for each channel corresponding to l″ k The tensors are normalized separately to obtain LayerNorm(l″ k ). The normalized results are then fed into a fully connected network mixer. 2,k (.),like Figure 6 As shown, all mixers 2,k (LayerNorm(l″ k )) will output the results l″′ respectively k ∈R 72 Finally, they are superimposed in the form of channels to obtain the feature tensor feats of the latent space, and feats∈R 108*72 , where the ratio of 108 to 72 is the same as the 6D rotation transformation mentioned in the subsequent step S4033, r∈R 3*2 The ratio is consistent, which is 3 to 2. It can be understood that this dimension is not a limiting value and can be adaptively adjusted according to actual conditions and experimental conditions.

[0104] It should be noted that the mixer in the key information encoder m,n (.) represents a simple fully connected residual structure unit, which is used to extract and analyze the features of another dimension while keeping the features of one dimension of the input tensor unchanged. Specifically, mixer 1,j (.), mixer 2,k (.) Feature extraction is performed on the facial key point information under specific spatial information and the spatial information under specific facial key point information.

[0105] It should also be noted that this application is for mixer m,nThe specific composition structure of (·) is not limited, and can be as follows Figure 7 For example, with mixer 1,j Taking the network structure of (·) as an example, the operation process of the fully connected residual basic unit can be expressed as: input tensor l″1∈R 468 After that, after a linear layer nn.Linear(), we get an operation process tensor l∈R n , where the tensor length n can be customized according to actual conditions, and then after a layer of Gaussian error linear unit activation layer nn.GeLU() and a layer of linear layer nn.Linear(), the fully connected residual unit intermediate tensor l' is obtained mid ∈R m , where the tensor length m can be customized according to actual conditions, and then the intermediate tensor l′ mid With the input tensor l′1∈R 468 Splice to get l′ mid ∈R m+468 , and finally pass through a linear layer nn.Linear(), and finally output l″1∈R 108 .

[0106] S4032: Inputting the latent space feature representation into the cartoon character expression base coefficient prediction module of the cartoon character expression driving network model to perform mean dimensionality reduction processing to obtain the cartoon character expression base coefficient.

[0107] In this implementation, after the latent space feature representation feats is obtained in step S4031, the target latent space feature representation feats can be further input into the cartoon character expression base coefficient prediction module of the cartoon character expression driving network model (this application defines it as D coeff (·)) performs mean dimensionality reduction processing to output a set of expression base coefficients through two simple fully connected layers as the cartoon character expression base coefficients, and define it as w, and w∈R 52 ,w={w i |i=1, 2, ..., 52}, used to execute the subsequent step S404.

[0108] The reason why w is set to 52 dimensions in this embodiment is that the 52 preset blend shape parameter models are taken into consideration, namely, the 52 predefined facial shapes, including left eye blink (eyeBlinkLeft), left eye looking down (eyeLookDownLeft), left eye looking at the tip of the nose (eyeLookInLeft), left eye looking left (eyeLookOutLeft), left eye looking up (eyeLookUpLeft), left eye squinting (eyeSquintLeft), left eye wide open (eyeWideLeft), right eye blink (eyeBlinkRight), right eye looking down (eyeLookD right eye (eyeLookInRight), right eye looking to the left (eyeLookOutRight), right eye looking upward (eyeLookUpRight), right eye squinting (eyeSquintRight), right eye wide open (eyeWideRight), chin forward when pouting (jawForward), chin to the left when pouting (jawLeft), chin to the right when pouting (jawRight), chin down when opening mouth (jawOpen), mouth closed (mouthClose), slightly open mouth with lips open (mouthFunnel), pursed lips (mouthPucker), pouting lips to the left (mouth outhLeft), right pout (mouthRight), smile with left pout (mouthSmileLeft), smile with right pout (mouthSmileRight), press left lip down (mouthFrownLeft), press right lip down (mouthFrownRight), left lip back (mouthDimpleLeft), right lip back (mouthDimpleRight), left corner of mouth to the left (mouthStretchLeft), right corner of mouth to the right (mouthStretchRight), lower lip rolled inward (mouthRollLower), lower lip rolled up (mouthRoll lUpper), lower lip down (mouthShrugLower), upper lip up (mouthShrugUpper), lower lip pressed to the left (mouthPressLeft), lower lip pressed to the right (mouthPressRight), lower lip pressed to the lower left (mouthLowerDownLeft), lower lip pressed to the lower right (mouthLowerDownRight), upper lip pressed to the upper left (mouthUpperUpLeft), upper lip pressed to the upper right (mouthUpperUpRight), left eyebrow outward (browDownLeft), right eyebrow outward (browDownRight),Frown (browInnerUp), left eyebrow up to the left (browOuterUpLeft), right eyebrow up to the right (browOuterUpRight), cheek outward (cheekPuff), left cheek up and rotated (cheekSquintLeft), right cheek up and rotated (cheekSquintRight), nose wrinkled left (noseSneerLeft), nose wrinkled right (noseSneerRight), tongue stuck out (tongueOut). "w" represents 52 sets of feature motion data, which are used to drive these 52 predefined facial shapes to form facial animations corresponding to the target user's 3D face image in the input target image.

[0109] It should be noted that this embodiment does not limit the specific composition structure of the cartoon character expression base coefficient prediction module. In order to reduce the performance requirements of the terminal device processor and meet the real-time requirements of the terminal device as much as possible, the cartoon character expression base coefficient prediction module used in the pre-construction of the cartoon character expression driving network model in this application is a decoder constructed based on a fully connected network structure. The encoder D coeff The role of (·) is to represent the latent space features feats∈R 108*72 Perform decoding processing to obtain the cartoon character expression base coefficient w∈R 52 ,w={w i |i=1,2,...,52}. Among them, w i The specific value of is not limited, but the value range can be set between 0-1, where 0 means the corresponding facial shape is completely inactive, and 1 means the corresponding facial shape is fully activated. The closer the value is to 1, the more obvious the expression of the corresponding facial shape is. For example, mouthPressLeft, 0 means inactive, and 1 means fully activated.

[0110] like Figure 8 As shown, in the decoding process, for the latent space feature representation feats∈R 108*72 The process of mean dimensionality reduction is as follows: first, the latent space feature tensor feats∈R 108*72 Expressed as Among them, f i ∈R 72 , then for all f i , i=1,2,…108 take the mean, and we get and Then, after another linear layer nn.Linear() and a layer of Softmax() activation, the prediction result w∈R of the cartoon character's expression base coefficient w can be output. 52 .

[0111] S4033: Input the latent space feature representation into the rotation representation prediction module of the cartoon character expression driven network model for size transformation and convolution processing to obtain the rotation feature representation.

[0112] In this implementation, after the latent space feature representation feats is obtained in step S4031, the target can be further represented by the latent space feature feats and input into the rotation representation prediction module of the cartoon character expression driving network model (this application defines it as E mlp (.)) performs size transformation and simple convolution to obtain a 6D rotation transformation representation as the rotation feature representation, and defines it as r, and r∈R 3*2 , used to execute the subsequent step S404.

[0113] Among them, this application does not limit the specific format content of the rotation feature representation. The reason why this embodiment uses 6D rotation transformation representation as the rotation feature representation is that: the 6D rotation transformation representation can map the rotation changes corresponding to the 3D model of the cartoon character's facial animation to the 2D plane, and the 3D rotation matrix can also be identically mapped with the corresponding 3*3 rotation matrix representation through specific mathematical operation rules. In addition, the advantage of the 6D rotation transformation representation is that this representation method can maintain the continuity of rotation (emphasis - continuity) in deep learning, thereby improving the learning effect of the model. In deep learning, the 6D representation performs better than the traditional rotation representation in tasks such as rotation estimation and inverse kinematics. Some traditional rotation representation methods may cause discontinuities in neural network training, thereby affecting the performance of the model. Therefore, this embodiment takes advantage of the aforementioned advantages of the 6D rotation transformation representation and uses the 6D rotation transformation representation as the representation of the rotation movement of the cartoon digital human head (i.e., the rotation feature representation) in deep learning tasks.

[0114] It should be noted that this embodiment does not limit the specific composition structure of the rotation representation prediction module. In order to reduce the performance requirements of the terminal device processor and meet the real-time requirements of the terminal device as much as possible, the rotation representation prediction module used in the pre-construction of the cartoon character expression driving network model in this application is a decoder based on a convolutional structure. The encoder E mlp The role of (.) is to represent the latent space features feats∈R 108*72 Perform decoding processing to obtain the rotation feature representation r (such as 6D rotation transformation representation r∈R 3*3 ).

[0115] like Figure 9 As shown, in the decoding process, for the latent space feature representation feats∈R 108*72 First, perform feature size transformation (resize) operation to obtain the intermediate feature f″∈R 96*64, where the resize operation is a basic and important step in computer vision, which can effectively adjust the image size to meet the needs of different applications. The specific operation method is to interpolate the feature values corresponding to adjacent pixels in the feature map to achieve the purpose of changing the size of the feature map. Then, after two convolution operations (the convolution kernel size is set to 3, the feature block edge padding is 1, and the step size is set to 2), we get f″∈R 24*16 , and finally after another feature size transformation, we can get the 6D rotation transformation representation r∈R 3*2 .

[0116] S404: Using the cartoon character's expression base coefficients and rotation feature representation, combined with the preset blend deformation BlendShape parameter model, render the cartoon character's expression animation image corresponding to the target image; and continuously output the cartoon character's expression animation image to form a real-time facial expression driven animation of the cartoon digital human corresponding to the target user.

[0117] In this embodiment, the terminal device uses the cartoon character expression to drive the network model in step S403 to predict the cartoon character expression base coefficient w (such as w∈R 52 ) and the rotation feature representation r (ie, 6D rotation transformation representation r∈R 3*2 ), in order to achieve the ideal driving effect, the cartoon character expression base coefficient w can be further used to perform weighted summation on the preset blend deformation (BlendShape) parameter model (i.e., the predefined 52 facial shapes mentioned in the above step S4032), and the cartoon character expression model corresponding to the target image can be predicted, and then the rotation feature representation r (i.e., the 6D rotation transformation representation r∈R 3*2 ), the predicted cartoon character expression model corresponding to the target image is rotated and transformed to obtain the cartoon character expression model mapped to the 2D plane, and the cartoon character expression animation is rendered to obtain the cartoon character expression animation, and then the cartoon character expression animation is aligned with the target user's face according to the target user's facial area image to obtain the cartoon character expression animation image corresponding to the target image, and finally, all the obtained cartoon character expression animation images are connected in chronological order, and all the connected cartoon character expression animation images are continuously output to form the real-time facial expression driven animation of the cartoon digital human corresponding to the target user, thereby improving the target user's social experience.

[0118] Specifically, since each BlendShape parameter model corresponds to a specific facial deformation (such as frowning), when multiple BlendShape parameter models are combined with different weights, complex facial expressions can be generated. The BlendShapes parameter model is technically a set of dictionaries that store the user's facial expression feature motion factors, which contains a total of 52 sets of feature motion data (i.e., weight combination coefficients). Taking the 52 predefined facial shapes (i.e., 52 sets of expression feature 3D mesh models) mentioned in the above step S4032 as an example, these 52 facial shapes are combined with these motion factors (i.e., cartoon character expression base coefficients w∈R 52 ) can realize the driving from 2D face image to 2D or 3D character model. This embodiment uses the BlendShapes parameter model as the technical basis, and uses the cartoon character expression base coefficient w∈R 52 As well as the predefined 52 facial shapes (i.e., 52 sets of expression feature 3D mesh models), various expressions of cartoon characters can be realized through the weighted summation method. The specific formula is as follows:

[0119]

[0120] Among them, b represents the predicted cartoon character expression model; b0 represents the reference shape, which is usually the face shape with neutral expression. All BlendShapes parameter models are deformed based on this reference shape. i |i=1, 2, ..., 52} represents 52 predefined facial shapes (i.e., 52 sets of facial expression feature 3D mesh models), each shape representing a specific facial deformation, such as frowning, etc., see the introduction of step S4032 above for details. i |i=1, 2, ..., 52} represents the weight set of 52 predefined facial shapes (i.e., 52 sets of expression feature 3D mesh models), which is the cartoon character expression base coefficient w∈R predicted by step S403. 52 , the value range is usually between 0 and 1, indicating the influence of the corresponding type of facial shape on the final facial expression.

[0121] In this way, for any cartoon character, the mesh model corresponding to the predefined 52 facial shapes (i.e., 52 sets of expression feature 3D mesh models) and its reference shape mesh model {b i |i=0,1,2,...,52},combined with the cartoon character expression base coefficient w∈R predicted in step S403 52 ,w={w i |i=1,2,...,52}, the corresponding cartoon character expression animation mesh model driven by the 2D face image is predicted.

[0122] At the same time, based on the realization of the identity mapping between the 3D rotation 9*9 matrix and the 6D rotation change representation, the 6D rotation change representation, combined with the cartoon character expression animation 3D-mesh model, can simultaneously predict the rotation change of the character avatar and map it to the 2D plane. Therefore, the cartoon character expression driven network model in this application can predict and output the 6D rotation change representation r∈R corresponding to the target user's face. 3*2 , and then use this 6D rotation change to represent r∈R 3*2 , the predicted cartoon character expression model b is subjected to a 6D rotation transformation, and a cartoon character expression model mapped to a 2D plane can be obtained. Further, the cartoon character expression model is rendered according to relevant knowledge of computer graphics to obtain a cartoon character expression animation.

[0123] Finally, the target user's facial area can be cropped from the target image as mentioned in the above step S402 to obtain the target user's facial area image, and the rendered cartoon character expression animation can be aligned with the target user's face. In this way, after the target image frame containing the target user's facial information is captured by the front camera of the mobile terminal device in real time at a preset frequency (such as 30fps), the image data can be processed and processed to obtain the character expression animation frame by executing the process of the above steps S402-S404. Then, all the character expression animation frames are connected and continuously output, thereby realizing real-time high-precision expression driving of the cartoon digital human.

[0124] Next, this embodiment will introduce the training process of the cartoon character expression driven network model in detail. The specific training process may include the following steps AC:

[0125] Step A: Obtain a sample image containing sample user facial information.

[0126] It should be noted that, since this application adopts basic neural network operators that are friendly to mobile terminal device deployment when constructing the cartoon character expression driving network model, and utilizes a smaller number of parameters and a simple model structure (such as fully connected layers and convolutional structures), it greatly improves the reasoning efficiency of mobile devices, and at the same time does not place excessive demands on the computing performance and power consumption of the device, which leads to certain challenges in the training accuracy of the model. From an empirical point of view, the reasoning performance of the model depends to a large extent on the number of parameters of the model and the quality of the training data. Under the premise that the model structure is fixed, the volume and quality of the training data greatly affect the calculation quality of the model. The real-time expression driving method of cartoon digital humans proposed in this application uses image frames containing user facial information as the original input data, obtains the user's facial 3D key point information through FaceMesh network reasoning, and then uses the cartoon character expression driving network model to predict and output the cartoon character expression base coefficient w∈R 52 And the rotation feature representation (i.e. 6D rotation transformation representation) r∈R 3*2 Therefore, how to obtain high-precision training data (input data & output label data pairs) is particularly important.

[0127] To this end, this embodiment adopts an incompletely supervised training strategy, using a small amount of labeled data and a large amount of unlabeled data to train the cartoon character expression-driven network model, thereby improving the performance and generalization ability of the model.

[0128] Specifically, in order to build a cartoon character expression-driven network model, it is first necessary to obtain a small amount of raw data (i.e., a small number of sample images containing sample user facial information) as a basis. Then, through the subsequent step B, it is processed, treated, calculated, etc. to generate a large number of high-precision input and output data pairs as training data for model training. Among them, this embodiment does not limit the method and number of sample images to obtain. For example, the method of obtaining a small amount of raw data (i.e., a small number of sample images containing sample user facial information) can be summarized as follows:

[0129] (1) The baseline cartoon character mesh produced by the art engineer;

[0130] (2) 52 cartoon character expression base meshes created by art engineers;

[0131] (3) 8000 original collected 3D meshes of real faces under standard conditions.

[0132] Step B: Using the preset coefficient prior and rotation prior as constraints, use sample images to construct input and output data pairs consisting of the sample user's facial 3D key point information, the sample cartoon character's expression base coefficients, and the sample rotation feature representation as the training image dataset.

[0133] It should be noted that in order to achieve incomplete supervised training of the model, after obtaining a small amount of original data (i.e., a small number of sample images containing sample user facial information) through step A, a large number of input-output data pairs consisting of sample user's facial 3D key point information, sample cartoon character expression base coefficients and sample rotation feature representation can be further generated based on these small amounts of original data, and defined as landmarksori-coefficients&rotation matrix input-output data pairs as training image data sets to execute subsequent step C.

[0134] Specifically, in order to construct a training image dataset, we first need to determine the restrictions that the input and output data pairs that can be used as training images need to meet, that is, we need to first determine the label data prior, including coefficient prior and rotation prior. These two parts serve as restrictions, which can provide prior restrictions for the subsequent generation of sample cartoon character expression base coefficient labels and sample rotation feature representation labels in the label data of the input and output data pairs, so as to ensure that the generated input and output data pairs are labels that may appear in the actual situation of human facial expressions.

[0135] Among them, the role of the rotation prior is to limit the actual head rotation angle corresponding to all 6D rotation transformations to avoid the situation where the rotation angle is too large and causes distortion of the training data. The rotation prior is used as a constraint condition, and its specific restrictions are: first determine the mesh grid coordinates corresponding to the rotation center position, which can usually be the center of the head and neck, and then limit the possible rotation angle range of the head in the three dimensions of x, y, and z in actual situations, and define it as [-θ x ,θ x ],[-θ y ,θ y ],[-θ z ,θ z ], wherein the rotation angle in the x-direction represents the angle at which a person tilts his head left or right, the rotation angle in the y-direction represents the angle at which a person nods his head up and down, and the rotation angle in the z-direction represents the angle at which a person shakes his head left or right. It should be noted that this application does not limit the values of the rotation angle ranges of the above three dimensions, and they can be set according to actual conditions and empirical values. For example, the rotation angle ranges of the three dimensions of x, y, and z can be set to [-45, 45], [-45, 45], [-70, 70], etc., respectively. However, the rotation angle limit ranges of these three dimensions need to meet the following requirements: In actual situations, random head rotation can meet the realization of random facial expressions and movements, and no distorted rotation will occur.

[0136] Regarding the coefficient prior, different coefficient combinations, combined with the predefined 52 facial shape mesh expression bases, can be combined to form different facial expression shapes through a weighted summation method. However, since the weighted summation of randomly selected coefficient combinations can easily result in distorted facial expressions, this type of data is not helpful for network training and can have a negative impact. Therefore, to limit the authenticity of the generated coefficient labels, this application proposes the following constraints: First, the 52 facial shape mesh expression bases are grouped, and the facial area is divided into four regions: left eye, right eye, mouth, and cheek. Within each region, mutually exclusive groups of combinations are extracted for all facial shapes. Any two facial shapes in each mutually exclusive group cannot appear at the same time. That is, within all facial shapes in each mutually exclusive group, only 0 or 1 facial shape can have a non-zero coefficient. For example, "left-turning lip" and "right-turning lip" must exist in the same mutually exclusive group; they cannot be activated at the same time. A weighted combination of the two will result in distorted expressions.

[0137] Furthermore, after determining the label data prior including coefficient prior and rotation prior, the preset coefficient prior and rotation prior can be used as restriction conditions, and a small number of sample images obtained can be used to construct a large number of input-output data pairs consisting of sample user's facial 3D key point information, sample cartoon character expression base coefficients and sample rotation feature representation (i.e., landmarksori-coefficients&rotation matrix input-output data pairs) as training image data sets.

[0138] Specifically, 52 facial shape and expression meshes corresponding to each standard-state real-world face mesh can be generated. The deformation transfer technique employed is as follows: given the original mesh information S, the deformed mesh information S′ based on the original mesh, and the target mesh information T, the target mesh information T′′ can be calculated and generated after undergoing the same deformation as the original mesh. The two meshes are morphologically identical, with similar deformation relationships between S and S′′, and between T and T′′, and S and T do not need to have the same number of mesh vertices. Based on this, the standard-state mesh information of the cartoon character, the 52 facial expression shape meshes of the cartoon character, and the standard-state mesh information of 8,000 real-world faces are combined to generate 52 meshes corresponding to each sample user's face. In other words, 52 BlendShapes parametric models corresponding to each of the 8,000 real-world users can be generated.

[0139] Then, under the constraints of preset coefficient priors and rotation priors, a large number of random sample cartoon character expression base coefficient combinations and sample 6D rotation transformation representations can be randomly generated, and a random set of sample cartoon character expression base coefficient combinations and 52 groups of facial expression shapes corresponding to a random sample user are weightedly combined to obtain the real facial expression mesh information corresponding to the sample user. Then, 468 3D facial key point information is extracted from the corresponding mesh, and a 6D rotation transformation representation that satisfies the rotation prior is randomly specified. Its identity is mapped to a 3*3 rotation matrix, and the extracted 3D key point information is rotated to obtain the 3D key point information of the specific expression of the corresponding sample user at a certain head rotation angle. This information is used as the input of the initial cartoon character expression driven network model, and the 52 cartoon character expression base coefficient combinations and 6D rotation transformation representations are used as the output of the label driven network, that is, a set of landmarksori-coefficients&rotation Matrix input and output data pairs, and so on, by collecting random coefficients that meet the label prior and 6D rotation transformation representation collection multiple times for all 8,000 real human sample users, a large number of training data sets with guaranteed accuracy can be obtained.

[0140] Step C: Using the training image data set, the first target loss function, the second target loss function and the third target loss function to train the initial cartoon character expression driven network model to generate a cartoon character expression driven network model.

[0141] After obtaining a large number of training data sets with guaranteed accuracy through step B, the initial cartoon character expression-driven network model can be further trained based on the obtained training data set, the first target loss function, the second target loss function and the third target loss function. In the training process, the parameters of the initial cartoon character expression-driven network model can be continuously updated according to the changes in the function values of the first target loss function, the second target loss function and the third target loss function until the function values of the first target loss function, the second target loss function and the third target loss function meet the requirements, such as the weighted sum of the three reaches the minimum value and the change range is very small (basically unchanged), or reaches the preset maximum number of iterations (such as 10,000 times), then the update of the model parameters is stopped, the training of the initial cartoon character expression-driven network model is completed, and the cartoon character expression-driven network model is generated.

[0142] Among them, the first objective function is used to constrain the cartoon character expression base coefficient prediction module D in the cartoon character expression driven network model coeff (.) Output cartoon character expression base coefficient w∈R 52 , so that it is consistent with the sample cartoon character expression base coefficient in the training image dataset (using w GT ∈R52 The specific calculation formula is as follows:

[0143] l coeffs =||ww GT ||2 (4)

[0144] Among them, l coeffs Represents the first objective loss function.

[0145] The second objective function is used to constrain the rotation representation prediction module d in the cartoon character expression driven network model. rotation The rotation feature representation of the output of (.) r∈R 3*2 , so that it is consistent with the sample rotation feature representation in the training image dataset (using r GT ∈R 3*2 The specific calculation formula is as follows:

[0146] l rotation =||rr GT ||2 (5)

[0147] Among them, l rotation Represents the second objective loss function.

[0148] In addition, the third objective function is used to simultaneously constrain the cartoon character expression base coefficient prediction module D in the cartoon character expression driven network model. coeff (.) Output cartoon character expression base coefficient w∈R 52 and rotation representation prediction module D rotation The rotation feature representation of the output of (.) r∈R 3*2 , so that the two are respectively compared with the sample cartoon character expression base coefficients (w GT ∈R 52 ) and sample rotation feature representation (r GT ∈R 3×2 ) as close as possible, specifically, firstly, according to the coefficient weighting and combined with the different facial shapes of the cartoon characters, the predicted mesh grid state of the cartoon character's facial expression (expressed by b) and the label mesh grid state (expressed by b) are calculated respectively. GT The calculation formula is as follows:

[0149]

[0150] l recon =||bb GT ||2 (8)

[0151] Among them, l recon Represents the third objective loss function.

[0152] An optional implementation method is to perform weighted summation on the first objective function, the second objective function, and the third objective function to form a total target loss function, and define it as L, to pre-train the initial cartoon character expression driven network model, and in the training process, the model parameters of the initial cartoon character expression driven network model can be continuously updated according to the change of the function value of the total target loss function L, until the function value of L meets the requirements, such as reaching the minimum value and the change amplitude is very small (basically unchanged), and the model converges. Then stop updating the model parameters, complete the training of the initial cartoon character expression driven network model, and generate the cartoon character expression driven network model. Among them, the specific calculation formula of L is as follows:

[0153] L=λ1*l coeffs +λ2*l rotation +λ3*l rec on (9)

[0154] Among them, λ1, λ2, and λ3 all represent preset coefficients, which can be set according to actual conditions and empirical values. This application does not limit the specific values of the three. For example, λ1, λ2, and λ3 can be set to 0.3, 0.3, 0.6, etc. respectively.

[0155] In terms of optimizer, this application uses the Adam optimizer with an initial learning rate of lr = 1e-5 and eps = 1e-8. Training is performed for 100 epochs, with the learning rate decayed by 0.1 every 30 epochs. Other common optimizers are also available, including but not limited to SGD and RMSprop.

[0156] To facilitate understanding of the real-time expression driving method for cartoon digital human proposed in this application, this application also provides a schematic diagram of the overall implementation process of real-time expression driving for cartoon digital human, as shown in FIG. Figure 8As shown in the figure, specifically: a video stream containing facial information of the target user's face is collected by the front camera of a mobile terminal device (such as a mobile phone) at a preset frequency (such as 30fps), and is parsed frame by frame into a target image containing facial information of the target user's face, which is input into the FaceMesh network, and 468 3D facial landmarks are estimated in real time as 3D key point information of the target user's face, which is then input into the cartoon character expression driven network model. The model is a fully connected-convolutional hybrid neural network structure based on the MLP-Mixer architecture and the basic structural unit of the fully connected neural network, including a key information encoder, a cartoon character expression base coefficient prediction module and a rotation representation prediction module. The key information encoding module extracts features from the target user's facial 3D key point information to obtain the latent space feature representation of the target user's facial information, which is then input into the cartoon character expression base coefficient prediction module and the rotation representation prediction module respectively to obtain a set of cartoon character expression base coefficients and rotation feature representations (i.e., 6D rotation transformation representations) calculated based on the target user's facial key point information. Finally, the output of the cartoon character expression-driven network model (i.e., cartoon character expression base coefficients and rotation feature representations (i.e., 6D rotation transformation representations)) is combined with the cartoon character's corresponding expression base (i.e., the predefined 52 facial shapes) to render an expression animation image corresponding to the cartoon character and the target user's face RGB.

[0157] As can be seen, this application uses a mobile terminal device-friendly neural network infrastructure, which facilitates the quantitative deployment of various deployment tools. At the same time, the cartoon character expression-driven network model constructed by this application achieves high-speed inference and terminal-side implementation. Moreover, this model uses a method of automatically annotated data to expand the training data set and is trained using an incompletely supervised training strategy, ultimately achieving high-real-time and high-precision expression driving for digital humans on mobile terminal devices.

[0158] In addition, the present application also provides an electronic device (i.e., a terminal device). For the hardware structure and software framework of the electronic device, please refer to Figure 2 and Figure 3 The electronic device includes a memory and a processor, wherein the memory stores a computer program and the processor is used to call and execute the computer program to implement the cartoon digital human real-time expression driving method provided in the above description.

[0159] The present application also provides a computer-readable storage medium in an embodiment, on which a computer program is stored. When the computer program is executed by a processor of a terminal device, it is used to implement the real-time expression driving method of a cartoon digital human provided in the above description.

[0160] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for driving real-time expression of cartoon digital human, characterized in that: Applied to a terminal device, the terminal device being equipped with a camera, the method comprising: Acquiring a target image containing target user's facial information through the camera at a preset frequency; Input the target image into the facial mesh FaceMesh network to perform 3D reconstruction of the target user's facial image to obtain 3D key point information of the target user's face; Inputting the target user's facial 3D key point information into a pre-built cartoon character expression driven network model to predict the cartoon character's expression basis coefficients and rotation feature representation; the cartoon character expression driven network model is based on a fully connected neural network and convolutional structure, and is obtained through incompletely supervised training using training images and a target loss function; The expression base coefficients and rotation feature representation of the cartoon character are used in combination with a preset blend deformation BlendShape parameter model to render a cartoon character expression animation image corresponding to the target image; and the cartoon character expression animation image is continuously output to form a real-time facial expression driven animation of a cartoon digital human corresponding to the target user.

2. The method according to claim 1, characterized in that Inputting the target image into the facial mesh FaceMesh network to perform 3D reconstruction of the target user's facial image to obtain 3D facial key point information of the target user includes: cropping a target user's facial region from the target image to obtain a target user's facial region image; Performing an alignment process of rotating and resizing the target user's facial region image to obtain an aligned target user's facial region image; and calculating an inverse rotation transformation matrix representation corresponding to the target user's facial region image; Inputting the aligned target user's facial region image into the facial mesh FaceMesh network to perform 3D reconstruction of the target user's facial image to obtain 468 facial key point information of the target user; The 3D facial key point information of the target user is calculated using the inverse rotation transformation matrix representation corresponding to the target user's facial region image and the 468 facial key point information of the target user.

3. The method according to claim 1, characterized in that The cartoon character expression driven network model includes a key information encoder, a cartoon character expression base coefficient prediction module, and a rotation representation prediction module; the 3D key point information of the target user's face is input into the pre-built cartoon character expression driven network model to predict the cartoon character expression base coefficient and rotation feature representation, including: Inputting the target user's facial 3D key point information into a pre-built key information encoder of a cartoon character expression-driven network model for feature extraction to obtain a latent space feature representation; Inputting the latent space feature representation into the cartoon character expression base coefficient prediction module of the cartoon character expression driving network model to perform mean dimensionality reduction processing to obtain the cartoon character expression base coefficient; The latent space feature representation is input into the rotation representation prediction module of the cartoon character expression driven network model for size transformation and convolution processing to obtain the rotation feature representation.

4. The method according to claim 3, characterized in that The key information encoder is an encoder constructed based on a fully connected network structure; the cartoon character expression base coefficient prediction module is a decoder constructed based on a fully connected network structure; and the rotation representation prediction module is a decoder constructed based on a convolutional structure.

5. The method according to claim 2, characterized in that The method of rendering a cartoon character expression animation image corresponding to the target image by using the cartoon character expression base coefficient and the rotation feature representation in combination with a preset blend deformation BlendShape parameter model includes: Using the cartoon character expression base coefficients, a weighted summation is performed on a preset blend deformation BlendShape parameter model to predict a cartoon character expression model corresponding to the target image; Using the rotation feature representation, a rotation transformation process is performed on the predicted cartoon character expression model corresponding to the target image to obtain the cartoon character expression model mapped to a 2D plane, and the cartoon character expression model is rendered to obtain a cartoon character expression animation; According to the target user's facial region image, the cartoon character expression animation is aligned with the target user's face to obtain a cartoon character expression animation image corresponding to the target image.

6. The method according to claim 1, characterized in that The continuously outputting the cartoon character expression animation image to form a real-time facial expression driven animation of the cartoon digital human corresponding to the target user includes: The cartoon character expression animation images are connected in sequence, and all connected cartoon character expression animation images are continuously output to form a real-time facial expression driven animation of a cartoon digital human corresponding to the target user.

7. The method according to claim 1, characterized in that The target loss function includes a first target loss function, a second target loss function, and a third target loss function; the cartoon character expression driven network model is constructed as follows: Obtain a sample image containing sample user facial information; Using the preset coefficient prior and rotation prior as constraints, the sample images are used to construct input and output data pairs consisting of the sample user's facial 3D key point information, the sample cartoon character's expression base coefficients, and the sample rotation feature representation as a training image dataset; The initial cartoon character expression driven network model is trained using the training image data set, the first objective loss function, the second objective loss function and the third objective loss function to generate a cartoon character expression driven network model.

8. The method according to claim 6, characterized in that The first objective loss function is used to constrain the cartoon character expression base coefficients output by the cartoon character expression driven network model so that they are as close as possible to the sample cartoon character expression base coefficients in the training image data set; the second objective loss function is used to constrain the rotation feature representation output by the cartoon character expression driven network model so that it is as close as possible to the sample rotation feature representation in the training image data set; the third objective loss function is used to simultaneously constrain the cartoon character expression base coefficients and rotation feature representations output by the cartoon character expression driven network model so that the two are as close as possible to the sample cartoon character expression base coefficients and sample rotation feature representations in the training image data set, respectively.

9. The method according to any one of claims 1 to 8, characterized in that The rotation feature representation is a 6D rotation transformation representation.

10. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor is configured to call and execute the computer program to implement the method according to any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by an electronic device, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Virtual social contact method based on Avatar expression transplantation

    CN110135215A

  • Virtual expression generation method and device

    CN111798551A

  • Method and device for driving expression of virtual character

    CN114821734A

  • Face parameter determination method and device, electronic equipment and storage medium

    CN115512417A

  • Virtual image driving method and device, electronic equipment and storage medium

    CN115631295A

Cited By

  • Expression base generation and hybrid driving method and system based on topological consistency and sparse constraint

    CN121883674A