Animation generation device and method

CN116152402BActive Publication Date: 2026-08-07HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2022-12-30
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]通常在生成风格迁移动画时,可以根据用户需求生成卡通、拟人、写实、超写实等不同的动画风格,同时,风格迁移对象的形象塑造、驱动效果和应用场景部署也会对生成效果产生影响,由于在由真实人像的源域迁移至二维风格迁移对象的目标域过程中,在源域与目标域的风格差距较大时,导致生成的风格迁移对象的形变程度也会过大,影响后续驱动虚拟数字人的观赏效果,用户体验不佳

Benefits of technology

[0023]As can be seen from the above technical solutions, the animation generation device and method provided in this application, through a controller, performs style transfer on a target face image to obtain a transferred image corresponding to the target face image; performs key point detection on the transferred image to obtain facial key points of the transferred image; determines driving key points and driving anchor points of the transferred image based on the facial key points and the corner points of the transferred image; obtains texture coordinates of the transferred image based on the coordinates of the driving key points and the coordinates of the driving anchor points; performs triangulation on the driving key points and the driving anchor points to obtain multiple texture triangles of the transferred image; and drives the transferred image based on the vertex coordinate sequence corresponding to the driving speech obtained in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image. Compared with the prior art, the generated style transfer animation has poor effect and low viewing quality. This application achieves better stretching and driving of the transferred image by triangulating the style transfer image to obtain multiple texture triangles, thereby improving the effect of the generated animation and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152402B_ABST
    Figure CN116152402B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an animation generation device and method, and relate to the technical field of display. The animation generation device and method comprise: firstly performing style transfer on a target face image to obtain a transfer image corresponding to the target face image, then performing key point detection on the transfer image to obtain face key points of the transfer image, determining driving key points and driving anchor points of the transfer image according to the face key points and corner points of the transfer image, then obtaining texture coordinates of the transfer image and a plurality of texture triangles of the transfer image; finally driving the transfer image according to a vertex coordinate sequence corresponding to a real-time obtained driving voice, the texture coordinates and the plurality of texture triangles to generate a style transfer animation corresponding to the target face image. Embodiments of the present application are used to improve the conversion effect of generating a style transfer animation through a target portrait.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of display technology, and more particularly to an animation generation device and method. Background Technology

[0002] In the field of style transfer animation technology, artificial intelligence technology is used to reconstruct and drive the animation, enabling it to evolve from a digital appearance to an interactive behavior, and then to an intelligent concept. Style transfer animation is based on artificial intelligence and combines technologies such as speech synthesis, speech recognition, semantic understanding, and image vision. In modern society, in various forms of human-computer interaction and smart terminal experience, images and vision are the most direct carriers of communication and sensory transmission.

[0003] When generating style transfer animations, different animation styles such as cartoon, anthropomorphic, realistic, and hyper-realistic can be generated according to user needs. At the same time, the image shaping, driving effect, and application scenario deployment of the style transfer object will also affect the generated effect. When the style difference between the source domain and the target domain of the two-dimensional style transfer object is large during the migration from the source domain of a real human image to the target domain, the deformation degree of the generated style transfer object will be too large, which will affect the viewing effect of the subsequent driving of the virtual digital human and result in a poor user experience. Summary of the Invention

[0004] This application provides an exemplary embodiment of an animation generation device and method for improving the generation effect of style transfer animation.

[0005] The technical solutions provided in this application are as follows:

[0006] Firstly, an animation generation device is provided, comprising:

[0007] The controller is configured as follows:

[0008] Perform style transfer on the target face image to obtain the transferred image corresponding to the target face image;

[0009] Key point detection is performed on the migrated image to obtain the facial key points of the migrated image;

[0010] Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image;

[0011] The texture coordinates of the migrated image are obtained based on the coordinates of the driving key points and the coordinates of the driving anchor points.

[0012] The driving key points and the driving anchor points are triangulated to obtain multiple texture triangles of the migration image;

[0013] The transfer image is driven by the vertex coordinate sequence corresponding to the driving speech acquired in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image.

[0014] Secondly, an animation generation method is provided, including:

[0015] Perform style transfer on the target face image to obtain the transferred image corresponding to the target face image;

[0016] Key point detection is performed on the migrated image to obtain the facial key points of the migrated image;

[0017] Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image;

[0018] The texture coordinates of the migrated image are obtained based on the coordinates of the driving key points and the coordinates of the driving anchor points.

[0019] The driving key points and the driving anchor points are triangulated to obtain multiple texture triangles of the migration image;

[0020] The transfer image is driven by the vertex coordinate sequence corresponding to the driving speech acquired in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image.

[0021] Thirdly, the present invention provides a computer-readable storage medium, comprising: storing a computer program on the computer-readable storage medium, wherein the computer program, when executed by a controller, implements the animation generation method as shown in the second aspect.

[0022] Fourthly, the present invention provides a computer program product, comprising: when the computer program product is run on a computer, causing the computer to implement the animation generation method as shown in the second aspect.

[0023] As can be seen from the above technical solutions, the animation generation device and method provided in this application, through a controller, performs style transfer on a target face image to obtain a transferred image corresponding to the target face image; performs key point detection on the transferred image to obtain facial key points of the transferred image; determines driving key points and driving anchor points of the transferred image based on the facial key points and the corner points of the transferred image; obtains texture coordinates of the transferred image based on the coordinates of the driving key points and the coordinates of the driving anchor points; performs triangulation on the driving key points and the driving anchor points to obtain multiple texture triangles of the transferred image; and drives the transferred image based on the vertex coordinate sequence corresponding to the driving speech obtained in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image. Compared with the prior art, the generated style transfer animation has poor effect and low viewing quality. This application achieves better stretching and driving of the transferred image by triangulating the style transfer image to obtain multiple texture triangles, thereby improving the effect of the generated animation and enhancing the user experience. Attached Figure Description

[0024] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0025] Figure 1 A scene architecture diagram of the animation generation method provided in an embodiment of the present invention is shown;

[0026] Figure 2 A hardware configuration block diagram of the control device provided in an embodiment of the present invention is shown;

[0027] Figure 3 A hardware configuration block diagram of the animation generation device provided in an embodiment of the present invention is shown;

[0028] Figure 4 A schematic diagram of the network architecture of the animation generation device provided in an embodiment of the present invention is shown;

[0029] Figure 5 A flowchart illustrating the steps of the animation generation method provided in an embodiment of the present invention is shown.

[0030] Figure 6 A schematic diagram of an animation generation method provided by an embodiment of the present invention is shown;

[0031] Figure 7 A schematic diagram of another animation generation method provided by an embodiment of the present invention is shown;

[0032] Figure 8 A flowchart illustrating the steps of another animation generation method provided by an embodiment of the present invention is shown;

[0033] Figure 9 A schematic diagram of the animation generation method provided in an embodiment of the present invention is shown;

[0034] Figure 10 A flowchart illustrating the steps of another animation generation method provided by an embodiment of the present invention is shown;

[0035] Figure 11 An architecture diagram of the animation generation method provided in an embodiment of the present invention is shown;

[0036] Figure 12 An architectural diagram of another animation generation method provided by an embodiment of the present invention is shown;

[0037] Figure 13 This diagram illustrates the overall framework of another animation generation method provided by an embodiment of the present invention. Detailed Implementation

[0038] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0039] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0040] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0041] Figure 1 This is a schematic diagram of the scene architecture for the animation generation method provided in an embodiment of this application. Figure 1 As shown in the embodiment of this application, the scene architecture includes: a server 400 and an animation generation device 300.

[0042] The animation generation device 300 provided in this application embodiment can have various implementation forms, such as smart speakers, televisions, refrigerators, washing machines, air conditioners, smart curtains, routers, set-top boxes, mobile phones, personal computers (PCs), smart TVs, laser projection equipment, monitors, electronic bulletin boards, wearable devices, in-vehicle devices, electronic tables, etc.

[0043] In some embodiments, a user can operate the animation generation device 300 through a smart device 200 or a control device 100. When the animation generation device 300 receives an audio signal, it can communicate with the server 400. The animation generation device 300 can communicate with the server 400 through a local area network (LAN) or a wireless local area network (WLAN).

[0044] Server 400 can be a server that provides various services, such as a server that supports audio data collected by animation generation device 300. The server can analyze and process the received audio and other data, and feed back the processing results (such as endpoint information) to the terminal device. Server 400 can be a server cluster or multiple server clusters, and can include one or more types of servers.

[0045] The animation generation device 300 can be either hardware or software. When the animation generation device 300 is hardware, it can be various electronic devices with sound acquisition capabilities, including but not limited to smart speakers, smartphones, televisions, tablets, e-book readers, smartwatches, media players, computers, AI devices, robots, smart vehicles, etc. When the animation generation device 300 is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules (e.g., used to provide sound acquisition services) or as a single software program or software module. No specific limitations are made here.

[0046] It should be noted that the animation generation method provided in this application embodiment can be executed by server 400, animation generation device 300, or both server 400 and animation generation device 300. This application does not limit the execution method in this way.

[0047] Figure 2 A hardware configuration block diagram of an animation generation device 300 according to an exemplary embodiment is shown. Figure 2The animation generation device 300 shown includes at least one of the following: a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, an audio processor, RAM, ROM, and a first to an nth interface for input / output.

[0048] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The animation generation device 300 can establish the transmission and reception of control signals and data signals through the communicator 220 and the server 400.

[0049] User interface 280 can be used to receive external control signals.

[0050] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0051] The sound acquisition device can be a microphone, also known as a "microphone" or "voice transducer," which can be used to receive the user's voice and convert the sound signal into an electrical signal. The animation generation device 300 can be equipped with at least one microphone. In some embodiments, the animation generation device 300 can be equipped with two microphones, which, in addition to acquiring sound signals, can also perform noise reduction. In other embodiments, the animation generation device 300 can also be equipped with three, four, or more microphones, enabling sound signal acquisition, noise reduction, sound source identification, and directional recording functions, etc.

[0052] Furthermore, the microphone can be built into the animation generation device 300, or it can be connected to the animation generation device 300 via wired or wireless means. Of course, this embodiment does not limit the location of the microphone on the animation generation device 300. Alternatively, the animation generation device 300 may not include a microphone, meaning the microphone is not located within the animation generation device 300. The animation generation device 300 can connect an external microphone (also called a microphone) via an interface (such as USB interface 130). This external microphone can be fixed to the animation generation device 300 using an external fastener (such as a camera bracket with a clip).

[0053] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the animation generation device 300.

[0054] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.

[0055] In some embodiments, the operating system of the smart device is, for example, the Android system. Figure 3 As shown, the animation generation device 300 can be logically divided into an application layer (referred to as "application layer") 31, a kernel layer 32, and a hardware layer 33.

[0056] Among them, such as Figure 3 As shown, the hardware layer may include Figure 2 The controller 250, communicator 220, detector 230, etc., are shown. Application layer 31 includes one or more applications. These applications can be system applications or third-party applications. For example, application layer 31 includes a speech recognition application, which can provide an animation generation interface and services for connecting the animation generation device 300 to the server 400.

[0057] The kernel layer 32 serves as a software middleware between the hardware layer and the application layer 31, and is used to manage and control hardware and software resources.

[0058] In some embodiments, kernel layer 32 includes a detector driver, which is used to send the speech data collected by detector 230 to a speech recognition application. For example, when the speech recognition application in animation generation device 300 is started and a communication connection is established between animation generation device 300 and server 400, the detector driver is used to send the user-input speech data collected by detector 230 to the speech recognition application. Subsequently, the speech recognition application sends query information containing this speech data to intent recognition module 102 in the server. Intent recognition module 102 is used to input the speech data sent by animation generation device 300 into the intent recognition model.

[0059] To clearly illustrate the embodiments of this application, the following description is provided in conjunction with... Figure 4 This application describes a speech recognition network architecture provided in its embodiments.

[0060] See Figure 4 As shown, Figure 4 This is a schematic diagram of an animation generation network architecture provided in an embodiment of this application. Figure 4 In this embodiment, the animation generation device receives input information and outputs the processing result of that information. The speech recognition module deploys a speech recognition service to convert audio into text; the semantic understanding module deploys a semantic understanding service to perform semantic parsing on the text; the business management module deploys a business instruction management service to provide business instructions; the language generation module deploys a language generation service (NLG) to convert instructions instructing the animation generation device to execute into text language; and the speech synthesis module deploys a speech synthesis (TTS) service to process the text language corresponding to the instructions and send it to a speaker for playback. In one embodiment, Figure 4 The architecture shown can contain multiple entity service devices with different business services deployed, or one or more entity service devices can combine one or more functional services.

[0061] In some embodiments, the following describes the basis Figure 4 The process of processing information from the input animation generation device in the architecture shown is described with an example, taking voice commands input via voice as an example:

[0062] [Speech Recognition]

[0063] The animation generation device can perform noise reduction and feature extraction on the audio of the voice command after receiving a voice input. The noise reduction process may include steps such as removing echoes and environmental noise.

[0064] [Semantic understanding]

[0065] Using acoustic and language models, natural language understanding is performed on the identified candidate text and associated contextual information. The text is parsed into structured, machine-readable information, including business domain, intent, slots, and other semantic information. An executable intent confidence score is obtained, and the semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.

[0066] [Business Management]

[0067] Based on the semantic parsing results of the voice command text, the semantic understanding module issues execution instructions to the corresponding business management module to execute the operation corresponding to the voice command, complete the user's request for this operation, and provide feedback on the execution result of the operation corresponding to the voice command.

[0068] In some embodiments, the animation generation device 300 performs style transfer on a target face image via a controller 250 to obtain a transferred image corresponding to the target face image; performs keypoint detection on the transferred image to obtain facial keypoints; determines driving keypoints and driving anchor points of the transferred image based on the facial keypoints and corner points of the transferred image; obtains texture coordinates of the transferred image based on the coordinates of the driving keypoints and the coordinates of the driving anchor points; performs triangulation on the driving keypoints and the driving anchor points to obtain multiple texture triangles of the transferred image; and drives the transferred image based on the vertex coordinate sequence corresponding to the driving speech acquired in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image.

[0069] In some embodiments, the animation generation device 300 performs style transfer on a target face image via a controller 250. The method for obtaining the transferred image corresponding to the target face image can be as follows: Based on the target face image and a style transfer model, a transferred image corresponding to the target face image is generated. The style transfer model is a generator in a Generative Adversarial Network (GAN), comprising an encoder, an auxiliary classifier, an attention module, and a decoder. The encoder is used to downsample the input image of the GAN to obtain a first downsampled image, process the first downsampled image through convolutional residual blocks to obtain residual features, and perform instance normalization (IN) on the residual features to obtain encoded features. The auxiliary classifier is used to obtain the weight coefficients of each channel of the encoded features. The attention module is used to obtain attention features based on the weight coefficients of each channel of the encoded features and the encoded features. The decoder is used to perform IN or layer normalization (LN) on the attention features to obtain normalized features and upsample the normalized features to obtain the output image of the generator.

[0070] In some embodiments, the animation generation device 300 is further configured via the controller 250 to, before generating a transfer image corresponding to the target face image based on the target face image and the style transfer model, acquire an original image set, the original image set including multiple sample face images and multiple sample transfer images; acquire facial key points of each image in the original image set; perform rotation correction on each image in the original image set based on the facial key points of each image to acquire a corrected image set; scale the facial bounding boxes corresponding to each image in the corrected image set to a preset size, and extract the image content within the scaled facial bounding boxes to acquire a cropped image set; the facial bounding box corresponding to any image is the positive circumscribed direction of the facial key points of that image; set the background of each image in the cropped image set to a preset color to acquire a sample image set; and train the GAN based on the sample image set to acquire the style transfer model.

[0071] In some embodiments, the animation generation device 300 performs key point detection on the migrated image through the controller 250. The method for obtaining the facial key points of the migrated image can be: performing key point detection on the migrated image based on a facial key point detection model to obtain the facial key points of the migrated image; wherein, the facial key point detection model is a model obtained by training a preset machine learning model based on sample data, and the sample data includes: multiple sample face images and facial key point information corresponding to each sample face image.

[0072] In some embodiments, the facial key points of the migration image in the animation generation device 300 include 68 key points. The controller is specifically configured to: select 20 driving key points from the mouth key point, left eye key point, and right eye key point among the facial key points; set 8 driving anchor points based on the mouth key point, chin key point, nose key point, left eye key point, and right eye key point among the facial key points; and set 4 driving anchor points based on the corner points of the migration image.

[0073] In some embodiments, the controller 250 of the animation generation device 300 is specifically configured to: normalize the coordinates of the driving key points and the coordinates of the driving anchor points to obtain the texture coordinates of the migration image; the controller is further configured to: process each coordinate value in the vertex coordinate sequence into a value within [-1, 1].

[0074] In some embodiments, the controller 250 of the animation generation device 300 is specifically configured to acquire a first offset sequence corresponding to the driving voice in real time; convert the first offset sequence into a second offset sequence corresponding to the migration image according to a preset scaling parameter and a preset offset; update the coordinates of the driving key points and the coordinates of the driving anchor points according to the second offset sequence, and acquire the vertex coordinate sequence corresponding to the driving voice.

[0075] In some embodiments, the controller 250 of the animation generation device 300 is specifically configured to convert the first vertex offset into the second vertex offset according to the preset scaling parameters, the preset offset, and the following formula:

[0076]

[0077] in, This is the i-th offset in the second offset sequence. Let be the i-th offset in the first offset sequence, scale be the preset scaling parameter, and sh be the preset offset.

[0078] In some embodiments, the controller 250 of the animation generation device 300 is further configured to: determine a target frequency based on the frame rate of the style transfer animation; and acquire the vertex coordinate sequence based on the target frequency.

[0079] Figure 5 An exemplary flowchart of the animation generation method provided in an embodiment of this application is shown, such as... Figure 5 As shown, the animation generation method provided in this application includes the following steps:

[0080] S501. Perform style transfer on the target face image to obtain the transferred image corresponding to the target face image.

[0081] In some embodiments, the style transfer of the target face image is a conversion between images in two different domains. Specifically, it involves providing a style image, converting any image into this style, and preserving the content of the original image as much as possible.

[0082] S502. Perform key point detection on the migrated image to obtain the facial key points of the migrated image.

[0083] In some embodiments, the key point detection is actually processing the relationship description of the image. The facial key points can represent point information, position information, and association information. Facial key points can be obtained by detection, and the pose of the face can be calculated based on them. Each key point in the face can represent the type features of the face. For example, the key points of the eyes can represent the shape of the eyes and the position information of the eyes on the face.

[0084] S503. Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image.

[0085] In some embodiments, the corner points of the migration image are located at the four corners of the vertex of the migration image, used to fix the overall stretching effect of the migration image, as shown in the reference. Figure 6 As shown, there are four corner points A, B, C, and D at the four corners of the vertex of the migration image 600. The positions of points A, B, C, and D are the corner points of the migration image.

[0086] S504. Obtain the texture coordinates of the migration image based on the coordinates of the driving key points and the coordinates of the driving anchor points.

[0087] In some embodiments, the coordinates of the driving key point and the coordinates of the driving anchor point can be represented as Ref[a1,b1,c1], Ref[a2,b2,c2], etc., where a, b, and c represent position parameters of a coordinate.

[0088] S505. Triangulate the driving key points and the driving anchor points to obtain multiple texture triangles of the migration image.

[0089] In some embodiments, the triangulation of the driving key points and the driving anchor points involves using the Delaunay triangulation algorithm to generate a set of triangles from a given set of planar points. For example, in finite element simulation, ray tracing rendering, and other calculations, it is necessary to convert the geometric model into triangular mesh data, i.e., "triangular mesh generation".

[0090] For example, refer to Figure 7As shown, this is a schematic diagram of the formation process of multiple texture triangles in the migrated image. The coordinate parameters of the driving key points and the driving anchor points are input using the subdiv.gettrianglelist method. Based on the Delauney partitioning empty circle characteristics and the rule of maximizing the minimum angle, the points are connected into triangles and the triangle index is output. The Delauney triangulation is unique and closest to the regularized triangulation network no matter where it starts to be constructed. This can optimize the subsequent end-side stretching effect and improve stretching efficiency.

[0091] S506. Drive the transfer image according to the vertex coordinate sequence corresponding to the driving speech acquired in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image.

[0092] In some embodiments, the method of obtaining the driving voice can be user voice recognized by intelligent technologies such as speech semantic recognition, or it can be directly obtained from a cloud voice library based on the current scene content.

[0093] In some embodiments, the method for acquiring the vertex coordinate sequence corresponding to the driving speech in real time includes the following steps A to D:

[0094] Step A: Obtain the first offset sequence corresponding to the driving voice in real time.

[0095] In some embodiments, the first offset sequence corresponding to the driving speech can be obtained by transmitting speech-driven keypoint offsets from the cloud. For example, when there are 32 keypoints, the first offset sequence can be represented as: Delta[A1, B1, C1] N Delta[A2, B2, C2] N ...Delta[A 32 B 32 C 32 ] N .

[0096] Step B: Convert the first offset sequence into the second offset sequence corresponding to the migrated image according to the preset scaling parameters and preset offset.

[0097] In some embodiments, the preset scaling parameter is a floating-point number, and the preset offset can be represented as [x_,y_,z_].

[0098] Step C: Convert the first vertex offset to the second vertex offset according to the preset scaling parameters, the preset offset, and the following formula:

[0099]

[0100] in, This is the i-th offset in the second offset sequence. Let be the i-th offset in the first offset sequence, scale be the preset scaling parameter, and sh be the preset offset.

[0101] In some embodiments, the preset scaling parameters and the preset offset can be manually adjusted multiple times according to different cartoon characters.

[0102] Step D: Update the coordinates of the driving key point and the driving anchor point according to the second offset sequence to obtain the vertex coordinate sequence corresponding to the driving speech.

[0103] As can be seen from the above technical solutions, the animation generation device and method provided in this application, through a controller, performs style transfer on a target face image to obtain a transferred image corresponding to the target face image; performs keypoint detection on the transferred image to obtain facial keypoints of the transferred image; determines driving keypoints and driving anchor points of the transferred image based on the facial keypoints and the corner points of the transferred image; obtains texture coordinates of the transferred image based on the coordinates of the driving keypoints and the coordinates of the driving anchor points; performs triangulation on the driving keypoints and the driving anchor points to obtain multiple texture triangles of the transferred image; and drives the transferred image based on the vertex coordinate sequence corresponding to the driving speech obtained in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image. Compared with the prior art, the generated style transfer animation has poor effect and low viewing quality. This application, by performing triangulation on the style transfer image, makes the obtained transferred image have a better stretching effect, thereby making the generated animation effect better and improving the user experience.

[0104] As an extension and refinement of the above embodiments, Figure 8 An exemplary flowchart of the animation generation method provided in an embodiment of this application is shown, such as... Figure 8 As shown, the animation generation method provided in this application includes the following steps:

[0105] S801. Perform style transfer on the target face image to obtain the transferred image corresponding to the target face image.

[0106] S802. Perform key point detection on the migrated image based on the facial key point detection model to obtain the facial key points of the migrated image.

[0107] The facial landmark detection model is a model obtained by training a preset machine learning model based on sample data. The sample data includes: multiple sample facial images and facial landmark information corresponding to each sample facial image.

[0108] The facial landmarks in the migrated image described in S802 above include 68 landmarks, as shown in the reference. Figure 9 The diagram shown illustrates 68 facial landmarks in a migrated image. The method for obtaining the facial landmarks in the migrated image includes the following steps one and two:

[0109] Step 1: Select 20 driving key points from the mouth key point, left eye key point, and right eye key point of the face key points.

[0110] In some embodiments, refer to Figure 9 The image shown is a schematic diagram of the location of 68 facial key points in the transferred image. Eight key points are selected from the mouth key points of the 68 facial key points, and six key points are selected from the left eye key points and six key points from the right eye key points. These 20 key points are used as driving key points.

[0111] Step 2: Set 8 driving anchor points based on the mouth key point, chin key point, nose key point, left eye key point and right eye key point of the face key points, and set 4 driving anchor points based on the corner points of the migration image.

[0112] In some embodiments, eight key points are selected from facial key points, including the mouth key point, chin key point, nose key point, left eye key point, and right eye key point, as the driving anchor points to control the stretching of the surrounding area caused when the key points of the driving part are eliminated. For example, refer to the above. Figure 7 As shown, the positions of 32 key points, including driving keypoints and driving anchor points, are referenced. Figure 7 shown in (1).

[0113] It should be noted that the selection of key points and the setting of anchor points should follow the following two principles to ensure the stretching effect:

[0114] Principle 1: Fine-tune the key points of the eyes and mouth to ensure the symmetry of the overall coordinates of the key points and their enclosure of the area.

[0115] Principle 2: Different images require different anchor point positions, and the result of the triangulation algorithm should be the standard.

[0116] S803. Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image.

[0117] S804. Obtain the texture coordinates of the migration image based on the coordinates of the driving key points and the coordinates of the driving anchor points.

[0118] In the above S804, the method for obtaining the texture coordinates of the migrated image based on the coordinates of the driving keypoint and the coordinates of the driving anchor point further includes:

[0119] The coordinates of the driving key points and the driving anchor points are normalized to obtain the texture coordinates of the migrated image.

[0120] In some embodiments, the texture coordinate value range is 0 to 1, so the coordinates of the driving key points and the coordinates of the driving anchor points need to be normalized to the [0, 1] space. For example, the scaling method can be to divide the coordinates of the driving key points and the coordinates of the driving anchor points by the pixel width of the stretched image (square).

[0121] S805. Triangulate the driving key points and the driving anchor points to obtain multiple texture triangles of the migration image.

[0122] S806. Drive the transfer image according to the vertex coordinate sequence corresponding to the driving voice acquired in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image.

[0123] In the above S806, it is also included that: each coordinate value in the vertex coordinate sequence is processed into a value within [-1, 1].

[0124] As an extension and refinement of the above embodiments, Figure 10 An exemplary flowchart of the animation generation method provided in an embodiment of this application is shown, such as... Figure 10 As shown, the animation generation method provided in this application includes the following steps:

[0125] S1001. Generate a transfer image corresponding to the target face image based on the target face image and the style transfer model.

[0126] The style transfer model is a generator in a Generative Adversarial Network (GAN), which includes an encoder, an auxiliary classifier, an attention module, and a decoder. In some embodiments, refer to... Figure 11 As shown, the generator 1100 in the Generative Adversarial Network (GAN) includes:

[0127] Encoder 1101 is used to downsample the input image of the GAN to obtain a first downsampled image, process the first downsampled image through a convolutional residual block to obtain residual features, and perform instance normalization (IN) on the residual features to obtain encoded features.

[0128] Auxiliary classifier 1102 is used to obtain the weight coefficients of each channel of the encoded feature.

[0129] Attention module 1103, weight coefficients for each channel of the encoded feature, and attention features obtained from the encoded feature.

[0130] Decoder 1104 is used to perform IN or LN layer normalization on the attention features to obtain normalized features and to upsample the normalized features to obtain the output image of the generator.

[0131] The style transfer model also includes a discriminator in a Generative Adversarial Network (GAN), referring to... Figure 12 As shown, the discriminator 1200 in the Generative Adversarial Network (GAN) includes: encoder 1201, auxiliary classifier 1202, attention module 1203, and classifier 1204. The discriminator has a basically the same structure as the generator, except that the generator and discriminator use different loss functions when training the classifier.

[0132] Before generating the transferred image corresponding to the target face image based on the target face image and the style transfer model, it is also necessary to obtain the style transfer model. The method for obtaining the style transfer model includes the following steps 1 to 6:

[0133] Step 1: Obtain the original image set, which includes multiple sample face images and multiple sample migration images.

[0134] Step 2: Obtain the facial key points of each image in the original image set.

[0135] Step 3: Perform rotation correction on each image in the original image set based on the facial key points of each image to obtain a corrected image set.

[0136] Step 4: Scale the face bounding boxes corresponding to each image in the corrected image set to a preset size, and extract the image content within the scaled face bounding boxes to obtain a cropped image set; the face bounding box corresponding to any image is the positive circumscribed direction of the face key points of that image.

[0137] Step 5: Set the background of each image in the cropped image set to a preset color to obtain a sample image set.

[0138] Step 6: Train the GAN based on the sample image set to obtain the style transfer model.

[0139] In some embodiments, when obtaining a style transfer model, it is also necessary to set a loss function to calculate the difference between the forward computation result of the neural network in each iteration and the true value, so as to guide the next training step in the right direction. In this embodiment, there are a total of four loss functions, namely Adversarial loss, Cycle loss, Identity loss and CAM loss.

[0140] Specifically, a function Gs->t is trained that maps the source neighborhood Xs to the target neighborhood Xt. Training is performed using mismatched samples taken from both neighborhoods. Our framework consists of Generators Gs->t and Gt->s, and two Discriminators Ds and Dt. The attention modules within the Discriminators help the Generator focus on the regions most important for generating realistic images; the Generator's attention focuses on the key parts that distinguish the source neighborhood from the other neighborhoods.

[0141] Let x∈{Xt,Gs->t(Xs)} represent samples from the target domain or generated samples, D t It is a function that includes an encoder EDt, a classifier CDt, and an auxiliary classifier ηDt. ηDt(x) and Dt(x) are used to identify whether x comes from Xt or Gs->t(Xs).

[0142] Adversarial loss: Calculated using mean squared error (MSE) to match the distribution of the translated image with the distribution of the target image. The expression for this loss function is as follows:

[0143]

[0144] Cycle loss: Cycle loss is designed to mitigate the mode collapse problem by applying a periodic consistency constraint to the generator. Given an image x∈Xs, after sequential transformations of x from Xs to Xt and from Xt to Xs, the image should be successfully transformed back to the original domain. The loss function is expressed as follows:

[0145]

[0146] Identity loss: The L1 regression loss function is used to ensure that the color distributions of the input and output images are similar, applying an identity consistency constraint to the generator. Given an image x∈Xt, the image should not change after transforming X using Gs→t. The expression for this loss function is as follows:

[0147]

[0148] CAM loss: The generator uses Binary Cross Entropy Loss (BCEloss) for calculation, while the discriminator uses MSE loss for calculation. The maximum difference between the target domain and the source domain is calculated using information from the auxiliary classifier.

[0149] The loss function expression in Gnerator is as follows:

[0150]

[0151] The loss function expression in Discriminator is as follows:

[0152]

[0153] Where ηs(x) represents the probability that x comes from Xs, ηD t (x) indicates that x comes from D. t The probability of.

[0154] S1002. Perform key point detection on the migrated image based on the facial key point detection model to obtain the facial key points of the migrated image.

[0155] The facial landmark detection model is a model obtained by training a preset machine learning model based on sample data. The sample data includes: multiple sample facial images and facial landmark information corresponding to each sample facial image.

[0156] S1003. Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image.

[0157] S1004. Obtain the texture coordinates of the migration image based on the coordinates of the driving key points and the coordinates of the driving anchor points.

[0158] S1005. Triangulate the driving key points and the driving anchor points to obtain multiple texture triangles of the migration image.

[0159] S1006. Drive the transfer image according to the vertex coordinate sequence corresponding to the driving speech acquired in real time, the texture coordinates, and the multiple texture triangles to generate a style transfer animation corresponding to the target face image.

[0160] As an extension and refinement of the above embodiments, refer to Figure 13 The diagram shown illustrates the overall framework of the animation generation method, which includes:

[0161] The overlay key point module 1301 is used to obtain the vertex coordinate sequence.

[0162] Processing module 1302 is used to obtain the vertex coordinate sequence corresponding to the driving voice, the texture coordinates, and the multiple texture triangles, drive the migration animation, and control the audio-visual alignment.

[0163] The processing module includes an OpenGL ES renderer, a Vertex Array Object (VAO) area, a Vertex Buffer Object (VBO) area, a Verter processor, and a Shader. In some embodiments, the audio-visual alignment is controlled by the speed of the input vertex coordinates (rendering speed). The generated vertex coordinate sequence has a frame rate of 80 FPS (frames per second), so the rendering speed based on the Open Graphics Library (OpenGL) is 12.5 ms / frame. Generally, the vertex coordinate sequence can be selected every other frame, for example, selecting one frame every four frames, resulting in 20 FPS and a rendering speed of 50 ms / frame.

[0164] The driving module 1303 is used to acquire driving audio and control the superimposed key point module 1301 to generate vertex coordinate sequences in real time.

[0165] Playback speed control module 1304 is used to control the playback speed of animation.

[0166] The audio and video synchronization playback module 1305 is used for synchronous playback of audio and video.

[0167] Audio track module 1306 is used to acquire the audio and video of the animation to be played.

[0168] In some embodiments, this application provides a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the animation generation method described in any of the above embodiments.

[0169] In some embodiments, this application provides a computer program product that, when run on a computer, enables the computer to implement the animation generation method described in the second aspect or any embodiment of the second aspect.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0171] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. An animation generation device, characterized in that, include: The controller is configured as follows: Perform style transfer on the target face image to obtain the transferred image corresponding to the target face image; Key point detection is performed on the migrated image to obtain the facial key points of the migrated image; Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image; The texture coordinates of the migrated image are obtained based on the coordinates of the driving key points and the coordinates of the driving anchor points. The driving key points and the driving anchor points are triangulated to obtain multiple texture triangles of the migration image; The transfer image is driven by the vertex coordinate sequence corresponding to the driving speech, the texture coordinates, and the multiple texture triangles acquired in real time, so as to generate a style transfer animation corresponding to the target face image. The controller is specifically configured as follows: The first offset sequence corresponding to the driving voice is acquired in real time; The first offset sequence is converted into the second offset sequence corresponding to the migrated image according to preset scaling parameters and preset offset. Update the coordinates of the driving keypoint and the driving anchor point according to the second offset sequence to obtain the vertex coordinate sequence corresponding to the driving voice.

2. The animation generation device according to claim 1, characterized in that, The controller is specifically configured as follows: Based on the target face image and the style transfer model, a transfer image corresponding to the target face image is generated; The style transfer model is a generator in a Generative Adversarial Network (GAN). The generator includes an encoder, an auxiliary classifier, an attention module, and a decoder. The encoder downsamples the input image of the GAN to obtain a first downsampled image, processes the first downsampled image using convolutional residual blocks to obtain residual features, and performs instance normalization (IN) on the residual features to obtain encoded features. The auxiliary classifier obtains the weight coefficients of each channel of the encoded features. The attention module obtains attention features from the weight coefficients of each channel of the encoded features and the encoded features. The decoder performs IN or layer normalization (LN) on the attention features to obtain normalized features and upsamples the normalized features to obtain the output image of the generator.

3. The animation generation device according to claim 2, characterized in that, The controller is also configured to: Before generating the transfer image corresponding to the target face image based on the target face image and the style transfer model, an original image set is obtained, which includes multiple sample face images and multiple sample transfer images. Obtain facial key points from each image in the original image set; Based on the facial key points of each image, each image in the original image set is rotated and corrected to obtain a corrected image set; The face bounding boxes corresponding to each image in the corrected image set are scaled to a preset size, and the image content within the scaled face bounding boxes is extracted to obtain a cropped image set; the face bounding box corresponding to any image is the positive circumscribed direction of the face key points of that image; Set the background of each image in the cropped image set to a preset color to obtain a sample image set; The GAN is trained based on the sample image set to obtain the style transfer model.

4. The animation generation device according to claim 2, characterized in that, The controller is specifically configured as follows: The key points of the migrated image are detected based on the facial key point detection model to obtain the facial key points of the migrated image. The facial landmark detection model is a model obtained by training a preset machine learning model based on sample data. The sample data includes: multiple sample facial images and facial landmark information corresponding to each sample facial image.

5. The animation generation device according to claim 2, characterized in that, The facial landmarks in the migrated image include 68 landmarks, and the controller is specifically configured as follows: Twenty driving key points are selected from the mouth key point, left eye key point, and right eye key point of the facial key points; Eight driving anchor points are set based on the mouth, chin, nose, left eye, and right eye key points of the face key points, and four driving anchor points are set based on the corner points of the migration image.

6. The animation generation device according to claim 1, characterized in that, The controller is specifically configured to: normalize the coordinates of the driving key points and the coordinates of the driving anchor points to obtain the texture coordinates of the migrated image; The controller is further configured to process each coordinate value in the vertex coordinate sequence into a value within the range [-1, 1].

7. The animation generation device according to claim 1, characterized in that, The controller is specifically configured as follows: The first offset sequence is converted into the second offset sequence according to the preset scaling parameters, the preset offset, and the following formula: in, This is the i-th offset in the second offset sequence. Let be the i-th offset in the first offset sequence. The preset scaling parameter, This is the preset offset.

8. The animation generation device according to any one of claims 1-7, characterized in that, The controller is also configured to: Determine the target frequency based on the frame rate of the style transfer animation; The vertex coordinate sequence is obtained based on the target frequency.

9. An animation generation method, characterized in that, include: Perform style transfer on the target face image to obtain the transferred image corresponding to the target face image; Key point detection is performed on the migrated image to obtain the facial key points of the migrated image; Based on the facial key points and the corner points of the migrated image, determine the driving key points and driving anchor points of the migrated image; The texture coordinates of the migrated image are obtained based on the coordinates of the driving key points and the coordinates of the driving anchor points. The driving key points and the driving anchor points are triangulated to obtain multiple texture triangles of the migration image; The transfer image is driven by the vertex coordinate sequence corresponding to the driving speech, the texture coordinates, and the multiple texture triangles acquired in real time, so as to generate a style transfer animation corresponding to the target face image. The method further includes: The first offset sequence corresponding to the driving voice is acquired in real time; The first offset sequence is converted into the second offset sequence corresponding to the migrated image according to preset scaling parameters and preset offset. Update the coordinates of the driving keypoint and the driving anchor point according to the second offset sequence to obtain the vertex coordinate sequence corresponding to the driving voice.

Citation Information

Patent Citations

  • A method and apparatus for generating an animation

    CN109377539A

  • Data processing system, object detection method and device

    CN113591872A