Character animation generation method, device and electronic equipment

By performing two animation diffusions on the target character image, a character animation with both facial and body movements is generated, solving the problems of high cost and poor versatility in existing technologies, and realizing efficient and automated character animation generation.

CN118864677BActive Publication Date: 2025-10-24SHENZHEN TUYI SHIBEI TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411032428.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-10-24
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

Existing character animation generation methods are costly and lack versatility, making it difficult to simultaneously achieve facial and body motion performance, thus failing to meet the practical application needs of most scenarios.

Method used

By acquiring the target character and its facial region from the original image, the facial mask image and the target character image are determined, and two animation diffusions are performed: the first diffusion generates a reference image related to facial movements, and the second diffusion generates a sequence of movements related to body movements. Animation post-processing is then performed to generate a character animation that combines facial and body movement expressions.

Benefits of technology

It enables rapid generation of character animations, reduces production costs, improves versatility, meets the actual needs of various scenarios, and eliminates the need for manual production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864677B_ABST
    Figure CN118864677B_ABST
Patent Text Reader

Abstract

The application provides a role animation generation method and device, electronic equipment, a storage medium and a computer program product, and relates to the technical field of computer vision. The method comprises the following steps: obtaining an original image; determining a face mask image of a target role and a target role image based on the target role and a face region thereof in the original image; taking the face mask image and the target role image as inputs, performing first animation diffusion related to a face action, and obtaining a plurality of reference images; each reference image corresponds to one face action; obtaining an action reference sequence, and driving each reference image to perform second animation diffusion related to a body action by using the action reference sequence, thereby obtaining a plurality of action sequences; each action sequence corresponds to one reference image; and performing animation post-processing on each action sequence, thereby obtaining a role animation of the target role. The application solves the problems of high cost and poor universality of role animation generation in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and in particular, the present application relates to a role animation generation method and device, an electronic device, a storage medium and a computer program product. BACKGROUND

[0002] At present, role animation has been widely applied in the fields of electronic games, Internet advertisements, short video animations, virtual anchors and the like.

[0003] The existing role animation mainly includes two generation methods: the first generation method is based on traditional animation software technology, and the second generation method is based on generative AI (Artificial Intelligence) technology. However, both generation methods have some shortcomings, for example, in order to continuously play to achieve coherent role animation effect, a large amount of time and labor is required, or it is difficult to realize a role with both facial motion performance and body motion performance, resulting in that the actual application demand in most scenes cannot be met.

[0004] From the above, the existing role animation generation still has the problems of high cost and poor universality, which need to be solved. SUMMARY

[0005] The present application provides a role animation generation method, device, electronic device, storage medium and computer program product, which can solve the problem of high cost and poor universality of role animation generation in related technologies. The technical solution is as follows:

[0006] According to an aspect of the present application, a role animation generation method comprises: obtaining an original image; the original image contains at least one target role with a front full body; determining a face mask image of the target role and a target role image based on the target role and its face region in the original image; taking the face mask image and the target role image as input, performing a first animation diffusion related to facial motion to obtain a plurality of reference images; each reference image corresponds to a facial motion; obtaining a motion reference sequence and driving each reference image to perform a second animation diffusion related to body motion using the motion reference sequence to obtain a plurality of motion sequences; each motion sequence corresponds to a reference image; performing animation post-processing on each motion sequence to obtain a role animation of the target role.

[0007] In an example embodiment, the first-time animation diffusion related to facial motion is performed on the face mask image and the target character image as input to obtain a plurality of reference images, including: determining the position of the face region of the target character in the face mask image based on the visible region and the mask region in the face mask image; and performing local redrawing on the target character in the target character image according to the determined position and different facial motions to obtain a plurality of reference images.

[0008] In an example embodiment, the local redrawing on the target character in the target character image according to the determined position and different facial motions to obtain a plurality of reference images includes: obtaining facial motion prompt words; the facial motion prompt words are used to describe different facial motions; inputting the face mask image and the target character image into an image diffusion model, and guiding the image diffusion model to generate different facial motions for the target character using the facial motion prompt words to output a plurality of reference images; the reference image refers to an image of the target character performing the facial motion.

[0009] In an example embodiment, the second-time animation diffusion related to body motion is driven by the action reference sequence to obtain a plurality of action sequences, including: inputting each reference image into a video diffusion model; and using the video diffusion model to infer each reference image from a static image to a dynamic video according to a plurality of reference actions in the action reference sequence to obtain a plurality of action sequences; the action sequence refers to a video of the target character performing the facial motion while performing a plurality of reference actions.

[0010] In an example embodiment, the animation post-processing is performed on each action sequence to obtain a character animation of the target character, including: sampling each action sequence in a set manner to obtain a plurality of frame blocks respectively belonging to different action sequences; and splicing the plurality of frame blocks belonging to different action sequences to obtain the character animation of the target character; the character animation refers to a video of the target character performing a plurality of reference actions while alternately performing different facial motions.

[0011] In an example embodiment, the determining the face mask image and the target character image of the target character based on the target character and the face region of the target character in the original image comprises: determining a position of the face region of the target character in the original image in response to a first circle selection operation triggered for the face region of the target character in the original image; and performing marking processing on the face region and other regions of the target character in the original image based on the determined position, to obtain the face mask image.

[0012] In an example embodiment, the determining the face mask image and the target character image of the target character based on the target character and the face region of the target character in the original image further comprises: determining a position of the target character in the original image in response to a second circle selection operation triggered for the target character in the original image; and performing image segmentation on the target character and other target characters in the original image based on the determined position, to obtain the target character image.

[0013] In an example embodiment, after the original image is obtained, the method further comprises: performing screening processing on the original image; the screening processing comprises at least one of resolution detection, target character detection, and front full-body detection of the target character; and deleting the original image that does not meet a set condition, so that the original image that meets the set condition performs the step of determining the face mask image and the target character image of the target character based on the target character and the face region of the target character in the original image.

[0014] In an example embodiment, the face region comprises a mouth region, and the face action comprises a first mouth action and a second mouth action; the plurality of reference images comprise a first reference image in which the target character performs the first mouth action, and a second reference image in which the target character performs the second mouth action; the plurality of action sequences comprise a first action sequence in which the target character simultaneously performs the first mouth action and a plurality of reference actions, and a second action sequence in which the target character simultaneously performs the second mouth action and the plurality of reference actions; and the character animation refers to a video in which the target character simultaneously performs the plurality of reference actions on the basis of alternately performing the first mouth action and the second mouth action.

[0015] According to an aspect of the present application, a role animation generation apparatus comprises: an original image acquisition module configured to acquire an original image, wherein the original image comprises at least one target role having a front full body; an input image determination module configured to determine a face mask image of the target role and a target role image based on the target role and a face region of the target role in the original image; a first animation diffusion module configured to perform a first animation diffusion related to a face action by taking the face mask image and the target role image as inputs, to obtain a plurality of reference images, wherein each of the reference images corresponds to a face action; a second animation diffusion module configured to acquire a motion reference sequence, and drive each of the reference images to perform a second animation diffusion related to a body action by using the motion reference sequence, to obtain a plurality of motion sequences, wherein each of the motion sequences corresponds to one of the reference images; and an animation post-processing module configured to perform animation post-processing on each of the motion sequences, to obtain a role animation of the target role.

[0016] According to an aspect of the present application, an electronic device comprises at least one processor and at least one memory, wherein the memory stores a computer program, and the computer program is executed by the processor to implement the role animation generation method as described above.

[0017] According to an aspect of the present application, a storage medium stores a computer program, and the computer program is executed by one or more processors to implement the role animation generation method as described above.

[0018] According to an aspect of the present application, a computer program product comprises a computer program, and the computer program is executed by one or more processors to implement the role animation generation method as described above.

[0019] The technical scheme provided by the present application has the beneficial effects that:

[0020] In the technical solution, the original image containing at least one target character with a front full body is acquired to determine a face mask image and a target character image of the target character based on the target character and the face region thereof in the original image, and then the face mask image and the target character image are used for first animation diffusion related to face actions, and then the plurality of reference images corresponding to different face actions obtained through the first animation diffusion are used in combination with the motion reference sequence for second animation diffusion related to body actions, and finally the character animation of the target character is obtained through animation post-processing on the plurality of motion sequences corresponding to different reference images obtained through the second animation diffusion. In the above process, through the two times of animation diffusion, the character animation can be quickly generated, the dependence on manual production is avoided, and a large amount of time and labor is not consumed, the target character can have both face action performance and body action performance, has richer character animation effect, and is beneficial to meet the actual application requirements in most scenes, thereby effectively solving the problems of high cost and poor universality of character animation generation in the related art. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0022] Figure 1 is a schematic diagram according to the implementation environment involved in the present application;

[0023] Figure 2 is a hardware structure diagram of a computer device according to an exemplary embodiment;

[0024] Figure 3 is a flowchart of a character animation generation method according to an exemplary embodiment;

[0025] Figure 4a is a schematic diagram of a target character image according to an exemplary embodiment;

[0026] Figure 4b is a schematic diagram of a face mask image according to an exemplary embodiment;

[0027] Figure 5a is a schematic diagram of a first reference image according to an exemplary embodiment;

[0028] Figure 5b is a schematic diagram of a second reference image according to an exemplary embodiment;

[0029] Figure 6a is a schematic diagram of an action reference sequence according to an example embodiment;

[0030] Figure 6b is a schematic diagram of a first action sequence according to an example embodiment;

[0031] Figure 6c is a schematic diagram of a second action sequence according to an example embodiment;

[0032] Figure 7 is a schematic diagram of a character animation according to an example embodiment;

[0033] Figure 8 is a schematic diagram of implementing frame-by-frame splicing of a first action sequence and a second action sequence according to an example embodiment;

[0034] Figure 9 is Figure 3 is a flowchart of step 330 in a corresponding embodiment in one embodiment;

[0035] Figure 10 is Figure 3 is a flowchart of step 350 in a corresponding embodiment in one embodiment;

[0036] Figure 11 is a schematic diagram of implementing local redrawing of a target character by calling an image diffusion model according to an example embodiment;

[0037] Figure 12 is a schematic diagram of implementing inference from a static image to a dynamic video by calling a video diffusion model according to an example embodiment;

[0038] Figure 13 is a schematic diagram of a specific implementation of a character animation generation method in an application scenario;

[0039] Figure 14 is a structural block diagram of a character animation generation apparatus according to an example embodiment;

[0040] Figure 15 is a structural block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION

[0041] Embodiments of the present application are described in detail below with reference to the accompanying drawings. The embodiments described below are examples for explaining the present application and should not be construed as limiting the present application.

[0042] It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It is further understood that the terms "comprise" and "comprising" and the like mean that there are the recited features, integers, steps, or components, but do not preclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. It is further understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements can be present. Further, "connected" or "coupled" as used herein can include wirelessly connected or wirelessly coupled. As used herein, the term "and / or" includes all combinations of one or more of the associated listed items.

[0043] As mentioned above, the existing role animation mainly includes two generation methods: the first generation method is based on traditional animation software technology, and the second generation method is based on generative AI technology.

[0044] Specifically, the generation method based on traditional animation software technology is usually to use traditional animation software to draw each frame of the whole character of the role separately, and then realize the coherent role animation effect (such as facial action performance and body action performance) by continuous playing. However, this method relies too much on animators and requires a lot of time and labor, resulting in very high production cost of role animation.

[0045] In order to reduce the production cost of role action, the generation method based on generative AI technology is introduced later, which mainly uses diffusion model to automatically generate coherent role animation effect. Although this method does not rely on animators compared with the generation method based on traditional animation software technology, it can save the production cost of role animation, but the generated role animation effect is not rich enough. It can only generate the body action performance of the role, but cannot generate the facial action performance of the role, or it can only generate the facial action performance of the role, but cannot generate the body action performance of the role, so it cannot meet the actual application requirements of most scenes.

[0046] As can be seen from the above, the related art still has the defects of high cost and poor universality of role animation generation.

[0047] To this end, the present application provides a role animation generation method, which can not only quickly generate role animation, avoid relying on animators, thereby reducing the production cost of role animation, but also enable the target role to have facial motion performance and body motion performance, have richer role animation effect, be beneficial to meet the actual application requirements in most scenes, and effectively improve the generality of role animation generation, accordingly, the role animation generation method is applicable to a role animation generation device, the role animation generation device can be deployed on an electronic device, the electronic device can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a notebook computer, a server, etc.; the electronic device can also be a portable mobile electronic device, for example, the electronic device includes a smart phone, a tablet computer, etc.

[0048] In order to make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0049] Figure 1 A schematic diagram of an implementation environment for a role animation generation method. It should be noted that this implementation environment is only an example adapted to the present application and should not be considered as providing any limitation on the use range of the present application.

[0050] The implementation environment includes a user terminal 110 and a server 130.

[0051] Specifically, the user terminal 110 can be an electronic device with image acquisition function, which includes but is not limited to camera equipment such as camera, camera, camcorder, etc., and computer equipment such as desktop computer, notebook computer, server, etc., and can also be a portable mobile electronic device carried by a user, for example, the electronic device includes a smart phone, a tablet computer, etc. Based on this, using the user terminal 110, at least one target role with a front full body can be captured at any time to obtain an original image, without the participation of professional animators, a large amount of role animation materials can be obtained, which is beneficial to save the expensive labor cost of role animation production.

[0052] The server 130 can be a computer device such as a desktop computer, a notebook computer, a server, etc., and can also be a computer cluster composed of multiple servers, or even a cloud computing center composed of multiple servers. The server 130 is used to provide background services, for example, the background services include but are not limited to automatic role animation generation services, etc.

[0053] The service end 130 and the user end 110 are pre-established a network communication connection through wired or wireless and the like, and the data transmission between the service end 130 and the user end 110 is realized through the network communication connection. The transmitted data includes but is not limited to: the original image containing at least one target role with a front full body, the role animation of the target role and the like.

[0054] In an application scenario, through the interaction between the user end 110 and the service end 130, the user end 110 can shoot and collect the original image for at least one target role with a front full body, and then upload the original image to the service end 130 to request the service end 130 to provide the automatic role animation generation service.

[0055] For the service end 130, after receiving the original image uploaded by the user end 110, the automatic role animation generation service is called to generate the corresponding role animation for the target role in the original image. Specifically, based on the target role and the face region thereof in the original image, the face mask image of the target role and the target role image are determined; the face mask image and the target role image are taken as inputs to perform the first animation diffusion related to the face action, to obtain a plurality of reference images corresponding to different face actions; the action reference sequence is obtained, and each reference image is driven by the action reference sequence to perform the second animation diffusion related to the body action, to obtain a plurality of action sequences corresponding to different reference images; the animation post-processing is performed on each action sequence to obtain the role animation of the target role. In this way, the role animation production is highly automated, completely without the participation of professional animators, and the target role can have both face action performance and body action performance, has richer role animation effect, and is beneficial to meet the actual needs in most scenarios, thereby effectively solving the problems of high cost and poor universality of role animation generation in related technologies.

[0056] Taking the user end 110 as an example of a smart phone, after the service end 130 generates the role animation of the target role, the role animation of the target role sent by the service end 130 based on the network connection can be received, and the role animation of the target role can be further displayed to the user through the screen of the smart phone. For the user, the whole process only needs to upload the original image, which is highly automated and easy to use, and is beneficial to improve the user's production experience of the role animation.

[0057] Of course, in other application scenarios, the role animation generation service provided by the service end 130 can also be completed independently by the user end 110, at this time, after the user end 110 obtains the original image, the twice animation diffusion can be continued for the target role in the original image, and finally the role animation of the target role is obtained, which is not a specific limitation.

[0058] Figure 2 is a hardware structure diagram of a computer device according to an exemplary embodiment. The computer device is suitable for Figure 1 the user terminal 110 and the server 130 in the illustrated implementation environment.

[0059] It should be noted that the computer device is only an example suitable for the present application, and should not be considered as providing any limitation on the scope of use of the present application. The computer device should also not be interpreted as requiring reliance on or must have Figure 2 one or more components in the exemplary computer device 200.

[0060] The hardware structure of the computer device 200 can vary greatly due to different configurations or performance, such as Figure 2 As shown, the computer device 200 includes a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.

[0061] Specifically, the power supply 210 is configured to provide operating voltage for each hardware device on the computer device 200.

[0062] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. For example, it is configured to Figure 1 interact between the user terminal 110 and the server 130 in the illustrated implementation environment.

[0063] Of course, in other examples suitable for the present application, the interface 230 can further include at least one serial-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, etc., as Figure 2 shown, which is not specifically limited herein.

[0064] The memory 250, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon include an operating system 251, an application program 253, and data 255, etc., and the storage mode can be temporary storage or permanent storage.

[0065] The operating system 251 is configured to manage and control each hardware device on the computer device 200 and the application program 253, so as to realize the operation and processing of the central processing unit 270 on the mass data 255 in the memory 250, which can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0066] The application 253 is essentially a computer program formed by computer-readable instructions that perform at least one specific task based on the operating system 251, and may include at least one module ( Figure 2 Each module may include corresponding computer-readable instructions. For example, the character animation generating device may be considered as an application 253 deployed on the computer device 200.

[0067] Data 255 can be photos, pictures, etc. stored in a disk, or it can be a facial mask image and target character image of the target character, a first reference image, a second reference image, a first action sequence, a second action sequence, a character animation, an action reference sequence, facial action prompt words, various models, etc., stored in the memory 250.

[0068] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read the computer program stored in the memory 250, thereby performing operations and processing on the massive data 255 in the memory 250. For example, the character animation generation method is completed by the central processing unit 270 reading the application program 253 stored in the memory 250.

[0069] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0070] See also Figure 3 The embodiment of the present application provides a method for generating character animation, which is applicable to an electronic device, which may be Figure 1 The server side in the implementation environment shown can also be Figure 2 The user end in the implementation environment is shown, and the hardware structure of the electronic device can be as follows Figure 2 shown.

[0071] In an exemplary embodiment, the facial region includes a mouth region, and the facial action includes a first mouth action and a second mouth action. The multiple reference images include a first reference image of the target character performing the first mouth action, and a second reference image of the target character performing the second mouth action. The multiple action sequences include a first action sequence in which the target character simultaneously performs the first mouth action and the multiple reference actions, and a second action sequence in which the target character simultaneously performs the second mouth action and the multiple reference actions. Character animation refers to a video in which the target character simultaneously performs the multiple reference actions while alternating between the first mouth action and the second mouth action.

[0072] In the following method embodiment, for the convenience of description, the execution subject of each step of the method is taken as an electronic device for illustration, but this does not constitute a specific limitation.

[0073] As shown in the method Figure 3 may include the following steps:

[0074] Step 310, obtaining an original image.

[0075] The original image contains at least one target character with a front full body.

[0076] It can be understood that the original image containing at least one target character with a front full body is used as a character animation material, and then the target character refers to any object that can be used as a character animation material, for example, the object can be a real person, a virtual character image, a cartoon character image, etc.

[0077] The original image can be obtained by Figure 1 The user terminal in the implementation environment. For example, the user terminal captures the real person in the environment to obtain the original image, or the user terminal captures the cartoon character image appearing on the TV to obtain the original image, or the user downloads the virtual character image he likes from the Internet with the help of the user terminal to obtain the original image. It should be noted that the shooting can be single shooting, and can also be continuous shooting. For single shooting, multiple photos can be obtained, that is, multiple static images (hereinafter also referred to as images). For continuous shooting, a video is obtained, which can also be understood as a dynamic video containing multiple static images. In this embodiment, the original image refers to a static image, that is, the original image can be one of multiple photos, or an arbitrary static image in a video.

[0078] Regarding the acquisition of the original image, it can be derived from the image obtained by the user terminal in real time, or from the image obtained by the user terminal in a historical period of time. Then, for the electronic device, after the user terminal captures or collects the original image, the original image obtained by the user terminal in real time can be timely processed to generate a character animation, or the original image stored in the electronic device in advance can be processed to generate a character animation according to the actual needs of the user. No specific limitation is made here.

[0079] The inventor realizes that for some original images not containing the target role, the target role in the original image cannot be detected at all, or some original images containing the target role do not have a full-body front view, which may result in poor role animation effect due to not being a full-body front view of the target role, and even some unclear original images may result in the face region of the target role being unable to be determined during role animation production. Therefore, these original images do not need to be subjected to subsequent role animation production. For this purpose, in a possible implementation, in order to improve the efficiency of role animation production and improve the role animation effect, before step 330 is performed, the original image is subjected to screening processing to delete the original image that does not meet the set condition. The screening processing includes at least one of resolution detection, target role detection, and full-body front view detection of the target role.

[0080] Specifically, the resolution detection is to ensure that the width and height of the original image meet the condition of a high-quality image, that is, w(I)≥w θ ,h(I)≥h θ , where I represents the original image, w(I), h(I) represents the width and height of the original image, w θ ,h θ represents the width and height of the high-quality image, and can also be regarded as a screening parameter of the resolution detection, which can be flexibly adjusted according to the actual needs of the user, and is not limited herein. At this time, the set condition is that the width and height of the original image exceed the width and height of the high-quality image. Therefore, for the original image whose width and height exceed the width and height of the high-quality image, it is considered that the original image meets the set condition through the resolution detection, and step 330 can be continued to be performed.

[0081] The target role detection is to call a target detection model to perform target detection on the original image to obtain a target detection result of the original image, which is used to indicate whether the original image contains the target role. At this time, the set condition is that the original image contains at least one target role. Therefore, only when the original image contains at least one target role, it is considered that the original image meets the set condition through the target role detection, and step 330 can be continued to be performed. The target detection model is a machine learning model that is trained and has the ability to perform target detection on the original image. For example, the machine learning model can be a YOLOV8 model.

[0082] The front full-body detection of the target role is to call a pose detection model to perform pose detection on the target role in the original image to obtain a pose detection result of the target role, and the pose detection result is used to indicate whether the pose of the target role in the original image is a front full-body. At this time, the condition is that the pose of the target role in the original image is a front full-body. When the pose of the target role in the original image is a front full-body, it can be considered that the original image is an original image that meets the set condition through the front full-body detection of the target role, and then step 330 can be continued. The pose detection model is a machine learning model that is trained and has the ability to perform pose detection on the target role in the original image. For example, the machine learning model can be a GhostNet model.

[0083] In this way, the original image that meets the set condition can continue to determine the face mask image of the target role and the target role image, avoiding subsequent processing of the original image that does not meet the set condition. This not only helps to improve the efficiency of role animation generation, but also helps to improve the role animation effect.

[0084] Step 330, based on the target role and the face region of the target role in the original image, determining the face mask image of the target role and the target role image.

[0085] The target role image refers to an image containing a front full-body of the target role, and the face mask image refers to an image in which the face region of the target role can be significantly distinguished from other regions. For example, the face mask image can be a binary image, and the region with a pixel value of 0 in the binary image represents other regions of the target role, and the region with a pixel value of 255 represents the face region of the target role.

[0086] For example, Figure 4a shows a schematic diagram of the target role image in an embodiment, Figure 4b shows a schematic diagram of the face mask image in an embodiment. In Figure 4a , the target role image only contains one target role, and the pose of the target role is a front full-body, and in Figure 4b , the face mask image contains a visible region and a mask region, wherein the white visible region represents the face region of the target role, and the visible region especially refers to the mouth region of the target role, and the black mask region represents other regions of the target role.

[0087] In a possible implementation, the face mask image of the target role can be determined by the user in a circled manner, and can also be implemented by using a model, which is not specifically limited here and can be selected according to actual needs of the user in different application scenarios. For example, for an application scenario with high precision requirements, the user can determine the face mask image of the target role in a circled manner, and for an application scenario with high efficiency requirements or high automation requirements, the user can determine the face mask image of the target role by calling a model. Taking determination of the mouth region of the target role by calling a model as an example, after obtaining the original image, first, a target detection model is used to perform target detection on the original image to determine a region where the target role is located in the original image, then a face detection model is used to perform face detection on the region where the target role is located in the original image to determine a face region of the target role in the original image, and then a semantic segmentation model is used to segment the mouth region from the face region, and finally, based on the segmented mouth region, the values of each pixel included in the mouth region in the original image are marked as 255, and the values of each pixel included in other regions in the original image are marked as 0, to finally obtain the face mask image of the target role, which can also be understood as the mouth mask image of the target role.

[0088] Similarly, for an original image containing multiple target roles, in a possible implementation, the user can also select a circled manner or a model calling manner to determine the target role image according to actual needs of the application scenario, so that subsequent role animation generation is performed in units of the target role image corresponding to each target role. Taking determination of the target role image by calling a model as an example, after obtaining the original image, first, a target detection model is used to perform target detection on the original image to determine regions where each target role is located in the original image, and then a semantic segmentation model is used to segment each target role region, to obtain the target role image corresponding to each target role.

[0089] Step 350, taking the face mask image and the target role image as inputs, performing first animation diffusion related to the face motion to obtain multiple reference images.

[0090] Each reference image corresponds to a kind of face motion, and it can also be considered that the reference image refers to an image in which the target role performs the face motion.

[0091] In a possible implementation, on the basis that the facial action includes a first mouth action and a second mouth action, the reference image includes a first reference image and a second reference image. The first reference image refers to an image in which the target character performs the first mouth action, and the second reference image refers to an image in which the target character performs the second mouth action. It is explained herein that the first reference image and the second reference image are also static images, and compared with the target character in the original image, the target character in the first reference image performs the first mouth action, and the target character in the second reference image performs the second mouth action. The first mouth action and the second mouth action are considered as facial action expressions, which can be mouth opening, mouth closing, pursing the lips, smiling, and the like. It is worth mentioning that the facial action expression is not limited to the mouth region, but can also be the eye region, and then the facial action expression can also refer to eye opening, eye closing, squinting, crying, and the like. The process of generating character animation is described in detail in each embodiment of the present application by taking the facial action expression implemented in the mouth region as an example, but this does not constitute a specific limitation.

[0092] In this embodiment, the purpose of the first animation diffusion is to enable the target character in the static image to perform a facial action, that is, to enable the target character in the static image to have a facial action expression. In a possible implementation, the first animation diffusion can be implemented by performing local redrawing on the target character in the static image. The local redrawing can be implemented by using an image diffusion algorithm, which includes but is not limited to anything-4.0-inpainting, SD-XL-Inpainting0.1, PowerPaint, and the like.

[0093] For example, Figure 5a The first reference image in which the target character performs the first mouth action is shown in a schematic diagram of one embodiment, Figure 5b The second reference image in which the target character performs the second mouth action is shown in a schematic diagram of one embodiment. As Figure 5a and 5b shown, the first mouth action in the first reference image refers to a mouth closing action, and the second mouth action in the second reference image refers to a mouth opening action, compared with Figure 4a the target character shown in FIG. 6A, it can also be understood that, after the first animation diffusion, the target character changes from not having a facial action expression to being able to perform the first mouth action and the second mouth action, so as to enable the target character to have a facial action expression. Of course, in this embodiment, since the first mouth action is a mouth closing action, Figure 4a the target character shown in FIG. 6B actually also performs the mouth closing action, which is equivalent to that the first reference image and the target character image have no difference, however, the two do not have difference in all cases, if the first mouth action is a pursing the lips action or other mouth action different from the mouth closing action, the target character image to the first reference image will inevitably change.

[0094] At step 370, the action reference sequence is obtained, and the action reference sequence is used to drive each reference image to respectively perform a second animation diffusion related to the body action, to obtain a plurality of action sequences.

[0095] Each action sequence corresponds to a reference image, and can also be considered as a video in which the target character performing the facial action also performs a plurality of reference actions at the same time.

[0096] In a possible implementation, the reference images include a first reference image and a second reference image, the first action sequence is a video in which the target character performing the first mouth action also performs a plurality of reference actions at the same time, and the second action sequence is a video in which the target character performing the second mouth action also performs a plurality of reference actions at the same time. That is, the first action sequence and the second action sequence are essentially a dynamic video, which can be composed of a plurality of static images, and the target character in each static image performs a mouth action and a reference action at the same time. It can also be considered that the target character in the first action sequence and the second action sequence has both facial action performance and body action performance.

[0097] First of all, the action reference sequence is a dynamic video composed of a series of static images containing different reference actions, in other words, the action reference sequence is composed of a plurality of static images in the dynamic video, and each static image corresponds to a reference action.

[0098] In a possible implementation, the action reference sequence can be represented by a skeleton sequence. The skeleton sequence includes a plurality of skeleton images, each of which is used to describe a plurality of key points of an object performing different actions. The object refers to any object that can perform actions, which can be a person, a robot, an animal, etc. Taking a person as an example, the human body key points used to construct the skeleton sequence can include but are not limited to: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left knee, right knee, left ankle, right ankle, etc.

[0099] Taking the skeleton sequence constructed by the human body key points as an example, Figure 6a shows a schematic diagram of the action reference sequence containing different reference actions in an embodiment, in Figure 6a each reference action in the action reference sequence can be a different dance action, and it should be understood that these different dance actions are respectively composed of human body key points in different positions in the static images.

[0100] As for the acquisition of the action reference sequence, it can be acquired from a pre-made and stored action reference sequence in the electronic device. In the process of generating the character animation, the user can select an arbitrary action reference sequence from the pre-stored action reference sequences to participate in the second animation diffusion, so as to meet the personalized demand of generating the character animation. Alternatively, the electronic device can actively push an arbitrary action reference sequence to the user to participate in the second animation diffusion, so as to further improve the automation degree of generating the character animation. The present embodiment is not limited in this regard.

[0101] Secondly, the purpose of the second animation diffusion is to enable the target character in each static image contained in the dynamic video to perform the reference action contained in the action reference sequence, that is, to enable the target character in the dynamic video to have a body action performance. It should be understood that after two times of animation diffusion, the target character in the dynamic video can not only perform facial actions but also perform body actions, that is, the target character in the dynamic video has both facial action performance and body action performance.

[0102] In a possible implementation, the second animation diffusion is implemented by inferring the dynamic video from the static image under the guidance of the action reference sequence. The inference from the static image to the dynamic video can be implemented by using a video diffusion algorithm, which includes but is not limited to Moore-AnimateAnyone, MusePose, etc.

[0103] For example, Figure 6b Fig. 1 shows a schematic diagram of a target character simultaneously performing a first mouth action and a first action sequence of different reference actions in an embodiment, Figure 6c Fig. 2 shows a schematic diagram of a target character simultaneously performing a second mouth action and a second action sequence of different reference actions in an embodiment. Figure 6a under the guidance of the action reference sequence shown in Fig. 1, the target character in the first reference image shown in Fig. 1 performs the first mouth action and the different reference action, Figure 5a based on the first reference image shown in Fig. 1 containing the target character performing the closed-mouth action, as shown in Fig. 1, Figure 6b as shown in Fig. 1, the target character in the first action reference sequence not only performs the closed-mouth action but also performs the different reference action, and under the guidance of the action reference sequence shown in Fig. 2, Figure 5b based on the second reference image shown in Fig. 2 containing the target character performing the open-mouth action, as shown in Fig. 2, Figure 6cAs shown, the target character in the second action reference sequence not only performs the mouth opening action but also performs different reference actions. It is worth mentioning that the guidance of the action reference sequence actually means that the reference actions contained in each static image in the action reference sequence are respectively mapped to the first reference image and the second reference image to form a plurality of first reference images and a plurality of second reference images in which the target character performs different reference actions, and then the plurality of first reference images and the plurality of second reference images form the first action sequence and the second action sequence respectively, so that the reference actions contained in each static image in the action reference sequence have a one-to-one correspondence with the reference actions contained in each static image in the first action reference sequence / second action reference sequence.

[0104] In step 390, each action sequence is post-processed to obtain a character animation of the target character.

[0105] The character animation refers to a video in which the target character simultaneously performs a plurality of reference actions on the basis of alternately performing different facial actions.

[0106] In a possible implementation, on the basis that the facial action includes a first mouth action and a second mouth action, the character animation refers to a video in which the target character simultaneously performs a plurality of reference actions on the basis of alternately performing the first mouth action and the second mouth action.

[0107] Still taking the foregoing example as an example, Figure 7 a schematic diagram of the character animation in one embodiment is shown, in Figure 7 In the foregoing example, the target character not only alternately performs the mouth closing action and the mouth opening action, but also performs different reference actions, so it can be seen that the character animation generated through steps 310 to 390 can make the target character have both facial action performance and body action performance, and the character animation effect is more rich.

[0108] In a possible implementation, the post-processing can include at least one of video super-resolution processing, video background removal processing, and video splicing processing.

[0109] The video super-resolution processing refers to performing super-resolution operations on each action sequence to obtain a plurality of high-definition action sequences, so as to fully guarantee that the character animation has a high-definition effect.

[0110] The purpose of the video background removal processing is to make each action sequence only retain the target character and a transparent background, so as to facilitate subsequent further synthesis of the user based on the character animation, and further help to improve the user's character animation production experience. The video background removal processing can be implemented by calling a matting model, and through the calling of the matting model, the background area other than the target character in each static image contained in the plurality of action sequences is set to be transparent.

[0111] The video splicing process may include the following steps: sampling each action sequence according to a set method to obtain a number of frame blocks belonging to different action sequences; splicing the several frame blocks belonging to different action sequences to obtain the character animation of the target character.

[0112] In a possible implementation, the setting method includes an alternate frame sampling method. For example, Figure 8 As shown, the first action sequence and the second action sequence are first divided into n frame blocks, respectively. Then, according to the alternate frame sampling method, the first action sequence is sampled to obtain the first first frame block, the third first frame block, the fifth first frame block, ..., the nth first frame block, and the second action sequence is sampled to obtain the second second frame block, the fourth second frame block, the sixth second frame block, ..., the n-1th second frame block. Finally, these sampled frame blocks are sequentially spliced ​​together to obtain the character animation of the target character, specifically: the first first frame block, the second second frame block, the third first frame block, the fourth second frame block, the fifth first frame block, the sixth second frame block, ..., the n-1th second frame block, the nth first frame block. Of course, n can be flexibly adjusted according to the actual needs of the user to ensure a smoother and more delicate character animation effect. Although the splicing here is based on the first action sequence and the second action sequence corresponding to two facial actions, in other embodiments, the splicing can also be based on different action sequences corresponding to multiple facial actions. For example, the character animation can be spliced ​​together by the first action sequence corresponding to the closing action, the second action sequence corresponding to the opening action, the third action sequence corresponding to the pouting action, and the fourth action sequence corresponding to the smiling action. This does not constitute a specific limitation.

[0113] Through the above process, through two animation diffusions, not only can character animation be generated quickly, avoiding dependence on animators to produce animation frame by frame, without consuming a lot of time and labor, it can effectively save expensive time and labor costs, and is conducive to improving the efficiency of character animation production, but also can make the target character have both facial and body movement performances, with richer character animation effects, and better able to meet the actual application needs in most scenes, thereby effectively solving the problems of high cost and poor versatility of character animation generation in related technologies.

[0114] See also Figure 9 In an exemplary embodiment, step 330 may include the following steps:

[0115] Step 331 : In response to a first circle selection operation triggered on the facial region of the target character in the original image, a position of the facial region of the target character in the original image is determined.

[0116] In this embodiment, the face mask image is determined based on the user's selection.

[0117] To facilitate the user's selection of the face region of the target character in the original image, a selection entry is provided for the user, and the user can then complete the selection of the face region of the target character in the original image through the selection entry.

[0118] Taking a user's smartphone as an example, the smartphone is deployed with a client having a character animation generation function, which at least includes uploading of an original image, selection of a face region of a target character, and the like. With the running of the client, an auxiliary page for assisting in character animation generation is displayed to the user in the smartphone, and the auxiliary page displays a selection entry, for example, the selection entry can be a "brush", and the user can slide the "brush" to directly select the face region of the target character in the original image. For the smartphone, the sliding operation can be detected to determine the position of the face region of the target character in the original image. The sliding operation can be regarded as a first selection operation triggered for the face region of the target character in the original image.

[0119] It is explained that the operation type of the first selection operation will be different with different input components configured by the electronic device. For example, if the electronic device is a smartphone and the configured input component includes a touch screen, the first selection operation can be a gesture operation such as clicking or sliding. If the electronic device is a desktop computer and the configured input component includes a mouse and a keyboard, the first selection operation can be a mechanical operation such as single-clicking, double-clicking, or dragging. This is not a specific limitation.

[0120] At step 333, based on the determined position, the face region of the target character and other regions in the original image are marked respectively to obtain a face mask image of the target character.

[0121] The marking refers to setting different values for pixels in different regions where the target character is located in the original image, so that the face region of the target character in the face mask image can be significantly distinguished from other regions. For example, the value of the pixel in the face region of the target character in the original image is set to 255, i.e., the face region of the target character in the face mask image is a white region. The value of the pixel in other regions of the target character in the original image is set to 0, i.e., the other regions of the target character in the face mask image are black regions.

[0122] Thus, through the marking, the face region of the target character and other regions in the original image are segmented, thereby laying a foundation for uniquely determining the position of the target character in the target character image for subsequent local redrawing of the target character.

[0123] In addition, if the original image contains multiple target roles, the user can also identify the multiple target roles in the original image through the user's selection, and then generate the role animation for each target role identified by the user, so as to ensure the accuracy of the identification.

[0124] In an example embodiment, step 330 can further include the following steps:

[0125] Step 335, in response to the second selection operation triggered for the target role in the original image, determining the position of the target role in the original image.

[0126] Similarly, similar to the first selection operation, the second selection operation can also be a gesture operation such as clicking, sliding, and can also be a mechanical operation such as single-click, double-click, and dragging, which will not be described here.

[0127] Step 337, based on the determined position, performing image segmentation on the target role in the original image to obtain a target role image.

[0128] That is, for an original image containing multiple target roles with a full front body, after image segmentation, more than one target role image can be obtained, each of which contains only one target role with a full front body, in other words, the role animation generation process in the embodiments of the present application is carried out in units of a single target role.

[0129] Of course, in other embodiments, for an original image containing multiple target roles, the automatic identification of multiple target roles in the original image can also be realized through an image segmentation model, and then the subsequent role animation generation is performed for each target role identified automatically, so as to ensure efficient automatic identification, and the present embodiment is not limited thereto.

[0130] With the cooperation of the above embodiments, the face mask image of the target role and the target role image are determined, which serves as the basis for the first animation diffusion, so that the generation of the role animation can be realized.

[0131] Please refer to Figure 10 In an example embodiment, step 350 can include the following steps:

[0132] Step 351, based on the visible area and the mask area in the face mask image, determining the position of the face area of the target role in the face mask image.

[0133] The visible area is used to describe the face area of the target role, and the mask area is used to describe other areas of the target role. Please refer to Figure 4bThe visible region is white and the mask region is black, that is, the value of each pixel in the visible region is set to 255 and the value of each pixel in the mask region is set to 0, so that the mouth region of the target character and other regions can be distinguished in the face mask image.

[0134] Based on this, after obtaining the face mask image of the target character, the position of the face region of the target character can be uniquely determined based on the visible region and the mask region in the face mask image, thereby laying a foundation for subsequent local redrawing of the target character in the target character image. That is, the same position in the target character image can be uniquely determined through the position of the face region of the target character in the face mask image, and then the local region where the target character is located at the same position is redrawn.

[0135] Step 353, according to the determined position and different facial actions, the target character in the target character image is locally redrawn to obtain a plurality of reference images.

[0136] In this embodiment, the local redrawing of the target character in the target character image based on the position of the face region of the target character in the face mask image, especially the redrawing of the local region of the target character in the target character image based on the position of the face region of the target character in the face mask image, is realized by calling an image diffusion model. The image diffusion model is a machine learning model that has been trained and has the ability to locally redraw the target character in the target character image.

[0137] Specifically, as shown in Figure 11 The process of calling the image diffusion model to realize the local redrawing of the target character can include the following steps:

[0138] First, obtain the facial action prompt word.

[0139] The facial action prompt word is used to describe different facial actions. For example, the facial action prompt word can be "close your mouth" and "open your mouth". Of course, in other embodiments, if more different facial actions need to be generated for the target character, the facial action prompt word can be modified to guide the image diffusion model to generate more different facial actions.

[0140] Second, input the face mask image and the target character image into the image diffusion model, and use the facial action prompt word to guide the image diffusion model to generate different facial actions for the target character, and output a plurality of reference images of the target character performing different facial actions.

[0141] Under the action of the above embodiment, under the guidance of the facial action prompt word, the image diffusion model can automatically generate facial actions matching the facial action prompt word for the target character, so that the target character has facial action performance.

[0142] Referring to Figure 12 In an example embodiment, step 370 can include the following steps:

[0143] First, input each reference image into the video diffusion model.

[0144] Second, use the video diffusion model to respectively infer each reference image from a static image to a dynamic video according to the plurality of reference actions in the action reference sequence, to obtain a plurality of action sequences.

[0145] In this embodiment, the inference from a static image to a dynamic video is achieved by calling the video diffusion model, which is a machine learning model trained and having the ability to infer from a static image to a dynamic video.

[0146] Still taking the above example, referring to Figure 5a to Figure 6c , through the calling of the video diffusion model, based on the different reference actions in the action reference sequence Figure 6a , the different reference actions in the action reference sequence are respectively mapped to the first reference image Figure 5a and the second reference image Figure 5b , forming a plurality of first reference images of the target character performing different reference actions and a plurality of second reference images of the target character performing different reference actions, and further forming the first action sequence Figure 6b and the second action sequence Figure 6c respectively composed of the plurality of first reference images and the plurality of second reference images. It can be seen that Figure 6a , the reference actions contained in each static image in the action reference sequence Figure 6b have a one-to-one correspondence with the reference actions contained in each static image in the first action reference sequence Figure 6c / the second action reference sequence, in addition, the target character in each static image in the first action reference sequence / the second action reference sequence also simultaneously performs different mouth actions corresponding to the first reference image Figure 5a / the second reference image Figure 5b .

[0147] In the above process, under the guidance of the action reference sequence, each reference image can be automatically inferred from a static image to a corresponding action sequence by using the video diffusion model, so that the target character not only has facial action performance, but also has body action performance.

[0148] Figure 13 is a specific implementation diagram of a role animation generation method in an application scenario. In this application scenario, the video diffusion model Figure 1The user terminal in the shown implementation environment independently completes the role animation generation, which can be a desktop computer, a notebook computer, a smart phone, a tablet computer, or the like. The target role is a cartoon character, the original image is a cartoon character image, the facial action includes a mouth action, the mouth action includes an open-mouth action and a close-mouth action, and thus the cartoon character mouth mask image and the cartoon character image are formed, the reference image corresponding to one mouth action includes a first reference image and a second reference image, and the action sequence corresponding to one reference image includes a first action sequence and a second action sequence.

[0149] As shown in Figure 13 With the running of the client with the role animation generation function, the user can upload the cartoon character image to the user terminal by means of the client. At this time, after obtaining the cartoon character image on the user terminal, the user can be provided with automatic role animation generation service.

[0150] Specifically, first, the resolution of the uploaded cartoon character image is detected, the cartoon character is recognized, and the T-Pose front full body is detected. Only when the cartoon character image passes the resolution detection, the cartoon character recognition, and the T-Pose front full body detection, can it enter the next step of "user interactive drawing", that is, the user is prompted to circle the mouth area of the cartoon character in the cartoon character image. Otherwise, the user is prompted to upload the cartoon character image again.

[0151] The user will circle the mouth area of the cartoon character in the cartoon character image that passes the above detection by means of the client. At this time, the user terminal can detect the circle operation, and thus the cartoon character mouth mask image and the cartoon character image are obtained by marking the mouth area and other areas of the cartoon character in the cartoon character image.

[0152] Then, after the first animation diffusion, that is, "picture diffusion", the cartoon character in the cartoon character image is locally redrawn in the mouth area, and the first reference image of the cartoon character performing the close-mouth action and the second reference image of the cartoon character performing the open-mouth action are obtained, so that the cartoon character can have facial action performance, especially facial speaking performance.

[0153] After the second animation diffusion, i.e., "video diffusion", two reference images are generated into two action video sequences under the driving of the action reference sequence, specifically, a first action sequence is generated from the first reference image, i.e., a video in which the cartoon character performs a closed-mouth action and simultaneously performs different reference actions, and a second action sequence is generated from the second reference image, i.e., a video in which the cartoon character performs an open-mouth action and simultaneously performs different reference actions, so that the cartoon character also has a body action performance. Of course, the body action performance depends on the reference actions in the action reference sequence, which can be arbitrarily selected by the user from a plurality of pre-provided action reference sequences, or automatically pushed to the user by the user terminal.

[0154] Finally, through video super-resolution, background removal by matting, video frame splicing and other animation post-processing, the cartoon character with both facial action performance and body action performance is outputted, which is no longer as single as the cartoon character in the related art.

[0155] In the application scenario, in a first aspect, no professional animator is needed to participate, and the user only needs to upload the picture and enclose the mouth area of the character in the picture, which is simple and easy to use, can effectively enrich the role animation generation materials, and can greatly save the labor cost of expensive role animation production; in a second aspect, during the role animation generation process, various detections, twice animation diffusion, and animation post-processing are automatically completed by the user terminal, which has high automation degree, can quickly generate role animation, does not need to rely on the animator to produce animation frame by frame, can save a lot of time cost of role animation production, and is beneficial to improve the role animation production efficiency; in a third aspect, the character in the role animation can have both facial speaking performance and body action performance, has more rich role animation effect, can meet the actual application needs of most scenes, and is beneficial to improve the universality of the role animation.

[0156] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps has no strict sequence limitation, and they can be executed in other orders. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0157] The following is an embodiment of the device of the present application, which can be used to execute the role animation generation method involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the role animation generation method involved in the present application.

[0158] Please refer to Figure 14 In the embodiment of the present application, a role animation generation device 900 is provided, which includes but is not limited to: an original image acquisition module 910, an input image determination module 930, a first animation diffusion module 950, a second animation diffusion module 970, and an animation post-processing module 990.

[0159] The original image acquisition module 910 is configured to acquire an original image, wherein the original image contains at least one target role with a front full body.

[0160] The input image determination module 930 is configured to determine a face mask image of the target role and a target role image based on the target role and the face region thereof in the original image.

[0161] The first animation diffusion module 950 is configured to take the face mask image and the target role image as inputs, perform a first animation diffusion related to facial motion, and obtain a plurality of reference images. Each reference image corresponds to a facial motion.

[0162] The second animation diffusion module 970 is configured to acquire a motion reference sequence and drive each reference image to perform a second animation diffusion related to body motion by using the motion reference sequence, and obtain a plurality of motion sequences. Each motion sequence corresponds to a reference image.

[0163] The animation post-processing module 990 is configured to perform animation post-processing on each motion sequence to obtain a role animation of the target role.

[0164] In an exemplary embodiment, the first animation diffusion module 950 is further configured to perform the following steps: determining the position of the face region of the target role in the face mask image based on the visible region and the mask region in the face mask image; and performing local redrawing on the target role in the target role image according to the determined position to obtain a first reference image and a second reference image.

[0165] In an exemplary embodiment, the first animation diffusion module 950 is further configured to perform the following steps: acquiring facial motion prompt words; the facial motion prompt words are used to describe different facial motions; inputting the face mask image and the target role image into an image diffusion model, and guiding the image diffusion model to generate different facial motions for the target role by using the facial motion prompt words, and outputting to obtain a plurality of reference images; the reference image refers to an image of the target role performing a facial motion.

[0166] In an example embodiment, the second animation diffusion module 970 is further configured to perform the following steps: inputting each reference image into a video diffusion model; and using the video diffusion model, respectively performing inference from a static image to a dynamic video on each reference image according to the plurality of reference actions in the action reference sequence to obtain a plurality of action sequences, wherein the action sequence refers to a video in which the target character performs the plurality of reference actions simultaneously.

[0167] In an example embodiment, the animation post-processing module 990 is further configured to perform the following steps: sampling each action sequence in a set manner to obtain a plurality of frame blocks belonging to different action sequences; and splicing the plurality of frame blocks belonging to different action sequences to obtain a character animation of the target character, wherein the character animation refers to a video in which the target character performs the plurality of reference actions simultaneously on the basis of alternately performing different facial actions.

[0168] In an example embodiment, the input image determination module 930 is further configured to perform the following steps: in response to a first circle selection operation triggered in the original image for the facial region of the target character, determining the position of the facial region of the target character in the original image; and based on the determined position, performing marking processing on the facial region and other regions of the target character in the original image respectively to obtain a facial mask image.

[0169] In an example embodiment, the input image determination module 930 is further configured to perform the following steps: in response to a second circle selection operation triggered in the original image for the target character, determining the position of the target character in the original image; and based on the determined position, performing image segmentation on the target character and other target characters in the original image to obtain a target character image.

[0170] In an example embodiment, the character animation generation apparatus 900 is further configured to perform the following steps: performing screening processing on the original image, wherein the screening processing includes at least one of resolution detection, target character detection, and front full-body detection of the target character; and deleting original images that do not meet a set condition, so that the original images that meet the set condition perform the steps of determining the facial mask image of the target character and the target character image based on the target character and the facial region thereof in the original image.

[0171] In an example embodiment, the face region includes a mouth region, and the facial action includes a first mouth action and a second mouth action; the plurality of reference images includes a first reference image in which the target character performs the first mouth action and a second reference image in which the target character performs the second mouth action; the plurality of action sequences includes a first action sequence in which the target character simultaneously performs the first mouth action and a plurality of reference actions, respectively, and a second action sequence in which the target character simultaneously performs the second mouth action and the plurality of reference actions, respectively; and the character animation is a video in which the target character simultaneously performs the plurality of reference actions on the basis of alternately performing the first mouth action and the second mouth action.

[0172] It should be noted that the character animation generation apparatus provided in the above embodiments is only used as an example to illustrate the division of the above functional modules during the generation of the character animation, and in actual applications, the above functions can be completed by different functional modules according to the needs, that is, the internal structure of the character animation generation apparatus is divided into different functional modules to complete all or part of the above-described functions.

[0173] In addition, the character animation generation apparatus and the character animation generation method provided in the above embodiments belong to the same concept, and the specific manner in which each module performs the operation has been described in detail in the method embodiments, which will not be described here.

[0174] Please refer to Figure 15 In the embodiments of the present application, an electronic device 4000 can include a smart phone, a tablet computer, a desktop computer, a notebook computer, a server, etc.

[0175] In Figure 15 The electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0176] The data interaction between the processor 4001 and the memory 4003 can be realized through at least one communication bus 4002. The communication bus 4002 can include a channel for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect, Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture, Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience, Figure 15 In the above embodiments, only one thick line is used to represent the communication bus 4002, but it does not mean that there is only one bus or only one type of bus.

[0177] Optionally, the electronic device 4000 can further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that the transceiver 4004 is not limited to one in actual application, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0178] The processor 4001 can be a CPU (Central Processing Unit, central processing unit), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure content of the present application. The processor 4001 can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.

[0179] The memory 4003 can be a ROM (Read Only Memory, read only memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory, random access memory) or other types of dynamic storage devices that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory, electrically erasable programmable read only memory), a CD-ROM (Compact Disc Read Only Memory, compact disc read only memory) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage medium or other magnetic storage devices, or any other medium capable of carrying or storing computer programs in the form of instructions or data structures and capable of being accessed by the electronic device 4000, but not limited thereto.

[0180] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002.

[0181] The computer program is executed by one or more processors 4001 to implement the role animation generation method in each of the above embodiments.

[0182] In addition, a storage medium is provided in the embodiments of the present application, and the storage medium stores a computer program. The computer program is executed by one or more processors to implement the role animation generation method.

[0183] In the embodiments of the present application, a computer program product is provided, which includes a computer program. The computer program is executed by one or more processors to implement the role animation generation method.

[0184] The above only describes some embodiments of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, some improvements and refinements can be made, which should also be considered as the protection scope of the present application.

Claims

1. A method of character animation generation, characterized by, The method comprises: obtaining an original image; the original image contains at least one target character with a front full body; determining a face mask image and a target character image of the target character based on the target character and the face region thereof in the original image; performing first animation diffusion related to facial action on the face mask image and the target character image as input to obtain a plurality of reference images; obtaining an action reference sequence, and inputting each reference image into a video diffusion model; using the video diffusion model, performing inference from a static image to a dynamic video on each reference image according to a plurality of reference actions in the action reference sequence to obtain a plurality of action sequences; the action sequence refers to a video in which the target character performing the facial action simultaneously performs a plurality of reference actions; performing animation post-processing on each action sequence to obtain a character animation of the target character; wherein the first animation diffusion related to facial action on the face mask image and the target character image as input to obtain a plurality of reference images comprises: obtaining a facial action prompt word; the facial action prompt word is used to describe different facial actions; inputting the face mask image and the target character image into an image diffusion model, and using the facial action prompt word to guide the image diffusion model to generate different facial actions for the target character, and outputting to obtain a plurality of reference images; the reference image refers to an image in which the target character performs the facial action.

2. The method of claim 1, wherein, The first animation diffusion related to facial action on the face mask image and the target character image as input to obtain a plurality of reference images comprises: determining the position of the face region of the target character in the face mask image based on the visible region and the mask region in the face mask image; according to the determined position and different facial actions, performing local redrawing on the target character in the target character image to obtain a plurality of reference images.

3. The method of claim 1, wherein, The animation post-processing on each action sequence to obtain a character animation of the target character comprises: sampling each action sequence in a set manner to obtain a plurality of frame blocks respectively belonging to different action sequences; splicing a plurality of frame blocks belonging to different action sequences to obtain a character animation of the target character; the character animation refers to a video in which the target character simultaneously performs a plurality of reference actions on the basis of alternately performing different facial actions.

4. The method of claim 1, wherein, The determination of the face mask image and the target character image of the target character based on the target character and the face region thereof in the original image comprises: in response to a first circle selection operation triggered for the face region of the target character in the original image, determining the position of the face region of the target character in the original image; based on the determined position, performing marking processing on the face region and other regions of the target character in the original image respectively to obtain the face mask image.

5. The method of claim 1, wherein, After obtaining the original image, the method further comprises: The original image is filtered; the filtering process includes at least one of resolution detection, target character detection, and front full-body detection of the target character; The original image that does not meet the set condition is deleted, and the original image that meets the set condition is used to determine the face mask image and the target character image of the target character based on the target character and the face region thereof in the original image.

6. The method according to any one of claims 1 to 5, characterized in that, The face region includes a mouth region, and the face action includes a first mouth action and a second mouth action; The plurality of reference images include a first reference image in which the target character performs the first mouth action and a second reference image in which the target character performs the second mouth action; The plurality of action sequences include a first action sequence in which the target character simultaneously performs the first mouth action and a plurality of reference actions, and a second action sequence in which the target character simultaneously performs the second mouth action and a plurality of reference actions; The character animation refers to a video in which the target character simultaneously performs a plurality of reference actions based on alternately performing the first mouth action and the second mouth action.

7. A character animation generating device, characterized in that: The device comprises: An original image acquisition module for acquiring an original image; the original image contains at least one target character with a front full-body; An input image determination module for determining a face mask image and a target character image of the target character based on the target character and the face region thereof in the original image; A first animation diffusion module for performing a first animation diffusion related to a face action using the face mask image and the target character image as input to obtain a plurality of reference images; A second animation diffusion module for obtaining an action reference sequence and inputting each reference image into a video diffusion model; using the video diffusion model, each reference image is inferred from a static image to a dynamic video according to a plurality of reference actions in the action reference sequence to obtain a plurality of action sequences; the action sequence refers to a video in which the target character performing the face action simultaneously performs a plurality of reference actions; An animation post-processing module for performing animation post-processing on each action sequence to obtain a character animation of the target character; The first animation diffusion module is also used to obtain a face action prompt word; the face action prompt word is used to describe different face actions; the face mask image and the target character image are input into an image diffusion model, and the face action prompt word is used to guide the image diffusion model to generate different face actions for the target character to output a plurality of reference images; the reference image refers to an image in which the target character performs the face action.

8. An electronic device comprising at least one processor and at least one memory, wherein, The memory stores a computer program, and the computer program is executed by the processor to implement the character animation generation method of any one of claims 1-6.

Citation Information

Patent Citations

  • Expression creation method and device

    CN114140564A

  • Virtual character animation generation method and device, storage medium and terminal

    CN114219878A

  • Virtual object face driving and model training method and device, equipment and medium

    CN117911588A