Character animation generation method and device, equipment, storage medium and program product

By generating character animation from a single image of the target person and adjusting mouth movements in conjunction with call scene information, the problem of insufficient flexibility in facial expressions and body movements in existing technologies is solved, achieving high-quality, personalized animation generation and improving user satisfaction.

CN120976376APending Publication Date: 2025-11-18HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510622406.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2025-05-14
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing voice-driven technologies mainly focus on the generation and adjustment of the mouth area, resulting in poor flexibility in the generated character animations and making it difficult to meet the diverse needs of different users in different scenarios.

Method used

By acquiring dynamic information about a person and images of the target person, the system drives the generation of character animations from a single image of the target person. It also allows users to control and edit facial expressions and body movements, and adjust mouth movements in conjunction with call information in the call scenario to generate animations that meet the needs of the scene.

Benefits of technology

It enables users to control and edit facial expressions and body movements in character animations, generating high-quality and more expressive character animations that can meet users' personalized needs and improve the flexibility and naturalness of animation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976376A_ABST
    Figure CN120976376A_ABST
Patent Text Reader

Abstract

The invention provides a character animation generation method and device, equipment, a storage medium and a program product, and belongs to the field of artificial intelligence. The method comprises the following steps: driving a character picture of a target character into a first character animation, wherein the expression change of the target character in the first character animation is consistent with the character expression change reflected by the expression information and / or the limb action change of the target character is consistent with the limb action change reflected by the limb action information. And in response to a correction operation for the first character animation, performing expression driving on the face of the target character in the first character animation according to the corrected expression information and / or performing limb driving on the limbs of the target character in the first character animation according to the corrected limb action information to obtain a second character animation. The character animation is generated for the single character picture, controllability and editability of character expressions and / or limb actions in the character animation are achieved, and personalized expression requirements of different users in different scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202410634187.4, filed May 16, 2024, entitled "Character Animation Generation Method and Device, Computer Device", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of artificial intelligence, and in particular to a character animation generation method and device, a computer device, a storage medium, and a program product. BACKGROUND

[0003] With the rapid development of computer graphics, artificial intelligence (AI), and other technologies, applications represented by the metaverse and artificial intelligence generative content (AIGC) are increasingly appearing in the production of professional generated content (PGC) and user generative content (UGC). The most mature and already in use is voice-driven technology. Voice-driven technology is a technology that controls the operation of a device or application program by inputting voice signals. For example, voice-driven single photo or voice-driven 3-dimension (3D) digital human technology is used to generate character animation. In the example of voice-driven single photo, a user can control a static photo through voice instructions, such as making the character in the photo "speak". In the example of voice-driven 3D digital human, a user can upload their own voice information or select the corresponding audio from a pre-set audio library, and then based on the content of the audio, use deep generative networks (DGN) to generate and adjust the mouth of a pre-set 3D digital human image, so that the digital human's mouth shape matches the audio content. For example, when the audio content is "hello", the 3D digital human's mouth shape is also adjusted to say "hello".

[0004] However, the above voice-driven technology mainly focuses on the generation and adjustment of the mouth area, resulting in poor flexibility of the generated character animation, which is difficult to meet the diverse needs of different users in different scenarios. SUMMARY

[0005] The present application provides a character animation generation method, device, equipment, storage medium and program product, which can generate character animation that meets the diverse needs of different users in different scenarios.

[0006] In a first aspect, a method for generating a character animation is provided. The method comprises: obtaining character dynamic information and a character picture of a target character, the character dynamic information comprising expression information and / or body action information, the expression information being used to reflect a change in expression of the character, and the body action information being used to reflect a change in body action of the character; driving the character picture into a first character animation according to the character dynamic information, and displaying the first character animation, wherein the change in expression of the target character in the first character animation is consistent with the change in expression of the character reflected by the expression information, and / or the change in body action of the target character in the first character animation is consistent with the change in body action of the character reflected by the body action information; in response to a correction operation on the first character animation, obtaining correction information, the correction information comprising correction expression information and / or correction body action information; performing expression driving on a face of the target character in the first character animation according to the correction expression information, and / or performing body action driving on a body of the target character in the first character animation according to the correction body action information, to obtain a second character animation.

[0007] In the first character animation, the change in expression of the target character is consistent with the change in expression of the character reflected by the expression information, including that the expression category of the target character in the first character animation is the same as the expression category of the character reflected by the expression information. In the case that the expression category of the character reflected by the expression information has multiple categories, the change in expression of the target character in the first character animation is consistent with the change in expression of the character reflected by the expression information, further including that the transition order between different expression categories of the target character in the first character animation is consistent with the transition order of the expression categories of the character reflected by the expression information. In the first character animation, the change in body action of the target character is consistent with the change in body action of the character reflected by the body action information, including that the body action type of the target character in the first character animation is the same as the body action type of the character reflected by the expression information. In the case that the body action type of the character reflected by the body action information has multiple types, the change in body action of the target character in the first character animation is consistent with the change in body action of the character reflected by the body action information, further including that the switching order between different body action types of the target character in the first character animation is consistent with the switching order of the body action types of the character reflected by the body action information. The mouth action of the character in the character animation matches the audio information, meaning that the mouth action of the character in the character animation is the same as the mouth action of the character when speaking the content of the audio information.

[0008] It can be seen that the application drives a single picture of a target character to generate a character animation, and the expression and / or body movement of the target character in the character animation is controllable and editable. Compared with the related art, the expression and / or body movement of the character animation is controllable and editable, the user can edit and optimize the expression and / or body movement of the target character in the character animation, so that the expression and / or body movement of the target character in the finally generated second character animation can meet the personalized needs of the user, greatly improving the user's satisfaction with the animation effect. Since the user can finely control the expression and / or body movement, the character animation can be more lively, natural, and more consistent with the actual situation and the user's expectations. Compared with the single and unnatural expression and movement that may exist in the related art, the technical solution of the application can generate a character animation with high quality and more expressive, improving the overall animation effect.

[0009] Optionally, the expression information includes expression description text and / or an expression video of the first character. The body movement information includes body movement description text and / or a body movement video of the second character.

[0010] Optionally, the character dynamic information includes expression information, and driving the character picture to the first character animation according to the character dynamic information includes: driving the face in the character picture according to the expression information to obtain an expression animation, and the expression change of the target character in the expression animation is consistent with the expression change of the character reflected by the expression information. The expression change of the target character in the first character animation adopts the expression animation.

[0011] In this implementation, the user can only edit the expression in the character animation, so that the character expression can meet the user's needs.

[0012] Optionally, the character dynamic information includes body movement information, and driving the character picture to the first character animation according to the character dynamic information includes: driving the body in the character picture according to the body movement information to obtain a body animation, and the body movement change of the target character in the body animation is consistent with the body movement change of the character reflected by the body movement information. The body movement change of the target character in the first character animation adopts the body animation.

[0013] In this implementation, the user only edits the body movement in the character animation, so that the character body movement can meet the user's needs.

[0014] Optionally, the character dynamic information comprises expression information and body action information. The driving the character picture into the first character animation according to the character dynamic information comprises: driving the face in the character picture according to the expression information to obtain an expression animation, wherein the expression change of the target character in the expression animation is consistent with the expression change of the character reflected by the expression information; driving the body in the character picture according to the body action information to obtain a body animation, wherein the body action change of the target character in the body animation is consistent with the body action change of the character reflected by the body action information; and fusing the expression animation and the body animation to obtain the first character animation. In the first character animation, the expression change of the target character is adopted from the expression animation, and the body action change of the target character is adopted from the body animation.

[0015] In this implementation, the user edits the expression and the body action in the character animation respectively, so that the expression and the body action of the character can meet the user's demand.

[0016] Optionally, the driving the face in the character picture according to the expression information to obtain the expression animation comprises: driving the face in the character picture according to the expression information to obtain a plurality of expression segments; sorting the plurality of expression segments according to a first sorting instruction to obtain sorted expression segments; and inserting an expression transition frame between a first expression segment and a second expression segment if a first expression label corresponding to the last frame of the first expression segment is different from a second expression label corresponding to the first frame of the second expression segment. The first expression segment is adjacent to the second expression segment and precedes the second expression segment. The expression transition frame is used to smooth the expression transition from the first expression segment to the second expression segment.

[0017] The application realizes the smooth transition of the expression in the previous expression segment to the expression in the next expression segment by inserting the expression transition frame between the adjacent expression segments, which can prevent the expression from jumping and thus affect the expression animation effect.

[0018] Optionally, the sorting the plurality of expression segments according to the first sorting instruction comprises: displaying a first sorting interface, wherein the plurality of expression segments are displayed in the first sorting interface; and sorting the plurality of expression segments according to the first sorting instruction in response to receiving the first sorting instruction through the first sorting interface. The first sorting instruction is used to indicate the arrangement order of the plurality of expression segments.

[0019] The application realizes the user's custom arrangement of the plurality of expression segments by providing the user interactive interface.

[0020] Optionally, the limb driving according to the limb action information to obtain the limb animation comprises: driving the limbs in the character picture according to the limb action information to obtain a plurality of limb action clips; sorting the plurality of limb action clips according to a second sorting instruction to obtain sorted limb action clips; if a first limb action corresponding to a last frame of a first limb action clip in the sorted limb action clips is different from a second limb action corresponding to a first frame of a second limb action clip, inserting a limb transition frame between the first limb action clip and the second limb action clip. The first limb action clip and the second limb action clip are adjacent and the first limb action clip is before the second limb action clip. The limb transition frame is used to smooth the transition of the limb action from the first limb action clip to the second limb action clip.

[0021] The application realizes the smooth transition of the limb action from the previous limb action clip to the next limb action clip by inserting the limb transition frame between the adjacent limb action clips, which can prevent the jump of the limb action and affect the effect of the limb animation.

[0022] Optionally, the sorting the plurality of limb action clips according to the second sorting instruction comprises: displaying a second sorting interface, and the second sorting interface displays the plurality of limb action clips; and in response to receiving the second sorting instruction through the second sorting interface, sorting the plurality of limb action clips according to the second sorting instruction. The second sorting instruction is used to indicate the arrangement order of the plurality of limb action clips.

[0023] The application realizes the custom arrangement of the plurality of limb action clips by providing the user interaction interface.

[0024] Optionally, the method further comprises: in response to detecting a determination operation on the second character animation, saving the second character animation.

[0025] The application can obtain the second character animation meeting the personalized needs of the user by modifying and optimizing the first character animation. The saved second character animation can be added to the character animation set of the target character, so that the target character can use the voice driving based on the specific call information in the call scene.

[0026] Optionally, the obtaining the character dynamic information comprises: obtaining the input character dynamic information.

[0027] That is, the user can upload the character dynamic information to the computer device according to the needs, which improves the flexibility of generating the character animation.

[0028] Optionally, the obtaining the character dynamic information comprises: displaying a dynamic information selection interface, wherein a plurality of preset character dynamic information are displayed in the dynamic information selection interface; and in response to detecting a selection operation on a character dynamic information in the plurality of preset character dynamic information, obtaining the selected character dynamic information.

[0029] That is, the application can provide a plurality of preset character dynamic information, so that the user can select a character dynamic information for generating a character animation from the plurality of preset character dynamic information, thereby improving the user experience.

[0030] In a second aspect, a character animation generation method is provided, which comprises: obtaining call information of a target character in a call scene; determining an expression label and a body action label according to the call information; determining an initial character animation from a character animation set according to the expression label and / or the body action label, wherein the character animation set comprises at least one character animation corresponding to the target character; and adjusting a mouth action of the initial character animation according to the call information to obtain a target character animation, wherein the mouth action of the target character in the target character animation matches the call information.

[0031] As can be seen, the application determines an initial character animation from a character animation set in combination with call information of a target character in a call scene, and adjusts a mouth action of the initial character animation based on the call information to obtain a target character animation. Since the initial character animation is pre-constructed based on the requirements of the target character, the character animation meets the basic requirements of the expression and / or body action of the target character. For a specific call scene, the mouth action is generated and adjusted based on the initial character animation, so that the target character animation that meets the dual requirements of the call scene and the call information can be obtained.

[0032] As can be seen, the application realizes controllable and editable expression and / or body action in the offline stage, and determines an initial character animation from a pre-constructed character animation set and adjusts a mouth action in the online stage in combination with call information in a call scene. This two-stage combination makes the character animation generation method adaptable to different scenes and requirements, and has strong flexibility. Whether it is a fixed animation requirement or a dynamically changing call scene, the appropriate character animation can be quickly generated.

[0033] In a third aspect, a character animation generation apparatus is provided. The apparatus comprises a plurality of functional modules that interact with each other to implement the character animation method provided in the first aspect. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be combined or divided in any manner based on specific implementation.

[0034] In a fourth aspect, a character animation generation apparatus is provided. The apparatus comprises a plurality of function modules which interact to implement the character animation method provided in the second aspect. The plurality of function modules can be implemented based on software, hardware or a combination of software and hardware, and the plurality of function modules can be combined or divided based on specific implementation.

[0035] In a fifth aspect, a computer device is provided. The computer device comprises a processor and a memory. The memory is configured to store a computer program comprising program instructions. The processor is configured to invoke the computer program to implement the character animation generation method provided in the first aspect or the character animation generation method provided in the second aspect.

[0036] In a sixth aspect, a computer readable storage medium is provided. The computer readable storage medium stores instructions thereon. When the instructions are executed by a processor, the character animation generation method provided in the first aspect or the character animation generation method provided in the second aspect is implemented.

[0037] In a seventh aspect, a computer program product is provided. The computer program product comprises a computer program. When the computer program is executed by a processor, the character animation generation method provided in the first aspect or the character animation generation method provided in the second aspect is implemented.

[0038] In an eighth aspect, a chip is provided. The chip comprises programmable logic circuitry and / or program instructions. When the chip is running, the character animation generation method provided in the first aspect or the character animation generation method provided in the second aspect is implemented.

[0039] The technical effects obtained by the third aspect to the eighth aspect are similar to the technical effects obtained by the corresponding technical means in the first aspect and the second aspect, and thus will not be described here. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 FIG. 1 is a flow diagram of a character animation generation method provided by an embodiment of the present application;

[0041] Figure 2 FIG. 2 is a schematic diagram of a display interface provided by an embodiment of the present application;

[0042] Figure 3 FIG. 3 is a schematic diagram of driving a character picture based on an expression video provided by an embodiment of the present application;

[0043] Figure 4 FIG. 4 is a schematic diagram of another display interface provided by an embodiment of the present application;

[0044] Figure 5is a flowchart of another method for generating a character animation provided by an embodiment of the present application;

[0045] Figure 6 is an architecture diagram of a system for generating a character animation provided by an embodiment of the present application;

[0046] Figure 7 is a process diagram of generating a character animation provided by an embodiment of the present application;

[0047] Figure 8 is another process diagram of generating a character animation provided by an embodiment of the present application;

[0048] Figure 9 is still another process diagram of generating a character animation provided by an embodiment of the present application;

[0049] Figure 10 is a structure diagram of a device for generating a character animation provided by an embodiment of the present application;

[0050] Figure 11 is another structure diagram of a device for generating a character animation provided by an embodiment of the present application;

[0051] Figure 12 is a structure diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0053] Before explaining the method for generating a character animation provided by an embodiment of the present application in detail, some terms and application scenarios related to the embodiments of the present application will be introduced.

[0054] First, the terms related to the embodiments of the present application will be introduced.

[0055] 1. Audio driven dubbing: Audio driven dubbing is a technology for controlling the operation of a device or an application program by inputting a voice signal. In simple terms, audio driven dubbing is a data-driven, artificial intelligence algorithm for generating facial animation. The input of the algorithm is usually a facial video and a voice audio, and the main goal of the algorithm is to edit the mouth of the face in the input video to match the movements of the input voice audio.

[0056] 2. Digital human: A digital human refers to a digitalized character image (virtual character image) close to the image of a human being created by computer technology.

[0057] 3. User-generated content (UGC): UGC refers to content created by ordinary users or individual media, such as videos shared by users on video websites or video applications. UGC originated from the Internet, so UGC is mainly published and shared through the Internet.

[0058] 4. Professional-generated content (PGC): PGC refers to content created by professionals (such as traditional broadcasters). Compared with UGC, PGC is more professional in classification and has better content quality. Similarly, PGC is mainly published and shared through the Internet.

[0059] 5. Artificial intelligence-generated content (AIGC): AIGC refers to media content generated by artificial intelligence algorithms according to user instructions, such as text, voice, images or video.

[0060] Secondly, the application scenarios of the embodiments of the present application are introduced.

[0061] With the rapid development of computer graphics, artificial intelligence and other technologies, applications represented by the metaverse and AIGC are increasingly appearing in the production of PGC and UGC. As a typical AIGC, voice-driven character animation is widely used in various scenarios.

[0062] Currently, voice-driven technology is generally divided into the following two application examples:

[0063] Voice-driven single photo: Users can control a static photo through voice commands, such as making the person in the photo "open their mouth and speak". This technology may involve voice recognition and image processing, and by analyzing the voice content, the mouth of the person in the photo is fine-tuned to make it look like they are speaking.

[0064] Voice-driven 3D digital human technology: Users upload their own voice or select audio from a pre-set audio library, and the system will generate and adjust the mouth of the pre-set 3D digital human image according to the audio content, so that the digital human's mouth shape matches the audio content. For example, when the audio content is "hello", the digital human's mouth shape will be adjusted accordingly to say "hello".

[0065] Although voice-driven technology has made significant progress in some aspects, it still has limitations in the following two aspects:

[0066] Limited generation area: Current voice-driven technology mainly focuses on the generation and adjustment of the mouth area, but cannot comprehensively edit the entire driving avatar. This means that users cannot customize changes to other parts of the digital person (such as facial expressions, body posture, etc.) according to their own needs. For example, a user may want the digital person to have a specific expression when speaking or perform a specific choreography, but existing technology cannot meet this demand.

[0067] High operational difficulty: Although 3D digital people can perform limb choreography, this process involves complex modeling and binding knowledge. For ordinary users without relevant professional knowledge, these operation processes are relatively difficult and cannot be completed independently. For example, an ordinary user may want the digital person to dance a dance, but due to lack of modeling and binding knowledge, he cannot choreograph a set of smooth dance movements for the digital person.

[0068] In summary, although voice-driven technology has shown great potential in the metaverse and AIGC fields, it still has certain limitations in the flexibility and personalized expression of character animation, making it difficult to fully meet the diverse needs of different users in different scenarios.

[0069] To facilitate the improvement of the technical solution of the present application in the field of voice-driven generation of character animation, the implementation scheme of the related art using voice-driven technology to generate character animation will be briefly introduced first.

[0070] The related art one provides a scheme of voice-driven preset video. The key technology of this scheme is to change the mouth shape of the character in the preset video according to the input audio signal, so that the mouth shape of the character animation in the video matches the audio content. The implementation process of this scheme mainly includes the following steps 1-3:

[0071] Step 1: Face preprocessing.

[0072] First, the face in the preset video is detected, and the position of the face in the preset video is accurately located and cropped out. This step is to separate the face from the complex video background for subsequent processing. Then, the lower half of the cropped face (mainly the mouth area) is subjected to a mask operation (masking operation is like adding a "mask" to the image, covering the lower half of the face, and this part covered will be used for subsequent mouth shape generation); At the same time, other faces that have not been masked are randomly selected as reference frames (the part of the face that has not been masked is used to provide appearance, pose, etc. Information will be used as a reference to ensure that the generated face shape in the subsequent mouth shape generation can match the overall image and posture of the character in the preset video, thereby obtaining better generation effect).

[0073] Step 2: Audio signal processing.

[0074] First, the input audio signal is cut and sampled. Cutting is to divide the continuous audio signal into small audio segments, and sampling is to discretize the audio segments for computer processing and analysis. Then, features are extracted from the sampled audio signal, which contain important information about the mouth shape, such as the opening and closing degree of the mouth, the movement trend of the lips, etc. The extracted features will serve as an important reference for the generation network to generate a mouth shape that matches the audio content.

[0075] Step 3: Mouth synthesis.

[0076] The preprocessed face (including the masked face and the reference frame) and the extracted features from the audio signal are input into the pre-trained generation network. The generation network performs mouth synthesis based on these input information to generate a character animation consistent with the audio content.

[0077] It should be understood that the character animation obtained through the above method is equivalent to changing the mouth shape of the pre-set video to match the audio content.

[0078] However, the related art has the following two defects:

[0079] First, this scheme strongly depends on the pre-set video, and must be given a pre-set video to synthesize the mouth. The application has high limitations. In other words, if there is no pre-set video, the entire technical process cannot be started, greatly limiting its application scenarios and flexibility.

[0080] As an example, suppose an animation production team wants to create a mouth animation that matches a new audio content, but does not have suitable pre-set video materials. According to the requirements of the related art, they cannot work and have to find or create a pre-set video that meets the requirements, which undoubtedly increases the time cost and difficulty of the work.

[0081] Second, after giving the pre-set video, the above scheme can only drive the synthesis of the mouth and cannot modify the character in the pre-set video, such as expression editing, action arrangement, etc. This may result in the generated character animation based on the above scheme, where the audio content does not match the character action or expression in the pre-set video, affecting the realism and expressiveness of the character animation, resulting in poor character animation effect.

[0082] As an example, the character in the preset video is smiling, but the audio content input when generating the character animation is a sad monologue. According to the solution of the related art 1, only the mouth shape of the character can be modified to match the sad monologue audio, but the expression of the character is still smiling. The animation generated in this way will look very strange because the sad monologue is seriously inconsistent with the smiling expression, and the audience will feel unnatural when watching and cannot be well immersed in the atmosphere created by the animation.

[0083] The related art 2 provides a solution of voice-driven single photo. The key technology of the solution is to drive the single photo according to the input audio information, thereby generating a character animation. Compared with the above-mentioned related art 1, the solution of the related art 2 only needs a photo as input, which is the initial source of the character image in the subsequent generated character animation. The implementation process of the solution is as follows: first, the audio information is cut, sampled and feature extracted, the matching mouth shape information, head pose information and expression information are obtained according to the extracted audio features. Then, the single photo and the extracted mouth shape information, head pose information and expression information are input into the pre-trained generation network. The generation network will perform synthesis driving operation according to these input information, and finally output a piece of character animation. In the final output character animation, the character in the photo will present the corresponding mouth shape, head pose and expression change according to the audio information, and it looks like talking.

[0084] That is, the picture effect of each frame of image in the character animation generated by the above-mentioned solution is equivalent to changing the mouth shape, pose or expression of the character on the basis of the single photo.

[0085] Compared with the solution provided by the above-mentioned related art 1, the solution provided by the related art 2 only needs to provide a photo and does not need to provide a preset video, so the application limitation is lower. However, the related art 2 still has the following two defects:

[0086] First, the pose information and expression information extracted from the audio information have a one-to-many situation, that is, the same person may have different head pose actions when narrating the same content. For example, when telling an interesting story, some people may nod their heads while speaking, some people may shake their heads and move their brains, and some people may only slightly turn their heads. At the same time, due to the different speaking habits of each person, the facial expressions of each person will also differ when narrating the same content. For example, when expressing happy emotions, some people may laugh with their mouths open, some people may smile, and some people may only show joy in their eyes. These differences will cause the pose and expression information extracted by the above-mentioned solution based on the audio information to be unable to accurately determine the specific actions and expressions of the character in the generated video, and the driving effect has randomness, which is difficult to fully meet the specific needs of the user in a specific scenario.

[0087] As an example, assume that a user wants to generate a video for a formal business presentation scenario. In this scenario, the presenter needs to maintain a proper posture and a serious expression. However, when using the solution of related art II to generate the video, due to the inability of the audio information to accurately control the head posture and expression, the presenter in the generated video may make some actions that do not conform to the business presentation scenario during the speech, such as frequently tilting the head or making exaggerated expressions, which is inconsistent with the formal and serious presentation image expected by the user and cannot meet the needs in this specific scenario.

[0088] Secondly, similar to related art I, the main function of this solution is focused on lip modification, and although it can generate mouth movements that match the audio content according to the audio information, it is powerless for body movements. The user cannot perform secondary editing on the body movements of the character in the generated video, which limits the diversity and expressiveness of the video. In some scenarios that require rich body movements to enhance the expression effect, this technology is not up to the task.

[0089] As an example, a user wants to make a dance teaching video, which requires the character to demonstrate the corresponding body movements while explaining the dance movements. However, using related art II, although the character in the video can speak according to the explanation of the audio content, it cannot add dance movements to the character. The final generated video can only see the character "speaking" without any dance movements, which cannot meet the actual needs of the dance teaching video and reduces the teaching effect and appeal of the video.

[0090] In view of the emerging field of voice-driven generation of character animation, the embodiments of the present application comprehensively consider the shortcomings of existing solutions and provide a character animation generation solution to meet the diverse voice-driven needs of different users in different scenarios. In some embodiments, the technical solution provided by the present application can be divided into an offline asset construction stage and an online animation generation stage.

[0091] In the offline asset construction stage, the technical solution of the present application generates a character animation by driving a single picture of a target character based on dynamic information of the character and a character picture of the target character, and the expression and / or body movement of the target character in the character animation are controllable and editable. Compared with the related art, the expression and / or body movement of the target character in the character animation is controllable and editable, so that the expression and / or body movement of the target character in the finally generated character animation can meet the personalized needs of the user, greatly improving the user's satisfaction with the animation effect. Since the user can finely control the expression and / or body movement, the character animation can be more lively, natural, and more consistent with the actual situation and the user's expectations. Compared with the related art, which may have problems such as single expression and movement and unnaturalness, the technical solution of the present application can generate a character animation with high quality and more expressive, improving the overall animation effect.

[0092] In the online animation generation stage, the technical solution of the present application determines an initial character animation from a character animation set in combination with the call information of the target character in the call scene, adjusts the mouth movement of the initial character animation based on the call information, and obtains the target character animation. Since the initial character animation is pre-constructed based on the needs of the target character, the character animation meets the basic requirements of the expression and / or body movement of the target character. For a specific call scene, the initial character animation is generated and adjusted based on the mouth movement, and the target character animation that meets the dual requirements of the call scene and the call information can be obtained.

[0093] As can be seen, the embodiment of the present application realizes controllable and editable expression and / or body movement in the offline stage, and determines an initial character animation from a pre-constructed character animation set and adjusts the mouth movement in the online stage in combination with the call information in the call scene. This two-stage combination makes the character animation generation method adaptable to different scenes and requirements, and has strong flexibility. Whether facing fixed animation requirements or dynamic changing call scenes, the appropriate character animation can be quickly generated.

[0094] The character animation generation method provided by the embodiment of the present application can be applied to various application scenarios, including but not limited to mobile calls, virtual hosts, digital entertainment, human-computer interaction, and remote conferences.

[0095] During a mobile call, a user can want his / her virtual image to be more lively and interesting to increase the interest and interactivity of the call. Exemplarily, when a user uses a video call software to chat with a friend, a target character animation generated by the technical solution provided in the embodiments of the present application can be used as the virtual image of the user. During the chat, the technical solution provided in the embodiments of the present application can also adjust the expression and body movement of the virtual image in combination with the chat content of the user. For example, when telling a joke, the virtual image can make a laughing expression and a hand-waving movement to make the chat atmosphere more relaxed and pleasant.

[0096] In various online activities, programs or meetings, a virtual host can replace a traditional real host, which has the advantages of low cost and strong customizability. For example, in an online product launch, a virtual host can be generated by the technical solution provided in the embodiments of the present application. When introducing the features of the product, the virtual host can show a focused and serious expression, and cooperate with the corresponding gesture movement to enhance the appeal of the product introduction; when announcing important information, the virtual host can be switched to an excited and excited expression and movement to attract the attention of the audience.

[0097] In the field of digital entertainment, such as games, animation production, etc., rich and diverse character animations can improve the entertainment experience of users. Exemplarily, in a role-playing game, a player can use the technical solution provided in the embodiments of the present application to generate unique animation effects for the character created by the player. Moreover, the player can customize the expression and body movement of the character according to the personality and story background of the character. For example, a brave character will show a determined expression and decisive movement when fighting, while a gentle character will show a kind smile and elegant posture when communicating.

[0098] In a human-computer interaction scene, a character animation with lively expressions and movements can make the machine more humanized and improve the interaction experience between the user and the machine. Exemplarily, in an intelligent customer service system, a character animation generated by the technical solution provided in the embodiments of the present application can be used as the image of a customer service representative. When a user raises a question, the customer service animation can make a corresponding expression and movement response according to the type of the question and the emotion of the user. For example, for the dissatisfaction of the user, the customer service animation can show an apologetic expression and make a soothing movement to make the user feel valued and understood.

[0099] In a remote meeting, using vivid character animation can make up for the lack of face-to-face communication in remote communication, and enhance the participation and authenticity of the meeting. Illustratively, in a remote meeting of a multinational company, participants may come from different cultural backgrounds, and the character animation generated by the technical solutions provided by the embodiments of the present application can be used as a unified communication image. In the communication process, the technical solutions provided by the embodiments of the present application can also adjust the expressions and body movements of the corresponding character animation in combination with the speech content and emotions of each participant. For example, when making important suggestions, let the animation show a confident expression and a powerful gesture, so that other participants are more likely to understand and accept their own views.

[0100] The character animation generation method provided by the embodiments of the present application can be applied to a computer device, which can be a terminal or a server. The terminal can be any electronic product that can interact with the user through one or more of a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device, such as a personal computer (PC), a mobile phone, a smart phone, a personal digital assistant (PDA), a wearable device, a pocket PC (PPC), a tablet computer, a smart car machine, a smart television, a smart speaker, etc. The server can be a standalone server, a server cluster composed of multiple physical servers, or a distributed system, and can also be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc. Basic cloud computing services, or a cloud computing service center.

[0101] When the character animation generation method is applied to the server, the server can communicate with a specific user terminal, so that the user can send relevant information, instructions, or data, etc. to the server through the user terminal. When the character animation generation method is applied to the terminal, the user can input relevant information, instructions, or data, etc. through the display interface (or user interaction interface) of the terminal.

[0102] In some embodiments, if the technical solutions provided by the present application are divided into an offline asset construction stage and an online animation generation stage, the two stages can be executed in the same computer device or in different computer devices, and the embodiments of the present application do not limit this.

[0103] As an example, considering that the offline asset construction stage involves a significant amount of computation, and the online animation generation stage is related to the user's call scenario, the cloud platform can execute the offline asset construction stage scheme in this application embodiment to build a set of character animations tailored to the needs of different users. The user terminal then executes the online animation generation stage scheme in this application embodiment to call and optimize corresponding character animations from the character animation set, based on the call information of the target character in a specific call scenario, to obtain target character animations that match the expressions and / or actions in the call scenario and call information.

[0104] The cloud platform can communicate with multiple user terminals to build corresponding character animation sets for multiple users. This embodiment of the application does not limit the number of user terminals or the size of the constructed character animation sets; that is, the cloud platform can build a character animation set for multiple users, or it can build a corresponding character animation set for each user.

[0105] In this way, not only can the computing resources of the cloud platform be fully utilized to quickly and effectively construct the combination of character animation, but the quality of the final generated target character animation can also be guaranteed through end-to-end cloud collaboration in the call scenario. This not only takes into account the real-time requirements in the call scenario, but also improves the quality of character animation and user satisfaction.

[0106] As another example, the technical solution provided in this application embodiment can also be executed by a single user terminal. That is, the user terminal executes the offline asset construction phase, and then executes the online animation generation phase in the scenario of real-time user communication.

[0107] It should be noted that the application scenarios and implementation environments described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the emergence of new application scenarios and the evolution of implementation environments, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0108] Next, taking the offline asset construction stage as an example, we will explain in detail the character animation generation method provided in this application embodiment.

[0109] See Figure 1 , Figure 1 This is a flowchart illustrating a method for generating character animation according to an embodiment of this application. The method can be applied to a first computer device and includes the following steps.

[0110] Step 101: Obtain character dynamic information and a character picture of a target character, the character dynamic information including expression information and / or body action information, the expression information being used to reflect expression changes of the character, and the body action information being used to reflect body action changes of the character.

[0111] Optionally, the expression information includes expression description text and / or an expression video of the first character. That is, the expression information can be one of the expression description text and the expression video, or can include both the expression description text and the expression video.

[0112] The expression description text can be simple expression description words such as smiling, frowning, opening eyes, closing eyes, etc., can be complex expression combination description words such as surprised and smiling (for example, the eyes of the main character in the animation are suddenly widened, the mouth is slightly opened, the corners of the mouth are uplifted, and an expression of both surprise and joy is formed), angry with a little disdain (for example, the eyebrows of the main character are raised, the eyes are widened, the teeth are clenched to show anger, and the corners of the mouth are slightly tilted downward, and the eyes show disdain), etc., or can be a text describing an expression corresponding to a certain emotion, such as "a bright smile on the face, a joyful light in the eyes, the corners of the mouth uplifted, and the whole person exudes a happy atmosphere", "dark eyes, drooping mouth corners, tear marks on the cheeks, and the whole person exudes a lost and sad atmosphere", etc.

[0113] In a possible implementation, the first character and the target character are different characters, for example, the first character is a character in an animation, and the target character is a real person; or the first character and the target character can also be the same character in different images, for example, the first character and the target character are the same girl, but the first character is in a fashionable image with makeup and dress, and the expression is happy and laughing; and the target character is in a simple image without makeup and dress, and the face is expressionless.

[0114] When the first character and the target character are the same character in different images, or different characters, the technical solution of the embodiment of the present application can copy the expressions of other characters or other images to the face of the target character, and support flexible adjustment of the expression effect. Based on this, the technical solution of the embodiment of the present application can also be extended to any personalized character animation generation scene, for example, photo beautification, to achieve the effect of fine adjustment of the expression of the character in the photo.

[0115] In this way, based on these rich and personalized expression description texts and / or expression videos, the technical solution of the embodiment of the present application can generate character animations with more lively and diversified expressions.

[0116] Similarly, the body action information includes body action description text and / or a body action video of the second character. That is, the body action description information can be one of the body action description text and the body action video, or simultaneously include the body action description text and the body action video.

[0117] The body action description text can be a simple action description word such as waving, stamping feet, clapping, bending and bowing, or a complex body action combination description word such as jumping and cheering, or a description of a series of actions such as an action description text of an athlete or a dance action description text.

[0118] In a possible implementation, the second character and the target character are different characters, for example, the second character is a professional athlete and the target character is a student learning the sport, or the second character and the target character are the same character in different images, for example, the second character and the target character are the same boy, but the second character is in his teenage years and is lively and active, and the target character is in his youth and is stable and powerful.

[0119] When the second character and the target character are the same character in different images or different characters, the technical solution of the embodiment of the present application can copy the actions of other characters or other images to the body of the target character, and support flexible adjustment of the effect of body action. Based on this, the technical solution of the embodiment of the present application can also be extended to any personalized character animation generation scene, for example, in dynamic picture generation, to achieve the effect of fine adjustment of the action of the character in the photo.

[0120] In this way, based on these rich and personalized body action description texts and / or body action videos, the technical solution of the embodiment of the present application can generate more lively and diversified character animations.

[0121] It should be understood that, for the character picture of the target character, when the expression information and the body action information are both videos, the first character and the second character are the same character, or the same character in different images, or completely different characters. The embodiment of the present application does not limit this.

[0122] In an embodiment of the present application, the character picture of the target character at least includes a face of the target character. Optionally, the character picture of the target character further includes a limb of the target character, which includes one or more of a head, a torso or a limb. For example, the character picture of the target character is a full-body picture of the target character. In the case where the character picture only includes a face, the first computer device can drive the character picture with the expression information. In the case where the character picture includes a face and a limb, the first computer device can drive the character picture with the expression information and / or drive the character picture with the limb action information.

[0123] Optionally, the target character is a real person or a virtual character (e.g., a digital person). Correspondingly, the character picture of the target character can be a real character picture taken by a camera or an image created by a computer technology, and the present application does not limit the character picture.

[0124] In a possible implementation, the process of obtaining the character dynamic information in step 101 is that the computer device obtains the input character dynamic information. That is, the character dynamic information can be provided by a user.

[0125] As an example, the first computer device includes a display interface for interaction with the user, and the user can input the character dynamic information and the character picture of the target character through the display interface.

[0126] In another possible implementation, the process of obtaining the character dynamic information in step 101 is that the first computer device displays a dynamic information selection interface, and the dynamic information selection interface displays a plurality of preset character dynamic information. In response to detecting a selection operation on any one of the plurality of preset character dynamic information, the first computer device obtains the character dynamic information, that is, the preset character dynamic information selected by the selection operation is taken as the character dynamic information in step 101.

[0127] For example, the selection operation is to first click to select a certain character dynamic information, and then click a confirmation button, or double-click to confirm the selected certain character dynamic information, and the present application does not limit the implementation of the selection operation.

[0128] Optionally, the database includes a preset expression material library and / or a preset limb action material library. The expression material library includes a plurality of preset expression information, and the plurality of preset expression information includes text and video; similarly, the limb action material library includes a plurality of preset limb action information, and the plurality of preset limb action information includes text and video.

[0129] In actual application, the first computer device can present the expression material library and / or the body action material library provided by the database to the user, and the user can select appropriate expression material and / or body action material therefrom as the character dynamic information.

[0130] Optionally, considering that the expression material library and / or the body action material library accumulates limited material, and there can be a problem of not timely updating the material, when performing the step 101, the first computer device also supports the user uploading the character dynamic information by himself / herself, and / or supports the user selecting the character dynamic information from the preset character dynamic information provided by the database and performing secondary editing and optimization on the character dynamic information to obtain the character dynamic information required by the user.

[0131] As an example, Figure 2 is a display interface schematic diagram provided by an embodiment of the present application. As Figure 2 shown, the display interface includes a picture uploading control A, a dynamic information uploading control B, and a dynamic information selecting control C. The picture uploading control A is used for the user to upload a character picture. The dynamic information uploading control B is used for the user to upload character dynamic information. The dynamic information selecting control C is used for the user to select from multiple preset character dynamic information. The user can upload the character dynamic information by triggering the dynamic information uploading control B according to his / her own needs, or can select the character dynamic information required by himself / herself from the multiple preset character dynamic information by triggering the dynamic information selecting control C.

[0132] If the user neither uploads the character dynamic information by the dynamic information uploading control B nor selects the character dynamic information from the multiple preset character dynamic information, the first computer device can use the character dynamic information selected by the user last time for generating the character animation as the character dynamic information for this time by default.

[0133] It can be seen that, when generating the character animation, the embodiment of the present application only needs the user to provide a single character picture, and then can select the character dynamic information to arrange the expression and / or body action of the target character in the single picture, so as to realize the controllability and editability of the user on the character expression and / or body action. Compared with the one-to-many situation of the pose information and expression information extracted from the audio information in the related art, the character picture and the character dynamic information in the embodiment of the present application are both provided by the user, which can meet the actual needs of the user. Moreover, the embodiment of the present application is based on the single character picture provided by the user to perform the subsequent animation generation steps, and does not depend on the preset character image or preset video, so the flexibility is better, and the embodiment of the present application can meet the personalized character animation generation needs of different users.

[0134] Step 102: driving the character picture to a first character animation according to the character dynamic information, and displaying the first character animation, wherein the expression change of the target character in the first character animation is consistent with the expression change reflected by the expression information, and / or the body action change of the target character in the first character animation is consistent with the body action change reflected by the body action information.

[0135] The expression change of the target character in the first character animation is consistent with the expression change reflected by the expression information, including that the expression category of the target character in the first character animation is the same as the expression category reflected by the expression information. In the case that the expression category reflected by the expression information has multiple expression categories, the expression change of the target character in the first character animation is consistent with the expression change reflected by the expression information, further including that the transition order between different expression categories of the target character in the first character animation is consistent with the transition order of the expression categories reflected by the expression information.

[0136] As an example, assuming that the expression change reflected by the expression information in the character dynamic information is “first smile and then cry”, if the expression change of the target character in the first character animation is “first smile and then cry”, it means that the two are consistent; if the expression change of the target character in the first character animation is “first cry and then smile”, “always cry” or “always smile”, it means that the two are not consistent.

[0137] Similarly, the body action change of the target character in the first character animation is consistent with the body action change reflected by the body action information, including that the body action type of the target character in the first character animation is the same as the body action type reflected by the expression information. In the case that the body action type reflected by the body action information has multiple body action types, the body action change of the target character in the first character animation is consistent with the body action change reflected by the body action information, further including that the switching order between different body action types of the target character in the first character animation is consistent with the body action switching order reflected by the body action information.

[0138] As an example, assuming that the body action change reflected by the body action information in the character dynamic information is “first jump and then squat”, if the body action change of the target character in the first character animation is “first jump and then squat”, it means that the two are consistent; if the body action change of the target character in the first character animation is “jump and then stand”, “squat and then stand” or “first squat and then jump”, it means that the two are not consistent.

[0139] Since the character dynamic information includes the expression information and / or the body action information, the above step 102 of driving the character picture to the first character animation based on the character dynamic information corresponds to three cases, which will be introduced respectively.

[0140] In the first case, the character dynamic information includes expression information. In this case, the implementation process of driving the character picture to the first character animation according to the character dynamic information is as follows: driving the face of the character picture according to the expression information to obtain an expression animation, in which the expression change of the target character is consistent with the expression change reflected by the expression information. In the first character animation, the expression change of the target character adopts the expression animation.

[0141] That is, based on the expression information, the face of the target character in the originally static character picture is "moved" to show the dynamic change consistent with the expression information, and this dynamic effect is used for the expression presentation of the target character in the first character animation, so as to ensure that the expression change process of the target character in the first character animation is consistent with the change process shown by the expression information.

[0142] In a possible implementation, the implementation process of driving the face of the target character in the character picture according to the expression information includes the following steps 11-13:

[0143] Step 11: driving the face of the character picture according to the expression information to obtain a plurality of expression segments.

[0144] In a possible implementation, the implementation process of step 11 can be as follows: first, feature extraction and analysis are performed on the face of the character picture, for example, the shapes, positions and postures of key parts such as eyes, eyebrows and mouth are recognized. Then, the key parts are adjusted according to the expression information. If the expression information is "laughing happily", the mouth will be opened, the corners of the mouth will be raised, and the eyes will be narrowed into a curved shape. Through a series of mathematical operations and algorithm processing, the face of the character picture gradually presents an expression state consistent with the expression information, and this process is dynamic, as if a real person is making this expression, so as to generate an expression animation.

[0145] In the process of expression driving, the changes of the key parts of the face at different time points are recorded, and these changes are combined in a certain time sequence to form an expression animation. The expression animation can clearly show the whole process of the expression of the target character changing from the initial state to the state consistent with the expression information.

[0146] In another possible implementation, the implementation process of step 11 can be as follows: performing face detection on the character picture of the target character, and cropping the detected face to obtain a face image; inputting the face image and the expression information into a pre-trained expression animation generation model to obtain an expression sequence output by the expression animation generation model, that is, a plurality of expression segments.

[0147] The expression animation generation model can be trained based on an expression training sample in a supervised learning manner. The expression training sample can include a sample face image and sample expression information. Embodiments of the present application do not limit the construction / generation manner of the expression animation generation model, and aim to emphasize that the expression animation generation model can generate an expression sequence based on a face image and expression information.

[0148] As an example, in the case of expression information being an expression video, the expression animation generation model can be implemented based on a video driving framework, that is, after inputting the expression video and the face image into the expression animation generation model, the expression animation generation model drives the face image based on the expression video, thereby outputting an expression sequence.

[0149] As an example, in the case of expression information being an expression description text, the expression animation generation model can be trained based on an Arkit2Video model. In this case, the expression coefficients corresponding to the expression description text are first generated, and then the expression coefficients and the face image are input into the expression animation generation model to obtain the expression sequence output by the expression animation generation model. The expression coefficients are defined coefficients for describing facial expressions, such as an open-eye state represented as A and a closed-eye state represented as B, and (0.5A+0.5B) represents a half-open-eye state. Alternatively, the expression animation generation model can be trained based on a Text2Video model, in which case the expression description text and the face image can be directly input into the expression animation generation model to obtain the expression sequence output by the expression animation generation model.

[0150] It should be noted that embodiments of the present application do not limit the implementation manner of driving a single picture to obtain an expression animation using an expression video or an expression description text.

[0151] As an example, assuming that there is a front picture of a target person, and the expression of the target person in the picture is calm. The given expression information is a description text of “experience wide mouth and round eyes”, and the process of generating an expression animation (i.e., multiple expression segments) through expression driving can be as follows: first, the features of the eyes and mouth of the person in the picture are analyzed through face recognition technology to determine their positions, sizes and shapes. Then, according to the expression information, the outline of the mouth is expanded to show a wide-open state, and the pupils of the eyes are enlarged, and the distance between the upper and lower eyelids is increased to simulate the effect of round eyes. During the adjustment process, the changes of the mouth and eyes at different time points are gradually recorded, such as from 0 seconds, the mouth starts to slowly open, to 1 second, the mouth opens to a certain extent, and the eyes also start to gradually enlarge, to 2 seconds, the expression completely shows a surprised expression. Combining the changes of the facial features at different time points, an expression animation with a duration of 2 seconds is generated, which shows the expression change process of the target person from calm to surprise.

[0152] As another example, such as Figure 3 As shown, assuming the target person's image is a frontal photograph with a calm expression, and the given expression information is a video of the first person's expression, in which the first person gradually smiles with their eyes open, then their eyes crinkle into a smile, and finally return to an open-eyed smile, the process of generating facial animation (i.e., multiple expression segments) through expression-driven animation can be as follows: Using video analysis technology, extract feature data of key facial features (eyes, mouth, eyebrows, etc.) of the first person in each frame of the video. Then, match and compare these feature data with the facial features of the target person in the image. Next, according to the time sequence of each frame in the video, gradually apply the changes in the first person's facial features to the target person's face. For example, if the first person's eyes are fully open in the first frame of the video, then adjust the target person's mouth to a slightly upturned state; if the first person's eyes crinkle into a smile in the second frame of the video, i.e., the eyes are half-open or even closed, then adjust the target person's mouth and eyes accordingly. In this way, by integrating the facial feature changes of all frames, an facial animation that matches the facial expression changes in the video is generated, showing the process of the target person from opening their eyes and smiling - eyes curving into a smile (half-closed state) - gradually opening their eyes.

[0153] It should be understood that, Figure 3 Taking the first person's facial expression video as an example, which includes 3 frames, the actual process may include more frames. This application embodiment does not limit the number of frames or the duration of the facial expression video.

[0154] Step 12: Sort the multiple expression fragments according to the first sorting instruction for the multiple expression fragments to obtain the sorted expression fragments.

[0155] In one possible implementation, step 12 can be performed by: displaying a first sorting interface, which displays multiple emoticon fragments generated in step 11. In response to receiving a first sorting instruction through the first sorting interface, the multiple emoticon fragments are sorted according to the first sorting instruction.

[0156] The first sorting instruction is used to indicate the arrangement order of the multiple facial expression fragments. This arrangement order can be determined based on different needs, such as chronological order, emotional change logic, priority, etc., and this application embodiment does not limit this.

[0157] In the embodiments of the present application, the first computer device realizes the user's custom arrangement of the plurality of expression segments by providing a user interaction interface. The user can arrange the plurality of expression segments in a drag-and-drop manner, for example, arranging the plurality of expression segments in a left-to-right order, where the leftward position of an expression segment indicates a higher arrangement order. Alternatively, the user can input a serial number for each expression segment, such as a numerical serial number "1, 2, 3" or an alphabetical serial number "a, b, c". Alternatively, the user can click the plurality of expression segments in sequence, and the arrangement order of the plurality of expression segments is the same as the clicking order.

[0158] Step 13: If the last frame of the first expression segment corresponds to a first expression label that is different from a second expression label corresponding to the first frame of the second expression segment, an expression transition frame is inserted between the first expression segment and the second expression segment.

[0159] The first expression segment is adjacent to the second expression segment, and the first expression segment is before the second expression segment. The expression transition frame is used to smooth the expression transition from the first expression segment to the second expression segment. The expression transition frame inserted between the first expression segment and the second expression segment can be one image or a plurality of images.

[0160] As an example, the expression in the expression transition frame can be a neutral expression. Assuming that the last frame of the first expression segment corresponds to an expression label "happy" and the first frame of the second expression segment corresponds to an expression label "sad", one or more images of a "calm face" can be inserted between the first expression segment and the second expression segment as the expression transition frame.

[0161] As another example, the expression transition frame can also be generated by a linear interpolation method, a motion atlas method, a deep learning-based generative adversarial network, or a physics-based simulation method based on the last frame of the first expression segment and the first frame of the second expression segment.

[0162] In the embodiments of the present application, for the plurality of expression segments constituting the expression animation, a natural and smooth expression transition frame is generated between adjacent expression segments, realizing smooth transition of the expression in the previous expression segment to the expression in the next expression segment, which can prevent expression jump and thus affect the expression animation effect, and ensure the continuity and realism of the expression animation.

[0163] Optionally, after obtaining the expression animation, the first computer device can further display an expression animation playing interface for playing the expression animation.

[0164] In the second case, the character dynamic information includes limb action information. In this case, the implementation process of driving the character picture to the first character animation according to the character dynamic information is as follows: limb driving is performed on the limbs in the character picture according to the limb action information, and a limb animation is obtained, in which the limb action change of the target character is consistent with the limb action change reflected by the limb action information. The limb action change of the target character in the first character animation adopts the limb animation.

[0165] That is, based on the limb action information, the limbs of the target character in the originally static character picture are "moved" to exhibit dynamic changes consistent with the limb action information, and this dynamic effect is used for the limb action presentation of the target character in the first character animation, so as to ensure that the limb action of the target character in the first character animation is consistent with the action change exhibited by the limb action information.

[0166] In a possible implementation, the implementation process of performing limb driving on the limbs of the target character in the character picture according to the limb action information includes the following steps 21-23:

[0167] Step 21: Perform limb driving on the limbs in the character picture according to the limb action information, and obtain a plurality of limb action segments.

[0168] In a possible implementation process, the implementation process of step 21 can be as follows: first, feature extraction and analysis are performed on the limbs of the character in the character picture. The positions and angles of the key joints of the character limbs, such as the shoulder, elbow, wrist, hip, knee, and ankle, are identified. Then, according to the limb action information, the key joints are adjusted. If the limb action information is "raise the right arm, bend the elbow, and stretch the palm forward", the positions and angles of the right shoulder, right elbow, and right wrist in the character picture will be adjusted accordingly to simulate the action of raising the right arm and bending the elbow.

[0169] In the limb driving process, the changes of the key joints of the character limbs at different time points are recorded. Combining these changes in chronological order forms a limb animation. The limb animation can clearly show the whole process of the gradual change of the limbs of the target character from the initial state to the state consistent with the description of the limb action information.

[0170] In another possible implementation, the implementation process of step 21 can be as follows: performing limb detection on the character picture of the target character to obtain a limb image; inputting the limb image and the limb action information into a pre-trained limb animation generation model to obtain a limb action sequence output by the limb animation generation model, that is, a plurality of limb action segments.

[0171] The body animation generation model can be trained based on a body action training sample by using a supervised learning manner. The body training sample can include a sample body image and sample body action information. Embodiments of the present application do not limit the construction / generation manner of the body animation generation model, and aim to emphasize that the body animation generation model can generate a body action sequence based on a body image and body action information.

[0172] As an example, assuming that there is a front picture of a target person standing still, and the given body action information is a description text of "taking a left foot forward, slightly crouching the body, and stretching the hands forward and making a fist". The process of generating a body animation (i.e., multiple body action segments) by body driving is as follows: first, the positions of the key joints of the body of the person in the picture are analyzed by using a human pose estimation technology to determine the states of the left foot, the hip, the knee, and the hands. Then, the key joints are adjusted according to the body action description text. The left foot is moved forward by a certain distance, the hip and the knee are bent downward to make the body slightly crouched, and the hands are stretched forward and the palms are made to be in a fist shape. During the adjustment process, the changes of the key joints at different time points are recorded, such as from 0 seconds, the left foot starts to move forward, the body gradually crouches, and the hands start to stretch forward, to 1 second, the whole action is basically completed. Combining the body feature changes at different time points, a body animation with a duration of 1 second is generated, which shows the action change process of the target person from standing still to taking a left foot forward, crouching, and stretching hands to make a fist.

[0173] As another example, assuming that there is also a front picture of a target person standing still, and the given body action information is a video of an athlete (i.e., a second person in the present application) shooting a basket in a basketball game. The process of generating a body animation (i.e., multiple body action segments) by body driving is as follows: first, the motion trajectory data of the key joints of the athlete's body in the video are extracted by using motion capture technology, including the positions and angle changes of the shoulder, the elbow, the wrist, the hip, the knee, and the ankle at different frames. Then, the motion trajectory data are matched and compared with the key joints of the body of the target person in the picture. Next, the motion changes of the key joints of the athlete's body are gradually applied to the body of the target person according to the time sequence of each frame in the action video. For example, in the first frame of the video, the athlete only slightly bends the knee to prepare for jumping, so the knee of the target person is also adjusted to a slightly bent state; in the fifth frame of the video, the athlete has already jumped and stretched the arm to prepare for shooting, so the positions and angles of the knee, the hip, the shoulder, the elbow, and the wrist of the target person are adjusted accordingly. Integrating the body feature changes of all frames, a body animation consistent with the shooting action in the video is generated, which shows the whole process of the target person from standing to jumping and shooting.

[0174] The process of driving the character picture based on the body action video can refer to the above Figure 3 The process of driving the character picture based on the expression video is to drive the character picture based on the body action video to obtain the body animation.

[0175] Step 22: According to the second sorting instruction for the plurality of body action segments, the plurality of body action segments are sorted to obtain sorted body action segments.

[0176] In a possible implementation, the implementation process of step 22 can be: a second sorting interface is displayed, and the plurality of body action segments are displayed in the second sorting interface. According to the second sorting instruction, the plurality of body action segments are sorted in response to receiving the second sorting instruction through the second sorting interface.

[0177] The second sorting instruction is used to indicate the arrangement order of the plurality of body action segments. The arrangement order can be determined based on different requirements, such as time sequence, action change logic, action priority, etc., which is not limited in the embodiments of the present application.

[0178] In the embodiments of the present application, the first computer device realizes the user's self-arrangement of the plurality of body action segments by providing a user interaction interface. The way in which the user arranges the plurality of body action segments can refer to the way of arranging the plurality of expression segments in the first possible case, which will not be repeated here.

[0179] Step 23: If the last frame of the first body action segment in the sorted body action segment corresponds to a first body action different from a second body action corresponding to the first frame of the second body action segment, a body transition frame is inserted between the first body action segment and the second body action segment.

[0180] The first body action segment and the second body action segment are adjacent, and the first body action segment is before the second body action segment. The body transition frame is used to smooth the body action transition process from the first body action segment to the second body action segment. The body transition frame inserted between the first body action segment and the second body action segment can be one frame of image or can include multiple frames of image.

[0181] As an example, the character limb action in the limb transition frame can be an idle state determined based on the plurality of limb action clips. For example, the limb action with the highest frequency of occurrence in the plurality of limb action clips is taken as the idle state, or the idle state can also be determined through other algorithms. The idle state can be understood as a behavior mode of the character or object when waiting or stationary, such as posture switching, standing, or squatting, etc. The above-mentioned limb transition frame can also be referred to as a reset frame.

[0182] In the embodiments of the present application, for the plurality of limb action clips constituting the limb animation, the limb transition frame with smooth and natural action connection is generated between adjacent limb action clips, so as to realize smooth transition of the limb action in the previous limb action clip to the limb action in the next limb action clip, prevent the occurrence of limb action jump and thus affect the limb animation effect, and ensure the coherence and realism of the limb animation.

[0183] Optionally, after obtaining the limb animation, the first computer device can further display a limb animation playing interface, the limb animation playing interface being used for playing the limb animation.

[0184] In the third case, the character dynamic information includes expression information and limb action information. In this case, the implementation process of driving the character picture to the first character animation according to the character dynamic information is as follows: driving the face in the character picture according to the expression information to obtain an expression animation, the expression change of the target character in the expression animation being consistent with the character expression change reflected by the expression information; driving the limbs in the character picture according to the limb action information to obtain a limb animation, the limb action change of the target character in the limb animation being consistent with the limb action change reflected by the limb action information; and fusing the expression animation and the limb animation to obtain the first character animation.

[0185] The implementation process of generating the expression animation can refer to the first case described above, and the implementation process of generating the limb animation can refer to the second case described above, which will not be described herein again.

[0186] The expression change of the target character in the first character animation adopts the expression animation, and the limb action change of the target character in the first character animation adopts the limb animation. Based on this, the expression animation and the limb animation with the same number of frames are fused, and the first character animation is obtained.

[0187] For the third case described above, the generation process of the expression animation and the generation process of the limb animation can be decoupled, so that the user can independently edit the character expression and the character limb action, or can only edit and correct one of the expression and the limb action through subsequent steps 103-104, which can save time for the user.

[0188] Based on the above three cases, the first character animation generated by the embodiment of the application can be an animation (video) with dynamically changing expressions, an animation (video) with dynamically changing body movements, or an animation (video) with dynamically changing expressions and body movements.

[0189] Optionally, the first character animation can be saved as a reference character animation (or a preset character animation) and added to the character animation set of the target character, so that the target character can use it for voice-driven use in a call scene based on specific call information.

[0190] Step 103: In response to the correction operation on the first character animation, obtain correction information, which includes correction expression information and / or correction body movement information.

[0191] As an example, the correction information is user modification instruction information for the expression and / or body movement of a frame in the first character animation, or user modification instruction information for the expression and / or body movement of multiple frames in the first character animation.

[0192] As another example, the correction information is preset character dynamic information that is re-uploaded or re-selected by the user from the database. The correction expression information is expression information that is re-uploaded or selected by the user, and the correction body movement information is body movement information that is re-uploaded or selected by the user.

[0193] Figure 4 is another display interface schematic diagram provided by the embodiment of the application. As shown in Figure 4 , the display interface is an expression animation playing interface, which includes an expression animation playing window, a confirm button, and a correction button. When detecting a triggering operation of the user on the play button in the expression animation playing window, the computer device plays the expression animation. When detecting a triggering operation of the user on the confirm button, the computer device determines that a confirm operation on the expression animation is detected. When detecting a triggering operation of the user on the correction button, the computer device determines that a correction operation on the expression animation is detected, and then the computer device can switch to the display interface shown in Figure 2 , in which the user re-uploads or selects expression information (character dynamic information).

[0194] It should be understood that the body animation playing interface can also refer to the expression animation playing interface shown in Figure 4 , and the embodiment of the application will not be described here again.

[0195] In this way, the embodiment of the application provides an expression animation preview and correction function, and / or a body animation preview and correction function, thereby improving the flexibility of the user in making expression animations, and thus improving the user experience.

[0196] Step 104: performing expression driving on the face of the target character in the first character animation according to the corrected expression information, and / or performing body driving on the body of the target character in the first character animation according to the corrected body action information, to obtain a second character animation.

[0197] The implementation process of performing expression driving on the face of the target character in the first character animation according to the corrected expression information can refer to the first case in step 102, and the implementation logic is similar. The implementation process of performing body driving on the body of the target character in the first character animation according to the corrected body action information can refer to the second case in step 102, and the implementation logic is similar. The implementation process can also refer to the existing implementation manner in the related art, and the embodiments of the present application do not limit this, and will not be repeated here.

[0198] It should be noted that the above steps 103 and 104 can be executed repeatedly until the second character animation that meets the user's satisfaction is generated.

[0199] Optionally, in the case where the second character animation is obtained, in response to detecting a determination operation on the second character animation, the second character animation is saved.

[0200] The saved second character animation can be added to the character animation set of the target character, so that the target character can use the second character animation for voice driving based on specific call information in a call scene.

[0201] In summary, the embodiments of the present application can drive a single picture of a target character to generate a character animation, and the expression and / or body action of the target character in the character animation are controllable and editable. Compared with the related art, the user can control and edit the expression and / or body action in the character animation, so that the expression and / or body action of the target character in the finally generated character animation can meet the personalized needs of the user, greatly improving the user's satisfaction with the animation effect. Since the user can finely control the expression and / or body action, the character animation can be more lively, natural, and more consistent with the actual situation and the user's expectations. Compared with the related art, the technical solution of the present application can generate high-quality and more expressive character animations, and improve the overall animation effect.

[0202] Next, the online animation generation stage is taken as an example to explain the character animation generation method provided by the embodiments of the present application in detail.

[0203] Figure 5is a flowchart of a method for generating a character animation provided by an embodiment of the present application. The method can be applied in a second computer device. It should be understood that the second computer device can be the same device as the first computer device or a different device. Referring to Figure 5 The method includes the following steps.

[0204] Step 501: Obtain communication information of a target character in a communication scenario.

[0205] The target character can be a real character, an animated character, or a digital person. The present application does not limit the target character.

[0206] As an example, in a human-computer interaction communication scenario, the target character can be an animated character provided by the second computer device, such as a cartoon image of a cat or a dog.

[0207] As an example, in a mobile communication scenario, such as a scenario of chatting with a friend using a video communication software, the target character can be a real character participating in the communication.

[0208] In the present application, the communication information of the target character can be audio data of the target character or text corresponding to the communication content of the target character. The present application does not limit the form of the communication information, which can be audio or text.

[0209] As an example, the communication information is collected speech, or synthesized speech based on text, or from a pre-stored audio material library. In different application scenarios, the source of the communication information is different, such as in a mobile communication scenario, the communication information is real-time speech collected.

[0210] Step 502: Determine an expression label and a body movement label according to the communication information.

[0211] In one possible implementation, the implementation process of step 502 can be: inputting the communication information into a pre-trained expression recognition model to determine an expression label corresponding to the communication information through the expression recognition model; inputting the communication information into a pre-trained body movement recognition model to obtain a body movement label output by the body movement recognition model through the body movement recognition model.

[0212] Optionally, the expression recognition model and the body action recognition model can be implemented by a classification model, such as a support vector machine (SVM), a decision tree, a random forest, a neural network (such as a convolutional neural network (CNN), a recurrent neural network (RNN), and variants thereof, such as a long short-term memory network (LSTM) and a gated recurrent unit (GRU)), and the like, to classify expressions. Among them, the SVM is suitable for small sample data and can find the optimal classification hyperplane; the neural network is suitable for large-scale data and can learn complex feature representations.

[0213] Optionally, considering that a body action is often a continuous process, the above-mentioned body action recognition model can also combine a sequence model (such as an LSTM) to process the time sequence of the body action, so as to better capture the association between actions.

[0214] Optionally, the model (i.e., the expression recognition model and the body action recognition) training process can be: obtaining a large amount of call sample information and labeling it to label the expression label and the body action label corresponding to each call sample information, for example, labeling the call sample information containing "happy laughter" content as the "happy" expression label, and labeling the call sample information describing "slapping the table" as the corresponding body action label. Further, the labeled data is divided into a training set, a validation set, and a test set. The training set is used to train the model, and the parameters of the model are adjusted so that the model can learn the features and rules in the data; the validation set is used to optimize the model, and the optimal model parameters are selected; the test set is used to evaluate the performance of the model, such as accuracy, recall rate, and the like. After the training is completed, the trained model is obtained.

[0215] As an example, the result output by the expression recognition model can be a probability value of each expression label (in a plurality of preset expression labels), and the label with the maximum probability value is selected as the final expression label. Similarly, the result output by the body action recognition model can be a probability value of each body action label (in a plurality of preset body action labels), and the label with the maximum probability value is selected as the final body action label.

[0216] Of course, a neural network model can also be used to simultaneously perform expression analysis and body action analysis to determine the expression label and the body action label corresponding to the call information, and the embodiments of the present application do not limit this.

[0217] Therefore, the embodiment of the present application determines the expression label and the body action label based on the call information, so that the determined label is more comprehensive, and the subsequent screening of the initial character animation is facilitated.

[0218] It should be noted that although the present application determines the expression label and the body action label in combination with the call information, one label or both labels can be combined for screening when the initial character animation is screened in the following step 503. For example, all the character animations in the character animation set are in the form of "mugshot", that is, the character in the character animation only has expression dynamic change, and at this time, the screening can be performed only in combination with the expression label.

[0219] Step 503: determining an initial character animation from a character animation set according to the expression label and / or the body action label, the character animation set including at least one character animation corresponding to the target character.

[0220] Correspondingly, the character in the at least one character animation corresponding to the target character can be an image of a real character or an image of an animated character.

[0221] When the character animation set includes one character animation corresponding to the target character, the character animation can be directly determined as the initial character animation; or the character animation can be secondarily optimized in combination with the expression label and / or the body action label, and the optimized character animation can be determined as the initial character animation; in the case that the time delay is allowed, the corresponding initial character animation can also be generated in real time based on the expression label, the body action label and the call information.

[0222] It should be noted that the implementation process of generating the corresponding initial character animation in real time based on the expression label, the body action label and the call information can refer to the above-mentioned embodiments, and the implementation logic is similar, which will not be described herein. Figure 1

[0223] When the character animation set includes multiple character animations corresponding to the target character, each character animation can also carry at least one expression label and / or at least one body action label. In this case, one character animation with the highest label similarity can be selected from the multiple character animations by matching the labels based on the expression label and / or the body action label corresponding to the call information, and the character animation is determined as the initial character animation.

[0224] When there are multiple character animations with the same expression label and body action label corresponding to the call information, further screening can be performed according to other factors, such as the fluency of the character animation, the style and the fit degree of the current call scene, and the present embodiment does not limit this.

[0225] ​In some embodiments, considering that the initial character animation itself can include one or more frames of images, the above steps 502 and 503 can also be periodically performed. For example, the following step 504 is performed based on the conversation information every 30 seconds, calling an initial character animation.

[0226] As an example, in a scenario where a user is having a human-computer conversation with a cartoon character (i.e., a target character), suppose that based on the conversation information, it is determined that the expression label of the target character is "happy" and the body action label is "sitting". Suppose that the cartoon character corresponds to the following five character animations: animation 1: the character expression is "calm" and the body action is "standing"; animation 2: the character expression is "happy" and the body action is "jumping"; animation 3: the character expression is "happy" and the body action is "clapping"; animation 4: the character expression is "sad" and the body action is "sitting"; and animation 5: the character expression is "angry" and the body action is "stomping". In this case, the implementation process of determining the initial character animation is as follows: if only the expression label "happy" is included, then the multiple character animations corresponding to the target character are traversed, and it is found that the character expressions of animations 2 and 3 are both "happy". Further considering other factors (e.g., the clapping action of animation 3 can be more suitable for the current scenario), it is determined that animation 3 is the initial character animation. If only the body action label "sitting" is included, then the body actions of each animation are checked, and it is found that only the body action of animation 4 is "sitting", so it is determined that animation 4 is the initial character animation. If both the expression label "happy" and the body action label "clapping" are included, then the expression and body action of each animation are compared, and only animation 3 satisfies both conditions, so it is determined that animation 3 is the initial character animation.

[0227] Step 504: Adjusting the mouth action of the initial character animation according to the conversation information to obtain a target character animation, in which the mouth action of the target character matches the conversation information.

[0228] In the character animation, the mouth action of the character matches the audio information, which means that the mouth action of the character in the character animation is the same as the mouth action when the character speaks the content of the audio information.

[0229] Optionally, the second computer device mainly relies on deep learning and computer vision technology to implement the mouth action adjustment of the first character animation according to the conversation information. The second computer device can first detect and crop the face in the first character animation, further perform a mask operation on the lower half of the face (including the mouth region), and randomly select a face without mask as a reference frame (the face part without mask is used to provide appearance and pose information to ensure better generation effect). Wherein, the face with mask and the face as reference frame are used as face input information together. Secondly, the conversation information is cut, sampled and feature extracted, and the matching mouth shape information is obtained according to the extracted features. Finally, the face input information and the mouth shape information are input into the pre-trained generation network together to perform mouth synthesis, and the second character animation is obtained. The specific implementation technology for adjusting the mouth action of the character animation according to the conversation information is not limited in the embodiments of the application.

[0230] In summary, the embodiments of the application determine the initial character animation from the character animation set according to the conversation information of the target character in the conversation scene, adjust the mouth action of the initial character animation based on the conversation information, and obtain the target character animation. Since the initial character animation is pre-constructed based on the requirements of the target character, the character animation meets the basic requirements of the expression and / or limb action of the target character. For a specific conversation scene, the target character animation that meets the dual requirements of the conversation scene and the conversation information can be obtained by generating and adjusting the mouth action based on the initial character animation.

[0231] The order of the steps of the above-mentioned character animation generation method provided by the embodiments of the application can be appropriately adjusted, and the steps can also be appropriately increased or decreased according to the situation. Any person skilled in the art can easily think of changes within the technical range disclosed in the application, which should be covered within the protection scope of the application.

[0232] Based on the above Figures 1-5 The technical solutions shown are for easy understanding. Next, several exemplary embodiments of the character animation generation method provided by the application are explained and described taking the conversation scene as an example.

[0233] Figure 6 is a schematic diagram of the architecture of a character animation generation system provided by the embodiments of the application. As Figure 6 shown, the system architecture includes an input unit, a functional unit and an output unit.

[0234] The input unit interacts with the user to input (or obtain) the conversation information (such as audio information), character picture, expression information (optional) and limb action information (optional) from the user.

[0235] The expression information and the body action information are optional, and if the user does not upload the information, the related materials can be obtained from the database unit for modification and editing.

[0236] In a possible implementation, the functional units include, but are not limited to, an identification unit, a voice processing unit, an expression processing unit, a body action processing unit, and a voice driving unit.

[0237] When creating a character animation (or constructing a character animation set), the expression processing unit is configured to perform face detection, expression generation (for example, generating a plurality of expression segments in step 11 described above), segment splicing (for example, sorting a plurality of expression segments in step 12 described above), and smooth reorganization (for example, inserting an expression transition frame between two adjacent expression segments in step 13 described above). The body action processing unit is configured to perform reset frame selection (for example, selecting a body transition frame in step 23 described above), body action generation (for example, generating a plurality of body action segments in step 21 described above), segment splicing (for example, sorting a plurality of body action segments in step 22 described above), and smooth reorganization (for example, inserting a body transition frame between two adjacent body action segments in step 23 described above).

[0238] When the character animation needs to be voice driven, the identification unit is configured to identify and analyze the call information of the call scene to determine the expression label and / or the body action label corresponding to the call information. The voice processing unit is configured to perform audio cropping and feature extraction on the audio information. The voice driving unit is configured to drive the corresponding character animation (for example, the target character animation in step 504 described above) based on the audio information of the user.

[0239] The output unit is configured to output the character animation corresponding to the character picture (for example, a silent animation, that is, the mouth shape is not processed, but the expression and body action are edited by using the scheme of the present application), and can also output the target character animation (video) driven by the call information.

[0240] Optionally, please continue to refer to Figure 6 The system architecture further includes a database unit. The database unit is configured to provide an expression material library and a body action material library, so that the user can select and modify the corresponding materials to obtain the character dynamic information.

[0241] It should be understood that the above Figure 6 The specific content of the input unit, the functional unit, and the output unit shown in the figure is only an example, and more or fewer units can be included in actual applications, and the embodiments of the present application do not limit this.

[0242] Next, taking a call scene as an example, the possible implementation process of the technical scheme of the present application is exemplified.

[0243] Figure 7 is a schematic diagram of a character animation generation process provided by an embodiment of the present application. As shown in the figure, the production of a character animation (i.e., an asset) is independently completed by a user terminal, and in a process of real-time voice communication, the pre-produced character animation is driven by language in combination with specific communication information to obtain a target character animation. Figure 7

[0244] In the character animation production stage, the user terminal side provides a character picture, expression information (e.g., an expression video), and body movement information (e.g., a body movement video). An expression animation is generated according to the character picture and the expression information, and a body animation is generated according to the character picture and the body movement information. The generated expression animation, body animation, or character animation obtained by fusing the expression animation and the body animation (e.g., the first character animation shown in the foregoing embodiment) can be fed back to the user, and the user determines whether the generated animation meets the user's own needs. If the user is satisfied with the effect of the expression animation, the expression animation is used for character animation generation. If the user is not satisfied with the effect of the expression animation, an expression animation is generated according to the character picture and the re-provided expression information on the user terminal side, or the expression of the character in the expression animation is adjusted according to the expression update information provided by the user until the user is satisfied. The expression animation finally determined by the user (e.g., the second character animation shown in the foregoing embodiment) is saved and added to a character animation set. Similarly, if the user is satisfied with the effect of the body animation, the body animation is used for character animation generation. If the user is not satisfied with the effect of the body animation, a body animation is generated according to the character picture and the re-provided body movement information on the user terminal side, or the expression of the character in the expression animation is adjusted according to the expression update information provided by the user until the user is satisfied. The body animation finally determined by the user (e.g., the second character animation shown in the foregoing embodiment) is saved and added to a character animation set. Alternatively, a character animation is generated according to the expression animation and the body animation, and the character animation production stage ends.

[0245] It should be understood that the above steps can be performed multiple times to generate multiple character animations for the user to meet the user's expression needs in different scenarios, thereby facilitating the flexible invocation of the corresponding character animation in a real-time voice communication scenario.

[0246] In the real-time voice communication process, the real-time voice of the user collected from the user terminal side is subjected to emotion recognition to determine the corresponding expression label and body movement label. Then, based on the expression label and / or the body movement label, an initial character animation is selected from the pre-produced character animation set, and the initial character animation and the real-time voice are input into a voice driving model, the mouth synthesis is performed by the voice driving model, and finally a target character animation whose mouth movement matches the real-time voice of the user is output.

[0247] ​It can be seen that the person animation generation method provided in the embodiments of the present application can run on the terminal side, and supports user-defined preset assets (i.e., person animation sets) generation, so that the user can edit expressions and / or limb actions according to different scenes, thereby achieving controllable expression and / or limb conversation effect during real-time voice conversation.

[0248] Figure 8 is another person animation generation process schematic diagram provided in the embodiments of the present application. As shown in Figure 8 based on the person animation set constructed in advance by the user, the initial person animation can be determined from the person animation set according to the scene identifier of the current conversation scene in the conversation scene, and then the initial person animation and the real-time conversation information are input into the voice-driven model, the mouth synthesis is performed by the voice-driven model, and finally the target person animation whose mouth action matches the real-time conversation information is output.

[0249] Optionally, in the case where the time delay is allowed, the user can edit the expressions and / or limb actions of the selected initial person animation.

[0250] As an example, the scene identifier can be scene number, scene type, etc. For example, the scene identifier can be scene 1 - daily greeting.

[0251] It can be seen that in the case where the person animation set is constructed in advance, the initial person animation consistent with the current conversation scene can also be called in the conversation scene combined with the scene identifier, and the initial person animation is driven in real time by the real-time voice.

[0252] It should be understood that, Figure 7 and Figure 8 Two bases for screening the initial person animation from the person animation set are provided, in actual application, one of the two ways can be used to determine the initial person animation, or the two can be combined to determine the initial person animation, the embodiments of the present application do not limit this.

[0253] Figure 9 is another person animation generation process schematic diagram provided in the embodiments of the present application. As shown in Figure 9As shown, in some human-computer interaction, or large model (for example, large language model) question and answer scenarios, the large language model generates a reply text for a question input by a user in audio or text information. The reply text is subjected to sentiment recognition to determine its corresponding expression label and body action label. Then, based on the expression label and / or body action label, an initial character animation is selected from a pre-made character animation set, and the initial character animation and the reply text are input into a speech-driven model, and the speech-driven model performs mouth synthesis, and finally outputs a target character animation whose mouth action matches the reply text. Finally, the user's input question is replied to by playing or displaying the target character animation.

[0254] The character animation generation scheme provided by the embodiments of the present application can be coupled with any large language model in the industry to improve the flexibility of the output results of the large language model and increase the interestingness of the interaction with the user.

[0255] It should be understood that the above Figures 6-9 The scheme of the example and the above Figure 1 and Figure 5 The method embodiments belong to the same inventive concept and can be implemented by the combination of the related technical features in the above Figure 1 and Figure 5 method embodiments, and therefore the specific implementation process and benefits can refer to the foregoing description, and will not be repeated here.

[0256] Next, the character animation generation device provided by the embodiments of the present application will be introduced.

[0257] Figure 10 is a structural schematic diagram of a character animation generation device provided by an embodiment of the present application. As Figure 10 shown, the character animation generation device 1000 includes a first acquisition module 1001, an animation generation module 1002, a second acquisition module 1003, and an animation adjustment module 1004.

[0258] The first acquisition module 1001 is configured to acquire character dynamic information and a character picture of a target character, the character dynamic information including expression information and / or body action information, the expression information being used to reflect changes in character expressions, and the body action information being used to reflect changes in character body actions. For details of the implementation process, please refer to the related description of step 101 in the embodiments shown in Figure 1

[0259] ​The animation generation module 1002 is configured to drive the character picture into a first character animation according to the character dynamic information, and display the first character animation, wherein an expression change of a target character in the first character animation is consistent with a character expression change reflected by the expression information, and / or a body action change of the target character in the first character animation is consistent with a body action change reflected by the body action information. For details, please refer to the related description of step 102 in the embodiment shown in the following. Figure 1 The related description of step 103 in the embodiment shown in the following will not be repeated here.

[0260] The second acquisition module 1003 is configured to acquire correction information in response to a correction operation on the first character animation, wherein the correction information includes corrected expression information and / or corrected body action information. For details, please refer to the related description of step 103 in the embodiment shown in the following. Figure 1 The related description of step 104 in the embodiment shown in the following will not be repeated here.

[0261] The animation adjustment module 1004 is configured to perform expression driving on a face of the target character in the first character animation according to the corrected expression information, and / or perform body action driving on a body of the target character in the first character animation according to the corrected body action information, to obtain a second character animation. For details, please refer to the related description of step 104 in the embodiment shown in the following. Figure 1 The related description of step 104 in the embodiment shown in the following will not be repeated here.

[0262] Optionally, the expression information includes expression description text and / or an expression video of the first character; and the body action information includes body action description text and / or a body action video of the second character.

[0263] Optionally, the character dynamic information includes the expression information, and the animation generation module 1002 includes:

[0264] The expression animation generation unit is configured to perform expression driving on a face in the character picture according to the expression information, to obtain an expression animation, wherein an expression change of a target character in the expression animation is consistent with a character expression change reflected by the expression information, and the expression change of the target character in the first character animation adopts the expression animation.

[0265] Optionally, the character dynamic information includes the body action information, and the animation generation module 1002 includes:

[0266] The body action animation generation unit is configured to perform body action driving on a body in the character picture according to the body action information, to obtain a body action animation, wherein a body action change of a target character in the body action animation is consistent with a body action change reflected by the body action information, and the body action change of the target character in the first character animation adopts the body action animation.

[0267] Optionally, the character dynamic information includes the expression information and the body action information, and the animation generation module 1002 includes:

[0268] The expression animation generation unit is configured to drive the face in the character picture according to the expression information to obtain an expression animation, in which the expression change of the target character is consistent with the expression change reflected by the expression information.

[0269] The body animation generation unit is configured to drive the body in the character picture according to the body action information to obtain a body animation, in which the body action change of the target character is consistent with the body action change reflected by the body action information.

[0270] The animation fusion unit is configured to fuse the expression animation and the body animation to obtain a first character animation, in which the expression change of the target character in the first character animation is adopted from the expression animation, and the body action change of the target character in the first character animation is adopted from the body animation.

[0271] Optionally, the expression animation generation unit is specifically configured to drive the face in the character picture according to the expression information to obtain a plurality of expression segments; sort the plurality of expression segments according to a first sorting instruction for the plurality of expression segments to obtain sorted expression segments; if a first expression label corresponding to a last frame of a first expression segment in the sorted expression segments is different from a second expression label corresponding to a first frame of a second expression segment, insert an expression transition frame between the first expression segment and the second expression segment, the first expression segment and the second expression segment are adjacent and the first expression segment is before the second expression segment, and the expression transition frame is used to smooth the expression transition process from the first expression segment to the second expression segment.

[0272] Optionally, the character animation generation device 1000 further includes a display module.

[0273] The display module is configured to display a first sorting interface, and the first sorting interface displays the plurality of expression segments.

[0274] The expression animation generation unit is further configured to sort the plurality of expression segments according to a first sorting instruction in response to receiving the first sorting instruction through the first sorting interface, and the first sorting instruction is used to indicate the arrangement order of the plurality of expression segments.

[0275] Optionally, the body animation generation unit is specifically configured to: perform body driving on the body in the character picture according to the body action information to obtain a plurality of body action clips; perform sorting on the plurality of body action clips according to a second sorting instruction for the plurality of body action clips to obtain sorted body action clips; if a first body action corresponding to a last frame of a first body action clip in the sorted body action clips is different from a second body action corresponding to a first frame of a second body action clip, insert a body transition frame between the first body action clip and the second body action clip, the first body action clip and the second body action clip are adjacent and the first body action clip is before the second body action clip, and the body transition frame is used to smooth a body action transition process from the first body action clip to the second body action clip.

[0276] Optionally, the character animation generation apparatus 1000 further includes a display module.

[0277] The display module is configured to display a second sorting interface, and the second sorting interface displays the plurality of body action clips.

[0278] The body animation generation unit is further configured to, in response to receiving a second sorting instruction through the second sorting interface, sort the plurality of body action clips according to the second sorting instruction, and the second sorting instruction is used to indicate an arrangement order of the plurality of body action clips.

[0279] Optionally, the animation generation module 1002 is further configured to, in response to detecting a determination operation on the second character animation, save the second character animation.

[0280] Optionally, the first obtaining module 1001 is configured to obtain input character dynamic information.

[0281] Optionally, the character animation generation apparatus 1000 further includes a display module.

[0282] The display module is configured to display a dynamic information selection interface, and the dynamic information selection interface displays a plurality of preset character dynamic information.

[0283] The first obtaining module 1001 is configured to, in response to detecting a selection operation on character dynamic information in the plurality of preset character dynamic information, obtain the character dynamic information.

[0284] As to the apparatuses in the above-described embodiments, specific manners in which various modules perform operations have been described in details in the embodiments of the method, and will not be described in details here.

[0285] In summary, the embodiment of the present application drives the single picture of the target character to generate a character animation, and the expression and / or body movement of the target character in the character animation are controllable and editable. Compared with the related art, the expression and / or body movement of the target character in the character animation can be controlled and edited by the user, so that the expression and / or body movement of the target character in the finally generated character animation can meet the personalized needs of the user, greatly improving the satisfaction of the user to the animation effect. Since the user can finely control the expression and / or body movement, the character animation can be more vivid, natural, and more consistent with the actual situation and the user's expectation. Compared with the possible single and unnatural expression and movement of the related art, the technical solution of the present application can generate a character animation with high quality and more expressive, improving the overall animation effect.

[0286] Figure 11 is another structure schematic diagram of a character animation generation apparatus provided by the embodiment of the present application. As shown in Figure 11 the character animation generation apparatus 1100 comprises an information acquisition module 1101, a label determination module 1102, an animation determination module 1103, and a voice driving module 1104.

[0287] The information acquisition module 1101 is configured to acquire the call information of the target character in the call scene. For details, please refer to the related description of step 501 in the embodiment shown in Figure 5 .

[0288] The label determination module 1102 is configured to determine the expression label and the body movement label according to the call information. For details, please refer to the related description of step 502 in the embodiment shown in Figure 5 .

[0289] The animation determination module 1103 is configured to determine an initial character animation from a character animation set according to the expression label and / or the body movement label, wherein the character animation set comprises at least one character animation corresponding to the target character. For details, please refer to the related description of step 503 in the embodiment shown in Figure 5 .

[0290] The voice driving module 1104 is configured to adjust the mouth movement of the initial character animation according to the call information to obtain a target character animation, wherein the mouth movement of the target character in the target character animation matches the call information. For details, please refer to the related description of step 504 in the embodiment shown in Figure 5 .

[0291] In summary, the embodiment of the present application combines the call information of the target character in the call scene to determine the initial character animation from the character animation set, adjusts the mouth movement of the initial character animation based on the call information, and obtains the target character animation. Since the initial character animation is pre-constructed based on the requirements of the target character, the character animation meets the basic requirements of the expression and / or body movement of the target character. For a specific call scene, the mouth movement generation and adjustment are performed based on the initial character animation, and the target character animation that meets the dual requirements of the call scene and the call information can be obtained.

[0292] Therefore, the embodiment of the present application realizes controllable and editable expression and / or body movement in the offline stage, and determines the initial character animation from the pre-constructed character animation set and adjusts the mouth movement in the online stage. This two-stage combination makes the character animation generation method adaptable to different scenes and requirements, and has strong flexibility. Whether facing fixed animation requirements or dynamic changing call scenes, the appropriate character animation can be quickly generated.

[0293] It should be noted that: the character animation generation apparatus (i.e., the apparatus 1000 and the apparatus 1100) provided in the above embodiment is only exemplified by the division of the above functional modules when performing the steps of the corresponding technical solutions. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the character animation generation apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the character animation generation apparatus provided in the above embodiment and the corresponding character animation generation method embodiment (i.e., the embodiment shown in the Figure 1 or Figure 5 The specific implementation process is detailed in the method embodiment, which will not be described here.

[0294] Next, the computer device for executing the above character animation generation method provided by the embodiment of the present application is introduced.

[0295] Figure 12 is a hardware structure schematic diagram of a computer device provided by the embodiment of the present application. As Figure 12 shown, the computer device 1200 includes a processor 1201 and a memory 1202, and the memory 1201 and the memory 1202 are connected through a bus 1203. Figure 12 The processor 1201 and the memory 1202 are described independently. Alternatively, the processor 1201 and the memory 1202 are integrated together.

[0296] The memory 1202 is used to store computer programs, including an operating system and program codes. The memory 1202 is various types of storage media, such as a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), a flash memory, an optical memory, a register, an optical disc storage, a magnetic disc storage, or other magnetic storage devices.

[0297] The processor 1201 is a general-purpose processor or a special-purpose processor. The processor 1201 can be a single-core processor or a multi-core processor. The processor 1201 includes at least one circuit to perform the character animation generation method provided in the embodiments of the present application.

[0298] Optionally, the computer device 1200 further includes a network interface 1204 connected to the processor 1201 and the memory 1202 through the bus 1203. The network interface 1204 enables the computer device 1200 to communicate with other devices. The processor 1201 can interact with other devices through the network interface 1204.

[0299] Optionally, the computer device 1200 further includes an input / output (I / O) interface 1205 connected to the processor 1201 and the memory 1202 through the bus 1203. The processor 1201 can receive input commands or data through the I / O interface 1205. The I / O interface 1205 is used for the computer device 1200 to connect to input devices, such as a keyboard and a mouse. Optionally, in some possible scenarios, the network interface 1204 and the I / O interface 1205 are collectively referred to as a communication interface.

[0300] Optionally, the computer device 1200 further includes a display 1206 connected to the processor 1201 and the memory 1202 through the bus 1203. The display 1206 can be used to display intermediate results and / or final results, etc. generated by the processor 1201 executing the above method. In a possible implementation manner, the display 1206 is a touch display screen to provide a human-computer interaction interface.

[0301] The bus 1203 is any type of communication bus, for example, a system bus, for interconnecting the internal components of the computer device 1200. In the embodiment of the present application, the above-mentioned components in the computer device 1200 are interconnected by the bus 1203. Alternatively, the above-mentioned components in the computer device 1200 are interconnected by other connection manners, for example, the above-mentioned components in the computer device 1200 are interconnected by a logical interface in the computer device 1200.

[0302] The above-mentioned components can be respectively arranged on independent chips, or at least partially or entirely arranged on the same chip. Whether the components are arranged on independent chips or integrated on one or more chips depends on the product design requirement. The embodiment of the present application does not limit the specific implementation form of the above-mentioned components.

[0303] Figure 12 The computer device 1200 shown is only exemplary, and in the implementation process, the computer device 1200 further includes other components, which are not listed one by one herein.

[0304] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores instructions. When the instructions are executed by a processor, the steps of the character animation generation method provided by the embodiment of the present application are implemented.

[0305] The embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the character animation generation method provided by the embodiment of the present application are implemented.

[0306] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing related hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.

[0307] In the embodiment of the present application, the terms "first", "second" and "third" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance.

[0308] In the present application, the term "and / or" is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects have an "or" relationship.

[0309] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0310] The above only describes optional embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the concept and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for generating character animation, characterized in that, The method includes: Acquire dynamic information of a person and an image of the target person. The dynamic information of the person includes facial expression information and / or body movement information. The facial expression information is used to reflect changes in the person's facial expression, and the body movement information is used to reflect changes in the person's body movement. Based on the character dynamic information, the character image is driven into a first character animation, and the first character animation is displayed. The facial expression changes of the target character in the first character animation are consistent with the facial expression changes reflected by the facial expression information, and / or, the body movement changes of the target character in the first character animation are consistent with the body movement changes reflected by the body movement information. In response to a correction operation on the first character's animation, correction information is obtained, the correction information including corrected facial expression information and / or corrected body movement information; Based on the corrected facial expression information, the facial expression of the target character in the first character animation is driven, and / or, based on the corrected limb movement information, the limbs of the target character in the first character animation are driven, to obtain the second character animation.

2. The method according to claim 1, characterized in that, The facial expression information includes facial expression description text and / or a video of a first person's facial expression; the body movement information includes body movement description text and / or a video of a second person's body movement.

3. The method according to claim 1 or 2, characterized in that, The character dynamic information includes facial expression information, and the step of driving the character image into a first character animation based on the character dynamic information includes: Based on the expression information, the facial expressions in the character image are driven to produce expression animations. The expression changes of the target character in the expression animations match the expression changes of the character reflected by the expression information. The expression changes of the target character in the first character animations are achieved using the expression animations.

4. The method according to claim 1 or 2, characterized in that, The character dynamic information includes body movement information, and the step of driving the character image into a first character animation based on the character dynamic information includes: Based on the limb movement information, the limbs in the character image are driven to obtain a limb animation. The limb movement changes of the target character in the limb animation are consistent with the limb movement changes reflected by the limb movement information. The limb movement changes of the target character in the first character animation are adopted by the limb animation.

5. The method according to claim 1 or 2, characterized in that, The character dynamic information includes facial expression information and body movement information. The step of driving the character image into a first character animation based on the character dynamic information includes: Based on the expression information, the facial expressions in the image are driven to produce an expression animation, in which the expression changes of the target person in the expression animation are consistent with the expression changes of the person reflected by the expression information. Based on the limb movement information, the limbs in the person image are driven to obtain a limb animation, and the limb movement changes of the target person in the limb animation are consistent with the limb movement changes reflected by the limb movement information. The facial expression animation and the body animation are fused to obtain the first character animation, wherein the facial expression changes of the target character in the first character animation are achieved using the facial expression animation, and the body movement changes of the target character in the first character animation are achieved using the body animation.

6. The method according to claim 3 or 5, characterized in that, The step of driving facial expressions in the image based on the facial expression information to obtain facial animation includes: Based on the facial expression information, facial expressions are driven in the image of the person to obtain multiple facial expression fragments; According to the first sorting instruction for the plurality of facial expression fragments, the plurality of facial expression fragments are sorted to obtain sorted facial expression fragments; If the first expression tag corresponding to the last frame of the first expression segment in the sorted expression segments is different from the second expression tag corresponding to the first frame of the second expression segment, an expression transition frame is inserted between the first expression segment and the second expression segment. The first expression segment is adjacent to the second expression segment and the first expression segment is before the second expression segment. The expression transition frame is used to smooth the expression transition process from the first expression segment to the second expression segment.

7. The method according to claim 6, characterized in that, The step of sorting the plurality of facial expression fragments according to a first sorting instruction includes: The first sorting interface is displayed, and the first sorting interface displays the multiple emoticon fragments; In response to receiving the first sorting instruction through the first sorting interface, the plurality of facial expression fragments are sorted according to the first sorting instruction, wherein the first sorting instruction is used to indicate the arrangement order of the plurality of facial expression fragments.

8. The method according to claim 4 or 5, characterized in that, The step of driving the limbs in the character image based on the limb movement information to obtain limb animation includes: Based on the limb movement information, the limbs in the person image are driven to obtain multiple limb movement segments; According to the second sorting instruction for the plurality of limb movement segments, the plurality of limb movement segments are sorted to obtain sorted limb movement segments. If the first limb movement corresponding to the last frame of the first limb movement segment in the sorted limb movement segments is different from the second limb movement corresponding to the first frame of the second limb movement segment, a limb transition frame is inserted between the first limb movement segment and the second limb movement segment. The first limb movement segment is adjacent to the second limb movement segment and the first limb movement segment is before the second limb movement segment. The limb transition frame is used to smooth the limb movement transition process from the first limb movement segment to the second limb movement segment.

9. The method according to claim 8, characterized in that, The step of sorting the plurality of limb movement segments according to the second sorting instruction for the plurality of limb movement segments includes: A second sorting interface is displayed, which shows the multiple body movement fragments. In response to receiving the second sorting instruction through the second sorting interface, the plurality of limb movement segments are sorted according to the second sorting instruction, wherein the second sorting instruction is used to indicate the arrangement order of the plurality of limb movement segments.

10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: In response to the detection of a determination operation for the second character's animation, the second character's animation is saved.

11. The method according to any one of claims 1 to 10, characterized in that, The acquisition of character dynamic information includes: Obtain the input dynamic information of the character.

12. The method according to any one of claims 1 to 10, characterized in that, The acquisition of character dynamic information includes: The display screen shows a dynamic information selection interface, which displays multiple preset character dynamic information. In response to detecting a selection operation of the character dynamic information from the plurality of preset character dynamic information, the character dynamic information is acquired.

13. A method for generating character animation, characterized in that, The method includes: Obtain the call information of the target person in the call scenario; Based on the call information, determine facial expression tags and body language tags; Based on the facial expression tags and / or the body movement tags, an initial character animation is determined from a character animation set, the character animation set including at least one character animation corresponding to the target character; The initial character animation is adjusted based on the call information to obtain a target character animation, in which the mouth movements of the target character match the call information.

14. A character animation generation device, characterized in that, The device includes: The first acquisition module is used to acquire dynamic information of a person and a picture of the target person. The dynamic information of the person includes facial expression information and / or body movement information. The facial expression information is used to reflect changes in the person's facial expression, and the body movement information is used to reflect changes in the person's body movement. An animation generation module is used to drive the character image into a first character animation based on the character dynamic information, and display the first character animation, wherein the facial expression changes of the target character in the first character animation are consistent with the facial expression changes reflected by the facial expression information, and / or, the body movement changes of the target character in the first character animation are consistent with the body movement changes reflected by the body movement information. The second acquisition module is used to acquire correction information in response to a correction operation on the first character animation, the correction information including corrected facial expression information and / or corrected body movement information; An animation adjustment module is used to drive the facial expressions of the target character in the first character animation according to the corrected facial expression information, and / or to drive the limbs of the target character in the first character animation according to the corrected limb movement information, so as to obtain a second character animation.

15. The apparatus according to claim 14, characterized in that, The facial expression information includes facial expression description text and / or a video of a first person's facial expression; the body movement information includes body movement description text and / or a video of a second person's body movement.

16. The apparatus according to claim 14 or 15, characterized in that, The character dynamic information includes facial expression information, and the animation generation module includes: An expression animation generation unit is used to drive the facial expressions of the person in the image based on the expression information to obtain an expression animation. The expression changes of the target person in the expression animation are consistent with the expression changes of the person reflected by the expression information. The expression changes of the target person in the first character animation are generated using the expression animation.

17. The apparatus according to claim 14 or 15, characterized in that, The character dynamic information includes body movement information, and the animation generation module includes: The body animation generation unit is used to drive the limbs in the character image according to the body movement information to obtain body animation. The changes in the body movements of the target character in the body animation are consistent with the changes in body movements reflected by the body movement information. The changes in the body movements of the target character in the first character animation are adopted by the body animation.

18. The apparatus according to claim 14 or 15, characterized in that, The character dynamic information includes facial expression information and body movement information, and the animation generation module includes: An expression animation generation unit is used to drive the facial expressions of the person in the image based on the expression information to obtain an expression animation, wherein the expression changes of the target person in the expression animation are consistent with the expression changes of the person reflected by the expression information. The body animation generation unit is used to drive the limbs in the person image according to the body movement information to obtain body animation, wherein the changes in the body movements of the target person in the body animation are consistent with the changes in body movements reflected by the body movement information. An animation fusion unit is used to fuse the facial expression animation and the body animation to obtain the first character animation, wherein the facial expression changes of the target character in the first character animation are achieved using the facial expression animation, and the body movement changes of the target character in the first character animation are achieved using the body animation.

19. The apparatus according to claim 16 or 18, characterized in that, The facial expression animation generation unit is specifically used for: Based on the facial expression information, facial expressions are driven in the image of the person to obtain multiple facial expression fragments; According to the first sorting instruction for the plurality of facial expression fragments, the plurality of facial expression fragments are sorted to obtain sorted facial expression fragments; If the first expression tag corresponding to the last frame of the first expression segment in the sorted expression segments is different from the second expression tag corresponding to the first frame of the second expression segment, an expression transition frame is inserted between the first expression segment and the second expression segment. The first expression segment is adjacent to the second expression segment and the first expression segment is before the second expression segment. The expression transition frame is used to smooth the expression transition process from the first expression segment to the second expression segment.

20. The apparatus according to claim 19, characterized in that, The device further includes: a display module; The display module is used to display a first sorting interface, in which the multiple emoticon fragments are displayed; The facial expression animation generation unit is further configured to respond to receiving the first sorting instruction through the first sorting interface, and to sort the plurality of facial expression fragments according to the first sorting instruction, wherein the first sorting instruction is used to indicate the arrangement order of the plurality of facial expression fragments.

21. The apparatus according to claim 17 or 18, characterized in that, The limb animation generation unit is specifically used for: Based on the limb movement information, the limbs in the person image are driven to obtain multiple limb movement segments; According to the second sorting instruction for the plurality of limb movement segments, the plurality of limb movement segments are sorted to obtain sorted limb movement segments. If the first limb movement corresponding to the last frame of the first limb movement segment in the sorted limb movement segments is different from the second limb movement corresponding to the first frame of the second limb movement segment, a limb transition frame is inserted between the first limb movement segment and the second limb movement segment. The first limb movement segment is adjacent to the second limb movement segment and the first limb movement segment is before the second limb movement segment. The limb transition frame is used to smooth the limb movement transition process from the first limb movement segment to the second limb movement segment.

22. The apparatus according to claim 21, characterized in that, The device further includes: a display module; The display module is used to display a second sorting interface, in which the multiple body movement fragments are displayed; The limb animation generation unit is configured to, in response to receiving the second sorting instruction through the second sorting interface, sort the plurality of limb movement segments according to the second sorting instruction, wherein the second sorting instruction is used to indicate the arrangement order of the plurality of limb movement segments.

23. The apparatus according to any one of claims 14 to 22, characterized in that, The animation generation module is also used for: In response to the detection of a determination operation for the second character's animation, the second character's animation is saved.

24. The apparatus according to any one of claims 14 to 23, characterized in that, The first acquisition module is used for: Obtain the input dynamic information of the character.

25. The apparatus according to any one of claims 14 to 23, characterized in that, The device further includes: a display module; The display module is used to display a dynamic information selection interface, which displays multiple preset character dynamic information. The first acquisition module is used to acquire the character dynamic information in response to detecting a selection operation of the character dynamic information from the plurality of preset character dynamic information.

26. A character animation generation device, characterized in that, The device includes: The information acquisition module is used to acquire the call information of the target person in the call scenario; The tag determination module is used to determine facial expression tags and body movement tags based on the call information; An animation determination module is used to determine an initial character animation from a character animation set based on the facial expression tags and / or the body movement tags, wherein the character animation set includes at least one character animation corresponding to the target character; A voice-driven module is used to adjust the mouth movements of the initial character animation based on the call information to obtain a target character animation, wherein the mouth movements of the target character in the target character animation match the call information.

27. A computer device, characterized in that, include: Processor and memory; The memory is used to store computer programs, the computer programs including program instructions; The processor is configured to invoke the computer program to implement the method as described in any one of claims 1 to 12, or to implement the method as described in claim 13.

28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 12, or implement the method as described in claim 13.

29. A computer program product, characterized in that, It includes a computer program, which, when executed by a processor, implements the method as described in any one of claims 1 to 12, or implements the method as described in claim 13.