Animation generation methods, devices and electronic equipment

By using motion generation models and joint rigging technology, animations are automatically generated, solving the problem of high animation generation difficulty in existing technologies and achieving low-cost, high-efficiency animation production.

CN119131207BActive Publication Date: 2026-01-30SHENZHEN TUYI SHIBEI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411092736.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-01-30
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

Existing methods for generating animations from images are too difficult, require a lot of time and manpower, and are not user-friendly for the general public.

Method used

By acquiring the original image and the target mask image, the motion generation model is used to generate joint motion information, and the animated character is bound based on the joint points to automatically generate the target animation, avoiding the need to manually create motion files.

Benefits of technology

It reduces the difficulty and cost of animation production, generates smooth and natural animations, is easy to operate, does not require animator skills, and improves animation generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131207B_ABST
    Figure CN119131207B_ABST
Patent Text Reader

Abstract

This application provides an animation generation method, apparatus, and electronic device, relating to the field of animation production. The method includes: acquiring an original image and a target mask image; acquiring motion description information, which describes the expected actions of the animated character; inputting the motion description information into a motion generation model to obtain an motion file; extracting the animated character from the original image using the animated character region in the target mask image, and binding the joint points of the animated character with the joint motion information of each frame to generate a target animation; the target animation is used to display the animated character performing the actions described in the motion description information. This application solves the problem that image-based animation generation methods in related technologies are too difficult.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of animation production, and more specifically, to an animation generation method, apparatus, and electronic device. Background Technology

[0002] Doodles are an important form of expression in children's art. By bringing static doodles to life and turning them into animations, they have many practical applications, such as helping children's early education, enriching their imagination, analyzing their psychological growth, and cultivating their artistic talents.

[0003] Currently, doodle images can be converted into animation sequences through hand-drawing. However, this method involves manually drawing each frame of the animation sequence from the doodle image, which is a tedious and labor-intensive task. It requires artistic skills, a lot of time, and learning to operate complex animation software. The learning curve for many animation software programs is quite steep, such as Visme, OpenToonz, and Cinema 4D. Therefore, hand-drawing animation is not user-friendly for the general public, and the threshold and cost of animation production are significantly high.

[0004] As can be seen from the above, the method of generating animations based on images is too difficult and has become a problem that urgently needs to be solved. Summary of the Invention

[0005] This application provides an animation generation method, apparatus, electronic device, and storage medium, which can solve the problem that image-based animation generation methods in related technologies are too difficult. The technical solutions are as follows:

[0006] According to one aspect of this application, an animation generation method includes: acquiring an original image and a target mask image; the original image includes an animated character; the target mask image includes an animated character region and a background region; the animated character region is distinct from the background region; acquiring motion description information, the motion description information being used to describe the expected action of the animated character; inputting the motion description information into a motion generation model to obtain an motion file; the motion file includes at least one frame of joint motion information; the joint motion information being used to describe the motion of joint points corresponding to each action; using the animated character region in the target mask image, extracting the animated character from the original image, and binding the joint points of the animated character to the joint motion information of each frame to generate a target animation; the target animation being used to display the animated character performing the action described by the motion description information.

[0007] According to one aspect of this application, an animation generation apparatus includes: an image acquisition module for acquiring an original image and a target mask image; the original image includes an animated character; the target mask image includes an animated character region and a background region; the animated character region is distinct from the background region; an information acquisition module for acquiring motion description information, the motion description information describing the expected action of the animated character; a motion generation module for inputting the motion description information into a motion generation model to obtain a motion file; the motion file includes at least one frame of joint motion information; the joint motion information describing the motion of joint points corresponding to each action; and an animation generation module for extracting the animated character from the original image using the animated character region in the target mask image, and binding the joint points of the animated character to the joint motion information of each frame to generate a target animation; the target animation is used to display the animated character performing the action described by the motion description information.

[0008] According to one aspect of this application, an electronic device includes at least one processor and at least one memory, wherein computer-readable instructions are stored on the memory; the computer-readable instructions are executed by one or more of the processors to cause the electronic device to implement the animation generation method as described above.

[0009] According to one aspect of this application, a storage medium stores computer-readable instructions thereon, which are executed by one or more processors to implement the animation generation method described above.

[0010] According to one aspect of this application, a computer program product includes computer-readable instructions stored in a storage medium, wherein one or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, causing the electronic device to implement the animation generation method as described above.

[0011] The beneficial effects of the technical solution provided in this application are:

[0012] In the above technical solution, inputting motion description information into the motion generation model to generate motion files can avoid manually creating motion files, reduce labor costs, and lower the difficulty of animation production. By binding the joint points of the animated character in the original image with the joint motion information of each frame in the motion file, the target animation can be automatically generated. This animation generation method is easy to learn and does not require users to have animator skills. Compared with manual animation production, it can effectively reduce the labor cost and tool learning cost of animation production, and thus effectively solve the problem that the image-based animation generation method in related technologies is too difficult. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1a This is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0015] Figure 1b This is a flowchart illustrating an animation generation method according to an exemplary embodiment;

[0016] Figure 2 yes Figure 1b A flowchart of step 350 in one embodiment corresponds to the following example;

[0017] Figure 3 yes Figure 1b A flowchart of step 370 in one embodiment corresponds to the following example;

[0018] Figure 4 yes Figure 3 A flowchart of step 373 in one embodiment corresponds to the following example;

[0019] Figure 5 yes Figure 1b A schematic diagram of the target animation involved in the corresponding embodiment;

[0020] Figure 6 yes Figure 1b A flowchart of step 310 in one embodiment corresponds to the following example;

[0021] Figure 7 yes Figure 6 A flowchart of step 313 in one embodiment corresponds to the following example;

[0022] Figure 8 This is a schematic diagram illustrating the specific implementation of an animation generation method in an application scenario;

[0023] Figure 9 This is a structural block diagram of an animation generation apparatus according to an exemplary embodiment;

[0024] Figure 10 This is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation

[0025] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0026] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this disclosure means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0027] As mentioned earlier, doodle images can be transformed into animation sequences through handcrafted animation techniques. However, these techniques are not user-friendly for the general public, as the barriers to entry and costs for animation production are significantly high.

[0028] In addition, with the rapid development of AI technology, methods have emerged to automatically drive graffiti images to generate animations. However, these methods primarily utilize a series of computer vision technologies (including object detection, character segmentation, and pose estimation) to build an automated system that transforms graffiti images into animation sequences. These systems first identify the characters and their joints from the graffiti images, then use manually created BVH motion capture data to bind the characters and joints, generating animations of the graffiti characters. There are two main ways to manually create BVH motion capture data: one is to film live performances and extract BVH motion files from the video using tools like Rokoko or Deepmotion; the other is to create skeletal animations of the performance using Blender software and then export them as BVH motion files. These methods require significant manpower and time, and expanding the diversity of movements is not easy for users. Furthermore, the character segmentation used by these systems is mostly based on classic deep learning models such as Mask R-CNN and YOLOv8, which are not accurate enough for abstract graffiti images, leading to poor motion binding for some characters and generating unsatisfactory graffiti character animations.

[0029] As can be seen from the above, there is still a drawback in the related technologies: the method of generating animations based on images is too difficult.

[0030] Therefore, the animation generation method provided in this application can effectively improve the accuracy of animation generation. Accordingly, the animation generation method is applicable to an animation generation device, which can be deployed on an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, such as a desktop computer, a laptop computer, a server, etc.

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0032] Figure 1a A schematic diagram of the structure of an electronic device is shown according to an exemplary embodiment. The electronic device may be a desktop computer, a laptop computer, a server, etc.

[0033] It should be noted that this electronic device is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Furthermore, this electronic device should not be interpreted as requiring or depending on any specific feature. Figure 1a One or more components of the exemplary electronic device 2000 shown.

[0034] The hardware structure of electronic devices 2000 can vary significantly due to differences in configuration or performance, such as... Figure 1a As shown, the electronic device 2000 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.

[0035] Specifically, power supply 210 is used to provide operating voltage for various hardware devices on electronic device 2000.

[0036] Interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices.

[0037] Of course, in other examples adapted in this application, interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc. Figure 1a As shown, this does not constitute a specific limitation.

[0038] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it include the operating system 251, application programs 253, and data 255, etc., and the storage method can be temporary storage or permanent storage.

[0039] The operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 2000, so as to enable the central processing unit 270 to perform calculations and processing on the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0040] Application 253 is a computer-readable instruction based on operating system 251 that performs at least one specific task, and may include at least one module ( Figure 1a (Not shown), each module may contain computer-readable instructions for the electronic device 2000. For example, the animation generation device may be considered as an application 253 deployed on the electronic device 2000.

[0041] Data 255 can be photos, pictures, etc. stored on a disk, or it can be original images, target mask images, etc., stored in memory 250.

[0042] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer-readable instructions stored in the memory 250, thereby performing operations and processing on massive amounts of data 255 stored in the memory 250. For example, an animation generation method may be implemented by the central processing unit 270 reading a series of computer-readable instructions stored in the memory 250.

[0043] Furthermore, this application can also be implemented through hardware circuits or a combination of hardware circuits and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, or combination thereof.

[0044] Please see Figure 1b This application provides an animation generation method applicable to electronic devices, such as desktop computers, laptops, and servers.

[0045] In the following method embodiments, for ease of description, the execution subject of each step of the method is an electronic device, but this does not constitute a specific limitation.

[0046] like Figure 1b As shown, the method may include the following steps:

[0047] Step 310: Obtain the original image and the target mask image.

[0048] The original image includes animated characters, and the original image can be a doodle image, such as an image obtained from a child's doodle.

[0049] The target mask image is obtained by performing specific processing on the original image, such as background matting. The target mask image is used to represent specific areas, including the animation character area and the background area. Since the target mask image is a binary image, there is a very high contrast between the masked area (animation character area) and the non-masked area (background area). In other words, the animation character area is distinguished from the background area, and this contrast makes the masked area (animation character area) very obvious.

[0050] One possible implementation is to process the original image by inputting it into a background matting model, marking the pixels in the animated character area of ​​the original image as white and the pixels in the background area as black, in order to obtain the target mask image.

[0051] Step 330: Obtain action description information.

[0052] Among them, the action description information is used to describe the actions expected to be performed by the animated character.

[0053] Action description information can be obtained through user input, such as text input or language input. For example, if a user expects an animated character to perform a dance action, they can input "a person is dancing" to generate the corresponding action description information.

[0054] Step 350: Input the action description information into the action generation model to obtain the action file.

[0055] The motion file includes at least one frame of joint motion information, which describes the motion of the joint points corresponding to each motion.

[0056] For motion files, standard formats are usually used, such as BVH (Biovision Hierarchy) and FBX (Filmbox). These formats can store complex hierarchical structures of joint points and frame-by-frame joint motion information.

[0057] First of all, it should be noted that when an animated character makes a certain action, the corresponding joint points will also move in accordance with that action.

[0058] For example, the joints of an animated character include the shoulder, elbow, and wrist joints. When the animated character waves, the shoulder joint rotates forward and upward, raising the arm until it is parallel to or slightly higher than the ground. The shoulder joint also rotates left and right to drive the arm in a swinging motion. The elbow joint bends slightly, creating an angle between the forearm and upper arm, and adjusts appropriately with the shoulder joint's swing to maintain the arm's natural curve. The wrist joint makes subtle bending and rotation movements while the arm swings, increasing the flexibility and naturalness of the wave. In other words, the waving motion involves the coordinated movement of the shoulder, elbow, and wrist joints. These joints achieve a natural and continuous movement through coordinated motion. Therefore, by adjusting the position and angle of these joints frame by frame, a realistic waving animation can be achieved.

[0059] The action generation model can be a trained machine learning model capable of generating action files based on action description information. One possible implementation is the MoMask model, a generative model based on the transformer architecture. Since the MoMask model is pre-trained using English corpora, the action description information needs to be converted to English before being input into the MoMask model. This Chinese-to-English translation can be implemented using the Opus-MT model.

[0060] In one possible implementation, such as Figure 2 As shown, step 350 may include the following steps:

[0061] Step 351: Determine the number of animation frames for the animated character.

[0062] First, it should be noted that the action description information can also include the start time, end time, type, and details of the action. Therefore, the duration of the target animation can be preliminarily determined based on the action description information.

[0063] Furthermore, the number of motion frames affects the smoothness of the final generated target animation. Understandably, the higher the number of motion frames displayed per second, the smoother the target animation. Common frame rates (the number of motion frames displayed per second) include: 24 FPS: suitable for movies and animated shorts, providing sufficient smoothness; 30 FPS: suitable for television and most video games, providing even higher smoothness; 60 FPS: suitable for high frame rate video games and VR applications, providing the smoothest visual effects. Users can set the appropriate frame rate according to their needs. Therefore, based on the animation's duration and frame rate, the total number of frames in the target animation can be calculated: Motion Frames = Target Animation Duration (seconds) × Frame Rate (FPS)

[0064] Step 353: Input the action description information and the number of action frames into the action generation model for text-to-action generation processing to generate joint motion information that matches the number of action frames.

[0065] First, it should be noted that the motion generation model can convert motion description information into frame-by-frame joint motion information. The motion file contains joint motion information that stores the joint position and rotation information corresponding to each frame of motion. Specifically, the joint motion information can include joint position, joint rotation angle, and joint hierarchy.

[0066] In one possible implementation, the process of generating action files for the action generation model can be represented by formula (1), where P MoMask For the action generation model, p Opus This is the action description information, where n is the number of action frames, (m1, ..., m n} is p MoMask An action file generated based on action description information and action frame count, m n This represents the joint motion information in the nth frame.

[0067] {m1,...,m n}=P MoMask (p Opus (t prompt) ,n)...Formula (1)

[0068] For example, the process of generating a "waving" motion file through a motion generation model may include: creating a BVH file, setting the frame rate to 30fps, defining the joint hierarchy: shoulder joint -> elbow joint -> wrist joint, and writing joint position and rotation information frame by frame. Then, the generated joint motion information may be: frame 1: shoulder joint position (x1, y1, z1), rotation (rx1, ry1, rz1), frame 2: shoulder joint position (x2, y2, z2), rotation (rx2, ry2, rz2)...

[0069] Step 355: Obtain the motion file based on the joint motion information of each frame.

[0070] In summary, the joint motion information generated by the motion generation model for each frame ensures that the joint motion in each frame matches the motion described by the input motion description information, thereby enabling the subsequently generated target animation to accurately display the desired animation effect.

[0071] The above process eliminates the need for manual creation of motion files. The motion generation model solves the problem of motion file creation, reducing the manpower and time costs required to increase motion diversity. It can significantly reduce the manpower and learning costs of animation production and is characterized by its simplicity and user-friendliness.

[0072] Step 370: Using the animated character region of the target mask image, the animated character is extracted from the original image, and the joint points of the animated character are bound to the joint motion information of each frame to generate the target animation.

[0073] Among them, the target animation is used to display the actions described in the action description information of the animated character.

[0074] As mentioned earlier, the target mask image includes the animated character region, which can be marked as white in the target mask image. Therefore, the animated character in the original image can be extracted based on the animated character region in the target mask image.

[0075] It is understandable that when an animated character performs the action described in the action description information, the movement of its joint points should conform to the joint movement information of each frame in the action description information. Therefore, before this, it is necessary to determine the joint points of the animated character. Joint points can include shoulder joints, elbow joints, wrist joints, hip joints, knee joints, ankle joints, etc., which are not limited here.

[0076] It should be noted that the original image includes all the color, texture and edge information of the animated character, while the target mask image is a binary image that only contains the outline and region of the animated character and lacks detailed information. Therefore, determining the joint points based on the animated character in the original image can ensure the accuracy and reliability of the recognition.

[0077] Regarding the determination of joint points, in one possible implementation, before step 370, the following steps may also be included: performing target detection on the original image to determine the location region of the animated character in the original image and generating a positioning region; performing pose recognition on the animated character in the positioning region to obtain the joint points of the animated character in the original image.

[0078] First, it should be noted that object detection methods can be used to identify the position of animated characters in the original image and generate localization regions. The localization regions can be in the form of bounding boxes, which can be used to mark the location area of ​​the animated characters. The object detection method can be YOLO, Faster R-CNN, or SSD, and no particular method is specified here.

[0079] Furthermore, pose recognition methods can be used within the bounding box to detect joint points of the animated character. The pose recognition methods can be such as OpenPose, PoseNet, or DensePose, and are not limited here.

[0080] By recognizing poses, the coordinates of each joint point of the animated character can be determined, thus identifying the joint points of the animated character. This allows the joint points of the animated character to be subsequently bound to the joint motion information of each frame in the motion file to generate the target animation.

[0081] In one possible implementation, such as Figure 3 As shown, step 370 may include the following steps:

[0082] Step 371: Perform mesh texturing on the animated character, generate multiple meshes on the animated character, and associate the joint points of the animated character with the meshes.

[0083] First, it should be noted that mesh texturing can map an animated character onto a textured mesh, thereby generating multiple meshes on the animated character.

[0084] Furthermore, by associating the joint points of the animated character with the mesh, when the joint points move, only the corresponding mesh vertices need to be adjusted, which can achieve natural limb movements and deformations, making the target animation more vivid. Moreover, it eliminates the need to redraw every frame of the animated character, making the deformation and movements of the animated character more flexible.

[0085] One possible implementation is to use the Delaunay triangulation method to achieve mesh texturing. Specifically, the Delaunay triangulation method is used to generate 2D triangular meshes, which are then used to texturize the animated character. The triangular meshes are assigned to the corresponding joint points of the animated character, thereby completing the association between the joint points of the animated character and the mesh.

[0086] Step 373: Adjust the mesh associated with the joint points of the animated character according to the joint motion information, so that the animated character can perform the actions indicated by the motion description information.

[0087] It's understandable that when an animated character performs a certain action, its joint points also move accordingly. Furthermore, since joint motion information describes the movement of joint points for each action, when joint points move, it's only necessary to adjust the corresponding mesh vertices based on the joint motion information. It's important to note that moving mesh vertices ensures that the shape and posture of the animated character adjust naturally with the changes in joint points when performing a certain action. During the movement of joint points, complex deformations occur on the surface of the animated character; adjusting the mesh vertices adapts to these deformations, making the movements appear more realistic and fluid.

[0088] For example, when an animated character waves, the movement primarily involves the arm, especially the shoulder, elbow, and wrist joints. Joint motion information describes the trajectory and positional changes of the shoulder, elbow, and wrist joints during the wave. Based on this joint motion information, the corresponding mesh vertices are adjusted, including: shoulder joint movement, which pulls the associated mesh vertices, causing them to slightly lift and shift forward; elbow joint movement, where outward movement pulls the mesh vertices around the elbow, maintaining a reasonable deformation of the upper arm mesh shape while aligning with shoulder movement; and wrist joint movement, where the wrist's swing directly affects the mesh vertices in the hand and forearm areas. These vertices need to rotate and move according to the wrist's movement to maintain the arm's natural shape and fluidity. By adjusting the mesh vertices corresponding to each joint point, the mesh of the entire arm area changes with the joint point movement, resulting in a natural and fluid arm movement during the wave. Ultimately, the mesh changes perfectly match the joint point movements, making the animation look realistic and coherent.

[0089] In one possible implementation, such as Figure 4 As shown, step 370 may include the following steps:

[0090] Step 3731: Determine the target positions corresponding to each joint point of the animated character based on joint motion information.

[0091] Among them, each target position is the actual position of each joint point after the animated character completes the action described by the joint movement information.

[0092] Regarding the results of joint point motion, when the animated character completes the joint point motion indicated by the joint motion information in each frame, each joint point has a corresponding ending position, that is, the actual position of each joint point in the target animation of that frame. Therefore, the actual position of each joint point of the animated character in the target animation of that frame can be determined based on the joint motion information, thereby obtaining the target position corresponding to each joint point.

[0093] Step 3733: Translate the corresponding joint points according to each target position to update the mesh associated with each joint point.

[0094] For each frame of joint motion information, the target position corresponding to each joint point can be determined, and the associated mesh can be adjusted according to the target position of each joint point.

[0095] One possible implementation is to use the As-Rigid-As-Possible (ARAP) method to update the meshes associated with each joint point. The ARAP method can simultaneously adjust the positions of other mesh vertices, thereby keeping the local shape of the mesh as unchanged as possible, thus ensuring that the animation of the character is natural and realistic.

[0096] Through the above process, the joint points of the animated character can be bound to the joint motion information of each frame in the motion file, thereby generating a sequence of target animation frames corresponding to the action described in the motion description information. The generated sequence of target animation frames can then be combined into a continuous target animation, such as MP4 or GIF.

[0097] For example, Figure 5 The diagram illustrates a target animation, such as... Figure 5 As shown, the target animation displays GIF animation frames of an animated character dancing, with each frame representing a transition between different movements of the character's dance.

[0098] In conjunction with the above embodiments, inputting motion description information into the motion generation model to generate motion files avoids manual creation of motion files, reducing labor costs and simplifying animation production. By binding the joint points of the animated character in the original image with the joint motion information of each frame in the motion file, the target animation can be automatically generated. This animation generation method is easy to learn and does not require users to have animator skills. Compared with manual animation production, it can effectively reduce the labor costs and tool learning costs of animation production, thus effectively solving the problem that image-based animation generation methods in related technologies are too difficult. In addition, the above method generates target animations more efficiently, and compared with other current animation production methods, it can generate the desired animation effect in less time.

[0099] Please see Figure 6 In one exemplary embodiment, step 310 may include the following steps:

[0100] Step 311: Obtain the original image.

[0101] As mentioned earlier, the original image includes animated characters, and the original image can also be a doodle image.

[0102] Step 313: Perform background masking on the original image to obtain the target mask image.

[0103] First, it should be noted that background matting refers to separating the foreground object (animated character) from the background area of ​​an image. Through background matting, the animated character can be separated from the complex or irrelevant background, making the animated character more prominent and clear. When generating the target animation, you can focus more on the action details of the animated character without being disturbed by the background.

[0104] Furthermore, if the animated character has already undergone background removal when generating the target animation, it will be more convenient to edit and composite it later. The background can be easily replaced or modified without affecting the display effect of the animated character. For example, the animated character can be composited into different scenes, or special effects backgrounds can be added.

[0105] After performing background masking on the original image, the animated character area and the background area can be separated. The animated character area is marked as white, and the background area is marked as black, thus obtaining the target mask image.

[0106] In one possible implementation, such as Figure 7 As shown, step 313 includes the following steps:

[0107] Step 3131: Input the original image into the background matting model to obtain the first mask image.

[0108] The first mask image includes an animated character area and a background area that is distinct from the animated character area.

[0109] One possible implementation is to use the ViTMatte or PP-Matting models to perform background matting on the original image to obtain the target mask image. It's important to note that these techniques use semantic segmentation models like Mask R-CNN and YOLOv8 for background matting, performing pixel classification tasks—a form of hard segmentation. The predicted pixel values ​​for the person mask are discrete (e.g., values ​​of 0 or 1), resulting in obvious edges and an inability to handle edge details, leading to poor visual quality. In contrast, the ViTMatte and PP-Matting models are background matting models that perform regression tasks on the pixels in the original image—a form of soft segmentation. The predicted pixel values ​​for the target mask image are continuous and smooth (e.g., values ​​between [0,1]), thus preserving edge information more naturally and better handling of edge details.

[0110] Step 3133: Locate the animated character in the original image using the animated character region in the first mask image, and fill the located animated character to obtain the second mask image.

[0111] The animated character area in the second mask image is different from the animated character area in the first mask image.

[0112] It is understood that the first mask image includes an animated character area and a background area that is distinct from the animated character area. The animated character area is significantly different from the background area. Therefore, the animated character in the original image can be extracted using the animated character area in the first mask image.

[0113] Due to errors in the background masking process, a small portion of the pixel area in the first mask image becomes hollow and is marked as black pixels. Therefore, the animated character is filled in, and the area corresponding to the animated character is marked as black, thus obtaining the second mask image. The second mask image can fill in the hollow areas caused by the aforementioned errors and fill in the edge details of the animated character, thereby ensuring continuous and smooth edge pixels, and significantly improving the visual quality and naturalness of the subsequently generated target animation.

[0114] In one possible implementation, the animated character is filled using OpenCV to obtain a second mask image.

[0115] Step 3135: Select an image that meets the detection conditions from the first mask image and the second mask image as the target mask image.

[0116] The detection criteria are used to filter out more refined target mask images, thereby reducing the impact of errors caused by background masking.

[0117] In one possible implementation, step 3135 may include the following steps: calculating the contour area of ​​the animated character region in the first mask image and the second mask image respectively to obtain the first contour area and the second contour area; comparing the first contour area and the second contour area, and selecting the image with the larger contour area from the first mask image and the second mask image as the target mask image.

[0118] First, it should be noted that the first mask image is mask1, and the corresponding first contour area is area1. The second mask image is mask2, and the corresponding second contour area is area2. The detection condition is area2 - area1 > area2 / λ, where λ is an adjustable scaling parameter. If the condition is true, the second mask image mask2 is used as the target mask image; otherwise, if the condition is false, the first mask image mask1 is used as the target mask image. λ is set to 10 by default, which is estimated based on the area difference between the first mask image mask1 and the second mask image mask2 when background masking fails.

[0119] Under the above embodiments, the background matting model can improve the segmentation effect of animated characters, better distinguish animated characters from the background area, generate more detailed animated character outlines, and ensure that the action performance of the subsequently generated target animation is more natural.

[0120] Figure 8 This is a schematic diagram illustrating the specific implementation of an animation generation method in an application scenario. This application scenario is divided into an image processing workflow, an action file generation workflow, and an animation generation workflow.

[0121] like Figure 8 As shown, the image processing workflow includes the following steps:

[0122] In step 601, the original image is obtained.

[0123] In step 603, the original image is input into the background matting model to obtain the animated character after background matting and the first mask image mask1.

[0124] In step 605, the animated character is filled to obtain the second mask image mask2.

[0125] In step 607, the outline area of ​​the animated character region in the first mask image mask1 is calculated to obtain the first outline area area1.

[0126] In step 609, the outline area of ​​the animated character region in the second mask image mask2 is calculated to obtain the second outline area area2.

[0127] In step 611, the target mask image is selected from the first mask image and the second mask image using the detection conditions.

[0128] like Figure 8 As shown, the action file generation process includes the following steps:

[0129] In step 701, enter the text prompt.

[0130] In step 703, the text prompts are translated into English using a machine translation model to obtain action description information.

[0131] In step 705, the action description information is input into the action generation model to obtain the action file.

[0132] like Figure 8 As shown, the animation generation process includes the following steps:

[0133] In step 801, the original image and target mask image input to the image processing flow, as well as the action file input to the action file generation flow, are obtained.

[0134] In step 803, the target animation is generated based on the original image, the target mask image, and the motion file using the Animated Drawings algorithm framework.

[0135] The Animated Drawings algorithm framework includes Delaunay triangulation and ARAP algorithms. First, Delaunay triangulation is used to generate 2D triangular meshes, which are then used to texture the animated character. The triangular meshes are assigned to the corresponding joint points of the animated character. Then, when generating the target animation, the joint points are translated, and the ARAP algorithm is used to readjust the meshes associated with each joint point of the animated character, thereby binding the joint motion information of each frame of the motion file and making the animated character present the effect of performing actions.

[0136] In this application scenario, on the one hand, the method is easy to learn and the operation of generating the target animation is highly automated, requiring no animator skills from the user. Compared with manual animation production, it can effectively reduce the labor cost and tool learning cost of animation production. On the other hand, the method in this application scenario produces the target animation more efficiently. Compared with other current animation production methods, it can generate the desired doodle character animation effect in less time. Furthermore, the method in this application scenario does not require manual creation of motion files. It solves the problem of motion file creation through a text-to-motion generative model, reducing the labor and time costs required to increase the diversity of motion. In addition, the method in this application scenario improves the segmentation effect of the animated character through a background matting model. Compared with current AI-driven methods, it can better separate the animated character from the background area, generating a more detailed animated character outline and a more natural animation effect.

[0137] Compared to related technologies, this solution inputs motion description information into the motion generation model to generate motion files, avoiding manual creation of motion files, reducing labor costs, and simplifying animation production. By binding the joint points of the animated character in the original image to the joint motion information of each frame in the motion file, the target animation can be automatically generated. This animation generation method is easy to learn and does not require animator skills. Compared to manual animation production, it effectively reduces the labor and tool learning costs of animation production, thus effectively solving the problem of excessive difficulty in image-based animation generation methods in related technologies. Furthermore, the above method generates target animations more efficiently, achieving the desired animation effect in less time compared to other current animation production methods.

[0138] In addition, the background cutout processing method used in this solution can improve the segmentation effect of animated characters, better distinguish animated characters from the background area, and generate more detailed animated character outlines. Furthermore, it can also ensure that the action performance of the subsequently generated target animation is more natural.

[0139] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0140] The following are embodiments of the apparatus described in this application, which can be used to execute the animation generation method involved in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of the animation generation method involved in this application.

[0141] Please see Figure 9 This application provides an animation generation device 900, including but not limited to: an image acquisition module 910, an information acquisition module 930, an action generation module 950, and an animation generation module 970.

[0142] The image acquisition module 910 is used to acquire the original image and the target mask image; the original image includes the animated character; the target mask image includes the animated character area and the background area; the animated character area is distinct from the background area.

[0143] The information acquisition module 930 is used to acquire action description information, which describes the actions expected to be performed by the animated character.

[0144] The motion generation module 950 is used to input motion description information into the motion generation model to obtain motion files; the motion files include at least one frame of joint motion information; the joint motion information is used to describe the motion of the joint points corresponding to each motion.

[0145] The animation generation module 970 is used to extract the animated character from the original image using the animated character region in the target mask image, and bind the joint points of the animated character to the joint motion information of each frame to generate the target animation; the target animation is used to display the actions described by the action description information of the animated character.

[0146] It should be noted that the animation generation device provided in the above embodiments is only illustrated by the division of the above functional modules when generating animation. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the animation generation device will be divided into different functional modules to complete all or part of the functions described above.

[0147] Furthermore, the animation generation apparatus and the animation generation method provided in the above embodiments belong to the same concept, and the specific way in which each module performs its operation has been described in detail in the method embodiments, and will not be repeated here.

[0148] Please see Figure 1b This application provides an electronic device 4000, which may include: a desktop computer, a laptop computer, a server, etc.

[0149] exist Figure 1b In this context, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.

[0150] The data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. This communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 1b The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0151] Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0152] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0153] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program instructions or code in the form of instructions or data structures and accessible by the electronic device 400, but not limited thereto.

[0154] The memory 4003 stores computer-readable instructions, and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002.

[0155] The computer-readable instructions are executed by one or more processors 4001 to implement the animation generation methods in the above embodiments.

[0156] Furthermore, this application provides a storage medium storing computer-readable instructions that are executed by one or more processors to implement the animation generation method described above.

[0157] This application provides a computer program product including computer-readable instructions stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, thereby enabling the electronic device to implement the animation generation method described above.

[0158] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An animation generation method characterized by comprising: The method comprises the following steps: obtaining an original image; the original image comprises an animation character; respectively performing background matting processing and filling processing on the animation character in the original image to obtain a first mask image and a second mask image; the first mask image comprises an animation character region and a background region different from the animation character region; the animation character region in the second mask image is different from the animation character region in the first mask image; selecting an image meeting a detection condition from the first mask image and the second mask image as a target mask image; the target mask image comprises an animation character region and a background region; the animation character region is different from the background region; the target mask image has higher precision than the first mask image or the second mask image; obtaining action description information, the action description information being used to describe an action expected to be performed by the animation character; inputting the action description information into an action generation model to obtain an action file; the action file comprises at least one frame of joint motion information; the joint motion information is used to describe the motion of a joint part point corresponding to each action; extracting the animation character from the original image by using the animation character region in the target mask image, and binding the joint part point of the animation character with each frame of the joint motion information respectively to generate a target animation; the target animation is used to show the action described by the action description information performed by the animation character.

2. The method of claim 1, wherein, The method comprises the following steps: inputting the original image into a background matting model to obtain the first mask image; positioning the animation character in the original image by using the animation character region in the first mask image, and performing filling processing on the positioned animation character to obtain the second mask image.

3. The method of claim 1, wherein, The method comprises the following steps: respectively calculating the contour area of the animation character region in the first mask image and the second mask image to obtain a first contour area and a second contour area; comparing the first contour area and the second contour area, and selecting an image with a larger contour area from the first mask image and the second mask image as the target mask image.

4. The method of claim 1, wherein, The method comprises the following steps: determining the action frame number of the animation character; inputting the action description information and the action frame number into the action generation model to perform text-to-action generation processing and generate joint motion information meeting the action frame number; obtaining the action file based on each frame of the joint motion information.

5. The method of claim 1, wherein, Before the binding of the joint part point of the animation character with each frame of the joint motion information to generate the target animation, the method comprises the following steps: performing target detection on the original image to determine the position region of the animation character in the original image, and generating a positioning area. Perform pose recognition on the animation character based on the positioning area, to obtain the joint part points of the animation character in the original image.

6. The method of claim 1, wherein, The joint part points of the animation character are respectively bound with each frame of the joint motion information to generate a target animation, including: Perform mesh texturing processing on the animation character to generate a plurality of meshes on the animation character, and associate the joint part points of the animation character with the meshes; Adjust the meshes associated with the joint part points of the animation character according to the joint part point motion indicated by the joint motion information, so that the animation character performs the action indicated by the action description information.

7. The method of claim 6, wherein, The adjustment of the meshes associated with the joint part points of the animation character according to the joint part point motion indicated by the joint motion information includes: Determine target positions of each joint part point of the animation character based on the joint motion information; each target position is the actual position of the corresponding joint part point after the animation character performs the action described by the joint motion information; Translate the corresponding joint part point according to each target position to complete the update of the mesh associated with each joint part point.

8. An animation generation apparatus characterized by comprising: It includes: An image acquisition module is configured to acquire an original image; the original image includes an animation character; the animation character in the original image is subjected to background matting processing and filling processing respectively to obtain a first mask image and a second mask image; the first mask image includes an animation character region and a background region different from the animation character region; the animation character region in the second mask image is different from the animation character region in the first mask image; an image meeting a detection condition is selected from the first mask image and the second mask image as a target mask image; The target mask image includes an animation character region and a background region; the animation character region is different from the background region; The target mask image has higher precision than the first mask image or the second mask image; An information acquisition module is configured to acquire action description information, which is used to describe an action expected to be performed by the animation character; An action generation module is configured to input the action description information into an action generation model to obtain an action file; the action file includes at least one frame of joint motion information; The joint motion information is used to describe the joint part point motion corresponding to each action; An animation generation module is configured to extract the animation character from the original image using the animation character region in the target mask image, and bind the joint part points of the animation character with each frame of the joint motion information to generate a target animation; The target animation is used to show the animation character performing the action described by the action description information.

9. An electronic device, comprising: It includes: At least one processor and at least one memory, The memory has computer readable instructions stored thereon; The computer readable instructions are executed by one or more processors to enable the electronic device to implement the animation generation method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Animation generation method and device and electronic equipment

    CN117132687A