A method, apparatus, device, and readable medium for generating animations

By using the action library and feature vector matching technology of real action acquisition, the problem of animation distortion of the deep learning algorithm model output is solved, and more realistic and efficient animation generation is achieved, suitable for real-time applications.

CN114998487BActive Publication Date: 2025-07-08GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210495283.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-07-08
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

In the existing animation generation technology, the animation quality output by the deep learning algorithm model is affected by the data set scale and learning ability, resulting in distortion, poor user experience, and difficult to apply to real-time scenarios.

Method used

A pre-set action library is used to store action frames collected based on real action, match target action frames through external action signals, and efficient search is carried out based on feature vector libraries to generate animations with higher authenticity.

Benefits of technology

The generated animation is more realistic and accurate, suitable for real-time scenes, significantly improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998487B_ABST
    Figure CN114998487B_ABST
Patent Text Reader

Abstract

This specification provides a method, apparatus, device, and readable medium for generating an animation. The method includes: obtaining an external action signal, where the external action signal is used to represent a target action of a target animation; obtaining a target action frame that matches the external action signal from a specified action library, where the specified action library stores at least one animation frame, and the animation frame contains three-dimensional skeleton data generated based on a real action; processing the target action frame based on the external action signal to generate the target animation with the target action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of animation generation, and particularly to a method, apparatus, device, and readable medium for animation generation. Background Art

[0002] In the technical field of animation generation, a common technical means is to train a deep learning algorithm model so that the model can simulate animation actions according to the characteristics of the input actions and output the simulation results. After training, the deep learning algorithm receives input signals, and after the signal processing is completed, the synthesized animation is output. However, the animations synthesized by this method will have varying degrees of distortion depending on the scale and quality of the deep learning algorithm dataset and the learning ability of the algorithm model, resulting in a poor user experience. Summary of the Invention

[0003] To overcome the problems existing in the related art, this specification provides a method and apparatus for animation generation.

[0004] According to the first aspect of the embodiments of this specification, a method for animation generation is provided. The method includes:

[0005] Obtain an external action signal, where the external action signal is used to represent the target action of the target animation;

[0006] Obtain a target action frame that matches the external action signal from a specified action library, where the specified action library stores at least one animation frame, and the animation frame contains three-dimensional skeleton data generated based on real actions;

[0007] Process the target action frame based on the external action signal to generate the target animation with the target action.

[0008] According to the second aspect of the embodiments of this specification, an apparatus for animation generation is provided, including:

[0009] A signal acquisition module, configured to acquire an external action signal, where the external action signal is used to represent the target action of the target animation;

[0010] A matching module, configured to obtain a target action frame that matches the external action signal from a specified action library, where the specified action library stores at least one animation frame, and the animation frame contains three-dimensional skeleton data generated based on real actions;

[0011] An animation synthesis module, configured to process the target action frame based on the external action signal to generate the target animation with the target action.

[0012] According to the third aspect of the embodiments of this specification, a device is provided, including:

[0013] Processor;

[0014] A memory for storing instructions executable by the processor; wherein the processor is configured to perform the operations described in the method of the first aspect of this specification.

[0015] According to the fourth aspect of the embodiments of this specification, there is provided a computer-readable storage medium having computer instructions stored thereon, including:

[0016] When the instruction is executed by the processor, it implements the operations described in the first aspect of this specification.

[0017] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:

[0018] In the embodiments of this specification, the action library set in advance stores action frames collected according to real actions. When in use, an external action signal is obtained, and based on the external action signal, the action frame most similar to the external action signal is searched for in the action library, and then the action frame is adjusted with reference to the external action signal so that the action frame can contain the action features of the external action signal. In this way, the generated animation includes some features of the real action because it has the features of the action frame collected according to the real action, and because it includes the features of the target action, it will be more real and accurate.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0021] Figure 1A It is a schematic diagram of the prior art animation generation scheme.

[0022] Figure 1B It is a schematic diagram of an animation generation scheme proposed in this specification.

[0023] Figure 1C It is a flowchart of a method for generating an animation shown according to an exemplary embodiment of this specification.

[0024] Figure 2A It is an image collected from the action library shown according to an exemplary embodiment of this specification.

[0025] Figure 2B It is a schematic diagram of the feature vector library and the action library shown according to an exemplary embodiment of this specification.

[0026] Figure 3 It is a schematic diagram of a feature vector matching process shown in this specification according to an exemplary embodiment.

[0027] Figure 4 It is a schematic diagram of target frame processing shown in this specification according to an exemplary embodiment.

[0028] Figure 5 It is a block diagram of an animation generation device shown in this specification according to an exemplary embodiment.

[0029] Figure 6 It is an application flowchart in the scenario of virtual character control shown in this specification according to an exemplary embodiment.

[0030] Figure 7 It is a schematic diagram of an animation generation device shown in this specification according to an exemplary embodiment. Detailed implementation manners

[0031] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.

[0032] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0033] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0034] Currently, existing animation generation solutions on the market are as Figure 1AAs shown: The deep learning algorithm receives an input signal and outputs a corresponding synthetic animation, such as outputting an animation according to a given piece of music or video (RGB visual motion capture algorithm). The quality of the synthetic animation output by the model depends on two aspects: the scale and quality of the dataset and the learning ability of the model. Usually, there are inevitable losses in the rationality, naturalness, and accuracy of the actions output by the model, which often lead to animation distortion problems. At the same time, it cannot be guaranteed that the solution can be used in real-time scenarios, resulting in a poor user experience.

[0035] Next, the embodiments of this specification will be described in detail.

[0036] As Figure 1C shown, Figure 1C is a flowchart of a method for generating an animation shown in this specification according to an exemplary embodiment, including the following steps:

[0037] In step S101, an external action signal is obtained, and the external action signal is used to represent the target action of the target animation;

[0038] In step S102, a target action frame matching the external action signal is obtained from a specified action library, where the specified action library stores at least one animation frame, and the animation frame contains three-dimensional skeleton data generated based on real actions;

[0039] In step S103, the target action frame is processed based on the external action signal to generate a target animation with the target action.

[0040] In step S101, obtaining the external signal can be achieved by collecting the action signal through a visual motion capture device, such as shooting with an action camera, or obtaining the action signal from a bionic device, such as capturing motion instructions through a mouse, keyboard, VR device, etc. These motion instructions are usually used to control the movements of animated characters, such as moving and jumping.

[0041] In step S102, the specified action library can be pre-made. In one embodiment, high-quality action recording is performed using an inertial capture device or an optical capture device, and the captured images are as Figure 2A shown. The captured animations include running, walking, daily actions, sports actions, martial arts actions, etiquette actions, etc. Some data in the action library can use open-source high-quality actions, but it is necessary to ensure that the character joint topologies of all actions are consistent. Denote the number of frames in the saved basic action library as T, where the data of each frame of animation mainly includes the global translation and global rotation of the root joint, as well as the local translation and local rotation of other joints, denoted as y = [p r q rp, q], and calculate the relevant data of the high-quality action library as database, including motion data such as trajectories, joint positions, rotations, velocities, and angular velocities in the action library. The specific operations are as follows: (a) Project the motion data p r , q r of the root in the action library onto the xoz plane, that is, remove the translation and rotation components in the y direction. The translation data is denoted as The direction data is denoted as as the motion data of a virtual joint O, which is used to describe the motion trajectory of the entire character in the world coordinate system. Joint O serves as the new root joint of the entire topology, and the number of joints is J at this time; (b) Calculate the offset of the local position of each joint relative to its parent joint. Since the new joint O changes the hierarchical relationship of the joint topology, it is necessary to calculate the local translation of the root node relative to joint O, which is recorded together with the local translations of other joints as (c) The local rotation of each joint relative to its parent joint is represented by a quaternion (A quaternion is a point on a hypersphere in four-dimensional space, satisfying x 2 + y 2 + z 2 + w 2 = 1; There are 3 imaginary parts, and they are pairwise orthogonal. The real part is cosθ / 2, and the imaginary part is a unit axis multiplied by sinθ / 2. Each quaternion can correspond to a feature vector, that is, n.) The local rotation of the root node relative to joint O, which is recorded together with the local rotations of other joints as (d) Calculate the motion velocity of each joint relative to its parent joint The calculation formula is where f is the frame rate of the animation, and f is defaulted to 60 frames; (e) Calculate the angular velocity of each joint relative to its parent joint The calculation formula is The angular velocity is represented by axis-angle, which is convenient for subsequent kinematic calculations. The above motion data set is the database, denoted as

[0042] Optionally, after obtaining the external action signals, a first feature vector can be obtained based on these external action signals, and a target action frame matching the first feature vector can be obtained from a specified action library. In this embodiment, the above steps include: obtaining a second feature vector of the animation frames in the action library; establishing a feature vector library based on the second feature vector; obtaining a second feature vector matching the first feature vector in the feature vector library; and obtaining a target action frame in the action library based on the matched second feature vector. The above second feature vector includes position vectors, velocity vectors, trajectory vectors, etc. of some key nodes of the animation frame in the global coordinate system. The feature vector library is a unique feature vector library. The key nodes include the wrists of both hands and the ankles of both feet of the action model. The relationship between the feature vector library and the action library is as Figure 2B shown. In this embodiment, the above key nodes are mainly based on the {P o , Q o , P l , Q l , V l , W l} motion data in making a high-quality action library and combined with rigid body kinematics calculation. The steps are as follows: (a) Position vector calculation: First, the global positions of each joint are calculated by forward kinematics. Denote the global positions of all joints after forward kinematics as Then, the de-positioning and de-rotation transformation of the global positions are performed, that is The subscript j represents the jth joint, then the global position feature vector of the key joint is The dimension is 15. (b) Velocity vector calculation: Denote the global velocities of all joints calculated by rigid body kinematics as Perform the de-rotation transformation of the global velocities, that is Then the global velocity feature vector of the key joint is The dimension is 15. (c) Trajectory position vector calculation: The trajectory position vector consists of the trajectory data of three frames at the 20th, 40th, and 60th frames after the current frame. Similarly, the de-positioning and de-rotation transformation are performed, that is where n represents the integers 20, 40, 60, then the obtained trajectory position feature is The dimension is 9. (d) Similarly, calculate the de-centered orientation vector of the trajectory. The default orientation of the system is [0, 0, 1] T , then Then the trajectory orientation feature is The dimension is 9. Finally, the animation feature corresponding to the ith frame of the animation is The dimension is 48. Thus, a mapping between the feature and the animation y i is established. Denote the feature library as And normalize each dimension of the 48-dimensional feature, and save the normalized mean value and standard deviation.

[0043] Optionally, during matching, each second feature vector in the feature vector library can be divided into one of the feature vector sets, and the similarity of the second feature vectors in each feature vector set satisfies a first threshold. The step of obtaining the second feature vector that matches the first feature vector in the feature vector library includes: determining the feature vector set whose matching degree with the first feature vector reaches a second threshold; in the determined feature vector set, obtaining the second feature vector whose matching degree with the first feature vector reaches a third threshold. As an example, as Figure 3 shown Figure 3 is a schematic diagram of the matching process. During the matching process, let the normalized action feature vector obtained from the input signal through the animation system be F 48 , and it is necessary to find the one in the feature library that is closest to F 48 The corresponding animation frame y i is the animation that best matches the current state. The fast search of feature vectors is accelerated by using the axis-aligned bounding box algorithm, which is divided into two bounding boxes, a large one and a small one. The large bounding box is used for rough matching, and the small bounding box is used for precise matching. (a) First, the entire feature library is segmented. The length of the feature vector sequence in the large bounding box is 64, and the length of the feature vector sequence in the small bounding box is 16. Then the sets of large and small bounding box sequences after segmentation are and There are T / 64 and T / 16 respectively, and each large bounding box contains 4 small bounding boxes; (b) Calculate the upper and lower bounds of each bounding box. Taking the kth large bounding box as an example, it contains the animations from the ((k - 1)×64)th frame to the (k×64)th frame in the animation library. Therefore, the dimension of the feature matrix of the animations in the bounding box is 64×48. Traverse each frame of animation in the bounding box and each dimension of the feature, and obtain the maximum eigenvalue and the minimum eigenvalue at the corresponding dimension position, and finally obtain the upper and lower bound vectors Both dimensions are 48. Finally, the set of upper and lower bounds of all large bounding boxes obtained is and the set of upper and lower bounds of small bounding boxes is (c) Use the loss value of the previous match as the threshold, traverse all bounding boxes, and find the match for the action feature F 48 . First, search in the large bounding box , and calculate the L2 loss values 48 of F and respectively with the upper and lower bounds If both are greater than the threshold, it means that the best match does not exist in this bounding box, and then continue to search in the next bounding box ; Assume that the loss calculated in the bounding box or If it is less than the threshold, there is an optimal match in this large bounding box, and further search is performed in the 4 included small bounding boxes. Similarly, calculate the loss value in the small bounding box If both are greater than the threshold, continue the search in the subsequent bounding box If any one is less than the threshold, there is an optimal match in this small bounding box, and it is necessary to traverse the 16 included feature vectors to find the feature vector with the smallest loss value Take the corresponding animation frame y i as the optimal animation output, and record the loss value as the threshold for the next match. It should be noted that steps (a) and (b) are only calculated once when the system starts, while step (c) is executed every time a match is made. If the brute-force algorithm is used for animation feature search, the complexity is o(T), while the two-layer axis-aligned bounding box search algorithm of the present invention reduces the algorithm complexity to o(T / 64), greatly improving the search speed. In actual use, after each match, the subsequent animation frames will continue to be played until the preset number of frames (usually set to 10) or a new control signal is generated, and then the next round of matching will be performed. For example, if the i-th frame of the animation is matched this time, the subsequent i + 1, i + 2 frames will continue to be played until the i + 10-th frame or the user inputs a new signal, and then re-matching will be performed. It should be noted that in the above process, it is assumed that the normalized action feature vector F 48 , and the source of this vector can come from various algorithms. Taking the visual motion capture algorithm as an example below, the generation of the feature vector is briefly introduced. Usually, the data output by a frame of the visual motion capture algorithm is mainly the global translation and global rotation of the root joint, as well as the local rotation of other joints, denoted as y c =[p c q c q]. Since the character topology in the high-quality action library contains the length of each bone, the ratio of the actor's height to the character's height in the visual capture algorithm is used to scale the global translation p c of the root joint to obtain and combine the joint local translation into the motion capture data to obtain Combined with the motion state of the previous frame, the speed and angular velocity can be calculated, so as to obtain data similar to that in step S101 Based on rigid body kinematics, the global position feature vector and global velocity feature vector p 48 in F can be obtained g ,v g . Based on the historical data of the motion projection of the current root joint on the xoz plane, the trajectories of the subsequent 20, 40, and 60 frames can be simply predicted, and the trajectory position feature vector and trajectory orientation feature vector p 48 in F can be obtained o ,r​o , then the final F 48 is [p g v g p o r o , and it needs to be normalized based on the average value and standard deviation of the action library.

[0044] In step S103, there are various ways to process the target action frame. In this embodiment, the processing method is to fuse the above first feature vector with the target action frame. The part of the fusion here mainly modifies the motion state of the O joint of the character in the original animation and does not modify the motion data of other joints. Let the position and orientation of the O joint of the current character be p c , q c , the original animation matched is the i-th frame, and the corresponding position and orientation of the O joint are respectively The actually played animation should start from the (i + 1)-th frame. Therefore, the motion data in the (i + 1)-th frame of the original animation is updated. Then, the fused position and orientation are as shown in the following formula (1), and the animation data that the system should play is can completely represent an animation frame. The first two data represent the global position and global rotation of the fused O joint, and the last two data are the original local motion data of other joints (including the root joint) in the action library. As Figure 4 shown, the subsequent animations in the animation segment also need to be fused.

[0045]

[0046] The animation generation scheme proposed by the present invention is as Figure 1B shown. First, based on the high-quality action library, calculate the feature vector corresponding to each frame of the animation to form a feature library, and save the mapping between the feature library and the action library; then, based on the user's input signal, such as keyboard and mouse operations or actor actions, combine physical simulation or visual motion capture algorithms to calculate the action feature vector, and use an efficient real-time feature search algorithm to find the most matching unique feature vector in the feature library; finally, based on the unique feature vector, find the corresponding animation frame, which is the closest to the current action state and can be used as a substitute for the current action state to play, thereby generating reasonable and natural high-quality animations, which can be used in aspects such as driving optimization and game operations, significantly improving the user experience.

[0047] Corresponding to the embodiment of the foregoing method, this specification also provides an embodiment of the apparatus and the terminal to which it is applied.

[0048] As Figure 5 shown, Figure 5 is a block diagram of an animation generation device shown in this specification according to an exemplary embodiment. The device includes:

[0049] A signal acquisition module 501, configured to acquire an external action signal, where the external action signal is used to represent a target action of a target animation;

[0050] A matching module 502, configured to obtain a target action frame matching the external action signal from a specified action library, where the specified action library stores at least one animation frame, and the animation frame includes three-dimensional skeleton data generated based on a real action;

[0051] An animation synthesis module 503, configured to process the target action frame based on the external action signal to generate a target animation with the target action.

[0052] Wherein, the signal acquisition module 501 is further configured to: after acquiring the external action signal, obtain a first feature vector of the external action signal; the process of the matching module 502 obtaining a target action frame matching the external signal from the specified action library includes: obtaining a target action frame matching the first feature vector from the specified action library; the matching module 502 being configured to obtain a target action frame matching the first feature vector from the specified action library includes: establishing a feature vector library based on a second feature vector; obtaining a second feature vector matching the first feature vector in the feature vector library; and obtaining a target action frame in the action library based on the matched second feature vector.

[0053] In addition, the matching module 502 is further configured to divide each second feature vector in the feature vector library into one of the feature vector sets, where the similarity of the second feature vectors in each feature vector set satisfies a first threshold, and the step of the matching module 502 obtaining a second feature vector matching the first feature vector in the feature vector library includes:

[0054] Determining a feature vector set with a matching degree to the first feature vector reaching a second threshold; and obtaining a second feature vector with a matching degree to the first feature vector reaching a third threshold in the determined feature vector set.

[0055] The implementation processes of the functions and roles of each module in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0056] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0057] One application scenario of the animation generation solution provided in this specification is the control of virtual animation characters. In games or other scenarios that require controlling virtual characters, virtual characters are usually presented in the form of animated characters. When users control virtual characters to perform actions such as turning around and walking, it is necessary to make the virtual characters fully reflect the characteristics of the actions as much as possible, that is, to make the actions more realistic. The following takes the scenario where users operate virtual characters to perform actions as an example to elaborate an application example. The animation generation solution provided in this specification can be deployed on the client operated by users or on the server of the application. Deploying this solution on the client is as follows Figure 6 shown:

[0058] In step S601, the client obtains an external action signal through a client bionic device;

[0059] In step S602, the client obtains a target action frame that matches the external action signal from a specified action library. The specified action library stores at least one animation frame, and the animation frame contains three-dimensional skeleton data generated based on real actions;

[0060] In step S603, the client processes the target action frame based on the external action signal to generate a target animation with the target action.

[0061] In step S604, the client sends the generated target animation to the server, and the server renders the above target animation and synchronizes it to the current display screen of the client.

[0062] It should be noted that in step S602, the specified action library can be a pre-generated action library provided by the service provider or an open-source high-quality action library found by the user according to their own needs.

[0063] Correspondingly, as Figure 7 shown, this application also provides an image anti-shake device 70, including a processor 71; a memory 72 for storing executable instructions, and the memory 72 includes a computer program; wherein, the processor 71 is configured to:

[0064] Obtain an external action signal, where the external action signal is used to characterize the target action of the target animation;

[0065] Obtain a target action frame that matches the external action signal from a specified action library, where the specified action library stores at least one animation frame, and the animation frame contains three-dimensional skeleton data generated based on real actions;

[0066] Process the target action frame based on the external action signal to generate the target animation with the target action.

[0067] The processor 71 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0068] The memory 72 may include at least one type of storage medium, and the storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. Moreover, the device may cooperate with a network storage device that performs the storage function of the memory through a network connection. The memory 72 may be an internal storage unit of the device 70, such as the hard disk or memory of the device 70. The memory 72 may also be an external storage device of the device 70, such as a plug-in hard disk equipped on the device 70, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0069] Furthermore, the memory 72 may also include both an internal storage unit of the device 70 and an external storage device. The memory 72 is used to store computer programs and other programs and data required by the device. The memory 72 may also be used to temporarily store data that has been output or will be output.

[0070] Device 70 may be a computing device such as a desktop computer, a notebook, a handheld computer, and a cloud server. The device may include, but is not limited to, a processor 71 and a memory 72. Those skilled in the art can understand that Figure 7 merely examples of device 70, which do not constitute a limitation on device 70, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the device may further include input / output devices, network access devices, buses, etc.

[0071] The implementation processes of the functions and roles of the respective units in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0072] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0073] Those skilled in the art will readily conceive of other implementations of this specification after considering the specification and practicing the invention herein. This specification is intended to cover any variations, uses, or adaptations of this specification, which follow the general principles of this specification and include common general knowledge or conventional technical means in the technical field not claimed in this application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of this specification are pointed out by the following claims.

[0074] It should be understood that this specification is not limited to the exact structures described above and shown in the figures, and various modifications and changes can be made without departing from its scope. The scope of this specification is only limited by the appended claims.

[0075] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of protection of this specification.

Claims

1. A method for generating an animation, characterized in that, The method includes the steps of: Obtaining an external action signal, where the external action signal is used to represent a target action of a target animation; Obtaining a target action frame that matches the external action signal from a specified action library, where the specified action library stores at least one animation frame, and each of the animation frames contains three-dimensional skeleton data generated based on a real action. The three-dimensional skeleton data includes the global translation and global rotation of the root joint, as well as the local translation and local rotation of other joints. The global translation and global rotation of the root joint can generate motion data of a virtual joint O, and the motion data of the virtual joint O is used to describe the motion trajectory of the entire character in the world coordinate system; Based on the motion data of the virtual joint O of the external action signal and the motion data of the virtual joint O of the target action frame, updating the motion data of the virtual joint O of the animation frames after the target action frame to generate the target animation with the target action.

2. The method according to claim 1, characterized in that, Obtaining an external action signal, and obtaining a target action frame that matches the external action signal from a specified action library includes: Obtaining a first feature vector of the external action signal; obtaining a target action frame that matches the first feature vector from a specified action library.

3. The method according to claim 2, characterized in that, Obtaining a target action frame that matches the first feature vector from a specified action library includes: Obtaining a second feature vector of the animation frame in the action library; Establishing a feature vector library based on the second feature vector; Obtaining a second feature vector that matches the first feature vector in the feature vector library; Obtaining the target action frame in the action library based on the matched second feature vector.

4. The method according to claim 3, wherein Each of the second feature vectors in the feature vector library is divided into one of the feature vector sets, and the similarity of the second feature vectors in each feature vector set satisfies a first threshold; The step of obtaining a second feature vector that matches the first feature vector in the feature vector library includes: Determining the feature vector set whose matching degree with the first feature vector reaches a second threshold; In the determined feature vector set, obtaining a second feature vector whose matching degree with the first feature vector reaches a third threshold.

5. An apparatus for generating an animation, characterized in that, The device includes: A signal acquisition module, configured to acquire an external action signal, where the external action signal is used to represent a target action of a target animation; A matching module, configured to obtain a target action frame that matches the external action signal from a specified action library, where the specified action library stores at least one animation frame, and each of the animation frames contains three-dimensional skeleton data generated based on a real action. The three-dimensional skeleton data includes the global translation and global rotation of the root joint, as well as the local translation and local rotation of other joints. The global translation and global rotation of the root joint can generate motion data of a virtual joint O, and the motion data of the virtual joint O is used to describe the motion trajectory of the entire character in the world coordinate system; An animation synthesis module, configured to update the motion data of the virtual joint O in the animation frames after the target action frame based on the motion data of the virtual joint O in the external action signal and the motion data of the virtual joint O in the target action frame, so as to generate the target animation with the target action.

6. The apparatus according to claim 5, wherein the signal acquisition module is further configured to: after acquiring the external action signal, obtain a first feature vector of the external action signal; the process of the matching module obtaining the target action frame matching the external action signal from the specified action library includes: obtaining the target action frame matching the first feature vector from the specified action library.

7. The device according to claim 6, characterized in that, The matching module is configured to obtain the target action frame matching the first feature vector from the specified action library, including: obtaining a second feature vector of the animation frame in the action library; establishing a feature vector library based on the second feature vector; obtaining, in the feature vector library, a second feature vector matching the first feature vector; obtaining the target action frame in the action library based on the matched second feature vector.

8. The device according to claim 7, characterized in that, The matching module is further configured to divide each of the second feature vectors in the feature vector library into one of the feature vector sets, and the similarity of the second feature vectors in each feature vector set satisfies a first threshold; The step of the matching module obtaining, in the feature vector library, a second feature vector matching the first feature vector includes: determining the feature vector set whose matching degree with the first feature vector reaches a second threshold; obtaining, in the determined feature vector set, a second feature vector whose matching degree with the first feature vector reaches a third threshold.

9. An electronic device, characterized in that, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the operations described in the method according to any one of claims 1-4.

10. A computer-readable storage medium, on which computer instructions are stored, wherein: when the instructions are executed by a processor, the operations described in the method according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Animation image generation method and device and storage medium

    CN113345057A

  • Action driving method and device of virtual teacher system, equipment and storage medium

    CN114022645A