Interactive animation generation model training, interactive animation generation method and system
Through machine learning models, interactive animation generation models are trained, and interactive animations between characters and objects that conform to physical laws and human body habits are automatically generated, solving the problems of low efficiency and large errors in the existing methods, and achieving efficient animation production.
Patent Information
- Application Number
- CN202110883086.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-02
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-08-02
AI Technical Summary
The existing interactive animation production methods require a lot of manual intervention, which is costly and prone to calculation errors and motion problems that do not conform to human habits.
Through machine learning model training, interactive animation generation model is trained using sample input data and label data, bone motion parameters during the interaction between characters and objects are generated, and interactive animations are automatically generated.
The efficiency of interactive animation production is improved, the generated animations are in line with physical laws and human habits, and the cost of manual intervention and calculation is reduced.
Smart Images

Figure CN113570690B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to animation production, and in particular to an interactive animation generation model training, interactive animation generation method and system. Background Art
[0002] The application of animation production is becoming more and more extensive. For example, animation technology is often required in the production of cartoons, movies, games, advertisements and other works.
[0003] Currently, it is hoped to provide an efficient interactive animation production method. Summary of the Invention
[0004] One of the embodiments of this specification provides a method for training an interactive animation generation model. The method may include: obtaining a plurality of sample input data, wherein the sample input data includes initial skeletal state parameters of a character, spatial distribution parameters of a rigid object, and motion trajectory parameters of the rigid object, wherein the initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the character, and the motion trajectory parameters at least indicate the initial position, initial posture, target position, and target posture of the rigid object; obtaining a plurality of sample label data corresponding to the plurality of sample input data, wherein the sample label data includes skeletal motion parameters of the character, and the skeletal motion parameters indicate the position and posture of the one or more bones at at least two time points during the interaction between the character and the rigid object; and training an initial model using the plurality of sample input data and the plurality of sample label data to obtain the interactive animation generation model.
[0005] One of the embodiments of this specification provides an interactive animation generation method. The method may include: obtaining initial skeletal state parameters of a target character, spatial distribution parameters of a target rigid object, and motion trajectory parameters of the target rigid object, wherein the initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the target character, and the motion trajectory parameters at least indicate the initial position, initial posture, target position, and target posture of the target rigid object; inputting the initial skeletal state parameters of the target character, the spatial distribution parameters of the target rigid object, and the motion trajectory parameters of the target rigid object into an interactive animation generation model to obtain skeletal motion parameters of the target character output by the interactive animation generation model, wherein the skeletal motion parameters indicate the position and posture of one or more bones at at least two time points during the interaction between the target character and the target rigid object; and generating an interactive animation between the target character and the target rigid object based on the skeletal motion parameters of the target character.
[0006] One of the embodiments of this specification provides an interactive animation generation model training system. The system may include a sample input data acquisition module, a sample label data acquisition module, and a training module. The sample input data acquisition module can be used to acquire multiple sample input data, wherein the sample input data include the initial skeletal state parameters of the character, the spatial distribution parameters of the rigid object, and the motion trajectory parameters of the rigid object. The initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the character, and the motion trajectory parameters at least indicate the initial position, initial posture, target position, and target posture of the rigid object. The sample label data acquisition module can be used to acquire multiple sample label data corresponding to the multiple sample input data, wherein the sample label data include the skeletal motion parameters of the character, and the skeletal motion parameters indicate the position and posture of the one or more bones at at least two time points during the interaction between the character and the rigid object. The training module can be used to train an initial model using the multiple sample input data and the multiple sample label data to obtain the interactive animation generation model.
[0007] One embodiment of this specification provides an interactive animation generation system. The system may include an input parameter acquisition module, an output parameter acquisition module, and an interactive animation generation module. The input parameter acquisition module may be used to acquire initial skeletal state parameters of a target character, spatial distribution parameters of a target rigid object, and motion trajectory parameters of the target rigid object. The initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the target character, and the motion trajectory parameters at least indicate the initial position, initial posture, target position, and target posture of the target rigid object. The output parameter acquisition module may be used to input the initial skeletal state parameters of the target character, the spatial distribution parameters of the target rigid object, and the motion trajectory parameters of the target rigid object into an interactive animation generation model to obtain skeletal motion parameters of the target character output by the interactive animation generation model. The skeletal motion parameters indicate the position and posture of one or more bones of the target character at at least two time points during the interaction between the target character and the target rigid object. The interactive animation generation module may be used to generate an interactive animation between the target character and the target rigid object based on the skeletal motion parameters of the target character.
[0008] One embodiment of this specification provides an interactive animation generation model training device. The device includes a processor and a storage device, wherein the storage device is used to store instructions. When the processor executes the instructions, the interactive animation generation model training method described in any embodiment of this specification is implemented.
[0009] One embodiment of this specification provides an interactive animation generation device, which includes a processor and a storage device, wherein the storage device is used to store instructions, and when the processor executes the instructions, the interactive animation generation method described in any embodiment of this specification is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0011] Figure 1 is an exemplary flow chart of a method for training an interactive animation generation model according to some embodiments of this specification;
[0012] Figure 2 is an exemplary flow chart of a method for generating an interactive animation according to some embodiments of this specification;
[0013] Figure 3 is a schematic diagram of an exemplary structure of a neural network for generating interactive animation according to some embodiments of this specification;
[0014] Figure 4 is an exemplary module diagram of an interactive animation generation model training system according to some embodiments of this specification;
[0015] Figure 5 It is an exemplary module diagram of an interactive animation generation system according to some embodiments of this specification. DETAILED DESCRIPTION
[0016] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0017] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0018] As used in this specification, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprise" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0019] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0020] The movement (such as actions) of characters in three-dimensional animation is generally achieved using three-dimensional skeletal animation technology. Skeletal animation binds each vertex on the surface of the character's three-dimensional model to several bones, and the coordinates of the bones jointly determine the coordinates of the vertices on the model surface. The calculation of the movement of objects is generally divided into two types: rigid objects and non-rigid objects. This specification only discusses interactive animation for rigid objects. It should be understood that a rigid object refers to an object that does not deform (or the deformation is considered negligible). The state of a rigid object can be determined based on position and posture (also called "rotation"). For the sake of simplicity, the "objects" that appear below can specifically refer to rigid objects.
[0021] 3D animation involves a lot of interactions between characters and objects. Here are a few ways to create interactive animations between characters and objects.
[0022] In some embodiments, in animation software, animators can adjust the position and rotation parameters of the skeleton frame by frame to adjust the character to the desired state. Correspondingly, the position and rotation parameters of the object can also be adjusted frame by frame to ensure that the entire interactive animation conforms to the laws of physics. This method (denoted as Method 1) requires professional animators to spend a lot of time and effort to ensure the quality of the animation.
[0023] In some embodiments, motion capture technology can be used to capture and calculate the position and rotation parameters of the actor's skeleton using professional equipment. These parameters are then imported into the model through animation software. The animator then checks the model's motion effects and makes corrections when necessary. Although this method (denoted as Method 2) can improve animation production efficiency compared to Method 1, it introduces the cost of hiring actors and using motion capture venues and equipment. In addition, in actual applications, since the data often contains errors, it is almost impossible to use it without manual correction, that is, the cost of manual correction is often also included.
[0024] In some embodiments, the IK (Inverse Kinematics) algorithm is used to preset target points on the object model. Each target point corresponds to a designated touch point on the hand. The touch point then drives the hand to move to the target point, and the IK algorithm is used to reversely calculate the skeletal position and rotation value of the hand and arm. This method (denoted as method 3) is commonly used in high-precision 3D game production, and target points and hand touch points need to be set in advance for each model. Due to the use of pure spatial physical relationship calculations, the IK algorithm is prone to the problem of hand and object penetration. In addition, the calculated limb movement process is generally the shortest path, which is prone to movements that are not in line with human habits.
[0025] The embodiments of this specification also provide a method for automatically generating motion parameters required for a three-dimensional character to interact with an object using a machine learning model. After a user inputs one or more of the following information: the character's initial skeletal state, the object's spatial distribution, the object's initial position, initial posture, target position, and target posture, the interactive animation generation model can generate skeletal motion parameters for the character based on the user input. These skeletal motion parameters can indicate the position and posture of the skeleton at at least two time points during the interaction between the character and the object. Furthermore, an interactive animation between the character and the object can be generated based on the character's skeletal motion parameters.
[0026] The following is an explanation of the two stages: model training and model prediction.
[0027] Figure 1 This is an exemplary flow chart of the interactive animation generation model training method according to some embodiments of this specification. Process 100 can be executed by one or more processors. Process 100 corresponds to the model training stage, wherein the sample input data and the sample label data can be collectively referred to as training data. Figure 1 As shown, process 100 may include:
[0028] Step 110 , obtaining a plurality of sample input data. In some embodiments, step 110 may be implemented by the sample input data obtaining module 410 .
[0029] The sample input data may include initial skeletal state parameters of a character, spatial distribution parameters of an object, and motion trajectory parameters of the object. The initial skeletal state parameters may indicate the initial position and initial posture of one or more bones of the character, and the motion trajectory parameters may indicate at least the initial position, initial posture, target position, and target posture of the object.
[0030] In some embodiments, the one or more skeletons may include bones from the shoulder to the fingers. In practical applications, it is possible to focus only on the movement from the shoulder to the hand during the character's interaction with an object, ignoring the movement of other parts of the body during the interaction. This can significantly reduce the data collection cost and computational effort during the model training and prediction phases.
[0031] It should be noted that the bones mentioned in this specification can be derived from real biological bone structures (such as human bones) or customized (for example, modified from real biological bone structures or completely fictitious).
[0032] In some embodiments, in order to reduce the amount of data calculation, the model of the object (the expression of the object in the computing device) can be processed with low-polygon count polygons, and accordingly, the spatial distribution parameters of the object are obtained based on the model that has been processed with low-polygon count polygons. Low-polygon count polygons can refer to representing an object with polygons with fewer faces, which means that some of the concave and convex details of the object can be ignored. It can be understood that low-polygon count polygons can be performed in combination with actual needs. For example, the level of low-polygon count polygons can be adjusted according to the level of sophistication of the interactive animation. For another example, low-polygon count polygons can be performed according to the specific type of interactive animation. As an example only, when the interactive animation is to open the door by hand, only details that are strongly related to the door opening animation, such as the door handle, can be retained during low-polygon count polygons to reduce the amount of calculation.
[0033] In some embodiments, the spatial distribution of the object model can be limited to a cubic space (regular hexahedron) of a preset size, that is, a cubic space that does not exceed a preset size (volume). The cubic space of the preset size can be divided into multiple sub-cube spaces, and accordingly, the spatial distribution parameters of the object (hereinafter referred to as the object model parameters) can indicate the proportion of the part of the object model in each sub-cube space to the sub-cube space. For example only, a cubic space with a side length of 1m can be divided into 100 3 Sub-cube spaces, each sub-cube space is 1cm in size 3 Correspondingly, the model parameters of the object can be a 100*100*100 three-dimensional tensor (with 100*100*100 points / components), where each point / component can take a value between 0-1.
[0034] Step 120 , obtaining a plurality of sample label data corresponding to the plurality of sample input data. In some embodiments, step 120 may be implemented by the sample label data obtaining module 420 .
[0035] The sample label data may include skeletal motion parameters of the character, which may indicate the position and posture of one or more bones at at least two time points during the character's interaction with an object. It is understood that the skeletal motion parameters are supervisory signals for model training.
[0036] In some embodiments, one or more parameters mentioned in this specification (such as initial skeletal state parameters, skeletal motion parameters, and motion trajectory parameters of an object) can be represented by three-dimensional coordinates to represent the position (in three-dimensional space), and by a rotation quaternion to represent the rotation (i.e., posture) (in three-dimensional space). The rotation quaternion includes the elements (3 numbers) represented by the three-dimensional vector of the rotation axis and the rotation angle (1 number) around the rotation axis. In this way, if the position and rotation of an object (for example, a bone or object of a character) are regarded as the state of the object, 7 numerical values can be used to represent the state of the object. Specifically, the initial skeletal state parameters of the character can be or include an array of length N*(3+4), where N is the number of the one or more bones. The initial state and target state of the object (which can constitute the motion trajectory parameters of the object) can each be 7 numerical values. The sample label data or model output can be or include a frame sequence of T*N*(3+4), where N also represents the number of the one or more bones and T represents the number of frames. It can be understood that one frame represents a static picture (image) of the animation (video) at a single time point, that is, T represents the number of the at least two time points.
[0037] The object model parameters can be obtained by referring to physical objects of various sizes and / or shapes. The diversity / number of object models can be adjusted according to the task requirements. For example, to meet the task requirements, the number of object models can be more than 100.
[0038] The character's initial skeletal state parameters and skeletal motion parameters can be obtained by the animator completing the animation and then exporting the relevant parameters. The diversity / complexity of the interactive animations (actions) reflected by the initial skeletal state parameters and skeletal motion parameters can also be adjusted according to the task requirements. For example, in order to meet the task requirements, each object model may need at least one or more basic interactive actions (animations), such as one or more interactive actions such as picking up, lifting, grabbing, holding up, putting down, and throwing.
[0039] Should be understood that, about the acquisition of the initial skeletal state parameters and / or the skeletal motion parameters of role, can also refer to the interactive animation production method provided in the aforementioned embodiment.In certain embodiments, can be combined with the advantage of one or more interactive animation production methods, to obtain high-quality training data.Wherein, high-quality can refer to the interactive animation that the skeletal motion parameters reflect is natural and reasonable (as meeting physical laws).
[0040] Step 130 : Using the plurality of sample input data and the plurality of sample label data to train an initial model to obtain an interactive animation generation model. In some embodiments, step 130 may be implemented by the training module 430 .
[0041] The goal of training may include but is not limited to optimizing the error, that is, reducing the gap between the model output and the sample label data.
[0042] In some embodiments, the initial model used in training may include one or more of a linear regression model, a logistic regression (LR) model, a decision tree, a neural network, and the like.
[0043] In some embodiments, when the initial model used for training includes a neural network, the initial model can be trained based on a stochastic gradient descent algorithm to obtain an interactive animation generation model.
[0044] For the specific structure of the neural network used to generate interactive animations, please refer to Figure 3 and its related descriptions.
[0045] Figure 2 This is an exemplary flow chart of the interactive animation generation method according to some embodiments of this specification. Process 100 can be executed by one or more processors. Process 200 corresponds to the model prediction stage, wherein the interactive animation generation model can be a prediction model obtained by training the initial model according to process 100, and the target character and target object are the two interactive objects associated with the interactive animation to be predicted. Figure 1 As shown, process 100 may include:
[0046] Step 210 , obtaining the initial skeletal state parameters of the target character, the spatial distribution parameters of the target object, and the motion trajectory parameters of the target object. In some embodiments, step 210 may be implemented by the input parameter acquisition module 510 .
[0047] The initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the target character, and the motion trajectory parameters at least indicate the initial position, initial posture, target position, and target posture of the target object.
[0048] More details about the initial skeletal state parameters of the target character, the spatial distribution parameters of the target object, and the motion trajectory parameters of the target object can be found in the relevant description of step 110 and will not be repeated here.
[0049] In step 220 , the initial skeletal state parameters of the target character, the spatial distribution parameters of the target object, and the motion trajectory parameters of the target object are input into the interactive animation generation model to obtain the skeletal motion parameters of the target character output by the interactive animation generation model. In some embodiments, step 220 can be implemented by the output parameter acquisition module 520 .
[0050] The skeletal motion parameters indicate positions and postures of the one or more bones at at least two time points during the interaction between the target character and the target object.
[0051] More details about the target character's skeletal motion parameters can be found in the relevant description of step 110 and will not be repeated here.
[0052] Step 230 : Generate an interactive animation between the target character and the target object based on the skeletal motion parameters of the target character. In some embodiments, step 230 may be implemented by the interactive animation generation module 530 .
[0053] It can be understood that the skeletal motion parameters of the (target) character can reflect the actions of the (target) character during the interaction with the (target) object.
[0054] In some embodiments, an interactive animation between a target character and a target object can be generated based on the target character's skeletal motion parameters, and the target object may not be included in the interactive animation. Of course, a motion animation of the target object (i.e., excluding the character) can also be generated based on the target object's motion trajectory parameters, and the interactive animation excluding the target object and the motion animation of the target object can be merged to obtain an interactive animation that includes both the target character and the target object.
[0055] In some embodiments, an interactive animation between the target character and the target object may be generated based on the skeletal motion parameters of the target character and the motion trajectory parameters of the target object. The animation may include both the target character and the target object.
[0056] In some embodiments, an interactive animation between a target character and a target object can be generated based on the skeletal motion parameters of the target character, or based on the skeletal motion parameters of the target character and the motion trajectory parameters of the target object, by rendering or other means using an engine (the engine can be integrated into animation software or software related to animation).
[0057] It is worth noting that the interactive animation mentioned in this manual can only focus on the interactive actions after the character and the object begin to contact, that is, the actions before the character and the object interact can be ignored, for example, walking towards the object to be interacted with, bending over, standing up, squatting, and other preparatory actions before the interaction.
[0058] It should be noted that the above description of the relevant processes is for illustration and purpose only and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the processes under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.
[0059] Figure 3 This is a schematic diagram of an exemplary structure of a neural network for generating interactive animation according to some embodiments of this specification.
[0060] like Figure 3 As shown, the neural network 300 may include a first encoder 310 , a second encoder 320 , and a motion parameter decoder 330 .
[0061] The first encoder 310 can be used to extract the initial skeletal state parameters of the character to obtain the first feature vector. The second encoder 320 can be used to extract the spatial distribution parameters of the object ( Figure 3 The model parameters are referred to as model parameters in this paper to extract features and obtain the second eigenvector.
[0062] The motion parameter decoder 330 can be used to use the output of the current step (step, which can also be interpreted as "path") as part of the input of the next step, and predict the output of each step based on the input of the step. Figure 3 As shown, the motion parameter decoder 330 can predict the output out1 of the first step based on the input start (also called the start mark) of the first step, add the output out1 of the first step to the input of the second step, predict the output out2 of the second step based on the input of the second step, add the output out2 of the second step to the input of the third step, and so on, until the output end (also called the end mark) of the last step is obtained. The conditional parameters input to the motion parameter decoder 330 may include a combined feature vector formed by concatenating the first feature vector and the second feature vector. Figure 4 As shown in the dotted box, the character's skeletal motion parameters can be obtained based on the output of each step of the motion parameter decoder. In addition, the object's motion trajectory parameters (indicating the initial state and the target state, Figure 3 The start and end state parameters are referred to in the text as start and end state parameters) so that the interactive animation (action) reflected by the skeletal motion parameters output by the motion parameter decoder 330 can conform to / match the start and end state of the object, thereby generating a reasonable and natural interactive animation.
[0063] In some embodiments, the first encoder 310 may include a multi-layer perceptron for extracting skeletal features. In some embodiments, the second encoder 320 may include a three-dimensional convolutional neural network (3D CNN) for extracting object model features. In some embodiments, the motion parameter decoder 330 may include a Transformer model for generating skeletal motion parameters. As an example only, the Transformer model may have the same or similar structure as the Transformer model in the paper " Attention is all you need " published by the Google Machine Translation team at the NIPS (Neural Information Processing Systems, neural information processing systems) international conference in 2017.
[0064] Figure 4 This is an exemplary module diagram of the interactive animation generation model training system shown in some embodiments of this specification. Figure 4 As shown, the system 400 may include a sample input data acquisition module 410 , a sample label data acquisition module 420 and a training module 430 .
[0065] The sample input data acquisition module 410 can be used to acquire a plurality of sample input data. The sample input data includes initial skeletal state parameters of a character, spatial distribution parameters of an object, and motion trajectory parameters of the object. The initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the character, and the motion trajectory parameters indicate at least the initial position, initial posture, target position, and target posture of the object.
[0066] The sample label data acquisition module 420 can be used to acquire a plurality of sample label data corresponding to the plurality of sample input data, wherein the sample label data includes skeletal motion parameters of the character, the skeletal motion parameters indicating the position and posture of one or more bones at at least two time points during the interaction between the character and the object.
[0067] The training module 430 may be configured to train an initial model using the plurality of sample input data and the plurality of sample label data to obtain an interactive animation generation model.
[0068] For more details about the system 400 and its modules, please refer to Figure 1 and its related descriptions.
[0069] Figure 5 is an exemplary module diagram of an interactive animation generation system according to some embodiments of this specification. Figure 5As shown, the system 500 may include an input parameter acquisition module 510 , an output parameter acquisition module 520 and an interactive animation generation module 530 .
[0070] The input parameter acquisition module 510 can be used to obtain the initial skeletal state parameters of the target character, the spatial distribution parameters of the target rigid object, and the motion trajectory parameters of the target object. The initial skeletal state parameters indicate the initial position and initial posture of one or more bones of the target character, and the motion trajectory parameters indicate at least the initial position, initial posture, target position, and target posture of the target object.
[0071] The output parameter acquisition module 520 can be used to input the initial skeletal state parameters of the target character, the spatial distribution parameters of the target object, and the motion trajectory parameters of the target object into the interactive animation generation model to obtain the skeletal motion parameters of the target character output by the interactive animation generation model. The skeletal motion parameters indicate the position and posture of one or more bones at at least two time points during the interaction between the target character and the target object.
[0072] The interactive animation generation module 530 may be configured to generate an interactive animation between the target character and the target object based on the skeletal motion parameters of the target character.
[0073] For more details about the system 500 and its modules, please refer to Figure 3 and its related descriptions.
[0074] It should be understood that Figure 4 、 Figure 5 The system and its modules shown can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented by hardware, software, or a combination of software and hardware. Among them, the hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated hardware. Those skilled in the art will understand that the above-mentioned methods and systems can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the system and its modules of this specification. Not only can the hardware circuits such as ultra-large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc. be implemented, they can also be implemented using software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software (for example, firmware).
[0075] It should be noted that the above description of the system and its modules is for convenience of description only and does not limit this specification to the scope of the embodiments cited. It is understandable that for those skilled in the art, after understanding the principles of the system, it is possible to arbitrarily combine the various modules or form a subsystem connected to other modules without deviating from this principle. For example, in some embodiments, the sample input data acquisition module 410 and the sample label data acquisition module 420 can be separate modules or combined into one module. Such variations are all within the scope of protection of this specification.
[0076] The beneficial effects that may be brought about by the embodiments of this specification include but are not limited to: (1) using the interactive animation generation model obtained by training, the user inputs the initial skeletal state parameters of the target character, the spatial distribution parameters of the target object, and the motion trajectory parameters of the target object, and the system can automatically generate the interactive animation between the target character and the target object, which can greatly improve the production efficiency of the interactive animation; (2) by using high-quality training data, the prediction effect / quality of the model can be guaranteed; (3) the motion trajectory parameters indicating the start and end states of the object can be added to the input of each step of the motion parameter decoder 330, so that the interactive animation (action) reflected by the skeletal motion parameters output by the neural network (motion parameter decoder) can conform to / match the start and end states of the object, thereby generating a reasonable and natural interactive animation. It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects that may be produced may be any one or a combination of the above, or any other possible beneficial effects.
[0077] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit the embodiments of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and revisions to the embodiments of this specification. Such modifications, improvements, and revisions are suggested in the embodiments of this specification and remain within the spirit and scope of the exemplary embodiments of this specification.
[0078] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.
[0079] In addition, it will be understood by those skilled in the art that the various aspects of the embodiments of this specification may be illustrated and described by a number of patentable categories or situations, including any new and useful process, machine, product or combination of substances, or any new and useful improvements thereto. Accordingly, the various aspects of the embodiments of this specification may be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software may all be referred to as "data blocks," "modules," "engines," "units," "components," or "systems." In addition, the various aspects of the embodiments of this specification may be represented as a computer product located in one or more computer-readable media, the product including computer-readable program code.
[0080] A computer storage medium may include a propagated data signal embodying the computer program code, for example, in baseband or as part of a carrier wave. The propagated signal may be in a variety of forms, including electromagnetic, optical, or any suitable combination thereof. A computer storage medium may be any computer-readable medium other than a computer-readable storage medium that can be connected to an instruction execution system, apparatus, or device to communicate, propagate, or transfer the program for use. The program code on the computer storage medium may be transmitted via any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of these.
[0081] The computer program coding required for the operation of each part of the embodiment of this specification can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C language, VisualBasic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages such as Python, Ruby and Groovy, or other programming languages, etc. The program coding can be run completely on the user's computer, or run on the user's computer as an independent software package, or partly run on the user's computer and partly run on a remote computer, or run completely on a remote computer or processing equipment. In the latter case, the remote computer can be connected to the user's computer by any network form, such as a local area network (LAN) or a wide area network (WAN), or be connected to an external computer (such as by the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0082] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in the embodiments of this specification are not intended to limit the order of the processes and methods of the embodiments of this specification. Although the above disclosure discusses some of the invention embodiments that are currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing processing device or mobile device.
[0083] Similarly, it should be noted that, in order to simplify the presentation of the embodiments disclosed in this specification and thereby facilitate understanding of one or more invention embodiments, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not mean that the embodiments of this specification require more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.
[0084] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, and documents, cited in this specification is hereby incorporated by reference in its entirety. This includes application history documents that are inconsistent with or conflict with the content of this specification, as well as documents (currently or subsequently attached to this specification) that limit the broadest scope of the claims of this specification. It should be noted that if the descriptions, definitions, and / or terminology used in the accompanying materials are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or terminology used in this specification will control.
[0085] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of the embodiments of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A training method for an interactive animation generation model, wherein: include: Acquire a plurality of sample input data, the sample input data including initial skeletal state parameters of a character, spatial distribution parameters of a rigid object, and motion trajectory parameters of the rigid object, wherein the initial skeletal state parameters indicate an initial position and initial posture of one or more bones of the character, and the motion trajectory parameters indicate at least an initial position, an initial posture, a target position, and a target posture of the rigid object; Acquire a plurality of sample label data corresponding to the plurality of sample input data, the sample label data including skeletal motion parameters of the character, the skeletal motion parameters indicating positions and postures of the one or more bones at at least two time points during the process of the character interacting with the rigid object; Using the plurality of sample input data and the plurality of sample label data to train an initial model to obtain the interactive animation generation model, the initial model includes a neural network, and the interactive animation generation model is obtained based on a stochastic gradient descent algorithm; In which, the neural network includes a first encoder, a second encoder and a motion parameter decoder; the first encoder is used to extract features of the character's initial skeletal state parameters to obtain a first feature vector; the second encoder is used to extract features of the spatial distribution parameters of the rigid object to obtain a second feature vector; the motion parameter decoder is used to use the output of the current step as part of the input of the next step, and to predict the output of the step based on the input of each step; the conditional parameters input to the motion parameter decoder include a combined feature vector spliced from the first feature vector and the second feature vector, and the input of each step of the motion parameter decoder includes at least the motion trajectory parameters of the rigid object; the skeletal motion parameters are obtained based on the output of each step of the motion parameter decoder.
2. The interactive animation generation model training method according to claim 1, wherein: The one or more bones include bones from shoulders to fingers.
3. The interactive animation generation model training method according to claim 1, wherein: The spatial distribution parameters of the rigid object are obtained based on a model of the rigid object that has been processed by low-polygonization.
4. The interactive animation generation model training method according to claim 1 or 3, wherein: The spatial distribution of the model of the rigid object does not exceed a cubic space of a preset size, and the cubic space of the preset size is divided into multiple sub-cube spaces. The spatial distribution parameter of the rigid object indicates the proportion of the part of the model of the rigid object in each sub-cube space.
5. The interactive animation generation model training method according to claim 1, wherein: The initial skeletal state parameters include three-dimensional coordinates indicating the position of the one or more bones and a rotation quaternion indicating the posture of the one or more bones; The motion trajectory parameters of the rigid object include three-dimensional coordinates indicating a starting position and a target position of the rigid object, and a rotation quaternion indicating an initial attitude and a target attitude of the rigid object; The rotation quaternion includes elements represented by a three-dimensional vector of a rotation axis and a rotation angle around the rotation axis.
6. The interactive animation generation model training method according to claim 1, wherein: The first encoder includes a multilayer perceptron.
7. The interactive animation generation model training method according to claim 1, wherein: The second encoder includes a three-dimensional convolutional neural network.
8. The interactive animation generation model training method according to claim 1, wherein: The motion parameter decoder is a Transformer structure.
9. An interactive animation generation model training system, wherein: It includes a sample input data acquisition module, a sample label data acquisition module and a training module; The sample input data acquisition module is used to acquire a plurality of sample input data, wherein the sample input data includes initial skeletal state parameters of a character, spatial distribution parameters of a rigid object, and motion trajectory parameters of the rigid object, wherein the initial skeletal state parameters indicate an initial position and initial posture of one or more bones of the character, and the motion trajectory parameters indicate at least an initial position, an initial posture, a target position, and a target posture of the rigid object; The sample label data acquisition module is used to acquire a plurality of sample label data corresponding to the plurality of sample input data, wherein the sample label data includes skeletal motion parameters of the character, and the skeletal motion parameters indicate the position and posture of the one or more bones at at least two time points during the interaction between the character and the rigid object; The training module is used to train an initial model using the multiple sample input data and the multiple sample label data to obtain the interactive animation generation model, wherein the initial model includes a neural network, and the interactive animation generation model is obtained based on a stochastic gradient descent algorithm; In which, the neural network includes a first encoder, a second encoder and a motion parameter decoder; the first encoder is used to extract features of the character's initial skeletal state parameters to obtain a first feature vector; the second encoder is used to extract features of the spatial distribution parameters of the rigid object to obtain a second feature vector; the motion parameter decoder is used to use the output of the current step as part of the input of the next step, and to predict the output of the step based on the input of each step; the conditional parameters input to the motion parameter decoder include a combined feature vector spliced from the first feature vector and the second feature vector, and the input of each step of the motion parameter decoder includes at least the motion trajectory parameters of the rigid object; the skeletal motion parameters are obtained based on the output of each step of the motion parameter decoder.
10. An interactive animation generation model training device, wherein: The method comprises a processor and a storage device, wherein the storage device is used to store instructions. When the processor executes the instructions, the method according to any one of claims 1 to 8 is implemented.
11. A method for generating an interactive animation, wherein: include: Acquiring initial skeletal state parameters of a target character, spatial distribution parameters of a target rigid object, and motion trajectory parameters of the target rigid object, wherein the initial skeletal state parameters indicate an initial position and initial posture of one or more bones of the target character, and the motion trajectory parameters indicate at least an initial position, an initial posture, a target position, and a target posture of the target rigid object; Inputting initial skeletal state parameters of the target character, spatial distribution parameters of the target rigid object, and motion trajectory parameters of the target rigid object into an interactive animation generation model to obtain skeletal motion parameters of the target character output by the interactive animation generation model, wherein the skeletal motion parameters indicate the position and posture of one or more bones at at least two time points during the interaction between the target character and the target rigid object, the interactive animation generation model being obtained by training an initial model, the initial model including a neural network, and the interactive animation generation model being obtained based on a stochastic gradient descent algorithm; The neural network includes a first encoder, a second encoder, and a motion parameter decoder; the first encoder is used to extract features of the character's initial skeletal state parameters to obtain a first feature vector; the second encoder is used to extract features of the spatial distribution parameters of the rigid object to obtain a second feature vector; the motion parameter decoder is used to use the output of the current step as part of the input of the next step, and to predict the output of each step based on the input of each step; the conditional parameters input to the motion parameter decoder include a combined feature vector formed by splicing the first feature vector and the second feature vector, and the input of each step of the motion parameter decoder includes at least the motion trajectory parameters of the rigid object; the skeletal motion parameters are obtained based on the output of each step of the motion parameter decoder; An interactive animation between the target character and the target rigid object is generated based on the skeletal motion parameters of the target character.
12. The interactive animation generation method according to claim 11, wherein: Generating the interactive animation between the target character and the target rigid object based on the skeletal motion parameters of the target character includes generating the interactive animation between the target character and the target rigid object based on the skeletal motion parameters of the target character and the motion trajectory parameters of the target rigid object.
13. The interactive animation generation method according to claim 11, wherein: The interactive animation generation model is obtained by the interactive animation generation model training method according to any one of claims 1 to 8.
14. An interactive animation generation system, wherein: It includes input parameter acquisition module, output parameter acquisition module and interactive animation generation module; The input parameter acquisition module is used to acquire initial skeletal state parameters of a target character, spatial distribution parameters of a target rigid object, and motion trajectory parameters of the target rigid object, wherein the initial skeletal state parameters indicate an initial position and initial posture of one or more bones of the target character, and the motion trajectory parameters indicate at least an initial position, an initial posture, a target position, and a target posture of the target rigid object; The output parameter acquisition module is used to input the initial skeletal state parameters of the target character, the spatial distribution parameters of the target rigid object, and the motion trajectory parameters of the target rigid object into an interactive animation generation model to obtain skeletal motion parameters of the target character output by the interactive animation generation model, wherein the skeletal motion parameters indicate the position and posture of one or more bones at at least two time points during the interaction between the target character and the target rigid object, and the interactive animation generation model is obtained by training an initial model, wherein the initial model includes a neural network, and the interactive animation generation model is obtained based on a stochastic gradient descent algorithm; The neural network includes a first encoder, a second encoder, and a motion parameter decoder; the first encoder is used to extract features of the character's initial skeletal state parameters to obtain a first feature vector; the second encoder is used to extract features of the spatial distribution parameters of the rigid object to obtain a second feature vector; the motion parameter decoder is used to use the output of the current step as part of the input of the next step, and to predict the output of each step based on the input of each step; the conditional parameters input to the motion parameter decoder include a combined feature vector formed by splicing the first feature vector and the second feature vector, and the input of each step of the motion parameter decoder includes at least the motion trajectory parameters of the rigid object; the skeletal motion parameters are obtained based on the output of each step of the motion parameter decoder; The interactive animation generation module is used to generate an interactive animation between the target character and the target rigid object based on the skeletal motion parameters of the target character.
15. An interactive animation generating device, wherein: The method comprises a processor and a storage device, wherein the storage device is used to store instructions. When the processor executes the instructions, the interactive animation generation method according to any one of claims 11 to 13 is implemented.
Citation Information
Patent Citations
Virtual object pose control method and device and computer storage medium
CN112774203A