Physics-based controllable animal motion generation system in three-dimensional scene

Through the hierarchical architecture of high-level generative models and underlying motion controllers, combined with reinforcement learning and adversarial mechanisms, the problems of scene perception and physical constraints in animal motion generation are solved, and multimodal information alignment and physical authenticity are realized, which is suitable for fields such as virtual reality and movie animation.

CN120279148APending Publication Date: 2025-07-08BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510436340.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art lacks scene perception ability, user controllability and ignores physical law constraints in animal movement generation, resulting in limited generalization ability of the generated results under different scenarios and language instructions, making it difficult to generate diverse and physically reasonable motion sequences.

Method used

Using the hierarchical architecture of the high-level motion generation model and the underlying motion controller, the high-level model generates initial motion sequences through scene coding and language instructions. The underlying controller generates predicted motion sequences that conform to physical laws through physical constraints, combining reinforcement learning and adversarial mechanisms to achieve multimodal information alignment and physical authenticity.

Benefits of technology

It improves the diversity and authenticity of animal generative models, provides user-controllable interactive generation of animals and scenes, and is suitable for fields such as virtual reality and movie animation, reducing production costs and improving the flexibility and practicality of generative models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279148A_ABST
    Figure CN120279148A_ABST
Patent Text Reader

Abstract

The invention provides a controllable animal motion generation system based on physics in a three-dimensional scene, and the system comprises the steps: generating an initial motion sequence through a high-layer motion generation model, and guiding a bottom-layer motion controller through the initial motion sequence to generate a predicted physical motion sequence which accords with the constraint of a physical law; the authenticity and rationality of a generated result are improved; therefore, the method can further improve and enrich the ability of animal model generation, provides high-quality and low-cost materials for the fields of virtual reality, movie animation and the like, has important application value, provides new thinking and inspiration for generation tasks of digital human motion and other moving objects, and has important scientific significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of virtual reality, human - computer interaction and artificial intelligence, and particularly relates to a physics - based controllable animal motion generation system in a three - dimensional scene. Background Art

[0002] Artificial intelligence - generated content (AIGC) technology has achieved remarkable results in aspects such as image, video, three - dimensional model, and animation generation. It can automatically generate realistic three - dimensional scenes and characters, and generate rich human actions and interactions with the environment. These technologies empower industries such as virtual reality, film, and games, not only enhancing the user's immersion and interactivity but also significantly reducing production costs. Virtual animals are equally crucial in these applications, but technologies related to animal motion generation are still scarce, restricting the widespread application of virtual animals.

[0003] Animal motion generation technology has wide application value. It can be used in the human - computer interaction system of virtual animals, simulating the positive effect of interacting with animals with the help of virtual reality or augmented reality devices, providing an alternative for people who cannot raise animals and meeting their emotional needs; in animal behavior analysis, this technology can generate diverse training data to support research in ecology, biodiversity conservation, agriculture, animal husbandry, and medicine, improving the analysis performance in related fields; in addition, this technology can also provide automated animal action generation solutions for the film, game, and entertainment industries, reducing production costs and cycles, enhancing creative flexibility and work diversity, and promoting the innovative development of the industry.

[0004] In the past few years, significant progress has been made in the field of conditional human motion generation, with conditions including past actions, audio, action tags, natural language descriptions, objects, and three - dimensional scenes. In animal motion generation research, researchers have explored various types of animals, including quadruped animals, multi - legged animals, flying animals, etc. The related research mainly focuses on how to generate the motion gaits of animals according to given input conditions. Some research explores methods of gait transition to achieve the generation of gaits of multiple animals. In addition, some research is dedicated to exploring the joint modeling method of animal appearance and motion. The method based on neural representation and rendering can not only generate realistic appearance and high - quality actions but also achieve real - time rendering and interactive control of furry animals. Compared with human motion data, animal motion data is more difficult to obtain, and some research attempts to apply human motion prior knowledge to animal motion modeling to improve the efficiency of animal motion generation. For example, by combining human motion prior knowledge, diverse and realistic animal motion sequences are generated from complex text descriptions.

[0005] Traditional physics-based motion generation tasks aim to enable agents to learn how to execute relevant skills or complete relevant tasks (such as walking, jumping, running) in an environment with real physics simulation through the means of reinforcement learning, and have been widely applied in the fields of character animation generation and robot control. For physics-based human motion generation, previous methods learned the motion execution strategy from reference motion capture or video data through reinforcement learning, and some works incorporated adversarial mechanisms to make the executed motions more conform to the data distribution of the reference motions. Combining these learned prior policy models with control signals (such as trajectories, language descriptions, etc.) can be applied to a wide range of downstream tasks. In addition, another research direction focuses on how to combine motion generation models with physics-based constraints so that the motions generated by the models conform to the training data distribution and have physical rationality at the same time. Among them, some studies modeled physical constraints as differentiable loss functions during the learning process of the generation model to constrain the model to generate results that meet the corresponding constraints. Based on the Diffusion Model, some studies embedded the scene-aware objective equation and the physics-based motion projection module into the diffusion model respectively, and reduced phenomena such as floating, ground penetration, and foot slippage by constraining the denoising process.

[0006] Although extensive and mature research results have been achieved in the field of human motion generation, relatively little work has been done in animal motion generation, especially in the generation of animal-scene interactions, which is almost blank. Although the research on human motion generation can provide us with valuable experience, there are still two main problems: First, existing work has not yet clarified how to encode the scene and how to more efficiently integrate it into the motion space, and then construct a scene-aware motion generation algorithm. Second, limited by the scale of the dataset, the generalization ability of existing methods under different scenes and different language instructions is limited, making it difficult for the model to obtain satisfactory generation results in new scenes.

[0007] Existing physics-based motion generation mainly focuses on the field of human motion generation. How to generate animal motions with physical authenticity is still an unexplored field. Current physics-based methods mainly use the reinforcement learning framework to learn simple motions or a small number of motion skills, and it is difficult to model the distribution of large-scale and complex motion data, making the motion generation ability learned only by the reinforcement learning algorithm in the physics simulation environment relatively single and unable to generate richer and more diverse motion and interaction sequences. Considering that generation models such as the diffusion model have powerful data distribution modeling capabilities, but it is currently difficult to efficiently introduce physics-based laws into them. Therefore, how to combine generation models with powerful data modeling capabilities and physics-based reinforcement learning strategies for diverse and physically reasonable animal motion generation in the scene will be one of the key research directions in the future. Summary of the Invention

[0008] To solve the problems that the current animal motion generation models lack scene perception ability, lack user controllability, and ignore the constraints of physical laws, the present invention provides a physically-based controllable animal motion generation system in a three-dimensional scene, which can generate motion sequences with high physical authenticity when interacting with the environment, ensuring the rationality and realism of the generated motions.

[0009] A physically-based controllable animal motion generation system in a three-dimensional scene, comprising a high-level motion generation model and a low-level motion controller;

[0010] The high-level motion generation model is used to generate an initial motion sequence for the animal to be controlled according to the current three-dimensional scene where the animal to be controlled is located and the action instruction input by the user, which is used to interact with the current three-dimensional scene and is consistent with the description of the action instruction;

[0011] The low-level motion controller is used to perform sliding slicing on the initial motion sequence according to the set sliding window length and step size to obtain a plurality of motion segments;

[0012] The low-level motion controller is used to generate a predicted physical motion sequence that conforms to the constraints of physical laws for each motion segment, where each predicted physical motion sequence is the action finally executed by the animal to be controlled.

[0013] Furthermore, the high-level motion generation model includes a scene encoder, a text encoder, a scene affordance map diffusion module, and a motion affordance diffusion module; among them, the scene affordance map diffusion module is obtained by cascading T scene affordance map diffusion units, and the cascading numbers of the scene affordance map diffusion units decrease from T in order from front to back; the motion affordance diffusion module is obtained by cascading T motion affordance diffusion units, and the cascading numbers of the motion affordance diffusion units decrease from T in order from front to back;

[0014] The scene encoder is used to encode the current three-dimensional scene where the animal to be controlled is located to obtain scene features;

[0015] The text encoder is used to encode the action instruction input by the user to obtain action instruction features;

[0016] Each scene affordance map diffusion unit is used to perform an affordance map acquisition operation. Among them, the affordance map acquisition operation performed by any one scene affordance map diffusion unit includes the following steps:

[0017] Concatenate the scene features and the affordance noise of the scene along the feature dimension to obtain the fused scene noise feature; among them, the affordance noise used by the first scene affordance map diffusion unit is Gaussian noise, and the affordance noise used by the remaining scene affordance map diffusion units is the affordance map output by the scene affordance map diffusion unit cascaded in front of itself;

[0018] Encode the scene diffusion step length to obtain the step length feature; among them, the scene diffusion step lengths used by the respective scene affordance map diffusion units are their respective cascade numbers;

[0019] Concatenate the step length feature and the action instruction feature along the feature dimension to obtain the fused action step length feature;

[0020] Use the attention mechanism to extract features from the fused scene noise feature and the fused action step length feature to obtain the affordance map;

[0021] Each motion affordance diffusion unit is used to perform the operation of obtaining the initial motion sequence. Among them, the initial motion sequence obtained by the last motion affordance diffusion unit is used as the initial motion sequence finally output by the high-level motion generation model; at the same time, the operation of obtaining the initial motion sequence performed by any motion affordance diffusion unit includes the following steps:

[0022] Use the attention mechanism to extract features from the scene features, the action instruction features, the affordance map output by the last scene affordance map diffusion unit, the motion affordance noise, and the motion diffusion step length to obtain the initial motion sequence; among them, the motion affordance noise used by the first motion affordance diffusion unit is Gaussian noise, and the motion affordance noise used by the remaining motion affordance diffusion units is the initial motion sequence output by the motion affordance diffusion unit cascaded in front of itself; the motion diffusion step lengths used by the respective motion affordance diffusion units are their respective cascade numbers.

[0023] Furthermore, the affordance map C output by any scene affordance map diffusion unit is specifically:

[0024] C = max - pool(c1, c2, …, c L )

[0025]

[0026] Among them, L represents the number of actions included in the initial motion sequence, and c1 to c L are respectively the affordance sub - maps corresponding to the L actions in the initial motion sequence, l = 1, 2, …, L, and c l (n, k) is the affordance sub - map c lThe element in the n-th row and k-th column, where n represents the point cloud number in the three-dimensional scene point cloud corresponding to the current three-dimensional scene where the animal to be controlled is located, k represents the surface marker point number predefined on the surface of the SMAL model corresponding to the animal to be controlled, and c l (n,k) represents the regularized distance between the n-th point cloud and the k-th surface marker point, d(n,k) represents the distance between the n-th point cloud and the k-th surface marker point, and σ is a predefined standardized constant factor.

[0027] Furthermore, the underlying motion controller includes a slicing module, a motion encoder, and a motion strategy module;

[0028] The slicing module is used to perform sliding slicing on the initial motion sequence to obtain multiple motion segments;

[0029] The motion encoder is used to encode each motion segment into a hidden vector respectively;

[0030] The motion strategy module is used to determine the predicted physical motion sequence corresponding to each hidden vector respectively according to each hidden vector and the current skeleton state of the animal to be controlled.

[0031] Furthermore, the underlying motion controller further includes a conditional discriminator;

[0032] The conditional discriminator is used to, when training the underlying motion controller, judge whether the difference between the predicted physical motion sequence generated by the motion strategy module and the reference physical motion sequence is less than a set value according to the next skeleton state entered after the animal to be controlled executes the predicted physical motion sequence generated by the motion strategy module; if the judgment result is yes, the underlying motion controller at this time is used as the final underlying motion controller, if the judgment result is no, adjust the network parameters of the motion encoder, the motion strategy module, and the conditional discriminator according to the set loss function, then use the motion encoder with updated network parameters to re-obtain the hidden vector of the motion segment, use the motion strategy module with updated network parameters to re-generate the predicted physical motion sequence according to the re-obtained hidden vector, and then use the conditional discriminator with updated network parameters to judge the re-generated predicted physical motion sequence; and so on until the judgment result is yes.

[0033] Furthermore, the loss function used when training the underlying motion controller includes a motion encoder loss function, a motion strategy loss function, and a discriminant loss function;

[0034] Among them, the motion encoder loss function is as follows:

[0035]

[0036] Among them, represents the set of reference physical motion sequences, denotes performing an expected calculation operation in the set , M denotes a subset of the reference physical action sequence randomly sampled from the set , E(M) denotes the hidden vector obtained by encoding each reference physical action sequence in the subset M, M′ denotes another subset of the reference physical action sequence randomly sampled from the set , and E(M′) denotes the hidden vector obtained by encoding each reference physical action sequence in the subset M′; ∼ + denotes that the subsets M and M′ of the reference physical action sequence sampled from the set overlap, ∼ - denotes that the subsets M and M′ of the reference physical action sequence sampled from the set are independently sampled without overlap; L + denotes the first loss function when the subsets overlap, and L - denotes the second loss function when the subsets do not overlap;

[0037] Discriminative loss function is as follows:

[0038]

[0039] wherein, denotes the previous true skeleton state entered after the controlled animal executes the previous reference physical motion sequence Ⅰ in the set , denotes the current true skeleton state entered after the controlled animal executes the current reference physical motion sequence Ⅱ in the set , s denotes the previous predicted skeleton state entered after the controlled animal executes the predicted physical motion sequence Ⅰ generated by the motion policy module according to the reference physical motion sequence Ⅰ, s′ denotes the current predicted skeleton state entered after the controlled animal executes the predicted physical motion sequence Ⅱ generated by the motion policy module according to the reference physical motion sequence Ⅱ; z denotes the predicted hidden vector obtained by the motion encoder encoding any reference physical motion sequence; D(s, s′|z) denotes the discriminative probability obtained by the discriminator judging according to the predicted skeleton states s and s′ corresponding to the predicted hidden vector z, wherein the discriminative probability is a decimal between 0 and 1. When D(s, s′|z) is closer to 1, it indicates that the discriminator believes that the predicted physical motion sequence corresponding to the predicted hidden vector z is a "true" motion sequence; when D(s, s′|z) is closer to 0, it indicates that the discriminator believes that the predicted physical motion sequence corresponding to the predicted hidden vector z is a "false" motion sequence; denotes the discriminative probability obtained by the discriminator judging according to the true skeleton states and corresponding to the predicted hidden vector z, wherein the discriminative probability is a decimal between 0 and 1. When When it is closer to 1, it means that the discriminator believes that the reference physical motion sequence corresponding to the predicted latent vector z is a "true" motion sequence; when it is closer to 0, it means that the discriminator believes that the reference physical motion sequence corresponding to the predicted latent vector z is a "false" motion sequence; during the training process, the discriminator needs to learn to judge the predicted physical motion sequence corresponding to the predicted latent vector z as "false" and the reference physical motion sequence corresponding to the predicted latent vector z as "true"; w gp represents the set gradient penalty term coefficient, and sg(·) represents the gradient stop operation; represents the gradient of the model parameter θ of the motion policy module;

[0040] represents performing an expectation calculation operation in the set ;

[0041] d π (s, s′|z) represents the distribution that the predicted skeleton state corresponding to the predicted physical motion sequence generated by the policy π under the given predicted latent vector z conforms to, represents performing an expectation calculation operation under the distribution d π (s, s′|z);

[0042] represents the distribution that the true skeleton state corresponding to the reference physical action sequence in the subset M conforms to, represents performing an expectation calculation operation under the distribution ;

[0043] The motion policy loss function is as follows:

[0044]

[0045] where s t represents the previous predicted skeleton state s, and s t+1 represents the current predicted skeleton state s′; s0 represents the initial skeleton state; γ represents the discount factor; t represents the step size, and t = 1, 2,..., T; represents performing an expectation calculation operation in the set ; p(τ|π, z) represents the distribution that the complete predicted physical motion sequence τ generated by the policy π under the given predicted latent vector z conforms to; represents performing an expectation calculation operation under the distribution p(τ|π, z); r(s t , s t+1 , z) represents the reward function calculated according to the previous predicted skeleton state s t , the current predicted skeleton state s t+1 and the latent vector z, and there is r(s t , s t+1, z) = r(s, s′, z) = -log(1 - D(s, s′|z)).

[0046] Furthermore, the training process is carried out in a simulator with a physical simulation environment, and the physical law constraints include Newton's first law of motion, Newton's second law of motion, and Newton's third law of motion.

[0047] Advantages:

[0048] 1. The present invention provides a physically-based controllable animal motion generation system in a three-dimensional scene. First, an initial motion sequence is generated through a high-level motion generation model, and then the initial motion sequence is used to guide a low-level motion controller to generate a predicted physical motion sequence that conforms to physical law constraints, thereby improving the authenticity and rationality of the generation result. Therefore, the present invention can further enhance and enrich the capabilities of the animal generation model, provide high-quality and low-cost materials for fields such as virtual reality and film animation, has important application value, provides new ideas and inspiration for the generation tasks of digital human motion and other moving objects, and has important scientific significance.

[0049] 2. The present invention provides a physically-based controllable animal motion generation system in a three-dimensional scene, which divides the animal motion generation into two parts: high-level complex motion modeling and low-level motion controller learning, and adopts different learning strategies. At the high level, complex motion modeling is used to ensure the diversity and naturalness of the generated motion, and at the low level, physical constraints are applied for motion control to ensure the rationality and realism of the motion. That is to say, compared with the traditional single-stage animal motion generation model, the present invention divides the animal motion generation into two parts: high-level complex motion modeling and low-level motion controller learning, and applies different learning strategies at different levels. Through this mechanism, the generated result not only conforms to the complex motion distribution but also satisfies the low-level physical constraints, ensuring the truthfulness and rationality of the generated result. This hierarchical and multi-strategy generation framework has substantial innovation.

[0050] 3. The present invention provides a physically-based controllable animal motion generation system in a three-dimensional scene, which combines language instructions and scene information to generate the motion of animals. Through scene affordances as intermediate conditions, it aligns the multi-modal information of language, scene, and action to achieve user-controllable generation. This motion generation model conditioned on multi-modal signals improves the flexibility and practicality of the generation model, especially suitable for interactive applications. That is to say, compared with the traditional animal motion generation that relies on a single-modal signal, the present invention first combines language instructions and scene information to generate animal motion, aligns the multi-modal information of language, scene, and action in the generation framework to achieve user-controllable generation, which provides key technical support for interactive applications, significantly improves the flexibility and practicality of the generation model, and the construction and learning of this animal motion generation model conditioned on multi-modal signals have significant innovation.

[0051] 4. The present invention provides a physics-based controllable animal motion generation system in a three-dimensional scene. In a physical simulation environment, by using a reinforcement learning framework and introducing an adversarial mechanism, it learns animal motion control that conforms to physical laws and data distribution. In addition, domain randomization is used to enhance the robustness and generalization of the control strategy. That is to say, compared with the traditional method that only relies on learning motion generation from data, the present invention, based on a physical simulation environment, uses the framework of reinforcement learning to introduce an adversarial mechanism to learn physically reasonable and data-distribution-compliant animal motion control, and enhances the robustness and generalization of the control strategy through domain randomization. The construction and learning of such a physics-based underlying animal motion control strategy have significant innovations. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic block diagram of a physics-based controllable animal motion generation system in a three-dimensional scene provided by the present invention;

[0053] Figure 2 It is a schematic block diagram of a high-level motion generation model provided by the present invention;

[0054] Figure 3 It is a schematic block diagram of a low-level motion controller provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application.

[0056] The present invention relates to technological innovations in virtual reality (VR) and augmented reality (AR) environments, specifically to using generative artificial intelligence (GenAI) to generate physically realistic and controllable virtual animal interaction motion sequences in a three-dimensional scene, aiming to solve the following two technical problems:

[0057] (1) Implementing user-controllable animal motion generation in a three-dimensional scene. To achieve the purpose of interactive practicality, the controllability of the generation target is crucial. Therefore, the generation model needs to have the ability to complete the corresponding generation target according to user instructions. However, the generation task involving three modal information of language instructions, three-dimensional scenes, and animal motions faces the challenge of significant differences between modalities. The present invention aims to effectively encode, align, and fuse each modal information, model the joint distribution of animal motion and other conditional modalities, and achieve user-controllable animal and scene interaction motion generation.

[0058] (2) Introduce physical laws during generation to enhance the realism of animal movements. To simulate the movements of animals in three-dimensional space, the generation model not only needs to focus on the naturalness and coherence of actions but also must follow basic physical laws. For example, when a cat jumps onto a table, the generation model needs to consider physical factors such as gravity and friction to ensure that the speed, acceleration, and resulting trajectory of the movement conform to the physical laws of the real world. However, current related research mainly focuses on generating the apparent realism of animal actions while ignoring the constraints of physical laws on movements. Therefore, another technical problem solved by the present invention is to organically combine the generation model with physical laws so that the generated animal movements have a high degree of physical realism when interacting with the environment.

[0059] Specifically, as Figure 1 shown, a physically controllable animal movement generation system in a three-dimensional scene includes a high-level movement generation model and a low-level movement controller;

[0060] The high-level movement generation model is used to generate an initial movement sequence X = [x1,…,x L for the animal to be controlled according to the current three-dimensional scene where the animal to be controlled is located and the action instruction input by the user, which is used to interact with the current three-dimensional scene and is consistent with the action instruction description;

[0061] The low-level movement controller is used to perform sliding slicing on the initial movement sequence X = [x1,…,x L according to the set sliding window length w and step size u to obtain a plurality of movement segments M1 = [x1,…,x w , M2 = [x u ,…,x u+w-1 , M3 = [x 2u ,…,x 2u+w-1 ,…;

[0062] The low-level movement controller is used to generate a predicted physical movement sequence that conforms to the constraints of physical laws for each movement segment, where each predicted physical movement sequence is the action finally executed by the animal to be controlled.

[0063] In this form of sliding window, the high-level conditional movement generation model can guide the low-level physical-based movement strategy to generate a movement sequence with a high degree of physical realism when interacting with the environment.

[0064] It should be noted that the goal of the present invention is to generate a three-dimensional animal movement sequence that is consistent with the language description and has reasonable actions based on a physical engine; under a given three-dimensional scene and language description, the present invention denotes the three-dimensional scene as S ∈ R N×6, represents a scene point cloud with N points, and the information of each scene point cloud includes a set of three-dimensional coordinates and a set of RGB color values. The language description is a tokenized word sequence of length D, denoted as W = [w1, …, w D . The present invention uses the SMAL model to represent the animal pose and shape in each frame of the animal motion sequence, that is, for an animal motion sequence X = [x1, …, x L , are the SMAL model parameters. Among them, L is the length of the motion sequence, β is a parameter related to the animal body size, are the relative rotations of J non-root joints in the kinematic tree, and ν and o are the global displacement and rotation applied to the root joint respectively; for example, when constructing the SMAL model in the present invention, the pelvic joint point of the animal is used as the root joint, and then it extends outwards, with spinal joints, thigh joints, and shoulder joints, which are used as non-root joint points, and the global displacement and rotation of the non-root joint points relative to the root joint point in each frame of the image are used to characterize the pose and shape of the animal in each frame of the image. The SMAL model finally uses the linear skinning function to convert the parameter x i into a three-dimensional mesh model of an animal, that is to say, the parameter x i is equivalent to the bone parameters of the three-dimensional mesh model of the animal, and the linear skinning function is equivalent to the skin parameters of the three-dimensional mesh model of the animal.

[0065] Further, as Figure 2 shown, the high-level motion generation model includes a scene encoder, a text encoder, a scene affordance map diffusion module, and a motion affordance diffusion module; among them, the scene affordance map diffusion module is obtained by cascading T scene affordance map diffusion units, and the cascading numbers of each scene affordance map diffusion unit decrease from T in order from front to back; the motion affordance diffusion module is obtained by cascading T motion affordance diffusion units, and the cascading numbers of each motion affordance diffusion unit decrease from T in order from front to back;

[0066] The scene encoder is used to encode the current three-dimensional scene where the animal to be controlled is located to obtain scene features;

[0067] The text encoder is used to encode the action instruction input by the user to obtain action instruction features;

[0068] Each scene affordance map diffusion unit is used to perform an affordance map acquisition operation, where the affordance map acquisition operation performed by any one scene affordance map diffusion unit includes the following steps:

[0069] The scene features and scene affordance noises are concatenated according to the feature dimensions to obtain the scene noise fusion features; among them, the scene affordance noise used by the first scene affordance map diffusion unit is Gaussian noise, and the scene affordance noises used by the remaining scene affordance map diffusion units are the affordance maps output by the scene affordance map diffusion units cascaded in front of themselves;

[0070] The scene diffusion step lengths are encoded to obtain the step length features; among them, the scene diffusion step lengths used by each scene affordance map diffusion unit are their respective cascade numbers;

[0071] The step length features and the action instruction features are concatenated according to the feature dimensions to obtain the action step length fusion features;

[0072] The attention mechanism is adopted to extract features from the scene noise fusion features and the action step length fusion features to obtain the affordance map;

[0073] Each motion affordance diffusion unit is used to perform the initial motion sequence acquisition operation. Among them, the initial motion sequence obtained by the last motion affordance diffusion unit is used as the initial motion sequence finally output by the high-level motion generation model; at the same time, the initial motion sequence acquisition operation performed by any motion affordance diffusion unit includes the following steps:

[0074] The attention mechanism is adopted to extract features from the scene features, the action instruction features, the affordance map output by the last scene affordance map diffusion unit, the motion affordance noise, and the motion diffusion step lengths to obtain the initial motion sequence; among them, the motion affordance noise used by the first motion affordance diffusion unit is Gaussian noise, and the motion affordance noises used by the remaining motion affordance diffusion units are the initial motion sequences output by the motion affordance diffusion units cascaded in front of themselves; the motion diffusion step lengths used by each motion affordance diffusion unit are their respective cascade numbers.

[0075] It should be noted that the present invention uses scene affordance as an intermediate representation to promote language-guided animal motion generation in a three-dimensional scene. This scene affordance is calculated through the distance field between the predefined marker points on the surface of the SMAL model and the scene point cloud. Specifically, first, the l2 distance field d∈R between the given three-dimensional scene point cloud S∈R N×6 and K predefined marker points selected from the epidermis of the three-dimensional mesh model of the animal is calculated; d(n,k) represents the distance between the nth point of the scene point cloud and the kth SMAL surface marker point. Finally, the scene affordance map (Affordance Map) C is obtained by normalizing the distance field d and pooling in the temporal dimension; specifically, the affordance map C output by any scene affordance map diffusion unit is specifically: N×K ;

[0076] C = max - pool(c1, c2, …, c L )

[0077]

[0078] where L represents the number of actions included in the initial motion sequence, and c1 to c L are the available sub - maps corresponding to the L actions in the initial motion sequence respectively, l = 1, 2, …, L, and c l (n, k) is the element in the n - th row and k - th column of the available sub - map c l corresponding to the l - th action. n represents the point cloud number in the 3D scene point cloud corresponding to the current 3D scene where the animal to be controlled is located, and k represents the surface marker point number predefined on the surface of the SMAL model corresponding to the animal to be controlled. c l (n, k) represents the regularized distance between the n - th point cloud and the k - th surface marker point, d(n, k) represents the distance between the n - th point cloud and the k - th surface marker point, and σ is a predefined normalization constant factor.

[0079] In specific implementation, the scene affordance map is pre - calculated from the motion dataset of animal - scene interaction and jointly constitutes a training data pair with the original scene, language annotation, and motion sequence. On the one hand, the scene affordance under this definition precisely depicts the corresponding area through language description, and can also enhance the landing of the 3D scene in the case of limited data. On the other hand, this distance - based affordance map can better understand the scene geometry and assist in the generation of animal - scene interaction and model generalization.

[0080] It can be seen from this that the high - level motion generation model of the present invention includes two stages: scene affordance generation and conditional motion generation, as Figure 2 shown. In the first stage, the network architecture of Perceiver is adopted, with language description and 3D scene as inputs, to generate a reasonable scene affordance map. In the second stage, the network architecture of Transformer is adopted, with language description, 3D scene, and the scene affordance map generated in the first stage as inputs, to generate the final animal motion sequence. The present invention independently trains the generation models of the two stages using conditional diffusion models. Among them, the language description is encoded using the CLIP model, and the 3D scene is encoded using the PointNet++ pre - trained network. The models of the two stages share language and scene encoding, and at the same time adopt self - attention and cross - attention mechanisms, as well as the efficient fusion mechanism proposed below to fuse the features of each modality.

[0081] When using the attention mechanism for feature fusion, cross - attention is used to fuse the motion modality features and scene modality features Taking it as an example, the query, key, and value embedding vectors are calculated as Q = XW Q , K = YW K , V = YW V , where are learnable parameters in the attention mechanism. The motion modality features are updated as:

[0082]

[0083] where, is the attention weight matrix, and each element ω ij in it represents the attention weight between the i-th motion feature and the j-th scene feature.

[0084] The present invention combines scene affordance with the attention weight matrix calculated by the original attention mechanism to obtain a new attention weight matrix Ω′, and each element in it is calculated as:

[0085]

[0086] where, N ≡ N y is the number of scene point clouds. Since the scene affordance defined by the present invention itself describes the relationship between the points in the scene and the motion modality, this combination method can more efficiently fuse the features of the scene and the motion modality, thereby improving the effect of motion generation.

[0087] Furthermore, based on the existing animal motion dataset, the present invention learns a physics-based low-level animal motion controller π, which can control the animal skeleton model to execute corresponding actions in a simulator with a physical simulation environment, while ensuring that the executed actions conform to the action distribution in the reference action dataset. The present invention models this low-level motion controller as a conditional motion policy model π := π(a|s,z), and learns it through a conditional imitation learning objective function, that is

[0088]

[0089] where, D Js is the Jensen-Shannon divergence, d π (s,s′|z)| z=E(M) and are the state transition probability distributions of the motion controller π and the reference motion data respectively, z is the motion latent vector encoded by the encoder E from the reference motion sequence, s is the current state of the animal skeleton, s′ is the next state, is the current state of the real data.

[0090] The present invention defines the entire learning task as a Markov Decision Process, which is defined by Among them, the state and the state transition equation are determined by the physical simulation environment. The information contained in s includes the position and orientation of the root node of the animal skeleton, as well as the rotation angles and angular velocities of other joints. The initial state s0 can be obtained by transforming the initial parameters x0 of the SMAL model through forward kinematics. The motion policy π(a|s,z) calculates the execution action for each step according to the current state s and the motion latent vector z By executing a, the current state s / s t transfers to the next state s′ / s t+1 . The value function calculates the reward obtained for each step The goal of policy learning is to maximize the discounted reward

[0091] In addition, in order to make the final motion satisfy both physical rationality and action authenticity, the present invention uses an adversarial learning mechanism in the process of policy learning. Given a motion dataset The present invention aims to train an action encoder E, a discriminator D, and a conditional motion policy π(a|s,z) simultaneously. The "conditional" here refers to the latent vector z encoded by the encoder from the reference action . The specific content is as follows:

[0092] As Figure 3 shown, the underlying motion controller includes a slicing module, a motion encoder, a motion policy module, and a conditional discriminator;

[0093] The slicing module is used to perform sliding slicing on the initial motion sequence to obtain multiple motion segments;

[0094] The motion encoder is used to encode each motion segment into a latent vector; it should be noted that in the present invention, the motion encoder E can encode each motion segment into an action latent vector z to guide the conditional motion policy π(a|s,z) to generate corresponding actions, and each action latent vector z only guides the conditional motion policy π(a|s,z) to generate u actions;

[0095] The motion policy module is used to determine the predicted physical motion sequence corresponding to each latent vector according to each latent vector and the current skeleton state of the animal to be controlled;

[0096] The condition discriminator is used to determine whether the difference between the predicted physical motion sequence generated by the motion policy module and the reference physical motion sequence is less than a set value according to the next skeleton state entered after the to-be-controlled animal executes the predicted physical motion sequence generated by the motion policy module during the training of the underlying motion controller; if the judgment result is yes, the underlying motion controller at this time is used as the final underlying motion controller, and if the judgment result is no, the network parameters of the motion encoder, the motion policy module, and the condition discriminator are adjusted according to the set loss function, and then the motion encoder with updated network parameters is used to re-obtain the hidden vector of the motion segment, the motion policy module with updated network parameters is used to re-generate the predicted physical motion sequence according to the re-obtained hidden vector, and then the condition discriminator with updated network parameters is used to judge the re-generated predicted physical motion sequence; and so on until the judgment result is yes.

[0097] It should be noted that the present invention uses an end-to-end training method and combines the PPO algorithm to train the encoder E, the discriminator D, and the motion policy π(a|s,z) simultaneously, as Figure 3 shown. The present invention uses Isaac Gym as the simulation training environment and constructs a set of physics-based animal URDF models according to the definition of SMAL skeletons. During the training process, in order to further improve the robustness of the policy model and its generalization ability to different environments, the present application intends to introduce domain randomization, that is, during the training process, random parameter settings are made for the initialized scene, and these parameters include the parameters of the simulation environment (such as environmental friction, the height of objects in the scene, etc.) and the parameters of the animal model (such as the bone length and initial state in the animal model).

[0098] Furthermore, the training process is carried out in a simulator with a physical simulation environment, and the physical law constraints include Newton's first law of motion, Newton's second law of motion, and Newton's third law of motion; at the same time, the loss functions used during the training of the underlying motion controller include a motion encoder loss function, a motion policy loss function, and a discriminant loss function;

[0099] Among them, for a given reference action sequence the action M is encoded into a hidden vector using the action encoder E Considering that similar actions should have similar embedding vectors in the latent space and dissimilar actions should be as far away as possible in the latent space, the present invention uses the following two loss functions to train the encoder:

[0100]

[0101] Among them, represents the set of reference physical motion sequences, represents in the set perform the desired calculation operation, where M represents a subset of the reference physical action sequences randomly sampled from the set , E(M) represents the hidden vectors obtained by encoding each reference physical action sequence in the subset M, and M′ represents another subset of the reference physical action sequences randomly sampled from the set ; ~ + represents that the subsets M and M′ of the reference physical action sequences sampled from the set overlap, ~ - represents that the subsets M and M′ of the reference physical action sequences sampled from the set are independently sampled without overlap; L + represents the first loss function when the subsets overlap, and L - represents the second loss function when the subsets do not overlap;

[0102] Given an action hidden vector z, the discriminator needs to determine whether the input (s, s′) is from the reference action dataset or output by the policy π. At the same time, to enhance the robustness of the discriminator, for the from the reference dataset, the discriminator also needs to distinguish that it belongs to the corresponding hidden vector Z rather than a randomly sampled hidden vector In addition, adding gradient regularization constraints to the loss function can make the training more stable. The final discriminator loss function is as follows:

[0103]

[0104] where represents the previous true skeleton state entered after the controlled animal executes the previous reference physical motion sequence Ⅰ in the set , represents that the controlled animal executes the set The current true skeleton state entered after the current reference physical motion sequence II, S represents the previous predicted skeleton state entered after the predicted physical motion sequence I generated by the motion strategy module of the animal to be controlled according to the reference physical motion sequence I, and s′ represents the current predicted skeleton state entered after the predicted physical motion sequence II generated by the motion strategy module of the animal to be controlled according to the reference physical motion sequence II; z represents the predicted hidden vector obtained by the motion encoder encoding any reference physical motion sequence; D(s, s′|z) represents the discrimination probability obtained by the discriminator judging according to the predicted skeleton states s and s′ corresponding to the predicted hidden vector z, where the discrimination probability is a decimal between 0 and 1. When D(s, s′|z) is closer to 1, it means that the discriminator believes that the predicted physical motion sequence corresponding to the predicted hidden vector z is a "true" motion sequence; when D(s, s′|z) is closer to 0, it means that the discriminator believes that the predicted physical motion sequence corresponding to the predicted hidden vector z is a "false" motion sequence; represents the true skeleton state corresponding to the predicted hidden vector z by the discriminator and the discrimination probability obtained by making a judgment, where the discrimination probability is a decimal between 0 and 1. When is closer to 1, it means that the discriminator believes that the reference physical motion sequence corresponding to the predicted hidden vector z is a "true" motion sequence; when is closer to 0, it means that the discriminator believes that the reference physical motion sequence corresponding to the predicted hidden vector z is a "false" motion sequence; The discriminator needs to learn to discriminate the predicted physical motion sequence corresponding to the predicted hidden vector z as "false" and the reference physical motion sequence corresponding to the predicted hidden vector z as "true" during training; w gp represents the set gradient penalty term coefficient, and sg(·) represents the gradient stop operation; represents the gradient of the model parameter θ of the motion strategy module;

[0105] represents performing an expectation calculation operation in the set ;

[0106] d π (s, s′|z) represents the distribution that the predicted skeleton state corresponding to the predicted physical motion sequence generated by the policy π under the given predicted hidden vector z conforms to, represents performing an expectation calculation operation under the distribution d π (s, s′|z);

[0107] represents the distribution that the true skeleton state corresponding to the reference physical action sequence in the subset M conforms to, represents performing an expectation calculation operation under the distribution ;

[0108] That is to say, the present invention regards these reference motion sequences as positive samples ("true"), and the generated predicted motion sequences (s, s′, z) as negative samples ("false"). The role of the discriminator is actually binary classification, that is, it tries to learn to classify (s, s′, z) as negative samples (represented by the label y = 0), and classify as positive samples (represented by the label y = 1). The principle of the simplified classification loss function is loss = ylog(D(x)) + (1 - y)log(1 - D(x)), where x represents the input here. The input can be (s, s′, z), or Because the present invention knows that when the input is (s, s′, z), the label y is 0, and when the input is the label y is 1, so the formula loss = ylog(D(x)) + (1 - y)log(1 - D(x)) can be replaced by

[0109] Furthermore, the objective optimization function of the conditional motion policy is to maximize the cumulative adversarial imitation learning (GAIL) reward function r(s, s′, z) = -log(1 - D(s, s′|z)), so that the discriminator discriminates the motion generated by the policy as real motion data. That is, the motion policy loss function is as follows:

[0110]

[0111] where s t represents the previous predicted skeleton state s, s t+1 represents the current predicted skeleton state s′; s0 represents the initial skeleton state; γ represents the discount factor; t represents the step size, and t = 1, 2,..., T; represents the operation of performing the expectation calculation on the set ; p(τ|π, z) represents the distribution that the complete predicted physical motion sequence τ generated by the policy π conforms to under the given predicted hidden vector z; represents the operation of performing the expectation calculation under the distribution p(τ|π, z); r(s t , s t+1 , z) represents the reward function calculated according to the previous predicted skeleton state s t , the current predicted skeleton state s t+1 and the hidden vector z, and there is r(s t , s t+1 , z) = r(s, s′, z) = -log(1 - D(s, s′|z)).

[0112] It can be seen that the physics-based underlying motion controller of the present invention is a motion strategy model independent of scenes and language instructions. The model only needs to execute corresponding physics-based motion skills according to the given motion latent vectors. This underlying motion controller is mainly responsible for efficiently modeling physics-based motion skills without learning complex and rich conditional motion distributions; the high-level motion generation model is used to construct the joint distribution of three-dimensional scenes, language instructions, and animal motions. The model needs to have the ability of scene perception and understanding, and then generate motion sequences that are consistent with the language description and reasonable in the scene. The high-level generation model can be integrated as a separate algorithm into the final application, or combined with the underlying motion controller to generate physically reasonable controllable animal and scene interaction motion sequences.

[0113] In summary, the present invention aims to construct a generation model that can efficiently model the joint distribution of language, scenes, and animal motion data, realize user-controllable animal and scene interaction generation. At the same time, a physics-based underlying motion controller is constructed to make the generated motion conform to objective physical laws in the three-dimensional scene, improving the authenticity and reasonableness of the generation results. The present invention can further enhance and enrich the capabilities of the animal generation model, provide high-quality and low-cost materials for fields such as virtual reality and film animation, and has important application value. The present invention introduces a computational framework and computational model of physical laws in the generation process, providing new ideas and inspirations for similar generation tasks, such as the generation of digital human motions and other moving objects, and has important scientific significance.

[0114] Of course, the present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and deformations according to the present invention. However, these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A physics-based controllable animal motion generation system in a three-dimensional scene, characterized in that It includes a high-level motion generation model and a low-level motion controller; The high-level motion generation model is used to generate an initial motion sequence for the animal to be controlled, which is used to interact with the current three-dimensional scene and is consistent with the action instruction described by the user input, according to the current three-dimensional scene where the animal to be controlled is located and the action instruction input by the user; The low-level motion controller is used to perform sliding slicing on the initial motion sequence according to the set sliding window length and step size to obtain a plurality of motion segments; The low-level motion controller is used to generate a predicted physical motion sequence that conforms to the constraints of physical laws for each motion segment, where each predicted physical motion sequence is the action finally executed by the animal to be controlled.

2. The physical-based controllable animal motion generation system in a three-dimensional scene according to claim 1, wherein The high-level motion generation model includes a scene encoder, a text encoder, a scene affordance map diffusion module, and a motion affordance diffusion module; among them, the scene affordance map diffusion module is obtained by cascading T scene affordance map diffusion units, and the cascading numbers of each scene affordance map diffusion unit decrease from T in order from front to back; the motion affordance diffusion module is obtained by cascading T motion affordance diffusion units, and the cascading numbers of each motion affordance diffusion unit decrease from T in order from front to back; The scene encoder is used to encode the current three-dimensional scene where the animal to be controlled is located to obtain scene features; The text encoder is used to encode the action instruction input by the user to obtain action instruction features; Each scene affordance map diffusion unit is used to perform an affordance map acquisition operation. Among them, the affordance map acquisition operation performed by any one scene affordance map diffusion unit includes the following steps: Concatenate the scene features and the scene affordance noise according to the feature dimension to obtain a scene noise fusion feature; among them, the scene affordance noise used by the first scene affordance map diffusion unit is Gaussian noise, and the scene affordance noise used by the remaining scene affordance map diffusion units is the affordance map output by the scene affordance map diffusion unit cascaded in front of itself; Encode the scene diffusion step size to obtain a step size feature; among them, the scene diffusion step sizes used by each scene affordance map diffusion unit are their respective cascading numbers; Concatenate the step size feature and the action instruction feature according to the feature dimension to obtain an action step fusion feature; Adopt an attention mechanism to extract features from the scene noise fusion feature and the action step fusion feature to obtain an affordance map; Each motion affordance diffusion unit is used to perform an initial motion sequence acquisition operation. Among them, the initial motion sequence obtained by the last motion affordance diffusion unit is used as the initial motion sequence finally output by the high-level motion generation model; at the same time, the initial motion sequence acquisition operation performed by any one motion affordance diffusion unit includes the following steps: The attention mechanism is used to extract features from the scene features, action instruction features, affordance map output by the last scene affordance map diffusion unit, motion affordance noise, and motion diffusion step size to obtain an initial motion sequence; among them, the motion affordance noise used by the first motion affordance diffusion unit is Gaussian noise, and the motion affordance noise used by the remaining motion affordance diffusion units is the initial motion sequence output by the motion affordance diffusion unit cascaded in front of itself; the motion diffusion step sizes used by each motion affordance diffusion unit are their respective cascade numbers.

3. The physical-based controllable animal motion generation system in a three-dimensional scene according to claim 2, wherein The specific affordance map C output by any scene affordance map diffusion unit is as follows: C = max - pool(c1, c2, …, c L ) Among them, L represents the number of actions included in the initial motion sequence, c1 to c L are respectively the available submaps corresponding to the L actions in the initial motion sequence, l = 1, 2, …, L, c l (n, k) is the element in the n-th row and k-th column of the available submap c l corresponding to the l-th action, n represents the point cloud number in the point cloud corresponding to the current three-dimensional scene where the animal to be controlled is located, k represents the surface marker point number predefined on the surface of the SMAL model corresponding to the animal to be controlled, c l (n, k) represents the regularized distance between the n-th point cloud and the k-th surface marker point, d(n, k) represents the distance between the n-th point cloud and the k-th surface marker point, and σ is a predefined normalization constant factor.

4. A physically-based controllable animal motion generation system in a three-dimensional scene according to claim 1, characterized in that, The underlying motion controller includes a slicing module, a motion encoder, and a motion policy module; The slicing module is used to perform sliding slicing on the initial motion sequence to obtain multiple motion segments; The motion encoder is used to encode each motion segment into a hidden vector respectively; The motion policy module is used to determine the predicted physical motion sequence corresponding to each hidden vector according to each hidden vector and the current skeleton state of the animal to be controlled respectively.

5. The physical-based controllable animal motion generation system in a three-dimensional scene according to claim 4, characterized in that The underlying motion controller further includes a condition discriminator; The condition discriminator is used to judge whether the difference between the predicted physical motion sequence generated by the motion policy module and the reference physical motion sequence is less than a set value according to the next skeleton state entered after the animal to be controlled executes the predicted physical motion sequence generated by the motion policy module during the training of the underlying motion controller; if the judgment result is yes, the underlying motion controller at this time is used as the final underlying motion controller, if the judgment result is no, the network parameters of the motion encoder, the motion policy module, and the condition discriminator are adjusted according to the set loss function, then the hidden vector of the motion segment is re-obtained using the motion encoder with updated network parameters, the predicted physical motion sequence is regenerated using the motion policy module with updated network parameters according to the re-obtained hidden vector, and then the condition discriminator with updated network parameters is used to judge the regenerated predicted physical motion sequence; and so on, until the judgment result is yes.

6. The physical-based controllable animal motion generation system in a three-dimensional scene according to claim 4, wherein The loss function used during the training of the underlying motion controller includes a motion encoder loss function, a motion policy loss function, and a discriminant loss function; Among them, the motion encoder loss function is as follows: Among them, represents a set of reference physical motion sequences, represents performing an expected calculation operation in the set , M represents a subset of reference physical action sequences randomly sampled from the set , E(M) represents the hidden vectors obtained by encoding each reference physical action sequence in the subset M, M' represents another subset of reference physical action sequences randomly sampled from the set , and E(M') represents the hidden vectors obtained by encoding each reference physical action sequence in the subset M'; ~ + represents that the subsets M and M' of reference physical action sequences sampled from the set overlap, ~ - represents that the subsets M and M' of reference physical action sequences sampled from the set are independently sampled without overlap; L + represents the first loss function when the subsets overlap, and L - represents the second loss function when the subsets do not overlap; Discriminant loss function As follows: Among them, represents the previous true skeleton state entered after the previous reference physical motion sequence Ⅰ in the set of animals to be controlled ; represents the current true skeleton state entered after the current reference physical motion sequence Ⅱ in the set of animals to be controlled , s represents the previous predicted skeleton state entered after the predicted physical motion sequence Ⅰ generated by the motion strategy module of the animal to be controlled according to the reference physical motion sequence Ⅰ, and s′ represents the current predicted skeleton state entered after the predicted physical motion sequence Ⅱ generated by the motion strategy module of the animal to be controlled according to the reference physical motion sequence Ⅱ; z represents the predicted hidden vector obtained by the motion encoder encoding any reference physical motion sequence; D(s, s′|z) represents the discrimination probability obtained by the discriminator judging according to the predicted skeleton states s and s′ corresponding to the predicted hidden vector z, where the discrimination probability is a decimal between 0 and 1. When D(s, s′|x) is closer to 1, it means that the discriminator believes that the predicted physical motion sequence corresponding to the predicted hidden vector z is a "true" motion sequence; when D(s, s′|x) is closer to 0, it means that the discriminator believes that the predicted physical motion sequence corresponding to the predicted hidden vector x is a "false" motion sequence; represents the discrimination probability obtained by the discriminator judging according to the true skeleton state and corresponding to the predicted hidden vector x, where the discrimination probability is a decimal between 0 and 1. When is closer to 1, it means that the discriminator believes that the reference physical motion sequence corresponding to the predicted hidden vector x is a "true" motion sequence; when is closer to 0, it means that the discriminator believes that the reference physical motion sequence corresponding to the predicted hidden vector x is a "false" motion sequence; During the training process, the discriminator needs to learn to discriminate the predicted physical motion sequence corresponding to the predicted hidden vector z as "false" and the reference physical motion sequence corresponding to the predicted hidden vector z as "true"; w gp represents the set gradient penalty term coefficient, and sg(·) represents the gradient stop operation; represents the gradient of the model parameter θ of the motion strategy module. Indicates performing an expected calculation operation in the set ; d π (s, s′|z) represents the distribution that the predicted skeleton state corresponding to the predicted physical motion sequence generated by the policy π under the given predicted latent vector z conforms to. denotes performing an expectation calculation operation under the distribution d π (s, s′|z); represents the distribution that the true skeleton states corresponding to the reference physical action sequences in the subset M conform to, represents performing an expected calculation operation under the distribution ; The motion policy loss function is as follows: Among them, s t represents the previous predicted skeleton state s, and s t+1 represents the current predicted skeleton state s'; s0 represents the initial skeleton state; γ represents the discount factor; t represents the step size, and t = 1, 2, …, T; represents performing an expectation calculation operation on the set ; p(τ|π,z) represents the distribution that the complete predicted physical motion sequence τ generated by the policy π under the given predicted latent vector z conforms to; represents performing an expectation calculation operation under the distribution p(τ|π,z); r(s t , s t+1 , z) represents the reward function calculated according to the previous predicted skeleton state s t , the current predicted skeleton state s t+1 , and the latent vector z, and there is r(s t , s t+1 , z) = r(s, s', z) = -log(1 - D(s, s'|z)).

7. A physically-based controllable animal motion generation system in a three-dimensional scene according to claim 1, characterized in that, The training process is carried out in a simulator with a physical simulation environment, and the physical law constraints include Newton's first law of motion, Newton's second law of motion, and Newton's third law of motion.