Action editing method of role model and related device

By automatically outputting the pose information of all joints through the motion editing network, the problem of low efficiency in character model motion editing is solved, and efficient motion editing is achieved.

CN121330133APending Publication Date: 2026-01-13BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410924494.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing technologies, the motion editing process for character models requires users to manually adjust joint positions and axis angles multiple times, resulting in low efficiency, high labor costs, and a complex motion editing process.

Method used

By acquiring the joint and motion information of the original character model, and responding to the user's motion editing operation, the motion editing network automatically outputs the pose information of all joints to generate a character model that performs the target action, thus avoiding a complex manual adjustment process.

Benefits of technology

It simplifies the motion editing process, reduces labor costs, improves the efficiency of motion editing, and achieves automated motion editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330133A_ABST
    Figure CN121330133A_ABST
Patent Text Reader

Abstract

The invention discloses an action editing method of a role model and a related device, which can be applied to the fields of cloud technology, artificial intelligence, digital humans, virtual humans, games, augmented reality and the like. And acquiring joint information and motion information of the original role model. And in response to an action editing operation of the target object on the target joint in the original role model, obtaining attitude information of the target joint. According to the joint information, the motion information and the attitude information of the target joint, attitude prediction is carried out through an action editing network, attitude information of a second full-amount joint is output, and an original role model for executing the target action is generated according to the attitude information of the second full-amount joint. According to the method and the device, the attitude information of all the joints can be automatically obtained only by specifying the attitude information of a small number of joints by the target object, then the original role model for executing the target action is generated, action editing is completed, the complex and tedious manual attitude adjustment process is avoided, the action editing process is simplified, the labor cost is reduced, and the action editing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for editing the motion of a character model. Background Technology

[0002] Motion editing is a technology widely used in industries such as film, entertainment, and sports. It involves manipulating the skeleton of a character model to create, modify, and refine various movements of the character model, such as walking, jumping, and fighting, in order to achieve specific animation effects or meet specific needs.

[0003] The relevant technology mainly involves users (such as animators) editing motions using auxiliary software. Users control the skeleton to achieve corresponding movements by repeatedly setting the position and axis of the joints.

[0004] During manual motion editing by the user, because the character model has many bones—for example, a human model has nearly 30 main bones—each time motion editing is performed, the user needs to adjust the movements of all bones from scratch. This requires manual modification more than 30 times, and in actual operation, it can result in nearly 100 modifications. This leads to a large number of adjustments, low motion editing efficiency, high labor costs, and a complex motion editing process. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a method and related apparatus for editing the motion of a character model. This method automatically obtains the posture information of all joints by requiring only the user to specify the posture information of a small number of joints, thereby generating an original character model that performs the corresponding target action and completing the motion editing. This approach avoids the complex and lengthy process of manually adjusting postures, thus simplifying the motion editing process, reducing labor costs, and improving the efficiency of motion editing.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] On one hand, embodiments of this application provide a method for editing the actions of a character model, the method comprising:

[0008] Obtain the joint and motion information of the original character model;

[0009] In response to the target object's action editing operation on the target joint in the original character model, the pose information of the target joint is obtained. The pose information of the target joint is generated based on the action editing operation. The target joint is a specified part of the first full set of joints in the original character model.

[0010] Based on the joint information, the motion information, and the pose information of the target joint, pose prediction is performed through an action editing network to output the pose information of the second full joint. The second full joint is the full joint of the target character model that corresponds to the original character model.

[0011] The original character model for performing the target action is generated based on the posture information of the second full-length joint.

[0012] On one hand, embodiments of this application provide a motion editing device for a character model, the device comprising an acquisition unit, a determination unit, and a generation unit:

[0013] The acquisition unit is used to acquire the joint information and motion information of the original character model;

[0014] The acquisition unit is further configured to acquire the pose information of the target joint in response to the action editing operation of the target object on the target joint in the original character model. The pose information of the target joint is generated based on the action editing operation. The target joint is a specified joint in the first full set of joints of the original character model.

[0015] The determining unit is used to perform posture prediction through an action editing network based on the joint information, the motion information and the posture information of the target joint, and output the posture information of the second full joint, wherein the second full joint is the full joint of the target character model that corresponds to the original character model.

[0016] The generation unit is used to generate an original character model that performs the target action based on the posture information of the second full-length joint.

[0017] On one hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:

[0018] The memory is used to store computer programs and to transfer the computer programs to the processor;

[0019] The processor is configured to execute the method described in any of the foregoing aspects according to instructions in the computer program.

[0020] In one aspect, embodiments of this application provide a computer-readable storage medium for storing a computer program that, when executed by a processor, causes the processor to perform the methods described in any of the foregoing aspects.

[0021] On one hand, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the foregoing aspects.

[0022] As can be seen from the above technical solution, this application realizes the action editing of the character model through user interaction. When it is necessary to edit the action of the original character model, the joint information and motion information of the original character model can be obtained. Furthermore, the target object performs the action editing operation, interacting with the original character model by implementing the action editing operation, thereby specifying the pose information of a small number of joints in the original character model. At this time, in response to the target object's action editing operation on the target joint in the original character model, the pose information of the target joint can be obtained. The pose information of the target joint is generated based on the action editing operation. The target joint is a portion of the specified joints in the first full set of joints of the original character model, and this portion of specified joints is a small number of joints in the first full set of joints. Then, based on the joint information, motion information, and pose information of the target joint, pose prediction is performed through the action editing network, outputting the pose information of the second full set of joints. The second full set of joints is the full set of joints of the target character model that corresponds to the original character model. Thus, the pose information of the full set of joints can be automatically determined based on the pose information of a small number of joints. The pose information of the full set of joints can reflect the target action to be edited, and then the original character model that performs the target action can be generated based on the pose information of the second full set of joints. As can be seen, this application only requires the target object to specify a small number of joint posture information to automatically obtain the posture information of all joints, thereby generating the original character model that performs the corresponding target action and completing the action editing. This method avoids the complex and lengthy manual posture adjustment process, thereby simplifying the action editing process, reducing labor costs, and improving the efficiency of action editing. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 An example image illustrating an action editing method for related technologies;

[0025] Figure 2 An application scenario architecture diagram of a character model motion editing method provided in this application embodiment;

[0026] Figure 3 A flowchart illustrating a method for editing the motion of a character model, as provided in this application embodiment;

[0027] Figure 4 A schematic diagram of the processing flow of an action editing network provided for an embodiment of this application;

[0028] Figure 5 A schematic diagram illustrating one implementation scheme of the action editing network provided in this application embodiment;

[0029] Figure 6 This is a schematic diagram illustrating the definition of a feasible range of a joint, provided in an embodiment of this application.

[0030] Figure 7 A schematic diagram of T-poseation provided for an embodiment of this application;

[0031] Figure 8 A schematic diagram illustrating one implementation scheme of the model parameter estimation network provided in an embodiment of this application;

[0032] Figure 9 This is a schematic diagram of an optimized processing flow provided in an embodiment of this application;

[0033] Figure 10 A schematic diagram illustrating one implementation scheme of the post-processing network provided in an embodiment of this application;

[0034] Figure 11 This application provides a schematic diagram of an overall framework for motion editing of animation.

[0035] Figure 12 This is a schematic diagram illustrating a process for motion editing of an original character model in a keyframe, provided as an embodiment of this application.

[0036] Figure 13 The original character model provided for the embodiments of this application is a motion editing framework diagram of a general model;

[0037] Figure 14 An original character model provided for embodiments of this application is an action editing framework diagram of an SMPLX model;

[0038] Figure 15 Comparison diagrams of the effects of motion editing provided in the embodiments of this application;

[0039] Figure 16 This application provides a schematic diagram illustrating the effect of motion editing in an embodiment.

[0040] Figure 17 A structural diagram of a motion editing device for a character model provided in an embodiment of this application;

[0041] Figure 18 A structural diagram of a terminal provided in an embodiment of this application;

[0042] Figure 19 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation

[0043] The embodiments of this application will now be described with reference to the accompanying drawings.

[0044] To achieve motion editing, the relevant technology mainly involves users (such as animators) editing motions using auxiliary software. Users control the skeleton to achieve corresponding movements by repeatedly setting the position and axis of the joints.

[0045] The specific motion editing workflow is as follows: First, an existing animation or still frame animation is used. The animator performs keyframe animation editing and keyframe interpolation to complete the animation. These two steps of motion editing and keyframe animation compositing are usually done using auxiliary software, such as Maya. Users obtain keyframe animation by repeatedly setting the position and axis angle of the joints. During the manual adjustment process, because the main skeleton of the character model has nearly 30 bones, each manually created keyframe requires modification of 30+ bones. In fact, bone modifications can be achieved through joint modifications. In practice, this results in approximately 100 modifications.

[0046] For example, if you need to adjust a character model to perform a drunken fist move, the process is as follows:

[0047] Step 1: Start from the root node of the hip and adjust each bone one by one (edit 30+ times);

[0048] Step 2: If the curvature of the lumbar spine is found to be incorrect, readjust the lumbar and upper rib bones (edit 20+ times);

[0049] Step 3: The movement is complete, but you feel that your feet are not generating enough power, so you need to adjust the arc of your supporting foot (edit 25+ times);

[0050] Step 4: After adjusting, I found that there was clipping through the left hand and knee (edited 10+ times);

[0051] The effects achieved through motion editing using relevant technologies can be seen in [reference]. Figure 1 As shown, the number of edits provided by related technologies is approximately the number of bones multiplied by the number of iterations, resulting in a large number of adjustments, low motion editing efficiency, high labor costs, and a complex motion editing process.

[0052] To address the aforementioned technical issues, this application provides a method for editing the actions of a character model. This method enables action editing of the character model through user interaction. When action editing of the original character model is required, only a small amount of joint posture information needs to be specified for the target object. The action editing network, which has been pre-trained, can automatically output the posture information of all joints, thereby generating the original character model that performs the corresponding target action and completing the action editing. This approach avoids the complex and lengthy process of manually adjusting postures, thus simplifying the action editing process, reducing labor costs, and improving the efficiency of action editing.

[0053] It should be noted that the action editing method for character models provided in this application can be applied to fields such as cloud technology, artificial intelligence, digital humans, virtual humans, games, and extended reality. Specifically, it can be applied to various scenarios that require action editing, such as game production, animation production, film and television special effects, advertising design, and sports training. For example, in animation production, a typical scenario is to directly generate animations that perform various actions; another typical scenario is to obtain animation resources using various motion capture (Mocap) methods and then fine-tune the actions to generate an animation that performs new actions. Motion capture can also be abbreviated as motion capture.

[0054] The action editing method for the character model provided in this application embodiment can be executed by a computer device, which may be a server or a terminal. The server may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminals include, but are not limited to, vehicle terminals, smartphones, tablets, computers, intelligent voice interaction devices, smart home appliances, and aircraft.

[0055] like Figure 2 As shown, Figure 2 An application scenario architecture diagram of a character model action editing method is shown, which takes a computer device as the terminal as an example.

[0056] This application scenario may include a terminal 200, and the target object can use the terminal 200 to edit actions. For example, action editing software may be installed on the terminal 200, and the action editing software integrates the action editing method provided in the embodiments of this application. The target object may be the object performing action editing, such as a user.

[0057] This application embodiment enables action editing of a character model through user interaction. When action editing of the original character model is required, the terminal 200 can obtain the joint information and motion information of the original character model. The target object then performs action editing operations on the terminal 200. By performing these operations, the target object interacts with the original character model, thereby specifying the posture information of a small number of joints in the original character model.

[0058] The character model can be a virtual character created through processes such as 3D modeling, texturing, and skeletal rigging. This application embodiment does not limit the type of character model; for example, it can be a human character model, an animal character model, etc. This application embodiment mainly uses a human character model as an example for description. The original character model can be the character model targeted in this motion editing. Joint information refers to detailed information about nodes or connection points in the character model used to connect bones and allow relative bone movement. Joint information defines the connection method, range of motion, and possible motion limitations between bones. Motion information refers to a set of animation parameters such as position, rotation, and scaling of bones in the character model over time. Motion information records the dynamic changes of the character model during animation. Joint information and motion information can reflect the original character model's current action.

[0059] When an action editing operation is received, the terminal 200 can respond to the action editing operation of the target object on the target joint in the original character model and obtain the pose information of the target joint. The pose information of the target joint is generated based on the action editing operation. The target joint can be a portion of the specified joints in the first full set of joints of the original character model. This portion of the specified joints is a small number of joints in the first full set of joints, and is usually the joint targeted by the target object when performing the action editing operation.

[0060] Then, the terminal 200 can predict the pose through the motion editing network based on the joint information, motion information and pose information of the target joint, and output the pose information of the second full joint. The second full joint is the full joint of the target character model that corresponds to the original character model. Thus, the pose information of the full joint can be automatically determined based on the pose information of a small number of joints. The pose information of the full joint can reflect the target action to be edited. Then, the original character model that performs the target action can be generated based on the pose information of the second full joint.

[0061] It should be noted that the methods provided in this application's embodiments may involve artificial intelligence (AI) technology. AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0062] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0063] This application embodiment may also involve machine learning. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning. In this application embodiment, an action editing network can be trained using machine learning.

[0064] It should be noted that in the specific implementation of this application, the entire process may involve user information and other related data. When the above embodiments of this application are applied to specific products or technologies, separate consent or permission from the user is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0065] Next, taking a computer device as the terminal as an example, the action editing method for the character model provided in the embodiments of this application will be described in conjunction with the accompanying drawings. See also Figure 3 , Figure 3A flowchart of a method for editing the motion of a character model is shown. The method may include steps S301-S304, as detailed below:

[0066] S301. Obtain the joint information and motion information of the original character model.

[0067] In this embodiment, for motion editing, the terminal can load the original character model and then obtain the joint information and motion information of the original character model. The joint information and motion information reflect the current motion of the original character model.

[0068] It should be noted that motion editing can directly generate original character models performing various actions from original character models that have not yet performed any actions. In this case, the joint and motion information can reflect the original character model in its natural state when it is not currently performing any actions. Motion editing can also modify an original character model that is currently performing a certain action, thereby adjusting it to generate other original character models. In this case, the joint and motion information can reflect the original character model currently performing a certain action, such as raising a hand, jumping, or kicking.

[0069] Understandably, in some cases, the purpose of motion editing is not only to obtain the original character model performing the target action at a single moment, but more importantly, to obtain an animation of the original character model performing a series of target actions. In this case, motion editing can be performed on the original character model in an existing animation to generate the required final animation, in which the original character model performs a series of target actions. However, the existing animation includes multiple video frames, and not every video frame can reflect the original character model's current action. Therefore, to reduce the workload of motion editing, one possible implementation is to obtain the joint and motion information of the original character model by acquiring the source animation data of the original character model. The source animation data can represent the existing animation including the original character model, and then keyframes can be determined from the source animation data, and the joint and motion information of the original character model in the keyframes can be obtained.

[0070] The above method obtains the joint and motion information of the original character model from keyframes. Subsequent animation editing is also performed on the original character model within the keyframes, thus completing the keyframe depiction. Therefore, this type of motion editing can be called K-motion. In animation production, K-motion usually refers to the process of creating keyframe animation. Keyframe animation is a technique for creating animation by defining keyframes at specific points in time. Keyframes contain important state information from the existing animation, reflecting the original character model's current action; they are representative video frames within the existing animation.

[0071] The above method obtains the joint and motion information of the original character model from keyframes, and then uses the method provided in this application embodiment to perform motion editing only on keyframes, thereby eliminating the need to perform motion editing on every video frame, greatly reducing the amount of data for motion editing and improving motion editing efficiency.

[0072] S302. In response to the target object's action editing operation on the target joint in the original character model, obtain the pose information of the target joint, the pose information of the target joint is generated based on the action editing operation, and the target joint is a specified joint among the first full joints of the original character model.

[0073] After the terminal loads the original character model, the target object can perform action editing operations on the target joints in the original character model based on the target action. The target action can be the action that the original character model is expected to perform during action editing, such as walking, running, jumping, squatting, crawling, waving, raising an arm, high knee, etc. The target joint is a representative joint among the joints that perform the target action. In this embodiment, the target joint is a portion of the designated joints in the first full set of joints of the original character model, and the number of target joints is relatively small. The first full set of joints of the original character model can support various actions of the original character model. The first full set of joints can include, for example, the shoulder joint, elbow joint, wrist joint, hip joint, knee joint, ankle joint, neck joint, and hand joint. If the target action is a left leg bending and raising its leg, the target joint could be, for example, the left knee joint.

[0074] In one possible implementation, the target joint can be determined through an action editing operation, such as the joint targeted by the action editing operation. The action editing operation can be the target object dragging a specified joint, or it can be directly inputting the joint's identifier and corresponding posture information. For example, if the target action is a left leg bending and lifting its leg, and the action editing operation involves the target object dragging the left lower leg to a certain position, thereby changing the posture of the left knee joint, then the target joint could be the left knee joint.

[0075] Posture information can be information reflecting the joint's posture. In this embodiment, the joint can perform at least one of translation and rotation. Therefore, in one possible implementation, the posture information of the target joint can include at least one of the target joint's position information and axis angle information. The target joint's position information can reflect the target joint's translation, and the target joint's axis angle information can reflect the target joint's rotation.

[0076] S303. Based on the joint information, the motion information, and the pose information of the target joint, pose prediction is performed through an action editing network to output the pose information of the second full joint. The second full joint is the full joint of the target character model that corresponds to the original character model.

[0077] After obtaining joint information, motion information, and pose information of the target joint, the terminal can perform pose prediction through an action editing network to generate pose information for the second full set of joints. The second full set of joints is the full set of joints of the target character model that corresponds to the original character model. The target character model can be the original character model itself, or a character model after processing the original character model. Different processing methods may result in different target character models, which will be described in detail later.

[0078] In one possible implementation, the input to the motion editing network, in addition to pose information, may include the target joint's identifier, joint type, and gender. The identifier can be an ID, and the joint type can include fixed position, fixed global rotation, or eye-viewing direction. For example, considering pose information including position and axis-angle information, see [link to relevant documentation]. Figure 4 As shown, the input to the motion editing network can be joint attributes of non-specific length, including the target joint's position information, target joint's axis-angle information, target joint ID, joint type, and gender. The output of this motion editing network is the position and axis-angle information of the second full set of joints, where the position information can be represented by global translation (e.g., global translation of the root node, where the root node could be a hip joint). The resulting pose information of the second full set of joints can be used to drive the original character model to perform the target action.

[0079] It should be noted that, in one possible implementation, a controller can be bound to each joint, and the joint can be represented by the controller. Therefore, the position information of the joint can be represented by the position information of the controller, and the axis angle information of the joint can be represented by the axis angle information of the controller.

[0080] In one possible implementation, position information can be represented by position coordinates, which can be an Np*3 matrix, where Np represents the number of target joints with fixed positions. Axis angle information can be represented by a rotation matrix, denoted as Rot6d, which can be an NR*6 rotation matrix, where NR represents the number of target joints with fixed global rotation. The eye's viewing direction can be represented by an Nla*6 matrix, where Nla represents the number of target joints with fixed eye viewing direction, and Np + NR + Nla = N, where N represents the total number of joints.

[0081] The motion editing network can be a network that determines the pose information of all joints based on the pose information of a small number of joints. The motion editing network can be pre-trained. The embodiments of this application do not limit the network structure of the motion editing network, such as the structure of a transformer, and are not limited to schemes such as convolutional neural networks (CNN) and recurrent neural networks (RNN).

[0082] One implementation scheme for action editing networks can be found in [reference needed]. Figure 5 As shown, the input to the motion editing network is Np*3 position coordinates, padded with zeros to become Np*6, along with an NR*6 rotation matrix Rot6d and an Nla*6 eye viewing direction. These three inputs are combined into an N*6 input. The other two inputs are the N*1 target joint ID and the N*3 joint type. These five inputs are then merged into an N*11 matrix. The motion editing network structure can be a multi-layer fully connected network (FC) structure, such as four fully connected layers. After four fully connected layers, dimensional averaging or summing is performed on the joint dimensions of length N, followed by another fully connected layer. The output is a vector of total length 144, which is transformed into a 24*6 Rot6d to obtain the axis-angle information of the second full set of joints. The 24*6 Rot6d is then converted into a rotation matrix of 24 joints. Finally, forward kinetics (FK) is used to obtain the position information of the second full set of joints.

[0083] S304. Generate the original character model that performs the target action based on the posture information of the second full-length joint.

[0084] The posture information of the second full joint can reflect the target action. That is, the second full joint can be controlled according to the posture information of the second full joint to perform the target action. Therefore, after obtaining the posture information of the second full joint, the terminal can generate the original character model that performs the target action based on the posture information of the second full joint.

[0085] It should be noted that when the target character model is the original character model itself, the pose information of the second full joint is directly used to drive the original character model, thereby obtaining the original character model that performs the target action. When the target character model is a character model that has been processed from the original character model, in one possible implementation, the original character model that performs the target action can be obtained through redirection.

[0086] It should be noted that the above S301-S304 can be implemented using inverse kinetics (IK) systems.

[0087] As can be seen from the above technical solution, this application realizes the action editing of the character model through user interaction. When it is necessary to edit the action of the original character model, the joint information and motion information of the original character model can be obtained. Furthermore, the target object performs the action editing operation, interacting with the original character model by implementing the action editing operation, thereby specifying the pose information of a small number of joints in the original character model. At this time, in response to the target object's action editing operation on the target joint in the original character model, the pose information of the target joint can be obtained. The pose information of the target joint is generated based on the action editing operation. The target joint is a portion of the specified joints in the first full set of joints of the original character model, and this portion of specified joints is a small number of joints in the first full set of joints. Then, based on the joint information, motion information, and pose information of the target joint, pose prediction is performed through the action editing network, outputting the pose information of the second full set of joints. The second full set of joints is the full set of joints of the target character model that corresponds to the original character model. Thus, the pose information of the full set of joints can be automatically determined based on the pose information of a small number of joints. The pose information of the full set of joints can reflect the target action to be edited, and then the original character model that performs the target action can be generated based on the pose information of the second full set of joints. As can be seen, this application only requires the target object to specify a small number of joint posture information to automatically obtain the posture information of all joints, thereby generating the original character model that performs the corresponding target action and completing the action editing. This method avoids the complex and lengthy manual posture adjustment process, thereby simplifying the action editing process, reducing labor costs, and improving the efficiency of action editing.

[0088] It should be noted that in some cases, the original character model can be a biological character model, such as a human or animal character model. Due to physiological factors of the organism, such as skeletal structure, ligaments, and muscle strength, the maximum range of motion of a joint is limited. This limitation can be called the joint's feasible range, which can be considered a reasonable physiological range. When the target object inputs posture information to the target joint, the resulting action may be abnormal for some reason. In this case, to reduce the occurrence of abnormal actions, the posture information of the target joint can be obtained by acquiring the joint's feasible range and the posture information input by the target object to the target joint. Then, based on the joint's feasible range and the input posture information, the target joint's posture information can be determined.

[0089] Depending on the action editing operation, the method for obtaining the input posture information can vary. If the action editing operation involves the target object dragging a target joint, the input posture information can be obtained by automatically recognizing the posture information of the target joint when the target object stops dragging. If the action editing operation involves the target object directly inputting the identifier and corresponding posture information of the target joint, the input posture information can be obtained by directly acquiring the posture information input by the target object.

[0090] The following defines the feasible range of joints, such as... Figure 6 As shown in Figure (a), Figure 6 Figure (a) provides a joint chain where each circular node represents a joint, and the joints in the chain are fully free. The lines connecting two joints represent bones, r0, r1, ..., r n These represent the lengths of the corresponding bones. In one possible implementation, point A represents the passively fixed joint, and point B represents the actively fixed joint. The range of motion (i.e., the feasible range of the joint) of the actively fixed joint is a sphere or shell centered on the passively fixed joint, where the minimum radius of the shell can be:

[0091]

[0092] Among them, R min Let r represent the minimum radius of the spherical shell, i.e., the minimum feasible range of the joint; min indicates finding the minimum value; r0 represents the bone length between the actively fixed joint and adjacent joints; adjacent joints are the joints adjacent to the actively fixed node; r i represents the length of the bone between any two adjacent joints, and n represents the number of joints in the joint chain.

[0093] The maximum radius of the spherical shell can be:

[0094]

[0095] Among them, R max The minimum radius of the spherical shell, i.e., the maximum value of the feasible range of the joint, is represented by r. i represents the length of the bone between any two adjacent joints, and n represents the number of joints in the joint chain.

[0096] Based on the above method, the feasible range of each joint can be determined, thereby obtaining the feasible range of the target joint. It should be noted that... Figure 6 Figure (a) illustrates an example of a joint located on a joint chain. In some cases, multiple joint chains exist between joints, and the range of motion of an actively fixed joint may be determined by multiple passively fixed joints and joint chains. See also Figure 6As shown in Figure (b), when the joint represented by point B is an active fixation joint, it is located on three joint chains: a joint chain with point B as the active fixation joint and point C as the passive fixation joint, a joint chain with point B as the active fixation joint and point D as the passive fixation joint, and a joint chain with point B as the active fixation joint and point E as the passive fixation joint. The range of motion of point B is constrained by these three joint chains.

[0097] Based on this, in order to obtain Figure 6 In Figure (b), the feasible range of the actively fixed joint is determined by calculating the spherical shell for each joint chain as described above, and then intersecting the spherical shells corresponding to multiple joint chains. For example... Figure 6 As shown in Figure (b), the spherical shell shown in 601 is calculated based on the joint chain with point B as the active fixed joint and point C as the passive fixed joint. The spherical shell shown in 602 is calculated based on the joint chain with point B as the active fixed joint and point E as the passive fixed joint. The spherical shell shown in 603 is calculated based on the joint chain with point B as the active fixed joint and point D as the passive fixed joint. The feasible range of the active fixed joint shown in point B is obtained by intersecting the spherical shells shown in 601, 602 and 603.

[0098] The above method defines a reasonable joint feasible range, and then determines the posture information of the target joint based on the joint feasible range and the input posture information. This ensures that the determined posture information of the target joint conforms to the constraints of the joint feasible range, thereby reducing the occurrence of abnormal movements during motion editing, ensuring that the generated movements conform to the laws of natural movement, and improving the realism and credibility of the movements.

[0099] It is understandable that when determining the target joint's posture information based on the joint's feasible range and the input posture information, there are different relationships between the input posture information and the joint's feasible range. For example, the input posture information may be within the joint's feasible range or may be outside of it. Different relationships result in different determined target joint posture information. Therefore, in one possible implementation, determining the target joint's posture information based on the joint's feasible range and the input posture information could be as follows: If the input posture information is determined to be within the joint's feasible range, it is determined as the target joint's posture information. If the input posture information is determined to be outside the joint's feasible range, it is adjusted, and the adjusted posture information is determined as the target joint's posture information. The adjusted posture information must either be within the joint's feasible range or the excess value of the adjusted posture information beyond the joint's feasible range is less than a preset threshold.

[0100] As described above, joints can be bound to controllers. Therefore, when the posture information input to a target joint exceeds the feasible range of the joint, adjusting the input posture information can forcibly pull the controller bound to the target joint back to or near the feasible range of the joint. This allows the input of the motion editing network to be controlled within a reasonable physiological range, preventing unreasonable movements from occurring when the input of the motion editing network exceeds the reasonable physiological range.

[0101] The above method is based on the different relationships between the input posture information and the feasible range of the joint, and uses the corresponding appropriate method to determine the posture information of the target joint, thereby obtaining more reasonable posture information and avoiding the action editing network from generating unreasonable actions based on the posture information.

[0102] In this embodiment, the automatic generation of full-joint pose information based on a small number of joint pose information is the joint for motion editing. During motion editing, the original character model may be in a standard pose or performing a certain action; a standard pose is easier to edit than other actions. Therefore, when generating full-joint pose information based on a small number of joint pose information, the method for outputting the second full-joint pose information by predicting pose using the motion editing network based on joint information, motion information, and the pose information of the target joint can be to standardize the pose of the original character model based on the joint information and motion information, obtaining a standard pose original character model. The standard pose, as a preset posture used to bind the character model in animation, provides a unified benchmark for modelers and animators. In this standard pose, each part of the character model is in a relatively standard and easily observable state, facilitating subsequent motion editing. Then, based on the standard pose original character model, the character model representation parameters of the target character model are obtained. The character model representation parameters and the pose information of the target joint are input into the motion editing network, which performs pose prediction to output the second full-joint pose information.

[0103] Correspondingly, the original character model in the standard pose is different from the original character model. In order to ensure that the final generated target action is more natural, smooth and undistorted, the way to generate the original character model that performs the target action based on the pose information of the second full joint is to redirect the pose information of the second full joint to the original character model to obtain the original character model that performs the target action.

[0104] In one possible implementation, when the original character model is a human character model, the standard pose can be a T-pose, and pose standardization can be T-poseification. In the T-pose, the movable person's neck, waist, knees, and ankles are in a straight line, while the arms are extended horizontally, thus forming a standard T-shaped standing posture. This pose is named for its resemblance to the English letter "T". See also Figure 7 As shown, Figure 7 Figure (a) shows the original character model performing the current action. After T-poseation, the original character model in T-pose is obtained. The original character model in T-pose is as follows: Figure 7 As shown in Figure (b).

[0105] The above method standardizes the pose of the original character model, allowing for subsequent motion editing within a standardized pose. In this standardized pose, all parts of the character model are in a relatively standard and easily observable state, facilitating subsequent motion editing and producing natural and fluid animations, thus improving the accuracy of motion editing. Furthermore, the T-poseified original character model has better compatibility and portability. This means the model can be easily transferred between different software platforms while maintaining the consistency and stability of its movements.

[0106] Understandably, when creating an original character model, it's possible to create original character models with different body shape characteristics, such as height, weight, and build. Therefore, to efficiently and flexibly represent the original character model, we can obtain the character model representation parameters of the target character model. These parameters can represent original character models with different body shape characteristics. By adjusting these parameters, character models with different body shape characteristics can be quickly generated or modified, thus facilitating the representation and management of character models. These character model representation parameters can be represented by the Beta (β) parameter. Since body shape characteristics may include features in different dimensions, the β parameter can consist of multiple values, such as 10 values, each representing a single dimension of body shape characteristic.

[0107] There are several ways to obtain the character model representation parameters. In one possible implementation, the body shape features of the original character model can be disregarded, and its body shape features can be assumed to be a standard body shape, with the character model representation parameters for the standard body shape being 0. In this case, the way to obtain the character model representation parameters of the target character model based on the original character model in a standard pose can be to determine the original character model in a standard pose as the target character model and set the character model representation parameters of the target character model to 0.

[0108] In another possible implementation, in order to accurately represent the body features of the original character model in the standard pose and improve the accuracy of subsequent motion editing, the method of obtaining the character model representation parameters of the target character model based on the original character model in the standard pose can be to align the original character model in the standard pose with the zero-parameter model of the standard parameterized model to obtain the initial scaling factor of the original character model in the standard pose relative to the zero-parameter model. Then, the original character model in the standard pose is standardized based on the initial scaling factor to obtain the joint position information of the target character model. Finally, the character model representation parameters are determined based on the joint position information of the target character model.

[0109] The zero-parameter model is a standard parametric model, except that its character model representation parameter is 0. In this embodiment, the zero-parameter model can be represented as B0, and the original character model of the standard pose can be represented as A. This embodiment does not limit the type of standard parametric model. In one possible implementation, the standard parametric model can be a skinned multi-person linear model (SMPL) or an extended version of SMPL (SMPL eXpressive, SMPLX).

[0110] It should be noted that this application does not limit the method for alignment processing and determining the joint position information of the target character model. In one possible implementation, the alignment processing can be described by the following formula:

[0111]

[0112]

[0113] Among them, P src P represents the joint position information of the zero-parameter model. tgt P' represents the joint position information of the original character model in standard pose. src P represents src Decentralized joint position information, P' tgt P represents tgt Decentralized joint position information. (x) i ,y i ,z i ) represents P src The joint position information of the i-th joint, x i This represents the joint position information of the i-th joint along the x-axis of the spatial coordinate system, y i This represents the joint position information of the i-th joint along the y-axis in the spatial coordinate system, z. i This represents the joint position information of the i-th joint along the z-axis of the spatial coordinate system. (xi ',y i ',z i ') represents P tgt The joint position information of the i-th joint, x i ' represents the joint position information of the i-th joint along the x-axis of the spatial coordinate system, y i ' represents the joint position information of the i-th joint along the y-axis in the spatial coordinate system, z i ' represents the joint position information of the i-th joint along the z-axis in the spatial coordinate system. n represents the number of joints in the original character model with zero parameters and standard pose.

[0114] The formula for translation and scaling is as follows:

[0115]

[0116] Where s represents the initial scaling factor, and std() represents the function for calculating the standard deviation. P' src A matrix composed of the x-components of the position information of each joint. P' src A matrix composed of the y-components of the position information of each joint. P' src The matrix formed by the z-components of the joint position information. P' tgt A matrix composed of the x-components of the position information of each joint. P' tgt A matrix composed of the y-components of the position information of each joint. P' tgt The matrix is ​​composed of the z-components of the position information of each joint.

[0117] Then, the original character model in standard pose is normalized based on the initial scaling factor to determine the joint position information of the target character model. In one possible implementation, the normalization formula is as follows:

[0118]

[0119] Among them, P t ′ g ′ t This indicates the joint position information of the target character model.

[0120] Finally, based on the joint position information of the target character model, the character model representation parameters are determined. It should be noted that when determining the character model representation parameter β, a scaling factor can also be determined. This scaling factor can be called the intermediate scaling factor, denoted as s', and the final scaling factor can be s*s'.

[0121] This application embodiment can obtain more accurate character model representation parameters by performing parameter estimation, thereby using the character model representation parameters to express original character models with different body shape characteristics. By adjusting the character model representation parameters, character models with different body shape characteristics can be quickly generated or modified, thus facilitating the representation and management of character models.

[0122] It should be understood that the embodiments of this application do not limit the determination of character model representation parameters based on the joint position information of the target character model. In one possible implementation, the joint position information of the target character model can be input into a model parameter estimation network, thereby performing parameter estimation through the model parameter estimation network and outputting character model representation parameters. In addition, the model parameter estimation network can also output intermediate scaling factors.

[0123] It is understood that the model parameter estimation network can be pre-trained, and the training data used to train the model parameter estimation network can be provided by publicly available standard parameterized model (e.g., SMPL) datasets. This application does not limit the network structure of the model parameter estimation network. In one possible implementation, the model parameter estimation network can be a network composed of multiple fully connected layers. An implementation scheme of the model parameter estimation network can be found in [reference needed]. Figure 8 As shown. In Figure 8 In this model parameter estimation network, the input is the joint position information of the target character model, which can be, for example, an N*3 matrix. The joint position information of the target character model can be arranged according to the SMPL standard. The model parameter estimation network includes multiple fully connected layers, and the output is the character model representation parameters and intermediate scaling factors. For example, the multiple fully connected layers can be... Figure 8 The four fully connected layers shown have an output layer size of 10 + 1 = 11. The first ten values ​​are the character model representation parameters, and the last value is the intermediate scaling factor.

[0124] The above method outputs the character model representation parameters through a trained model parameter estimation network, achieving intelligent and automated determination of character model representation parameters and improving the efficiency of parameter determination. Furthermore, since the model parameter estimation network is trained on a large amount of data, it can learn complex patterns and features, thereby estimating the character model representation parameters more accurately.

[0125] In another possible implementation, the character model representation parameters can be obtained directly using an optimization algorithm, and the final scaling factor can also be obtained.

[0126] Specifically, when using the optimization algorithm, the joint position information of the target character model can be represented by the character model representation parameters to obtain the parameter representation value of the target character model. Then, optimization is performed with minimizing the parameter representation value and the obtained joint position information of the target character model as the optimization objective to obtain the character model representation parameters.

[0127] In one possible implementation, the optimization objective can be expressed as:

[0128]

[0129] Where min represents finding the minimum value, scale represents the final scaling factor, and J regressor Let M represent an M*N matrix that generates joint positions from skin vertices, where M is the number of joints and N is the number of vertices. J represents an N*3 matrix formed by all vertices of a standard parametric model. tgt The joint position information of the target character model is represented by an M*3 matrix.

[0130]

[0131] Where T0 represents the joint position information of the zero-parameter model, T i This represents the incremental position information of the joint, β. i The i-th value in the parameter represents the character model, and k represents the number of values ​​in the parameter.

[0132] The above method directly determines the character model representation parameters through optimization algorithms, which can automatically find the optimal solution, greatly saving time, reducing costs, and improving the efficiency of motion editing.

[0133] As described above, determining the pose information of all joints based on the pose information of a small number of joints is key to the motion editing method provided in this application embodiment, and the motion editing network is crucial for determining the pose information of all joints based on the pose information of a small number of joints. The accuracy of the motion editing network directly affects the accuracy of its output. When the implementation of S303 differs, a suitable method must be used to train the motion editing network to obtain a matching motion editing network. Therefore, in the implementation of S303, if parameter estimation based on a standard parametric model is required, the motion capture dataset of the standard parametric model can be used as training samples when training the motion editing network. The standard parametric model in the training process can be called the sample standard parametric model.

[0134] Specifically, this application also provides a training method for an action editing network. During training, a motion capture dataset of a sample standard parameterized model can be obtained first. The motion capture dataset includes sample pose information of the third full joint of the sample standard parameterized model. Then, sample pose information of sample joints (part of the third full joint) is randomly selected from the motion capture dataset. The sample pose information of the sample joints and the role model representation parameters of the sample standard parameterized model are then input into an initial network. The initial network performs pose prediction and outputs the predicted pose information of the third full joint. Subsequently, based on the difference between the predicted pose information of the third full joint and the sample pose information of the third full joint, the network parameters of the initial network are adjusted and repeatedly trained to obtain the action editing network.

[0135] To train the above motion editing network, the training data needs to be prepared first. Semantic redirection is used to redirect the actions of various character models to the standard parametric model to ensure correct body shape and pose redirection.

[0136] To facilitate subsequent use of the action editing network, after training the action editing network, the weight file of the action editing network can be saved, so that the weight file can be loaded during action editing to use the action editing network.

[0137] The embodiments of this application are trained on a standard parametric model, so the motion editing of any character model can be completed by retargeting, avoiding the problem of different initial poses and bone lengths of different bones in cross-model processes, and reducing the workload of retargeting.

[0138] It should be noted that in S304, an original character model for performing the target action can be generated based on the posture information of the second full-joint system. In some cases, if the generated posture information of the full-joint system is inaccurate, problems such as some joint postures exceeding reasonable physiological range, collisions, and clipping may occur. Therefore, to avoid some joint postures exceeding reasonable physiological range, ensure accurate limb contact with the ground, and avoid collisions and clipping, the method for generating the original character model for performing the target action based on the posture information of the second full-joint system can be to optimize the posture information of the second full-joint system to obtain optimized posture information, and then generate the original character model for performing the target action based on the optimized posture information.

[0139] It should be noted that the optimization of the pose information of the second full-length joint can be achieved through a post-processing network. The post-processing network can be pre-trained, and the network structure of the post-processing network is not limited in this embodiment. In one possible implementation, the post-processing network can be a CNN.

[0140] This application optimizes the posture information of the second full-range joint, which can accurately solve the problem of some joint postures exceeding the reasonable physiological range, ensuring accurate limb contact with the ground and avoiding problems such as collision and clipping. At the same time, the optimized processing method satisfies the problem of smooth transition of movements under different controller settings.

[0141] During optimization, various penalty items can be set, and then the penalty value of at least one penalty item can be calculated based on the pose information of the second full-joint. Then, the pose information of the second full-joint is iteratively optimized based on the penalty value of at least one penalty item until the iterative optimization stopping condition is met, and the optimized pose information is obtained.

[0142] At least one penalty item may include position loss, axis angle loss, joint physical constraint loss, head orientation loss, etc. Position loss may include position loss for fixed joints (the difference between the position information of the fixed joint determined by S303 and the actual position information of the controller) and position loss for non-fixed joints (the difference between the position information optimized for non-fixed joints and the position information determined by S303); axis angle loss may include local axis angle loss (the difference between the local axis angle information determined by S303 and the actual local axis angle information of the controller) and global axis angle loss.

[0143] In the embodiments of this application, the listed optimization penalty items can have many variations. For example, the constraint on the axis angle can be in the form of quaternions, rotation matrices, 5-dimensional rotation vectors, 6-dimensional rotation vectors, axis angles, Euler angles, partial Euler angles, etc., or it can be in the form of global axis angles and local axis angles.

[0144] The following introduces the definitions of several commonly used penalty items. For the position loss of a fixed joint, its expression can be as follows:

[0145]

[0146] Among them, E pos λ represents the position loss of the fixed joint. pos The weights represent the position loss, where i represents the i-th joint, and set pos This represents the set of joints that need to participate in optimizing the position loss. This indicates the actual location information of the controller. This represents the position information of the fixed joint determined by S303, i.e., the position information of the fixed joint output by the motion editing network, and |||| represents the calculated norm.

[0147] The expression for local axis angle loss can be given as follows:

[0148]

[0149] Among them, E local-rot λ represents the local axis angle loss. local-rot The weights represent the local axis angle loss, where i represents the i-th joint, and set l-rot This represents a set of joints containing local axis-angle information that needs to be constrained, where the order follows the standard skeleton order. `tr()` is a function that computes the trace of a matrix. This represents the actual local axis angle information of the controller. This indicates the local axis angle information determined by S303, i.e., the local axis angle information output by the action editing network.

[0150] The expression for the global axis-angle loss can be shown below:

[0151]

[0152] Among them, E global-rot λ represents the global axis angle loss. global-rot The weights represent the global axis-angle loss, where i represents the i-th joint, and set g-rot This represents the set of joints that require constraints on global axis angle information. This represents the actual global axis angle information of the controller. This represents the calculated global axis angle information.

[0153] In one possible implementation, the pose information of the second full-scale joint determined by S303 includes axis angle information and position information. The axis angle information is local axis angle information, and the position information is represented by the global translation of the root node. Therefore, during post-processing, in order to obtain the penalty value for subsequent penalty items, the position information and global axis angle information of the second full-scale joint can be calculated first based on the local axis angle information and the global translation of the root node. Then, the penalty value for at least one penalty item is calculated based on the position information and global axis angle information of the second full-scale joint.

[0154] After obtaining the penalty value for each penalty item, the iterative optimization process of the pose information of the second full joint based on the penalty value of at least one penalty item can involve weighting the values ​​according to specific weights, calculating the Jacobian matrix, approximating the Hessian matrix using the Jacobian matrix, weighting multiple Hessian matrices, and solving a quadratic equation. When solving the quadratic equation, the Gauss-Newton method (GS) or the Least Square Error Method (LSE) can be used, employing a preprocessed conjugate gradient method to iteratively optimize to a suitable penalty value, thereby obtaining the optimized pose information. During the iterative optimization process, the weights and iteration step size of different penalty items can be adjusted based on the penalty value obtained in the current iteration, thus proceeding to the next iteration, and so on, until the iterative optimization stopping condition is met, obtaining the optimized pose information. The iterative optimization stopping condition can be that the penalty value reaches a threshold, the penalty value stops changing, or the number of iterations reaches a preset number, etc. This application embodiment does not limit the iterative optimization stopping condition.

[0155] See Figure 9 As shown, after obtaining the global translation, position information, axis angle information, and controller head orientation of the second full-joint, the penalty value of the corresponding penalty item can be calculated based on these parameters. These penalty items may include, for example, position loss, axis angle loss, head orientation loss, and joint physical constraint loss. Then, based on the penalty values ​​of each penalty item, the optimizer uses the corresponding optimization algorithm for iterative optimization. During the iterative optimization process, in each iteration, it is determined whether the penalty value of the penalty item in this iteration reaches a threshold. If the threshold is reached, the iteration stops, and the optimized posture information is obtained. If the threshold is not reached, the global translation, position information, axis angle information, and controller head orientation of the second full-joint are updated, the penalty value of the corresponding penalty item is recalculated, and the next iterative optimization continues.

[0156] in, Figure 9 The optimizer can employ Sequential Quadratic Programming (SQP) and interior point methods. This application does not limit the optimization algorithm used by the optimizer; for example, it can also employ GS, Dogleg, or Alternating Direction Method of Multipliers (ADMM), etc.

[0157] It should be noted that when the pose information of the second full-scale joint determined by S303 includes axis angle information and position information, where the axis angle information is local axis angle information and the position information is represented by the global translation of the root node, one implementation scheme of the post-processing network can be found in [reference needed]. Figure 10 As shown. The input to the post-processing network is local axis-angle information, such as an N*6 rotation matrix Rot6d. This local axis-angle information, along with the global translation of the root node, is used to obtain global axis-angle information and position information via FK. The global axis-angle information can be an N*6 rotation matrix, and the position information can be an N*3 matrix. The position information and global axis-angle information are combined into an N*9 matrix, which is then input into the post-processing network. The network outputs a pose probability density, which is used to determine the optimized pose information.

[0158] The above embodiments have provided a detailed description of motion editing. Next, we will introduce the motion editing method for character models provided in this application embodiment, in conjunction with a practical application scenario. This application scenario involves obtaining animation resources using various motion capture methods, and then needing to fine-tune the motion to generate a new animation. The obtained animation resources can be called source animation, which includes the original character model.

[0159] In this application scenario, the overall framework for motion editing of animations can be found in [reference needed]. Figure 11 As shown. First, obtain the source animation data (see...). Figure 11 As shown in S1101), keyframes are characterized by the IK system user (see [reference]). Figure 11 As shown in S1102, this involves motion editing for specific frames (i.e., keyframes) in the source animation. When editing motion for any keyframe, the user sets the pose information of the target joints in the original character model. The IK system automatically estimates the rotation values ​​of all joints. The motion editing method provided in this application can quickly edit a user-expected keyframe. Then, the keyframe edited by the IK system is linked with other keyframes specified by the user in the source animation data for keyframe continuation (see [link to relevant documentation]). Figure 11 As shown in S1103, a new animation is generated. Keyframe continuity refers to defining changes in motion by setting multiple keyframes consecutively along the animation's timeline. Transition frames (also called intermediate frames) are automatically generated between these keyframes, making the changes in motion appear smooth and continuous. Automatic generation of transition frames based on keyframes can be achieved by interpolating keyframes to obtain transition frames. Next, physical post-processing is performed on the new animation (see [link to documentation]). Figure 11 As shown in S1104, the rhythm and physics of the new animation are enhanced through post-processing. Finally, the new animation is used to drive the original character model (see [reference]). Figure 11 (As shown in S1105).

[0160] When users are creating keyframes, they can edit the motion of the original character model within the keyframes. See [link / reference]. Figure 12 As shown, load the original character model (see...) Figure 12 As shown in S1201, the user performs motion editing operations on the original character model, inputting the pose information of the target joints (see...). Figure 12 As shown in S1202, the original character model is posed in a T-pose (see [reference]). Figure 12 As shown in S1203, the original character model in the T-pose is obtained. Parameter estimation is then performed using a model parameter estimation network (see [link]). Figure 12 As shown in S1204, the character model representation parameters are output. The character model representation parameters and the pose information of the target joints are input into the motion editing network, which then performs pose prediction (see [link to documentation]). Figure 12 As shown in S1205, output the pose information of the second full-length joint (see S1205). Figure 12 As shown in S1206). The pose information of the second full-length joint is optimized using a post-processing network (see [reference]). Figure 12 As shown in S1207, the optimized pose information is obtained. The optimized pose information is then redirected to the original character model (see [reference]). Figure 12 As shown in S1208, the original character model that performs the target action is generated.

[0161] In practical applications, the motion editing process will vary depending on the type of the original character model. The original character model can be an SMPLX model or a general model. Motion editing framework diagrams for general and SMPLX models can be found separately. Figure 13 and Figure 14 As shown. In Figure 13 In this process, motion editing is performed on the original character model, which is of the general type. Specifically, the original character model is loaded, and the user performs motion editing operations on it, inputting the pose information of the target joints through user interaction, such as the position and axis angle information of the target joints. The original character model is then T-posed to obtain the T-pose original character model. During T-poseation, the original character model in the standard pose is aligned with the zero-parameter model of the standard parametric model to obtain the initial scaling factor of the original character model in the standard pose relative to the zero-parameter model. Then, based on the initial scaling factor, the original character model in the standard pose is standardized to obtain the joint position information of the target character model. Finally, based on the joint position information of the target character model, the character model representation parameters are determined.

[0162] During motion editing, the input to the motion editing network includes not only pose information (including position and axis-angle information), but also the target joint identifier, joint type, gender, head orientation, and character model representation parameters. The output of this motion editing network is the position and axis-angle information of the second full set of joints. The position information can be represented by global translation (e.g., global translation of the root node, where the root node could be a hip joint). After obtaining the pose information of the second full set of joints, optimization processing can be performed to obtain optimized pose information. This optimized pose information is then redirected to the original character model, driving the original character model to perform the target action.

[0163] exist Figure 14 In this study, motion editing is performed on an original character model of type SMPLX. The specific process is as follows: the original character model is loaded, and the user performs motion editing operations on it. The user inputs the pose information of the target joints through interactive input, which may include the target joint's position and axis-angle information. During motion editing, the input to the motion editing network, in addition to pose information (including position and axis-angle information), may also include the target joint's identifier, joint type, gender, head orientation, and character model representation parameters. The output of this motion editing network is the position and axis-angle information of the second full set of joints. The position information can be represented by global translation (e.g., global translation of the root node, where the root node could be a hip joint). With the obtained pose information of the second full set of joints, the original character model can be driven to perform the target action.

[0164] The effects achieved by the motion editing methods described above are a significant improvement over those provided by related technologies. See also... Figure 15 As shown, Figure 15 Figure (a) shows the effect of the motion editing method provided by the related technology. Figure 15 Figure (b) is an effect diagram of the action editing method provided in the embodiment of this application. By comparison, it can be found that when the character model needs to perform the action of bending the knee and raising the leg, the foot is parallel to the ground in the effect shown in Figure (a), while the foot hangs down naturally in the effect shown in Figure (b), which better meets the constraints of the reasonable physiological range of the human body and obtains a more accurate edited action.

[0165] In addition, the effect diagrams of the action editing method provided in the embodiments of this application can also be found in... Figure 16 As shown, with Figure 1 Compared to the effect diagrams of the motion editing methods provided by the related technologies shown, when performing a rear leg raise, Figure 1 A clipping issue occurred, with the raised leg appearing to pass through the body, which did not match the actual movement. Figure 16The effect shown avoids clipping issues. Furthermore, the motion editing method provided in this application can analyze the force exerted on the human body based on its physical condition, adjust the posture to achieve a balance between support force and gravity, and if balance cannot be achieved, plan joint rotation according to a dynamic equilibrium state to conform to the dynamic process. For example... Figure 16 In order to maintain balance, the human body adaptively leans forward, and Figure 1 When one leg is raised backward, the other leg remains upright, which does not meet the actual needs of human balance.

[0166] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.

[0167] Based on the character model motion editing method provided in the foregoing embodiments, this application also provides a character model motion editing device 1700. See also... Figure 17 As shown, the character model motion editing device 1700 includes an acquisition unit 1701, a determination unit 1702, and a generation unit 1703:

[0168] The acquisition unit 1701 is used to acquire the joint information and motion information of the original character model;

[0169] The acquisition unit 1701 is further configured to acquire the pose information of the target joint in response to the action editing operation of the target object on the target joint in the original character model. The pose information of the target joint is generated based on the action editing operation. The target joint is a specified joint in the first full set of joints of the original character model.

[0170] The determining unit 1702 is used to perform posture prediction through an action editing network based on the joint information, the motion information and the posture information of the target joint, and output the posture information of the second full joint, wherein the second full joint is the full joint of the target character model that has a corresponding relationship with the original character model.

[0171] The generation unit 1703 is used to generate an original character model that performs the target action based on the posture information of the second full-length joint.

[0172] In one possible implementation, the acquisition unit 1701 is used for:

[0173] Obtain the feasible range of the target joint, and obtain the posture information input by the target object for the target joint;

[0174] The pose information of the target joint is determined based on the feasible range of the joint and the input pose information.

[0175] In one possible implementation, the acquisition unit 1701 is used for:

[0176] If it is determined that the input posture information is within the feasible range of the joint, the input posture information is determined as the posture information of the target joint;

[0177] If it is determined that the input posture information exceeds the feasible range of the joint, the input posture information is adjusted, and the adjusted posture information is determined as the posture information of the target joint. The adjusted posture information is within the feasible range of the joint, or the excess value of the adjusted posture information beyond the feasible range of the joint is less than a preset threshold.

[0178] In one possible implementation, the determining unit 1702 is configured to:

[0179] Based on the joint information and the motion information, the original character model is standardized in posture to obtain an original character model with a standard posture.

[0180] Based on the original character model of the standard pose, obtain the character model representation parameters of the target character model;

[0181] The character model representation parameters and the pose information of the target joint are input into the motion editing network, and the pose prediction is performed through the motion editing network to output the pose information of the second full set of joints.

[0182] The generation unit 1703 is used for:

[0183] The pose information of the second full joint is redirected to the original character model to obtain the original character model that performs the target action.

[0184] In one possible implementation, the determining unit 1702 is configured to:

[0185] The original character model in the standard pose is determined as the target character model, and the character model representation parameter of the target character model is set to 0.

[0186] In one possible implementation, the determining unit 1702 is configured to:

[0187] Align the original character model of the standard pose with the zero-parameter model of the standard parameterized model to obtain the initial scaling factor of the original character model of the standard pose relative to the zero-parameter model.

[0188] The original character model in the standard pose is standardized based on the initial scaling factor to obtain the joint position information of the target character model;

[0189] Based on the joint position information of the target character model, the character model representation parameters are determined.

[0190] In one possible implementation, the apparatus further includes a training unit, the training unit being configured to:

[0191] Obtain the motion capture dataset of the sample standard parameterized model, wherein the motion capture dataset includes the sample pose information of the third full joint of the sample standard parameterized model;

[0192] The sample pose information of sample joints is randomly selected from the motion capture dataset, and the sample joints are a portion of the third full set of joints;

[0193] The sample pose information of the sample joint and the role model representation parameters of the sample standard parameterized model are input into the initial network, and the pose prediction is performed through the initial network to output the predicted pose information of the third full set of joints.

[0194] Based on the difference between the predicted pose information of the third full joint and the sample pose information of the third full joint, the network parameters of the initial network are adjusted to obtain the motion editing network.

[0195] In one possible implementation, the generation unit 1703 is used for:

[0196] The posture information of the second full-length joint is optimized to obtain the optimized posture information;

[0197] Based on the optimized posture information, an original character model is generated to perform the target action.

[0198] In one possible implementation, the generation unit 1703 is used for:

[0199] Calculate the penalty value for at least one penalty item based on the posture information of the second full-length joint;

[0200] The posture information of the second full joint is iteratively optimized based on the penalty value of the at least one penalty item until the iterative optimization stop condition is met, thereby obtaining the optimized posture information.

[0201] In one possible implementation, the acquisition unit 1701 is used for:

[0202] Obtain the source animation data of the original character model;

[0203] Keyframes are determined from the source animation data, and the joint information and motion information of the original character model in the keyframes are obtained.

[0204] As can be seen from the above technical solution, this application realizes the action editing of the character model through user interaction. When it is necessary to edit the action of the original character model, the joint information and motion information of the original character model can be obtained. Furthermore, the target object performs the action editing operation, interacting with the original character model by implementing the action editing operation, thereby specifying the pose information of a small number of joints in the original character model. At this time, in response to the target object's action editing operation on the target joint in the original character model, the pose information of the target joint can be obtained. The pose information of the target joint is generated based on the action editing operation. The target joint is a portion of the specified joints in the first full set of joints of the original character model, and this portion of specified joints is a small number of joints in the first full set of joints. Then, based on the joint information, motion information, and pose information of the target joint, pose prediction is performed through the action editing network, outputting the pose information of the second full set of joints. The second full set of joints is the full set of joints of the target character model that corresponds to the original character model. Thus, the pose information of the full set of joints can be automatically determined based on the pose information of a small number of joints. The pose information of the full set of joints can reflect the target action to be edited, and then the original character model that performs the target action can be generated based on the pose information of the second full set of joints. As can be seen, this application only requires the target object to specify a small number of joint posture information to automatically obtain the posture information of all joints, thereby generating the original character model that performs the corresponding target action and completing the action editing. This method avoids the complex and lengthy manual posture adjustment process, thereby simplifying the action editing process, reducing labor costs, and improving the efficiency of action editing.

[0205] This application also provides a computer device capable of executing a method for editing the actions of a character model. This computer device may be a terminal. Figure 18 This diagram illustrates the structure of a terminal according to an embodiment of this application. Figure 18 In this example, taking a smartphone as the terminal:

[0206] refer to Figure 18 A smartphone includes components such as: a radio frequency (RF) circuit 1810, a memory 1820, an input unit 1830, a display unit 1840, a sensor 1850, an audio circuit 1860, a Wi-Fi module 1870, a processor 1880, and a power supply 1890. The input unit 1830 may include a touch panel 1831 and other input devices 1832, the display unit 1840 may include a display panel 1841, and the audio circuit 1860 may include a speaker 1861 and a microphone 1862. It is understood that... Figure 18The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0207] The memory 1820 can be used to store software programs and modules. The processor 1880 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 1820. The memory 1820 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 1820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0208] The processor 1880 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1820 and by accessing data stored in the memory 1820. Optionally, the processor 1880 may include one or more processing units; preferably, the processor 1880 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1880.

[0209] In this embodiment, the processor 1880 in the smartphone can execute the action editing method for the character model provided in the various embodiments of this application.

[0210] The computer device provided in this application embodiment can also be a server. Please refer to [link / reference]. Figure 19 As shown, Figure 19The diagram illustrates the structure of a server 1900 provided in this embodiment. The server 1900 can vary significantly depending on its configuration or performance. It may include one or more processors, such as a Central Processing Unit (CPU) 1922, and a memory 1932, as well as one or more storage media 1930 (e.g., one or more mass storage devices) for storing application programs 1942 or data 1944. The memory 1932 and storage media 1930 can be temporary or persistent storage. The program stored in the storage media 1930 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1922 may be configured to communicate with the storage media 1930 and execute the series of instruction operations stored in the storage media 1930 on the server 1900.

[0211] Server 1900 may also include one or more power supplies 1926, one or more wired or wireless network interfaces 1950, one or more input / output interfaces 1958, and / or one or more operating systems 1941, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0212] In this embodiment, the central processing unit 1922 in the server 1900 can execute the action editing method for the character model provided in the various embodiments of this application.

[0213] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program for performing the action editing method for the character model described in the foregoing embodiments.

[0214] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.

[0215] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0216] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0217] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0218] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0219] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0220] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a terminal, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0221] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0222] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for editing the motion of a character model, characterized in that, The method includes: Obtain the joint and motion information of the original character model; In response to the target object's action editing operation on the target joint in the original character model, the pose information of the target joint is obtained. The pose information of the target joint is generated based on the action editing operation. The target joint is a specified part of the first full set of joints in the original character model. Based on the joint information, the motion information, and the pose information of the target joint, pose prediction is performed through an action editing network to output the pose information of the second full joint. The second full joint is the full joint of the target character model that corresponds to the original character model. The original character model for performing the target action is generated based on the posture information of the second full-length joint.

2. The method according to claim 1, characterized in that, The step of obtaining the pose information of the target joint includes: Obtain the feasible range of the target joint, and obtain the posture information input by the target object for the target joint; The pose information of the target joint is determined based on the feasible range of the joint and the input pose information.

3. The method according to claim 2, characterized in that, Determining the pose information of the target joint based on the feasible range of the joint and the input pose information includes: If it is determined that the input posture information is within the feasible range of the joint, the input posture information is determined as the posture information of the target joint; If it is determined that the input posture information exceeds the feasible range of the joint, the input posture information is adjusted, and the adjusted posture information is determined as the posture information of the target joint. The adjusted posture information is within the feasible range of the joint, or the excess value of the adjusted posture information beyond the feasible range of the joint is less than a preset threshold.

4. The method according to claim 1, characterized in that, The step of predicting the pose using an action editing network based on the joint information, the motion information, and the pose information of the target joint, and outputting the pose information of the second full-scale joint, includes: Based on the joint information and the motion information, the original character model is standardized in posture to obtain an original character model with a standard posture. Based on the original character model of the standard pose, obtain the character model representation parameters of the target character model; The character model representation parameters and the pose information of the target joint are input into the motion editing network, and the pose prediction is performed through the motion editing network to output the pose information of the second full set of joints. The step of generating the original character model for performing the target action based on the posture information of the second full-length joints includes: The pose information of the second full joint is redirected to the original character model to obtain the original character model that performs the target action.

5. The method according to claim 4, characterized in that, The process of obtaining the character model representation parameters of the target character model based on the original character model of the standard pose includes: The original character model in the standard pose is determined as the target character model, and the character model representation parameter of the target character model is set to 0.

6. The method according to claim 4, characterized in that, The process of obtaining the character model representation parameters of the target character model based on the original character model of the standard pose includes: Align the original character model of the standard pose with the zero-parameter model of the standard parameterized model to obtain the initial scaling factor of the original character model of the standard pose relative to the zero-parameter model. The original character model in the standard pose is standardized based on the initial scaling factor to obtain the joint position information of the target character model; Based on the joint position information of the target character model, the character model representation parameters are determined.

7. The method according to claim 6, characterized in that, The method further includes: Obtain the motion capture dataset of the sample standard parameterized model, wherein the motion capture dataset includes the sample pose information of the third full joint of the sample standard parameterized model; The sample pose information of sample joints is randomly selected from the motion capture dataset, and the sample joints are a portion of the third full set of joints; The sample pose information of the sample joint and the role model representation parameters of the sample standard parameterized model are input into the initial network, and the pose prediction is performed through the initial network to output the predicted pose information of the third full set of joints. Based on the difference between the predicted pose information of the third full joint and the sample pose information of the third full joint, the network parameters of the initial network are adjusted to obtain the motion editing network.

8. The method according to claim 1, characterized in that, The step of generating the original character model for performing the target action based on the posture information of the second full-length joints includes: The posture information of the second full-length joint is optimized to obtain the optimized posture information; Based on the optimized posture information, an original character model is generated to perform the target action.

9. The method according to claim 8, characterized in that, The optimization of the posture information of the second full-length joint to obtain optimized posture information includes: Calculate the penalty value for at least one penalty item based on the posture information of the second full-length joint; The posture information of the second full joint is iteratively optimized based on the penalty value of the at least one penalty item until the iterative optimization stop condition is met, thereby obtaining the optimized posture information.

10. The method according to any one of claims 1-9, characterized in that, The acquisition of joint and motion information of the original character model includes: Obtain the source animation data of the original character model; Keyframes are determined from the source animation data, and the joint information and motion information of the original character model in the keyframes are obtained.

11. A motion editing device for a character model, characterized in that, The device includes an acquisition unit, a determination unit, and a generation unit: The acquisition unit is used to acquire the joint information and motion information of the original character model; The acquisition unit is further configured to acquire the pose information of the target joint in response to the action editing operation of the target object on the target joint in the original character model. The pose information of the target joint is generated based on the action editing operation. The target joint is a specified joint in the first full set of joints of the original character model. The determining unit is used to perform posture prediction through an action editing network based on the joint information, the motion information and the posture information of the target joint, and output the posture information of the second full joint, wherein the second full joint is the full joint of the target character model that corresponds to the original character model. The generation unit is used to generate an original character model that performs the target action based on the posture information of the second full-length joint.

12. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the method according to any one of claims 1-10 according to instructions in the computer program.

13. A computer-readable storage medium for storing a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1-10.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1-10.