Method for generating virtual image, method and device for training deep learning model
By generating the feature sequence of motion trajectory and deep learning models, the motion generation of virtual images is optimized, and the problem of unnatural motion of virtual images is solved, and the generation of virtual images with high simulation is achieved.
Patent Information
- Application Number
- CN202311121138.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-09-01
AI Technical Summary
In the prior art, when generating the motion posture of virtual images, it is difficult to achieve high simulation degree with the real motion posture of humans, resulting in mechanical and unnatural motion distribution of virtual images in the motion conversion.
By generating motion trajectory feature sequences, combining the training methods of deep learning models, using bone point features and attention mechanisms, the generation process of motion feature sequences is optimized to ensure that the movement of virtual images conforms to the real movement laws of the human body.
It improves the naturalness of the movement of the virtual image, reduces the mechanical sense, achieves high simulation with human motion postures, and improves the generation efficiency and effect.
Smart Images

Figure CN117152208B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and in particular to technologies such as computer vision, augmented reality, virtual reality, deep learning, etc. Specifically, it relates to a method for generating a virtual image, a method for training a deep learning model, and an apparatus therefor. Background Art
[0002] The goal of digital virtual humans is to create a digital image close to the human image through computer graphics technology (CG), and endow it with a specific character identity setting, so as to shorten the psychological distance from people visually and bring more real emotional interaction to users.
[0003] With the in-depth application of digital virtual humans in various fields, people's requirements for the similarity between the movement postures of virtual humans and real human movement postures are getting higher and higher. Summary of the Invention
[0004] The present disclosure provides a method for generating a virtual image, a method for training a deep learning model, and an apparatus therefor.
[0005] According to one aspect of the present disclosure, there is provided a method for generating a virtual image, including: generating a motion trajectory feature sequence according to the bone point features of the initial posture of a target object and the bone point features of the target posture of the target object, where the motion trajectory feature sequence includes the bone point features of a plurality of transitional actions required for the target object to transform from the initial posture to the target posture; processing the bone point features of the initial posture, the motion trajectory feature sequence, and the bone point features of the target posture to generate an action feature sequence, where the action feature sequence is used to drive the virtual image to obtain the target posture through continuous posture changes from the initial posture; and rendering the action feature sequence to generate a virtual image.
[0006] According to another aspect of the present disclosure, there is provided a method for training a deep learning model, including: performing local masking on an initial sample action feature sequence to obtain a masked action feature sequence, where the initial sample action feature sequence includes a sample bone point feature sequence of a sample object from a sample initial posture through continuous posture changes to a sample target posture; inputting the masked action feature sequence into an encoding module of an initial model to obtain an encoded sample feature sequence; inputting the encoded sample feature sequence into an attention module of the initial model to obtain a fused feature sequence, where the fused feature sequence includes a predicted bone point feature sequence of the sample object from the sample initial posture through continuous posture changes to the sample target posture; obtaining a target loss value based on a target loss function according to the predicted bone point feature sequence and the sample bone point feature sequence; and adjusting the model parameters of the initial model based on the target loss value to obtain a trained deep learning model.
[0007] According to another aspect of the present disclosure, there is provided an apparatus for generating an avatar, including: a generation module, a processing module, and a rendering module. The generation module is configured to generate a motion trajectory feature sequence according to the bone point features of the initial pose of the target object and the bone point features of the target pose of the target object, where the motion trajectory feature sequence includes the bone point features of a plurality of transition actions that the target object needs to complete when transforming from the initial pose to the target pose. The processing module is configured to process the bone point features of the initial pose, the motion trajectory feature sequence, and the bone point features of the target pose to generate an action feature sequence, where the action feature sequence is used to drive the avatar to obtain the target pose through continuous pose changes from the initial pose. The rendering module is configured to render the action feature sequence to generate an avatar.
[0008] According to another aspect of the present disclosure, there is provided an apparatus for training a deep learning model, including: a masking module, an encoding module, an attention module, a loss calculation module, and an adjustment module. The masking module is configured to perform local masking on the initial sample action feature sequence to obtain a masked action feature sequence, where the initial sample action feature sequence includes a sample bone point feature sequence of a sample object that undergoes continuous pose changes from a sample initial pose to a sample target pose. The encoding module is configured to encode the masked action feature sequence to obtain an encoded sample feature sequence. The attention module is configured to process the encoded sample feature sequence based on an attention mechanism to obtain a fused feature sequence, where the fused feature sequence includes a predicted bone point feature sequence of the sample object that undergoes continuous pose changes from the sample initial pose to the sample target pose. The loss calculation module is configured to obtain a target loss value based on a target loss function according to the predicted bone point feature sequence and the sample bone point feature sequence. The adjustment module is configured to adjust the model parameters of the initial model based on the target loss value to obtain a trained deep learning model.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described above.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method described above.
[0011] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the method described above.
[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings are used to better understand the solution of the present disclosure and do not constitute a limitation to the present disclosure. Among them:
[0014] Figure 1 Schematically shows an exemplary system architecture to which the method and apparatus for generating an avatar according to an embodiment of the present disclosure can be applied;
[0015] Figure 2 Schematically shows a flowchart of the method for generating an avatar according to an embodiment of the present disclosure;
[0016] Figure 3 Schematically shows a schematic diagram of the method for generating an avatar according to an embodiment of the present disclosure;
[0017] Figure 4 Schematically shows a schematic diagram of generating a motion trajectory feature sequence according to an embodiment of the present disclosure;
[0018] Figure 5 Schematically shows a schematic diagram of generating a motion feature sequence according to an embodiment of the present disclosure;
[0019] Figure 6 Schematically shows a flowchart of the method for training a deep learning model according to an embodiment of the present disclosure;
[0020] Figure 7 Schematically shows a block diagram of the apparatus for generating an avatar according to an embodiment of the present disclosure;
[0021] Figure 8 Schematically shows a block diagram of the apparatus for training a deep learning model according to an embodiment of the present disclosure; and
[0022] Figure 9 Schematically shows a block diagram of an electronic device suitable for implementing the method for generating an avatar or the method for training a deep learning model according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] In related technologies, the following several methods are usually adopted to generate an action sequence for driving a virtual image to move:
[0025] 1. Each action is designed to be in a zero state at the start and end. The zero state can be pre-configured. For example, it can be a standard standing posture. The connection between any two actions will return to the zero state, enabling the virtual image to achieve a seamless transition effect when performing continuous actions. However, the virtual image generated in this way returns to the pre-configured posture after any action ends, which does not conform to the natural movement law of the human body.
[0026] 2. A predetermined number of transition actions, such as linear interpolation, are inserted at the positions where transitions are needed between two actions. However, the action distribution of the virtual image generated in this way does not conform to the real movement distribution of the human body, resulting in a mechanical feeling in the actions of the virtual image.
[0027] In view of this, the present disclosure provides a method for generating a virtual image. First, according to the bone point features of the initial posture and the target posture of the target object, a motion trajectory sequence is generated; then, based on the bone point features of the initial posture, the motion trajectory sequence, and the bone point features of the target posture, an action feature sequence capable of representing the continuous change of the posture of the target object is generated, thereby reducing the mechanical feeling of the virtual image changing from the initial posture to the target posture and enabling the completion of actions that conform to the real movement distribution of the human body.
[0028] Figure 1 An exemplary system architecture to which the method and device for generating a virtual image according to embodiments of the present disclosure can be applied is schematically shown.
[0029] It should be noted that Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the method and device for generating a virtual image can be applied may include a terminal device, but the terminal device can implement the method and device for generating a virtual image provided by the embodiments of the present disclosure without interacting with the server.
[0030] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0031] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0032] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.
[0033] Server 105 can be a server providing various services, such as a background management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The background management server can analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0034] It should be noted that the method for generating a virtual image provided in the embodiments of the present disclosure can generally be executed by terminal device 101, 102, or 103. Correspondingly, the device for generating a virtual image provided in the embodiments of the present disclosure can also be set in terminal devices 101, 102, or 103.
[0035] Alternatively, the method for generating a virtual image provided in the embodiments of the present disclosure can generally also be executed by server 105. Correspondingly, the device for generating a virtual image provided in the embodiments of the present disclosure can generally be set in server 105. The method for generating a virtual image provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105. Correspondingly, the device for generating a virtual image provided in the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103, and / or server 105.
[0036] For example, the terminal devices 101, 102, and 103 can obtain the skeletal point features of the initial pose and the skeletal point features of the target pose of the target object, and then send the skeletal point features of the initial pose and the skeletal point features of the target pose to the server 105. The server 105 processes the skeletal point features of the initial pose and the skeletal point features of the target pose to generate a motion trajectory feature sequence. Then, according to the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose, an action feature sequence is generated, and the action feature sequence is rendered to generate a virtual image. Alternatively, a server or a server cluster capable of communicating with the terminal devices 101, 102, 103 and / or the server 105 processes the skeletal point features of the initial pose and the skeletal point features of the target pose, and finally generates a virtual image.
[0037] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0038] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc., of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and it does not violate public order and good customs.
[0039] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.
[0040] Figure 2 Schematically shows a flowchart of a method for generating a virtual image according to an embodiment of the present disclosure.
[0041] As Figure 2 shown, the method includes operations S210 to S230.
[0042] In operation S210, according to the skeletal point features of the initial pose of the target object and the skeletal point features of the target pose of the target object, a motion trajectory feature sequence is generated.
[0043] In operation S220, the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose are processed to generate an action feature sequence.
[0044] In operation S230, the action feature sequence is rendered to generate a virtual image.
[0045] According to an embodiment of the present disclosure, the initial pose of the target object can represent the starting motion pose of the target object at the start time, and the target pose can represent the terminating motion pose of the target object at the end time. For example: the initial pose can be a sitting pose, and the target pose can be a standing pose.
[0046] According to an embodiment of the present disclosure, the bone point features are used to characterize the absolute positions of multiple bone points of a target object and the relative positions between multiple bone points on adjacent joints.
[0047] For example: The bone point features can be obtained based on the human bone tree in the SMPL (Skinned Multi-Person Linear Model). The human bone tree can include 23 joints and 1 root node. The root node can be the center point of the human hip bone. Each joint can include multiple bone points. The absolute position of the bone point can represent the coordinate value of the bone point on each joint, or the coordinate value of the center point of the human hip bone. The relative position between multiple bone points on adjacent joints can be the rotation angle between multiple bone points on adjacent joints. For example: The adjacent joints can be the elbow joint and the wrist joint. Each motion posture of the target object can be represented by the relative positions of the bone points between 23 joints and the absolute position of the root node.
[0048] According to an embodiment of the present disclosure, the motion trajectory feature sequence includes the bone point features of multiple transitional actions required for the target object to transform from the initial posture to the target posture.
[0049] For example: The initial posture can be the standard standing posture, and the target posture can be the sitting posture. The transitional actions required for the target object to change from standing to sitting can include multiple transitional postures such as bending down, bending the knees, and half-squatting. Then the motion trajectory feature sequence can include the bone point features in each transitional posture.
[0050] According to an embodiment of the present disclosure, each transitional action in the motion trajectory feature sequence is isolated, and the distribution of the transitional actions does not fully conform to the human motion distribution law. Therefore, by fusing the bone point features of the initial posture, the motion trajectory feature sequence, and the bone point features of the target posture, the obtained action feature sequence can characterize the continuous posture change process of the target object from the initial posture to the target posture.
[0051] For example: The bone point features of the initial posture, the motion trajectory feature sequence, and the bone point features of the target posture can be input into a trained deep learning model to output an action feature sequence for driving the virtual image to change from the initial posture to the target posture through continuous posture changes.
[0052] According to an embodiment of the present disclosure, a virtual image that can simulate the target object changing from the initial posture to the target posture through continuous posture changes can be generated by rendering the action feature sequence.
[0053] Figure 3 A schematic diagram schematically shows a method for generating a virtual image according to an embodiment of the present disclosure;
[0054] As Figure 3 shown, in Embodiment 300, based on the skeletal point features of the initial pose 301 and the skeletal point features of the target pose 302, and using the method of linear interpolation, the skeletal point features of a motion trajectory sequence 303 including a plurality of transitional motions, the initial pose, and the target pose can be obtained. The motion trajectory sequence 303 is input into the trained Transformer network 304, and after model refinement, an action feature sequence 305 for driving the virtual avatar to continuously change from the initial pose 301 to the target pose 302 is obtained.
[0055] According to an embodiment of the present disclosure, first, a motion trajectory sequence is generated based on the skeletal point features of the initial pose and the skeletal point features of the target pose of the target object; then, based on the skeletal point features of the initial pose, the motion trajectory sequence, and the skeletal point features of the target pose, an action feature sequence capable of characterizing the continuous change of the target object's pose is generated, thereby reducing the mechanical feeling of the virtual avatar changing from the initial pose to the target pose and enabling the completion of actions that conform to the true motion distribution of the human body.
[0056] In the related art, usually, the number of transitional motions from the initial pose to the target pose is determined based on prior experience, and then, based on the action effect of the generated virtual avatar, the number of transitional motions is adjusted. This method requires repeatedly generating the virtual avatar and then reversely adjusting the number of transitional motions, with a large number of repeated operations, resulting in a low efficiency of generating the virtual avatar.
[0057] For example: Generating a motion trajectory feature sequence according to the skeletal point features of the initial pose of the target object and the skeletal point features of the target pose of the target object may include the following operations: obtaining the number of transitional motions based on the skeletal point features of the initial pose and the skeletal point features of the target pose; and based on the number of transitional motions, processing the skeletal point features of the initial pose and the skeletal point features of the target pose to obtain the motion trajectory feature sequence.
[0058] According to an embodiment of the present disclosure, the number of transitional motions can represent the number of skeletal point features that need to be inserted between the skeletal point features of the initial pose and the skeletal point features of the target pose. For example: In the case where the difference between the initial pose and the target pose is large, more transitional motions are required to complete the smooth transition from the initial pose to the target pose. In the case where the difference between the initial pose and the target pose is small, only fewer transitional motions are required to complete the smooth transition from the initial pose to the target pose.
[0059] According to an embodiment of the present disclosure, the motion trajectory feature sequence can represent a sequence of transitional motions. For example: The number of transitional motions is I, where I is an integer greater than 1. In the sequence of transitional motions, the order of the transitional motions can be 1, 2,..., i,..., I in turn.
[0060] For example, for the $i$-th transitional motion, the bone point features of the $i$-th transitional motion can be generated according to Equation (1) based on the bone point features of the initial pose, the number of transitional motions, the order of the transitional motions, and the bone point features of the target pose.
[0061]
[0062] Among them, $f$ a represents the bone point features of the initial pose; $f$ b represents the bone point features of the target pose; $i$ represents the order of the transitional motions; $n$ represents the number of transitional motions.
[0063] In the case where it is determined that $j \lt I$, return to perform the operation of generating the bone point features of the $i$-th transitional motion, and increment $i$; generate the bone point features of the $(i + 1)$-th transitional motion. And so on, until in the case where it is determined that $i = I$, generate the bone point features of the $I$-th transitional motion. And based on the bone point features of the $I$ transitional motions, generate a motion trajectory feature sequence.
[0064] According to an embodiment of the present disclosure, based on the order of the transitional motions, between the initial pose and the target pose, according to the uniform or uniformly accelerated change of the bone point features, generate the bone point features of the transitional motions for characterizing the motion trajectory of the target object, which increases the action pose density between the initial pose and the target pose and reduces the difficulty of motion pose fusion.
[0065] Since the features of the bone points can represent any pose of the target object, therefore, based on the differences between the bone point features, the number of transitional motions can be determined more accurately.
[0066] According to an embodiment of the present disclosure, according to the bone point features of the initial pose, obtain the first positions of multiple bone points of the target object in the state of the initial pose; according to the bone point features of the target pose, obtain the second positions of multiple bone points of the target object in the state of the target pose; and according to the first positions of the multiple bone points and the second positions of the multiple bone points, obtain the number of transitional motions.
[0067] For example: based on the forward kinematics algorithm, the bone point features of the initial pose can be processed according to Equation (2) to obtain the first positions of multiple bone points of the initial pose. The first positions can represent the position coordinates of the bone points on each joint of the target object in the initial pose.
[0068] $p$ a $= FK(f$ a ) (2)
[0069] Among them, $f$ a represents the bone point features of the initial pose; $FK()$ represents the forward kinematics function; $p$a Represents the first position, which can be a 23x3-dimensional vector (23 represents the number of joints in the human bone tree, and 3 represents the rotation angles of the bone points of each joint relative to the adjacent joint along the x, y, and z axes).
[0070] Similarly, based on the forward kinematics algorithm, the bone point features of the target pose can be processed to obtain the second positions of multiple bone points of the target pose. The second position can represent the position coordinates of the bone points on each joint of the target object in the target pose.
[0071] According to an embodiment of the present disclosure, obtaining the number of transitional actions based on the first positions of multiple bone points and the second positions of multiple bone points may include the following operations: according to the identification of the bone points, obtaining the position difference of the target bone point based on the first position of the multiple bone points and the corresponding second positions of the multiple bone points; and obtaining the number of transitional actions according to the position difference of the target bone point and a predetermined transition parameter.
[0072] For example: according to the identification of the bone points, the first position of the bone point on the elbow joint in the initial pose can be compared with the second position of the bone point on the elbow joint in the target pose to obtain the position difference of the bone point on the elbow joint. By analogy, multiple sets of position difference values of the bone points on 23 joints and the root node in different poses can be obtained.
[0073] According to an embodiment of the present disclosure, the position differences of the bone points on 23 joints and the root node between the initial pose and the target pose can be sorted, and the bone point with the largest position difference is determined as the target bone point. For example: during the running of the target object, the position of the bone point on the ankle joint changes from the ground to a position higher than the knee, which may be the bone point with the largest change in position difference among all the bone points on the joints. Therefore, the bone point on the ankle joint can be determined as the target bone point.
[0074] For example: the number of transitional actions can be obtained according to the position difference of the target bone point and a predetermined transition parameter according to Equation (3):
[0075] n = κ·Max(||p a - p b ||2) (3)
[0076] Where n represents the number of transitional actions, p a represents the first position of the bone point in the initial pose, p b represents the second position of the bone point in the target pose; κ represents a predetermined transition parameter.
[0077] According to an embodiment of the present disclosure, the predetermined transition parameter is not unique, and different position difference values can correspond to different predetermined transition parameters. The mapping relationship between the transition parameter and the position difference value can be preconfigured, and the matching predetermined transition parameter is selected based on the position difference value of the target bone point.
[0078] Figure 4 A schematic diagram showing the generation of a motion trajectory feature sequence according to an embodiment of the present disclosure is schematically illustrated.
[0079] As Figure 4 shown, in Embodiment 400, the bone point features 411 of the initial pose can be processed based on the forward kinematics algorithm to obtain the first positions 413 of multiple bone points of the initial pose. The bone point features 412 of the target pose are processed based on the forward kinematics algorithm to obtain the second positions 414 of multiple bone points of the target pose. The first positions 413 of multiple bone points of the initial pose and the corresponding second positions 414 of multiple bone points of the target pose are used to calculate the position differences of the corresponding bone points in sequence, obtaining the position differences 415 of multiple bone points. The maximum value among the position differences 415 of multiple bone points is used as the position difference 416 of the target bone point. Then, the number of transition actions 417 is obtained according to the position difference 416 of the target bone point. Based on the linear interpolation method, according to the bone point features 411 of the initial pose, the bone point features 412 of the target pose, and the number of transition actions 417, an action trajectory feature sequence 418 is generated.
[0080] According to an embodiment of the present disclosure, by determining the matching number of transition actions through the differences in the positions of bone points between different poses, the number of adjustments to the action feature sequence of the virtual image can be effectively reduced, improving the generation efficiency of the virtual image.
[0081] Although the method based on linear interpolation can increase the density of transition actions and reduce the mechanical feeling of the virtual image movement, however, the actual movement change process of the human body is not a completely uniform or uniformly variable process.
[0082] In view of this, the bone point features of the initial pose, the motion trajectory feature sequence, and the bone point features of the target pose can be spliced to generate a to-be-processed bone point feature sequence; and based on the attention mechanism, the to-be-processed bone point feature sequence is processed to obtain an action feature sequence.
[0083] For example: The bone point features of the initial pose, the motion trajectory feature sequence, and the bone point features of the target pose can be spliced in the order of pose change to generate a to-be-processed bone point feature sequence.
[0084] For example, the sequence of skeleton point features to be processed can be encoded to obtain an encoded feature sequence; and based on the cross-attention mechanism, the encoded feature sequence can be processed to obtain an action feature sequence.
[0085] The sequence of skeleton point features to be processed can be encoded and positional encoding can be added to obtain an encoded feature sequence. The encoded feature sequence is input into a multi-head attention layer (Multi-Head Attention). Based on the attention mechanism, it can be used to focus on important information with high weights, ignore unimportant information with low weights, and can exchange information by sharing important information with other information, thereby realizing the transmission of important information, and obtaining an action feature sequence coupled with global information.
[0086] Figure 5 Schematically shows a schematic diagram of generating a motion feature sequence according to an embodiment of the present disclosure.
[0087] As Figure 5 shown, in Embodiment 500, the motion feature sequence is generated through two stages. In the first stage, based on linear interpolation, an action trajectory feature sequence 513 is inserted between segment 1511 and segment 2512 to obtain a sequence of skeleton point features 514 to be processed. In the sequence of skeleton point features 514 to be processed, each transitional action is not continuous. The sequence of skeleton point features 514 to be processed is input into a trained Transformer network 515. In the Transformer network 515, first, the sequence of skeleton point features 514 to be processed is encoded through a temporal convolutional network, and positional encoding (PE: position encoding) is added to obtain an encoded sequence of skeleton point features. The encoded sequence of skeleton point features passes through N multi-head attention layers (Multi-Head Attention), couples global information, and outputs an action feature sequence 516 with continuous pose changes.
[0088] According to an embodiment of the present disclosure, the spliced skeleton point features are fused based on the attention mechanism, so that the skeleton point features can couple the skeleton point feature information of the initial pose, transitional pose, and target pose, so as to realize continuous pose changes that conform to the human motion distribution.
[0089] Figure 6 Schematically shows a flowchart of a training method for a deep learning model according to an embodiment of the present disclosure.
[0090] As Figure 6 shown, the training method 600 may include operations S610 to S650.
[0091] In operation S610, perform local masking on the initial sample action feature sequence to obtain a masked action feature sequence.
[0092] In operation S620, input the masked action feature sequence into the encoding module of the initial model to obtain an encoded sample feature sequence.
[0093] In operation S630, input the encoded sample feature sequence into the attention module of the initial model to obtain a fused feature sequence.
[0094] In operation S640, based on the target loss function, obtain a target loss value according to the predicted skeletal point feature sequence and the sample skeletal point feature sequence.
[0095] In operation S650, based on the target loss value, adjust the model parameters of the initial model to obtain a trained deep learning model.
[0096] According to an embodiment of the present disclosure, the initial sample action feature sequence includes a sample skeletal point feature sequence of a sample object from a sample initial pose through continuous pose changes to a sample target pose. The fused feature sequence includes a predicted skeletal point feature sequence of the sample object from the sample initial pose through continuous pose changes to the sample target pose.
[0097] According to an embodiment of the present disclosure, the definition range of the sample skeletal point features is the same as that of the skeletal point features.
[0098] According to an embodiment of the present disclosure, performing local masking on the initial sample action feature sequence to obtain a masked action feature sequence may be performed as follows: dividing the initial sample action feature sequence to obtain multiple sample action feature sequence segments; and masking a continuous predetermined number of sample action feature sequence segments among the multiple sample action feature sequence segments to obtain a masked action sequence.
[0099] For example: the initial sample action feature sequence may be divided according to a sliding window of a predetermined size to obtain m-frame sample action feature sequence segments, and p segments among the m-frame sample action feature sequence segments may be masked to obtain a masked action sequence.
[0100] According to an embodiment of the present disclosure, input the masked action feature sequence into the encoding module of the initial model to obtain an encoded sample feature sequence. Then input the encoded sample feature sequence into the attention module of the initial model, and based on the cross-attention mechanism, perform feature fusion on the masked action sequence to obtain a fused feature sequence.
[0101] According to an embodiment of the present disclosure, by locally masking a complete continuous pose change sequence, the model can learn implicit rules that conform to human kinematic laws during the training process. Based on the attention mechanism, the global skeletal point feature information is coupled, so that the generated action feature sequence conforms to the human motion distribution.
[0102] According to an embodiment of the present disclosure, the fused feature sequence includes a predicted skeletal point feature sequence of a sample object from a sample initial pose through continuous pose changes to a sample target pose. The predicted skeletal point feature sequence includes a predicted rotation angle feature sequence of the skeletal points between adjacent associated joints and a predicted position feature sequence of the hip center point of the sample object.
[0103] According to an embodiment of the present disclosure, the initial sample action feature sequence includes a sample skeletal point feature sequence of a sample object from a sample initial pose through continuous pose changes to a sample target pose. The sample skeletal point feature sequence includes a sample rotation angle feature sequence of the skeletal points between adjacent associated joints and a sample position feature sequence of the hip center point of the sample object.
[0104] For example: Based on the target loss function, according to the predicted skeletal point feature sequence and the sample skeletal point feature sequence, obtaining a target loss value may include the following operations: Based on the first loss function, according to the sample rotation angle feature sequence and the predicted rotation angle feature sequence, obtaining a rotation angle loss value; based on the second loss function, according to the sample position feature sequence and the predicted position feature sequence, obtaining a center point position loss value; and according to the rotation angle loss value and the center point position loss value, obtaining the target loss value.
[0105] For example: The rotation angle loss value can be calculated using Equation (4):
[0106]
[0107] where, rx ij represents the rotation angle feature of the j-th sample skeletal point in the i-th pose; ry ij represents the predicted rotation angle feature of the j-th skeletal point in the i-th pose; J represents the number of skeletal points; ia represents the initial pose; ib represents the target pose.
[0108] For example: The center point position loss value can be calculated using Equation (5):
[0109]
[0110] where, tx i represents the sample center point position feature of the i-th pose; ty i represents the predicted center point position feature of the i-th pose; ia represents the initial pose; ib represents the target pose.
[0111] According to an embodiment of the present disclosure, since the human body movement posture can be represented by the vector of the bone point features, therefore, based on the loss of the rotation angle of the bone points between adjacent joints and the position loss of the center point, the accuracy of the model prediction can be improved.
[0112] Figure 7 A block diagram of a virtual image generation device according to an embodiment of the present disclosure is schematically shown.
[0113] As Figure 7 shown, the generation device 700 may include a generation module 710, a processing module 720, and a rendering module 730.
[0114] The generation module 710 is configured to generate a motion trajectory feature sequence according to the bone point features of the initial posture of the target object and the bone point features of the target posture of the target object, where the motion trajectory feature sequence includes the bone point features of a plurality of transition actions required for the target object to transform from the initial posture to the target posture.
[0115] The processing module 720 is configured to process the bone point features of the initial posture, the motion trajectory feature sequence, and the bone point features of the target posture to generate an action feature sequence, where the action feature sequence is used to drive the virtual image to obtain the target posture through continuous posture changes from the initial posture.
[0116] The rendering module 730 is configured to render the action feature sequence to generate a virtual image.
[0117] According to an embodiment of the present disclosure, the generation module 710 may include: a first acquisition sub-module and a first processing sub-module.
[0118] The first acquisition sub-module is configured to obtain the number of transition actions according to the bone point features of the initial posture and the bone point features of the target posture.
[0119] The first processing sub-module is configured to process the bone point features of the initial posture and the bone point features of the target posture based on the number of transition actions to obtain a motion trajectory feature sequence.
[0120] According to an embodiment of the present disclosure, the first acquisition sub-module may include: a bone point position acquisition unit for the initial posture, a bone point position acquisition unit for the target posture, and a transition number calculation unit.
[0121] The bone point position acquisition unit for the initial posture is configured to obtain the first positions of a plurality of bone points of the target object in the state of the initial posture according to the bone point features of the initial posture.
[0122] The skeletal point position acquisition unit for the target pose is configured to obtain the second positions of multiple skeletal points of the target object in the state of the target pose according to the skeletal point features of the target pose.
[0123] The transition number calculation unit is configured to obtain the number of transition actions according to the first positions of multiple skeletal points and the second positions of multiple skeletal points.
[0124] According to an embodiment of the present disclosure, the transition number calculation unit may include: a position difference acquisition subunit and a transition number calculation subunit.
[0125] The position difference acquisition subunit is configured to obtain the position difference of the target skeletal point according to the first position of multiple skeletal points and the corresponding second positions of multiple skeletal points according to the identifier of the skeletal point, where the target skeletal point represents the skeletal point with the largest position difference between the first position and the second position.
[0126] The transition number calculation subunit is configured to obtain the number of transition actions according to the position difference of the target skeletal point and a predetermined transition parameter.
[0127] According to an embodiment of the present disclosure, the number of transition actions is I, where I is an integer greater than 1, and the first processing sub-module includes: a first generation unit, a second generation unit, and a third generation unit.
[0128] The first generation unit is configured to generate the skeletal point features of the i-th transition action according to the skeletal point features of the initial pose, the number of transition actions, the order of the transition actions, and the skeletal point features of the target pose for the i-th transition action.
[0129] The second generation unit is configured to, when determining that i is less than I, return to execute the operation of generating the skeletal point features of the i-th transition action and increment i.
[0130] The third generation unit is configured to, when determining that i is equal to I, generate a motion trajectory feature sequence according to the skeletal point features of I transition actions.
[0131] According to an embodiment of the present disclosure, the processing module may include: a splicing sub-module and an attention sub-module.
[0132] The splicing sub-module is configured to splice the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose to generate a skeletal point feature sequence to be processed.
[0133] The attention sub-module is configured to process the skeletal point feature sequence to be processed based on an attention mechanism to obtain an action feature sequence.
[0134] According to an embodiment of the present disclosure, the attention sub-module may include: an encoding unit and an attention unit.
[0135] An encoding unit for encoding a sequence of skeletal point features to be processed to obtain an encoded feature sequence.
[0136] An attention unit for processing the encoded feature sequence based on a cross-attention mechanism to obtain an action feature sequence.
[0137] Figure 8 Schematically shows a block diagram of a training device for a deep learning model according to an embodiment of the present disclosure.
[0138] As Figure 8 shown, the training device 800 may include a masking module 810, an encoding module 820, an attention module 830, a loss calculation module 840, and an adjustment module 850.
[0139] The masking module 810 is configured to perform local masking on an initial sample action feature sequence to obtain a masked action feature sequence, where the initial sample action feature sequence includes a sample skeletal point feature sequence of a sample object from a sample initial pose through continuous pose changes to a sample target pose.
[0140] The encoding module 820 is configured to encode the masked action feature sequence to obtain an encoded sample feature sequence.
[0141] The attention module 830 is configured to process the encoded sample feature sequence based on an attention mechanism to obtain a fused feature sequence, where the fused feature sequence includes a predicted skeletal point feature sequence of a sample object from a sample initial pose through continuous pose changes to a sample target pose.
[0142] The loss calculation module 840 is configured to obtain a target loss value based on a target loss function according to the predicted skeletal point feature sequence and the sample skeletal point feature sequence.
[0143] The adjustment module 850 is configured to adjust the model parameters of the initial model based on the target loss value to obtain a trained deep learning model.
[0144] According to an embodiment of the present disclosure, the sample skeletal point feature sequence includes a sample rotation angle feature sequence of skeletal points between adjacent associated joints and a sample position feature sequence of the hip center point of the sample object; the predicted skeletal point feature sequence includes a predicted rotation angle feature sequence of skeletal points between adjacent associated joints and a predicted position feature sequence of the hip center point of the sample object.
[0145] According to an embodiment of the present disclosure, the loss calculation module may include: a rotation angle loss calculation sub-module, a center point loss calculation sub-module, and a second obtaining sub-module.
[0146] A rotation angle loss calculation sub-module, configured to obtain a rotation angle loss value based on a first loss function according to a sample rotation angle feature sequence and a predicted rotation angle feature sequence.
[0147] A center point loss calculation sub-module, configured to obtain a center point position loss value based on a second loss function according to a sample position feature sequence and a predicted position feature sequence.
[0148] A second obtaining sub-module, configured to obtain a target loss value according to the rotation angle loss value and the center point position loss value.
[0149] According to an embodiment of the present disclosure, the masking module may include: a dividing sub-module and a masking sub-module.
[0150] The dividing sub-module is configured to divide an initial sample action feature sequence to obtain a plurality of sample action feature sequence segments.
[0151] The masking sub-module is configured to mask a continuous predetermined number of sample action feature sequence segments among the plurality of sample action feature sequence segments to obtain a masked action sequence.
[0152] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0153] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0154] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0155] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0156] Figure 9FIG. 0 shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0157] As Figure 9 shown, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0158] A plurality of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0159] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the method for generating an avatar or the method for training a deep learning model. For example, in some embodiments, the method for generating an avatar or the method for training a deep learning model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method for generating an avatar or the method for training a deep learning model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the method for generating an avatar or the method for training a deep learning model in any other suitable way (e.g., by means of firmware).
[0160] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0161] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0163] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0164] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0165] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0166] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0167] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating an avatar, comprising: Generating a motion trajectory feature sequence according to the skeletal point features of the initial pose of the target object and the skeletal point features of the target pose of the target object, wherein the motion trajectory feature sequence includes the skeletal point features of multiple transitional actions required for the target object to transform from the initial pose to the target pose; the number of the transitional actions is determined by the degree of difference between the skeletal point features of the initial pose and the skeletal point features of the target pose, and the skeletal point features of the multiple transitional actions are generated according to the order of the transitional actions and the uniform or uniformly accelerated change of the skeletal point features; so as to reduce the difficulty of motion pose fusion; Using a deep learning model based on an attention mechanism to perform motion pose fusion processing on the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose, and generating an action feature sequence, wherein the action feature sequence is used to drive the avatar to obtain the target pose from the initial pose through continuous pose changes; wherein, the deep learning model is trained by performing local masking on a complete continuous pose change sequence and using a target loss function; wherein, the target loss function includes a first loss function constructed based on a sample rotation angle feature sequence and a predicted rotation angle feature sequence, and a second loss function constructed based on a sample position feature sequence and a predicted position feature sequence; and Rendering the action feature sequence to generate the avatar.
2. The method according to claim 1, wherein, The generating a motion trajectory feature sequence according to the skeletal point features of the initial pose of the target object and the skeletal point features of the target pose of the target object includes: Obtaining the number of the transitional actions according to the skeletal point features of the initial pose and the skeletal point features of the target pose; and Processing the skeletal point features of the initial pose and the skeletal point features of the target pose based on the number of the transitional actions to obtain the motion trajectory feature sequence.
3. The method according to claim 2, wherein The obtaining the number of the transitional actions according to the skeletal point features of the initial pose and the skeletal point features of the target pose includes: Obtaining the first positions of multiple skeletal points of the target object in the state of the initial pose according to the skeletal point features of the initial pose; Obtaining the second positions of multiple skeletal points of the target object in the state of the target pose according to the skeletal point features of the target pose; and Obtaining the number of the transitional actions according to the first positions of the multiple skeletal points and the second positions of the multiple skeletal points.
4. The method according to claim 3, wherein, The obtaining the number of the transitional actions according to the first positions of the multiple skeletal points and the second positions of the multiple skeletal points includes: Obtaining the position difference of a target skeletal point according to the first positions of the multiple skeletal points and the corresponding second positions of the multiple skeletal points according to the identifier of the skeletal point, wherein the target skeletal point represents the skeletal point with the largest position difference between the first position and the second position; and Obtaining the number of the transitional actions according to the position difference of the target skeletal point and a predetermined transition parameter.
5. The method according to claim 2, wherein The number of the transition actions is I, where I is an integer greater than 1. Based on the number of the transition actions, processing the skeletal point features of the initial pose and the skeletal point features of the target pose to obtain the motion trajectory feature sequence, including: For the i-th transition action, generating the skeletal point features of the i-th transition action according to the skeletal point features of the initial pose, the number of the transition actions, the order of the transition actions, and the skeletal point features of the target pose; When it is determined that i is less than I, return to execute the operation of generating the skeletal point features of the i-th transition action and increment i; and When it is determined that i is equal to I, generating the motion trajectory feature sequence according to the skeletal point features of I transition actions.
6. The method according to claim 1, wherein, Based on the attention mechanism, performing motion pose fusion processing on the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose to generate an action feature sequence, including: Concatenating the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose to generate a to-be-processed skeletal point feature sequence; and Based on the attention mechanism, performing motion pose fusion processing on the to-be-processed skeletal point feature sequence to obtain the action feature sequence.
7. The method according to claim 6, wherein, Based on the attention mechanism, performing motion pose fusion processing on the to-be-processed skeletal point feature sequence to obtain the action feature sequence, including: Encoding the to-be-processed skeletal point feature sequence to obtain an encoded feature sequence; and Based on the cross-attention mechanism, processing the encoded feature sequence to obtain the action feature sequence.
8. The method according to claim 1, wherein The skeletal point features are used to represent the absolute positions of multiple skeletal points of the target object and the relative positions between multiple skeletal points on adjacent joints.
9. A training method for a deep learning model, including: Performing local masking on an initial sample action feature sequence to obtain a masked action feature sequence, where the initial sample action feature sequence includes a sample skeletal point feature sequence of a sample object from a sample initial pose through continuous pose changes to a sample target pose; Encoding the masked action feature sequence to obtain an encoded sample feature sequence; Based on the attention mechanism, performing motion pose fusion processing on the encoded sample feature sequence to obtain a fusion feature sequence, where the fusion feature sequence includes a predicted skeletal point feature sequence of the sample object from the sample initial pose through continuous pose changes to the sample target pose; Based on a target loss function, obtaining a target loss value according to the predicted skeletal point feature sequence and the sample skeletal point feature sequence; and Based on the target loss value, adjusting the model parameters of an initial model to obtain a trained deep learning model; wherein, the target loss function includes: a first loss function constructed based on a sample rotation angle feature sequence and a predicted rotation angle feature sequence and a second loss function constructed based on a sample position feature sequence and a predicted position feature sequence.
10. The method according to claim 9, wherein The sample skeletal point feature sequence includes a sample rotation angle feature sequence of skeletal points between adjacent associated joints and a sample position feature sequence of the center point of the hip bone of the sample object; the predicted skeletal point feature sequence includes a predicted rotation angle feature sequence of skeletal points between adjacent associated joints and a predicted position feature sequence of the center point of the hip bone of the sample object; Based on the target loss function, obtaining a target loss value according to the predicted skeletal point feature sequence and the sample skeletal point feature sequence includes: Based on a first loss function, obtaining a rotation angle loss value according to the sample rotation angle feature sequence and the predicted rotation angle feature sequence; Based on a second loss function, obtaining a center point position loss value according to the sample position feature sequence and the predicted position feature sequence; and Obtaining the target loss value according to the rotation angle loss value and the center point position loss value.
11. The method according to claim 9, wherein, Performing local masking on the initial sample action feature sequence to obtain a masked action sequence includes: Dividing the initial sample action feature sequence to obtain a plurality of sample action feature sequence segments; and Masking a continuous predetermined number of sample action feature sequence segments among the plurality of sample action feature sequence segments to obtain the masked action sequence.
12. A virtual image generation device, comprising: A generation module, configured to generate a motion trajectory feature sequence according to the skeletal point features of the initial posture of the target object and the skeletal point features of the target posture of the target object, wherein the motion trajectory feature sequence includes the skeletal point features of a plurality of transition actions required for the target object to transform from the initial posture to the target posture; the number of the transition actions is determined by the degree of difference between the skeletal point features of the initial posture and the skeletal point features of the target posture, and the skeletal point features of the plurality of transition actions are generated according to the order of the transition actions and the uniform or uniformly accelerated change of the skeletal point features; so as to reduce the difficulty of motion pose fusion; A processing module, configured to perform motion pose fusion processing on the skeletal point features of the initial posture, the motion trajectory feature sequence, and the skeletal point features of the target posture by using a deep learning model based on an attention mechanism to generate an action feature sequence, wherein the action feature sequence is used to drive the virtual image to obtain the target posture through continuous posture changes from the initial posture; wherein the deep learning model is trained by performing local masking on a complete continuous posture change sequence and using a target loss function; wherein the target loss function includes a first loss function constructed based on a sample rotation angle feature sequence and a predicted rotation angle feature sequence and a second loss function constructed based on a sample position feature sequence and a predicted position feature sequence; and A rendering module, configured to render the action feature sequence to generate the virtual image.
13. The apparatus according to claim 12, wherein, The generation module includes: A first obtaining sub-module, configured to obtain the number of the transition actions according to the skeletal point features of the initial posture and the skeletal point features of the target posture; and The first processing sub-module is configured to process the skeletal point features of the initial pose and the skeletal point features of the target pose based on the number of the transition actions, so as to obtain the motion trajectory feature sequence.
14. The device according to claim 13, wherein, The first obtaining sub-module includes: The skeletal point position obtaining unit of the initial pose is configured to obtain the first positions of multiple skeletal points of the target object in the state of the initial pose according to the skeletal point features of the initial pose; The skeletal point position obtaining unit of the target pose is configured to obtain the second positions of multiple skeletal points of the target object in the state of the target pose according to the skeletal point features of the target pose; and The transition number calculating unit is configured to obtain the number of the transition actions according to the first positions of the multiple skeletal points and the second positions of the multiple skeletal points.
15. The apparatus according to claim 14, wherein, The transition number calculating unit includes: The position difference obtaining sub-unit is configured to obtain the position difference of the target skeletal point according to the first positions of the multiple skeletal points and the corresponding second positions of the multiple skeletal points according to the identifier of the skeletal point, where the target skeletal point represents the skeletal point with the largest position difference between the first position and the second position; and The transition number calculating sub-unit is configured to obtain the number of the transition actions according to the position difference of the target skeletal point and a predetermined transition parameter.
16. The apparatus according to claim 13, wherein The number of the transition actions is I, where I is an integer greater than 1, and the first processing sub-module includes: The first generating unit is configured to generate the skeletal point features of the i-th transition action according to the skeletal point features of the initial pose, the number of the transition actions, the order of the transition actions, and the skeletal point features of the target pose for the i-th transition action; The second generating unit is configured to return to execute the operation of generating the skeletal point features of the i-th transition action and increment i when it is determined that i is less than I; and The third generating unit is configured to generate the motion trajectory feature sequence according to the skeletal point features of I transition actions when it is determined that i is equal to I.
17. The apparatus according to claim 12, wherein, The processing module includes: The splicing sub-module is configured to splice the skeletal point features of the initial pose, the motion trajectory feature sequence, and the skeletal point features of the target pose to generate a to-be-processed skeletal point feature sequence; and The attention sub-module is configured to perform motion pose fusion processing on the to-be-processed skeletal point feature sequence based on an attention mechanism to obtain the action feature sequence.
18. The apparatus according to claim 17, wherein, The attention sub-module includes: The encoding unit is configured to encode the to-be-processed skeletal point feature sequence to obtain an encoded feature sequence; and The attention unit is configured to process the encoded feature sequence based on a cross-attention mechanism to obtain the action feature sequence.
19. The apparatus according to claim 12, wherein The skeletal point features are used to represent the absolute positions of multiple skeletal points of the target object and the relative positions between multiple skeletal points on adjacent joints.
20. A training device for a deep learning model, including: A masking module for locally masking an initial sample action feature sequence to obtain a masked action feature sequence, where the initial sample action feature sequence includes a sample skeleton point feature sequence of a sample object changing from a sample initial pose through continuous pose changes to a sample target pose; An encoding module for encoding the masked action feature sequence to obtain an encoded sample feature sequence; An attention module for performing motion pose fusion processing on the encoded sample feature sequence based on an attention mechanism to obtain a fused feature sequence, where the fused feature sequence includes a predicted skeleton point feature sequence of the sample object changing from the sample initial pose through continuous pose changes to the sample target pose; A loss calculation module for obtaining a target loss value based on a target loss function according to the predicted skeleton point feature sequence and the sample skeleton point feature sequence; and An adjustment module for adjusting model parameters of an initial model based on the target loss value to obtain a trained deep learning model; where the target loss function includes: a first loss function constructed based on a sample rotation angle feature sequence and a predicted rotation angle feature sequence and a second loss function constructed based on a sample position feature sequence and a predicted position feature sequence.
21. The device according to claim 20, wherein, The sample skeleton point feature sequence includes a sample rotation angle feature sequence of skeleton points between adjacent associated joints and a sample position feature sequence of the hip bone center point of the sample object; the predicted skeleton point feature sequence includes a predicted rotation angle feature sequence of skeleton points between adjacent associated joints and a predicted position feature sequence of the hip bone center point of the sample object; the loss calculation module includes: A rotation angle loss calculation sub-module for obtaining a rotation angle loss value based on the first loss function according to the sample rotation angle feature sequence and the predicted rotation angle feature sequence; A center point loss calculation sub-module for obtaining a center point position loss value based on the second loss function according to the sample position feature sequence and the predicted position feature sequence; and A second obtaining sub-module for obtaining the target loss value according to the rotation angle loss value and the center point position loss value.
22. The apparatus according to claim 20, wherein, The masking module includes: A dividing sub-module for dividing the initial sample action feature sequence to obtain a plurality of sample action feature sequence segments; and A masking sub-module for masking a continuous predetermined number of sample action feature sequence segments among the plurality of sample action feature sequence segments to obtain the masked action sequence.
23. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-11.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
25. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for driving virtual image and computer readable storage medium
CN111145322A
Clothing model driving method and device and storage medium
CN114119908A
Motion generation method and virtual character animation generation method
CN116392812A