Method, device and equipment for automatically binding target 3D object with digital human hand
By matching the geometric features and pose information of the target 3D object, the target gesture and simulated trajectory of the hand skeleton are generated. Combined with collision detection, the precise binding of the digital human hand to the 3D object is achieved, which solves the visual defects in the existing technology and improves the realism of the interaction and the efficiency of the application.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies are insufficient in terms of automation and binding effect in digital human hand binding, resulting in visual defects such as unnatural grip, floating props, and clipping between joints and objects, making it difficult to achieve large-scale application.
By determining the geometric features of the target 3D object, the target gesture required for hand bone gripping is generated, and pose information matching and displacement calculation are performed. Combined with the simulation trajectory of the hand joints and collision detection, the precise binding of the hand bones to the target 3D object is achieved.
It enhances the realism and visual consistency of digital humans interacting with 3D objects, solves the problems of unnatural grip and visual defects, and supports the large-scale deployment of digital humans in various scenarios such as virtual live streaming, film and television animation, and interactive games.
Smart Images

Figure CN121837468A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of digital human processing technology, and in particular to a method, apparatus, and device for automatically binding a target 3D object to a digital human hand. Background Technology
[0002] In fields such as digital content creation and virtual interaction, with the increasingly widespread application of digital human technology, skeletal binding and motion capture technologies have become the core support for enabling digital humans to achieve natural movement. However, existing technologies still have a series of significant shortcomings in practical applications, and their degree of automation and binding effect are insufficient to meet the requirements. Therefore, there is an urgent need for an efficient automatic binding method to improve the realism of interaction and application efficiency. Summary of the Invention
[0003] This disclosure provides a method, apparatus, and device for automatically binding a target 3D object to a digital human hand to solve or alleviate one or more technical problems in the prior art.
[0004] In a first aspect, this disclosure provides a method for automatically binding a target 3D object to a digital human hand, the method comprising: Determine the hand skeleton of the preset digital human, and the target 3D object to be bound to the hand skeleton; Based on the geometric features of the target 3D object, the target gesture required for the hand skeleton to hold the target 3D object is determined; Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones; Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object as it moves to the reference position; Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton; After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, collision detection is performed on each hand joint in the hand skeleton based on the hand closure simulation trajectory to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0005] Secondly, this disclosure provides a device for automatically binding a target 3D object to a digital human hand, comprising: A preprocessing unit is used to determine the hand skeleton of a preset digital human and a target 3D object to be bound to the hand skeleton; based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object; determine the current pose information of the target 3D object and the current reference position of the hand skeleton; based on the current pose information of the target 3D object and the current reference position of the hand skeleton, determine the displacement information of the target 3D object moving to the reference position; and generate a hand closure simulation trajectory for the hand joints in the hand skeleton according to the target gesture. The simulation processing unit is used to perform collision detection on each hand joint in the hand skeleton based on the hand closure simulation trajectory after the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, so as to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0006] Thirdly, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0007] Fourthly, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of the present disclosure.
[0008] Fifthly, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of the present disclosure.
[0009] The beneficial effects of the technical solution provided in this disclosure include at least the following: This disclosed solution first determines the target gesture required for the hand skeleton to hold the target 3D object based on the geometric features of the target 3D object. Then, it determines the pose information of the target 3D object and the reference position of the hand skeleton, calculates the displacement information, and generates the hand joint closure simulation trajectory. Finally, after the target 3D object moves to the reference position of the hand skeleton, it performs joint collision detection to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture. In this way, the automatic and accurate binding of the target 3D object and the digital human hand is achieved. In the above process, the matching of the target gesture driven by geometric features ensures the adaptability of the holding posture and the target 3D object, avoids the deviation of the basic gesture selection, and improves the visual rationality. At the same time, with the help of the precise alignment of pose information and the joint-by-joint collision detection mechanism, it effectively solves the visual defects commonly found in traditional binding, such as unnatural holding, prop floating, and clipping between joints and objects. It significantly improves the realism and visual consistency of the interaction between the digital human and the 3D object, and provides favorable support for the large-scale deployment of digital humans in multiple scenarios such as virtual live streaming, film and television animation, and interactive games.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.
[0012] Figure 1 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 1 ; Figure 2 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 2 ; Figure 3 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 3 ; Figure 4 This is an illustrative diagram of collision detection according to an embodiment of this application. Figure 1 ; Figure 5 This is an illustrative diagram of collision detection according to an embodiment of this application. Figure 2 ; Figure 6 This is an illustrative diagram of collision detection according to an embodiment of this application. Figure 3 ; Figure 7 This is an illustrative diagram of collision detection according to an embodiment of this application. Figure 4 ; Figure 8 An illustrative flow diagram of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 4 ; Figure 9 An illustrative flow diagram of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 5 ; Figure 10 This is a schematic diagram of the structure of an automatic binding device for a target 3D object and a digital human hand according to an embodiment of this application; Figure 11 This is a block diagram of an electronic device used to implement the method for automatically binding a target 3D object to a digital human hand according to embodiments of the present disclosure. Detailed Implementation
[0013] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0014] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0015] Commonly used digital human rigging and motion capture methods can be mainly divided into three categories: manual rigging and skinning using DCC tools, automatic or semi-automatic rigging driven by templates / rules, and traditional object / product rigging. Among them, manual rigging and skinning require multiple processes such as mesh cleaning and repair, manual skeleton building, and manual weighting, which are highly dependent on experienced artists; while template-driven automatic or semi-automatic rigging can reduce some manual operations, it still requires manual quality inspection and repair; traditional object rigging involves setting attachment points on the wrist or palm bones and completing prop alignment with simple scaling and rotation, without involving detailed hand motion simulation. However, these methods generally have many drawbacks. Not only are the rigging and skinning processes time-consuming and labor-intensive, but the quality is also difficult to unify across characters. The movements in high deformation areas are prone to distortion. There are also problems with the disconnect and low merging efficiency of the "rigging / skinning" and "motion capture / motion template" links. At the same time, the details of hand-held objects are poorly represented, and props are prone to floating and clipping. The understanding and alignment of the product is insufficient, the maintenance of spatial relationships is weak, and the technology has poor scalability and reusability, making it difficult to achieve large-scale application.
[0016] Based on this, the present invention provides a method for automatically binding a target 3D object to a digital human hand. It can optimize the coordination between the binding and motion capture links through the hand closure simulation trajectory, thereby achieving finger-by-finger closure simulation and joint-by-joint collision detection. This improves the naturalness and accuracy of prop holding, enhances the hand-holdability and spatial alignment of the product, explicitly models the spatial relationship between the digital human and the prop, reduces appearance repair costs, and thus improves the application efficiency and large-scale implementation capability of digital human technology.
[0017] Specifically, Figure 1 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0018] Furthermore, the method includes at least a portion of the following: such as... Figure 1 As shown, it includes: Step S101: Determine the hand skeleton of the preset digital human and the target 3D object to be bound to the hand skeleton.
[0019] Here, the target 3D object is a digital model of a physical product with standardized 3D model data, which includes at least geometric information such as mesh topology, vertex coordinates, and bounding box parameters. For example, in one example, it could be a 3D model of a handheld product such as a water cup, mobile phone, headphones, or book.
[0020] Here, the pre-defined hand skeleton of the digital human refers to the hand skeletal structure of the digital human that has completed the basic bone binding and skinning steps, and includes at least key hand joints such as the metacarpal bones, phalanges (proximal phalanges, middle phalanges, and distal phalanges), and wrist joint, as well as the hierarchical relationship of the skeletal chain. Furthermore, it may also include the degrees of freedom parameters of each hand joint (such as the range of rotation angles).
[0021] Step S102: Based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object.
[0022] Here, in one example, the geometric features of the target 3D object include, but are not limited to, the bounding box size of the target 3D object (i.e., the bounding box size of the 3D model, such as length, width, and height), the outline of the gripping area (such as the curvature of the handle, the diameter of the bottle), and appearance attributes (such as the shape type of the target 3D object, such as rod-shaped, block-shaped, sheet-shaped, cup-shaped, etc.). This provides data support for subsequently determining the hand-holding method.
[0023] Furthermore, in another example, the target gesture may specifically include four basic hand gestures, such as pinching, gripping, grasping, and supporting. This facilitates the selection of a natural grip posture that matches the attributes of the target 3D object from the basic hand gestures, based on the geometric features of the target 3D object. For example, a "grip" posture is used when holding a water cup, a "pinch" posture when pinching a pen, a "grasp" posture when grasping a book, and a "support" posture when holding a tray, thus ensuring that the pre-set digital human's hand skeleton accurately fits the gripping area of the target 3D object and conforms to visual rationality.
[0024] Step S103: Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones.
[0025] Step S104: Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object moving to the reference position.
[0026] Step S105: Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton.
[0027] It should be noted that the execution order of the steps of determining the displacement information of the target 3D object moving to the reference position (i.e., steps S103 and S104) and the step of generating the hand closure simulation trajectory (i.e., step S105) can be interchanged or performed simultaneously. This disclosure does not restrict the execution order of the two.
[0028] Step S106: After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, collision detection is performed on each hand joint in the hand skeleton based on the hand closure simulation trajectory to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0029] In this way, the disclosed solution can first determine the target gesture required for the hand skeleton to hold the target 3D object based on the geometric features of the target 3D object, then determine the pose information of the target 3D object and the reference position of the hand skeleton and calculate the displacement information, and generate the hand joint closure simulation trajectory. Finally, after the target 3D object moves to the reference position of the hand skeleton, joint collision detection is performed to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture. In this way, the automatic and accurate binding of the target 3D object and the digital human hand is realized. In the above process, the matching of the target gesture driven by geometric features ensures the adaptability of the holding posture and the target 3D object, avoids the deviation of the basic gesture selection, and improves the visual rationality. At the same time, with the help of the joint-by-joint collision detection mechanism, it effectively solves the visual defects such as unnatural holding, prop floating, and clipping between joints and objects that are common in traditional binding, significantly improving the realism and visual consistency of the interaction between digital humans and 3D objects, and providing favorable support for the large-scale deployment of digital humans in multiple scenarios such as virtual live broadcast, film and television animation, and interactive games.
[0030] Furthermore, in a specific example, the current pose information of the target 3D object can be obtained in the following manner; specifically, the determination of the current pose information of the target 3D object (e.g., step S103) can specifically include: Step S103-1-1: Render the target 3D object from multiple angles to identify the target viewpoint required to display the target 3D object.
[0031] For example, in one example, the multi-angle rendering refers to generating multiple rendering perspectives around the center point of the target 3D object within a preset spherical range at uniform angular intervals (such as every 15° interval). Each perspective corresponds to a 2D rendered image of the target 3D object, thereby identifying the target perspective required to display the target 3D object.
[0032] Furthermore, in one example, the target viewpoint can be the frontal viewpoint of the target 3D object. This facilitates the complete display of the core features of the target 3D object (such as the logo, front pattern, main functional areas, etc. of the product). In other words, the visible area of the core features of the target 3D object shown by the target viewpoint is the largest and the degree of occlusion is the lowest.
[0033] Step S103-1-2: Determine the object normal vector of the target 3D object from the target's perspective.
[0034] Here, the object normal vector can represent the current pose information of the target 3D object. For example, in one example, the outer surface normal vector can be determined based on the 2D rendered image corresponding to the identified frontal view. In this case, the outer surface normal vector can be used as the object normal vector of the target 3D object.
[0035] Alternatively, in another example, based on the 2D rendered image corresponding to the identified frontal view, the core feature plane of the target 3D object's front (such as the plane where the product logo is located, or the reference plane of the frontal pattern) can be extracted. The normal vector of this core feature plane can be calculated and determined as the frontal normal vector of the target 3D object, which can be denoted as N. prod At this time, the frontal normal vector N prod This can be used as the object normal vector of the target 3D object. Here, the direction of the obtained frontal normal vector is perpendicular to the core feature plane and points outward from the viewpoint. Its vector coordinates can be directly output through the coordinate system (such as the world coordinate system) of the 3D modeling software. In practical applications, if the front of the target 3D object is an irregular curved surface, the normal vector of the front can be approximated by fitting the tangent plane of the core feature region. Through the above method, the frontal normal vector (i.e., N) is obtained. prod It can accurately represent the current orientation of the target 3D object, and its vector parameters can be directly used as the core data of pose information. In this way, it provides data support for the subsequent hand skeleton to accurately grasp the target 3D object.
[0036] Thus, this disclosed solution provides a refined approach to determining the current pose information of a target 3D object. This approach involves multi-angle rendering of the target 3D object to identify the required target viewpoint, and further determining the object normal vector of the target 3D object under the target viewpoint. This object normal vector can then be used as the current pose information of the target 3D object. In this way, multi-angle rendering ensures that the target viewpoint can fully present the core features of the target 3D object, avoiding pose judgment errors caused by feature occlusion under a single viewpoint, thereby achieving accurate identification of the target 3D object's pose information. Simultaneously, by using the object normal vector to transform the pose information into quantifiable vector parameters, standardized data support is provided for the subsequent accurate alignment of the target 3D object with the hand skeleton reference position, significantly improving the accuracy of the pose information and thus enhancing the reliability and accuracy of the entire binding process.
[0037] Furthermore, in another example, the current reference position of the hand bones can be obtained in the following manner; specifically, determining the current reference position of the hand bones as described above (e.g., step S103) can specifically include: Step S103-2-1: Determine the line connecting the center of the palm and the specified finger joint in the hand bones.
[0038] Here, the center of the palm can specifically refer to the geometric center point of the metacarpal region in the hand skeleton model. Its coordinates can be obtained by extracting the vertex coordinates of different metacarpal bones in the metacarpal region and calculating the average value. It is understood that the present invention does not impose specific restrictions on the selection method of the center of the palm in the hand skeleton, as long as the center position of the palm can be determined.
[0039] It should be noted that when specifying finger joints, distinctive feature joints should be prioritized, such as the proximal interphalangeal joint and distal interphalangeal joint of the index finger, and the metacarpophalangeal joint of the thumb. In practical applications, the specified finger joints can be determined based on the target gesture, thus laying the foundation for ensuring that the hand bones accurately fit the gripping area of the target 3D object and conform to visual rationality.
[0040] Step S103-2-2: Based on the connection line, determine the current reference position of the hand bones.
[0041] For example, in one example, the line connecting the center of the palm to a specified finger joint in the hand bones can be denoted as the hand reference line, denoted as L. hand At this point, the hand reference line can represent the current reference position of the hand skeleton, which is also the baseline reference line for precise docking between the target 3D object and the hand skeleton. Furthermore, in practical applications, the endpoint coordinates (such as the center of the palm, a specified finger joint, etc.) and midpoint coordinates of the hand reference line can be used as key points for hand skeleton positioning. In this way, a precise positional reference is provided for subsequent operations such as precise gripping of the hand skeleton and the target 3D object, hand motion tracking, and virtual interaction.
[0042] Thus, this disclosed solution provides a refined method for determining the current reference position of the hand bones. It can determine the current reference position of the hand bones by identifying the line connecting the center of the palm and a specified finger joint. In this way, by using the feature point connection line as the core benchmark, the reference position of the hand bones can be quickly and accurately determined. This avoids complex bone model fitting operations, reduces computational complexity, and ensures the robustness of the reference position through the stable feature point combination of the palm center and finger joints. This lays the foundation for ensuring that the hand bones fit accurately and visually in accordance with the gripping area of the target 3D object.
[0043] Figure 2 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 2 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.
[0044] Furthermore, the method includes at least a portion of the following: such as... Figure 2 As shown, it includes: Step S201: Determine the hand skeleton of the preset digital human and the target 3D object to be bound to the hand skeleton.
[0045] Step S202: Based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object.
[0046] Step S203: Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones.
[0047] For details regarding pose information and reference positions, please refer to the above statements; they will not be repeated here.
[0048] Step S204: Align the current pose information of the target 3D object with the current reference position of the hand bone in three-dimensional space to obtain an alignment vector representing the alignment effect.
[0049] Here, in one example, the alignment process may refer to a spatial rotation alignment operation based on pose information and a reference position.
[0050] For example, in one instance, the current pose information of the target 3D object can be aligned with the current reference position of the hand bones in three-dimensional space. For instance, this can be achieved by rotating the frontal normal vector (i.e., N) of the target 3D object. prod ) and hand reference line (i.e., L) hand When the angle (denoted as △θ) between the vectors (for example, the direction vector pointing to the palm in the hand reference line) is adjusted to within the preset angle tolerance (denoted as 0±ε), the alignment operation can be considered complete. Furthermore, the aligned vector, that is, the vector that can reflect the alignment effect, can be used as the alignment vector.
[0051] It should be noted that in practical applications, the preset angle tolerance can be flexibly adjusted according to the binding accuracy requirements. For example, a smaller value can be taken in high-precision scenarios, while a larger value can be taken in fast preview scenarios, so as to achieve approximate collinearity between the two directions, thereby facilitating the rapid acquisition of the alignment vector that represents the alignment effect.
[0052] Step S205: If the geometric center and centroid of the target 3D object both meet visual requirements, determine the offset of the target 3D object towards the reference position based on the alignment vector. Here, the offset can be used as the displacement information.
[0053] It should be noted that the geometric center of the target 3D object is the spatial geometric midpoint of the target 3D object, and the centroid is the mass distribution center of the target 3D object.
[0054] For example, in one example, the geometric center and centroid of the target 3D object satisfying the visual requirements means that the relative positions of the geometric center and centroid of the target 3D object must conform to the laws of real physics and visual aesthetic standards. For example, in one example, the relative positional relationship between the geometric center and centroid of the target 3D object must satisfy that the distance between the geometric center and centroid (i.e., the deviation distance) is ≤ a preset threshold. At this time, it can be considered that the visual requirements are met. In this way, visual incongruity is avoided through visual requirements.
[0055] Step S206: Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton.
[0056] Step S207: After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, collision detection is performed on each hand joint in the hand skeleton based on the hand closure simulation trajectory to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0057] In this way, the disclosed solution aligns the pose information of the target 3D object with the reference position of the hand skeleton in three-dimensional space to obtain an alignment vector. Based on this alignment vector, and assuming that both the geometric center and centroid of the target 3D object meet visual requirements, the offset is determined. Thus, through three-dimensional alignment, precise matching between the target 3D object and the hand gripping reference is achieved, ensuring consistency in their posture directions. Simultaneously, by combining the visual requirements of the geometric center and centroid to determine the offset, the fitting accuracy between the target 3D object and the hand skeleton is guaranteed, while adhering to real physical laws and visual requirements. This effectively avoids problems such as center of gravity imbalance and visual inconsistencies, providing a precise positional reference for subsequent hand skeleton gripping simulation and collision detection, further improving the overall accuracy and realism of the automatic binding of the target 3D object and the digital human hand.
[0058] Figure 3 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 3 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 and Figure 2 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.
[0059] Furthermore, the method includes at least a portion of the following: such as... Figure 3 As shown, it includes: Step S301: Determine the hand skeleton of the preset digital human and the target 3D object to be bound to the hand skeleton.
[0060] Step S302: Based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object.
[0061] Step S303: Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones.
[0062] For details regarding pose information and reference positions, please refer to the above statements; they will not be repeated here.
[0063] Step S304: Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object moving to the reference position.
[0064] For details regarding displacement information, please refer to the above statements; they will not be repeated here.
[0065] Step S305: Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton.
[0066] Step S306: After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, each hand joint in the hand skeleton is driven to perform a preset movement based on the hand closed simulation trajectory.
[0067] In one example, the preset motion specifically refers to the continuous movement simulating the natural gripping action of a human hand, such as an opening-closing motion. Specifically, the hand first maintains an initial open posture matching the target gesture (ensuring the target 3D object can smoothly enter the gripping area), and then, along the simulated closing trajectory, drives the metacarpophalangeal joints, proximal interphalangeal joints, distal interphalangeal joints, and other joints to bend sequentially until a preset closing degree is reached. It should be noted that during the simulation, the motion process must adhere to the constraints of human skeletal kinematics, such as ensuring that the joint bending angle does not exceed physiological limits and that the movement speed matches the natural gripping rhythm. This lays the foundation for improving the overall accuracy and realism of the automatic binding between the target 3D object and the digital human hand, and avoiding visual inconsistencies.
[0068] Step S307: During the preset movement of the hand joint, detect whether the joints at each level of the hand joint collide with the collision bodies bound to each level of the joint, so as to perform collision detection for the joints at each level of the hand joint.
[0069] It should be noted that different joints and different levels of joints in the hand can be bound to collision bodies to detect whether a collision occurs during a preset movement.
[0070] For example, in one instance, a distance threshold can be used to determine whether a valid collision has occurred: When the shortest distance between the collider bound to the joint and the surface of the target 3D object is less than or equal to the preset contact threshold (which can be denoted as C), contact For example, C contact When the distance between the joint and the target 3D object is greater than the preset contact threshold (e.g., 1mm, etc.), it is considered a "valid collision". Conversely, when the shortest distance between the collider bound to the joint and the target 3D object surface is greater than the preset contact threshold, it is considered a "no valid collision". In this case, the joint can continue to move along the closed simulation trajectory of the hand.
[0071] Furthermore, after a joint is determined to have "occurred a valid collision," its movement stops. Further, if the joint is a parent joint and has a sub-joint, the parent joint stops moving, while the sub-joint continues to move along the hand's closed simulation trajectory. Then, when the shortest distance between the collider bound to the sub-joint and the surface of the target 3D object is determined to be less than or equal to a preset contact threshold, the sub-joint is considered to have "occurred a valid collision," and its movement stops. This process is repeated until collision detection for each level of joint is completed.
[0072] Step S308: If the collision detection is confirmed to be complete, end the collision detection.
[0073] It should be noted that in one example, collision detection is considered complete after all joints at each level have completed collision detection, at which point the collision detection ends.
[0074] Alternatively, in another example, to further enhance the realism of the automatic binding between the target 3D object and the digital human hand, achieving a "human-like" visual effect, coverage detection can be performed during collision detection. For instance, after driving the joints at various levels to perform preset movements based on the hand's closed simulation trajectory and detecting the collision results, the grip coverage rate (which can be denoted as P) can be determined based on the collision results. cover Furthermore, it can be determined whether the grip coverage is greater than or equal to the preset minimum coverage. For example, if the grip coverage is greater than or equal to the preset minimum coverage (which can be denoted as P0, e.g., P0=60%, etc., and can be set according to the actual scene requirements), collision detection is considered complete and can be terminated. Otherwise, if the grip coverage is less than the preset minimum coverage, the displacement information can be re-determined until the grip coverage is greater than or equal to the preset minimum coverage. In this way, it can ensure that the hand skeleton forms a stable grip on the target 3D object, avoiding the problem of unstable grip caused by only single-point contact. At the same time, it also further enhances the realism of the grip, thereby achieving a "human-like" visual effect.
[0075] For example, in one example, taking the target 3D object as a rectangular Bluetooth earphone case, the target gesture as a "four-finger wrapping + thumb pressing" grip gesture, and the preset movement as an opening-closing movement, collision detection is performed on the joints at various levels in the hand. Specifically, such as Figure 4 As shown, at this time, the hand bones are in the initial opening stage of the preset movement. The Bluetooth earphone box has moved to the reference position where the hand bones are located based on the determined offset, for example, to the center of the palm. At the same time, all finger joints are not close to the Bluetooth earphone box.
[0076] Furthermore, such as Figure 5As shown, based on the simulated trajectory of hand closure, the joints at various levels are driven to perform opening-closing movements. The hand skeleton enters the intermediate closing stage, and some joints in the hand are already in a bent state. The opening-closing movement continues, and after collision detection, the following results are obtained: Figure 6 The collision effect diagram shown depicts the hand skeleton holding the rectangular Bluetooth earphone case in a "four-finger wrapping + thumb pressing" grip posture. In this case, the relative positional relationship between each hand joint (such as various articulation points) and the target 3D object (i.e., the rectangular Bluetooth earphone case) can be recorded. This relative positional relationship adapts to the shape of the earphone case while ensuring a comfortable grip, providing a natural posture basis for subsequent dynamic interaction.
[0077] It should be noted that, because this disclosed solution can detect whether collisions occur between the joints at each level of the hand joint and the colliders bound to those joints through simulation experiments, and stops immediately after a collision occurs at the parent joint while the child joints continue to move, and because coverage detection is introduced during the collision process, it effectively avoids the occurrence of... Figure 7 The problem of clipping between the hand joints and the target 3D object is effectively solved, which improves the realism of the automatic binding between the target 3D object and the digital human hand, and achieves a "human-like" visual effect.
[0078] Step S309: Obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton grasps the target 3D object with the target gesture.
[0079] In this way, the disclosed solution drives the hand joints to perform preset movements based on the hand's closed simulation trajectory, and performs collision detection on the joints at each level during the movement. The collision detection ends after the conditions for completion are determined. Thus, the preset movements driven by the hand's closed simulation trajectory ensure that the hand gripping action conforms to the laws of human movement and avoids the incongruity of the movement posture. At the same time, combined with the real-time collision detection of the joints at each level, it realizes the fine control of the gripping process, effectively prevents the clipping problem between the hand joints and the target 3D object, and provides reliable collision constraint support for the accurate binding of the target 3D object and the digital human hand.
[0080] The following is a detailed scheme for obtaining the target video based on the relative positional relationship between each hand joint in the hand skeleton and the target 3D object, as well as the animation sequence of each bone joint point in the preset digital human.
[0081] Figure 8 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 4 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1 to 7 The relevant content of the method shown in any of the attached figures can also be applied to this example, and the relevant content will not be described again in this example.
[0082] Furthermore, the method includes at least a portion of the following: such as... Figure 8 As shown, it includes: Step S801: Determine the hand skeleton of the preset digital human and the target 3D object to be bound to the hand skeleton.
[0083] Step S802: Based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object.
[0084] Step S803: Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones.
[0085] For details regarding pose information and reference positions, please refer to the above statements; they will not be repeated here.
[0086] Step S804: Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object moving to the reference position.
[0087] For details regarding displacement information, please refer to the above statements; they will not be repeated here.
[0088] Step S805: Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton.
[0089] Step S806: After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, collision detection is performed on each hand joint in the hand skeleton based on the hand closure simulation trajectory to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0090] For relevant explanations regarding collision detection, please refer to the above statements; they will not be repeated here.
[0091] Step S807: Based on the relative positional relationship between each hand joint in the hand skeleton and the target 3D object, and the animation sequence of each bone joint in the preset digital human, perform rendering processing to obtain a target video of limb movement for the preset digital human.
[0092] Here, the hand skeleton of the preset digital human in the target video holds the target 3D object with the target gesture.
[0093] In practical applications, during the rendering process, the hand gripping constraints in the animation sequence of each skeletal joint in the digital human can be maintained. For example, if the original motion of the animation sequence may disrupt the gripping posture, hand constraints are added to the motion layer to ensure that the hand gripping posture remains unchanged.
[0094] Furthermore, before rendering, and in conjunction with the aforementioned animation sequence, anti-interference detection can be performed. For example, through the collision protection mechanism between upper limbs and 3D objects, the spatial overlap between the upper limbs (such as arms and wrists) of the preset digital human and the target 3D object can be detected in real time. If model interference occurs, the wrist angle and the pose of the target 3D object will be automatically adjusted. In this way, the incongruous effect of limbs penetrating the product can be effectively avoided.
[0095] In this way, the disclosed solution can perform rendering processing based on the relative positional relationship between the hand joints and the target 3D object, as well as the animation sequence of each skeletal joint in the preset digital human, to obtain a target video of the hand bones of the preset digital human holding the target 3D object with the target gesture to perform limb movements. In this way, the adaptability and coordination between different movements of the preset digital human are realized, thereby adapting to the application needs of multiple scenarios such as product display and virtual interaction.
[0096] Here, to further enhance the visual effects, the generated target video can be further optimized using a model. Specifically, the method also includes: The video generation model is invoked to process the target video based on the target video and target prompts indicating that the video effect of the target video should be optimized, so as to output an optimized target video.
[0097] Here, target cue words refer to text instructions used to clarify the direction, effect standards, and core requirements of video optimization. Their content can cover multiple dimensions such as improving physical realism, enhancing image quality, and calibrating spatiotemporal consistency. This guides the model to optimize in the indicated dimensions, effectively preventing excessive model divergence. This disclosed solution does not restrict the specific details of setting target cue words.
[0098] For example, in one example, the target prompt can be set to "enhance the physical realism of the video, enhance image quality, and fix rendering defects." After receiving the above instructions, the video generation model optimizes from the core dimensions, such as: optimizing the physical properties of elements such as cloth and hair, calibrating the consistency of lighting and shadows in the scene, thereby optimizing the physical realism of the rendered video; enhancing image quality and improving video accuracy by increasing resolution, optimizing anti-aliasing effects, and expanding dynamic range; automatically identifying and fixing defects such as blurring, glitches, clipping, and material distortion that occur during the rendering process; and finally outputting an optimized video that is physically realistic, has clear image quality, and consistent spatiotemporal effects, while fully preserving the original high-quality body movements and complex camera movements.
[0099] In this way, by calling the video generation model and combining it with target prompts, the target video is processed and an optimized target video is output. Thus, the target prompts accurately clarify the direction of video optimization, enabling the video generation model to achieve targeted optimization, thereby outputting a high-quality target video and significantly improving the visual presentation quality and adaptability of the target video.
[0100] The following presents an animation sequence of each skeletal joint in a preset digital human based on initial video data, and thereby obtains a refinement scheme for the target video.
[0101] Figure 9 This is an illustrative flowchart of a method for automatically binding a target 3D object to a digital human hand according to an embodiment of this application. Figure 5 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1 to 8 The relevant content of the method shown in any of the attached figures can also be applied to this example, and the relevant content will not be described again in this example.
[0102] Furthermore, the method includes at least a portion of the following: such as... Figure 9 As shown, it includes: Step S901: Based on the limb movement video contained in the initial video data (e.g., a video uploaded by a user), determine the skeletal animation sequence of the limb movement of the 3D skeleton corresponding to the target body contained in the initial video data.
[0103] Step S902: Based on the skeletal animation sequence, obtain the animation sequence of each skeletal joint in the preset digital human.
[0104] Step S903: Decouple the hand skeleton of the preset digital human from the animation sequence.
[0105] In other words, in this example, the animation sequence of each skeletal joint in the preset digital human can be obtained based on the initial video data input by the user, and then the hand bones can be decoupled. This lays the foundation for obtaining the target video of the preset digital human that matches the limb movement of the target body in the initial video data and whose hand bones hold the target 3D object with the target gesture.
[0106] Step S904: Determine the hand skeleton of the preset digital human and the target 3D object to be bound to the hand skeleton.
[0107] Step S905: Based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object.
[0108] Step S906: Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones.
[0109] For details regarding pose information and reference positions, please refer to the above statements; they will not be repeated here.
[0110] Step S907: Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object moving to the reference position.
[0111] For details regarding displacement information, please refer to the above statements; they will not be repeated here.
[0112] Step S908: Generate a simulated hand closure trajectory for the hand joints in the hand skeleton based on the target gesture.
[0113] Step S909: After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, collision detection is performed on each hand joint in the hand skeleton based on the hand closure simulation trajectory to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0114] For relevant explanations regarding collision detection, please refer to the above statements; they will not be repeated here.
[0115] Step S910: Based on the relative positional relationship between each hand joint in the hand skeleton and the target 3D object, and the animation sequence of each bone joint in the preset digital human, perform rendering processing to obtain a target video of limb movement for the preset digital human.
[0116] Here, the hand skeleton of the preset digital human in the target video grasps the target 3D object with the target gesture. Furthermore, the limb movements of the preset digital human in the target video match the limb movements of the 3D skeletons in the skeletal animation sequence; in other words, the limb movements of the preset digital human in the target video match the limb movements of the target body in the initial video data.
[0117] In this way, the disclosed solution can determine the skeletal animation sequence of the 3D skeleton corresponding to the target body for limb movement based on the initial video data, and then obtain the animation sequence of each bone joint in the preset digital human based on the skeletal animation sequence. In this way, the limb movement of the preset digital human is matched with the movement posture of the initial video data input by the user, thus meeting the user's customization needs. On the other hand, by decoupling the hand skeleton, it can independently adapt to the gripping and binding requirements of the target 3D object without adjusting the full-body animation sequence of the preset digital human, thereby reducing the complexity of animation adaptation and improving the adaptation efficiency of automatic binding between the preset digital human and the target 3D object. At the same time, it also improves the generation efficiency of 3D digital human videos and the adaptability of application scenarios.
[0118] Further, in a specific example, the skeletal animation sequence of the 3D skeleton corresponding to the target body performing limb movements can be obtained in the following manner; specifically, the above-mentioned determination of the skeletal animation sequence of the 3D skeleton corresponding to the target body performing limb movements based on the limb movement video contained in the initial video data (for example, step S901) can specifically include: Step S901-1: Extract key point information of the target body from each video frame of the initial video data.
[0119] Furthermore, in one example, the key point information of the target body can be obtained in the following manner; specifically, the extraction of key point information of the target body in each video frame of the initial video data (for example, step S901-1) described above can specifically include: inputting each video frame of the initial video data into a two-dimensional pose estimation model to obtain the key point information of the target body in each video frame. For example, if the target body is a human body, the human body key point information of each video frame can be output, such as the coordinate information of the human body key points.
[0120] Here, a two-dimensional pose estimation model refers to a model that can automatically identify and locate key points of a target (e.g., a human body) from an image or video frame, thereby outputting relevant information for each key point (e.g., human body key points), such as coordinate information. For example, in one example, the key point information may include the coordinate information of each skeletal joint of the target body in the video frame.
[0121] Furthermore, in one example, each video frame of the initial video data can be input into the two-dimensional pose estimation model in frame order, and the two-dimensional coordinates (such as (x, y)) corresponding to each key point of the target body in each video frame can be output, that is, the key point information.
[0122] In this way, the disclosed solution inputs each video frame of the initial video data into a two-dimensional pose estimation model to quickly obtain the key point information of the target body in each video frame. This comprehensively captures the details of the target body's limb movements, thereby achieving accurate extraction of key point information. This provides a reliable data foundation for the accurate generation of subsequent 3D skeletal animation sequences, and thus ensures the motion reconstruction quality of controllable digital human videos.
[0123] Step S901-2: Based on the key point information of the target body in each video frame and the temporal information between video frames, generate a skeletal animation sequence of 3D skeleton for limb movement.
[0124] Thus, this disclosed solution provides a refined method for obtaining a skeletal animation sequence of 3D skeletons for limb movement. It can first extract the key point information of the target body in each video frame of the initial video data, and then generate a skeletal animation sequence of 3D skeletons for limb movement based on the key point information and the temporal information between video frames. In this way, on the one hand, by accurately extracting key points frame by frame, a reliable data foundation is provided for the reconstruction of 3D skeleton movement, and the accurate capture of the details of the initial limb movement is achieved; on the other hand, by fusing the temporal information between video frames, the temporal rhythm of 3D skeleton movement and initial limb movement is ensured to be consistent, and the skeletal animation sequence is coherent and natural.
[0125] Furthermore, in a specific example, the skeletal animation sequence of limb movement by a 3D skeleton can be obtained in the following manner; specifically, the above-mentioned generation of the skeletal animation sequence of limb movement by a 3D skeleton based on the key point information of the target body in each video frame and the temporal information between video frames (for example, step S901-2) can specifically include: Step S901-2-1: Input the key point information of the target object in each video frame into the pose regression model to obtain the coordinate information of the 3D skeleton corresponding to the key point information of the target object in each video frame. In this way, the conversion from 2D to 3D is achieved.
[0126] Here, the pose regression model refers to a deep learning model that can directly predict the coordinates of key points in three-dimensional space based on the input key point information through a regression algorithm. In other words, the pose regression model can map the key point information in the input two-dimensional plane to three-dimensional space, and then predict the coordinate information of the 3D skeleton corresponding to the key point information, thus providing favorable support for the generation of subsequent skeletal animation sequences.
[0127] For example, in a specific example, the pose regression model can be a 3D pose regression model. In this case, the two-dimensional coordinate information (such as (x, y)) corresponding to each key point of the target body in the video frame can be input into the 3D pose regression model. Then, the three-dimensional coordinates (such as (x, y, z)) corresponding to each key point of the target body in the video frame can be output, that is, the coordinate information of the 3D skeleton. In this way, the conversion from 2D image to 3D model is realized.
[0128] Step S901-2-2: Based on the coordinate information of the 3D skeleton corresponding to the key point information of the target body in each video frame, and the temporal information between video frames, obtain the skeletal animation sequence of the 3D skeleton performing limb movements.
[0129] Here, the timing information between video frames can be specifically determined by the timestamps of each video frame, thereby ensuring that the motion trajectory of the 3D skeleton is consistent with the limb movement of the target body.
[0130] Furthermore, the skeletal animation sequence refers to a continuous sequence of 3D skeleton posture changes on the timeline. Each frame in the sequence can specifically include the skeletal hierarchy, the spatial position / rotation angle of each joint, and time information, thus providing data support for the generation of the subsequent target video.
[0131] For example, in one example, the coordinate information of the 3D skeleton corresponding to the key point information of the target body in each video frame can be processed by temporal smoothing constraint (e.g., applying a continuity and natural constraint mechanism to the 3D skeleton motion trajectory based on the temporal correlation between video frames), and the 3D skeleton coordinates of consecutive frames can be concatenated in the order of timestamps to generate a skeletal animation sequence of 3D skeletons performing limb movements.
[0132] In this way, the disclosed solution inputs the key point information of the target body in each video frame into the pose regression model to obtain the coordinate information of the corresponding 3D skeleton. Then, it combines the temporal information between video frames to generate a skeletal animation sequence of limb movement of the 3D skeleton. In this way, through the accurate mapping capability of the pose regression model, the conversion from two-dimensional key point information to three-dimensional skeleton coordinate information is realized, ensuring the accuracy of 3D skeleton spatial position restoration. On the other hand, by fusing the temporal information between video frames, the discrete 3D skeleton coordinates are connected in sequence, ensuring that the temporal rhythm of 3D skeleton movement is consistent with the limb movement in the initial video data. At the same time, it achieves a coherent and natural skeletal animation sequence.
[0133] Furthermore, in a specific example, the animation sequence of each skeletal joint in the preset digital human can be obtained in the following manner; specifically, the above-mentioned method of obtaining the animation sequence of each skeletal joint in the preset digital human based on the skeletal animation sequence (for example, step S902) can specifically include: Step S902-1: Eliminate the difference between the bone length in the skeletal animation sequence and the bone length of the preset digital human.
[0134] For example, in one instance, the difference between the bone length in the skeletal animation sequence and the bone length of the preset digital human can be eliminated by a scale normalization algorithm. The core of this algorithm is to establish a unified scale benchmark, thereby ensuring that the original motion posture features are retained after the bone length is scaled, so that the limb movements in the generated target video match the limb movements in the initial video data.
[0135] Step S902-2: Based on the skeletal animation sequence, the binding relationship between bones in the preset digital human, and the weight values of the association between each vertex in the preset digital human and the bones in the preset digital human, perform bone relocation processing on the skeletal joints in the preset digital human after eliminating the bone length difference, to obtain the animation sequence of the skeletal joints in the preset digital human.
[0136] For example, in one example, the above-obtained binding relationships between bones in the preset digital human and the weight values of the association relationships between each vertex in the preset digital human and the bones in the preset digital human (e.g., step S902-2) can specifically include: Step S902-2-1: Perform skeletal binding on the preset digital human to obtain the skeletal information of the preset digital human.
[0137] Here, the skeletal binding (also known as bone binding) refers to the process of establishing a connection between a preset skeletal system and a preset digital human's three-dimensional mesh model, thereby defining the skeletal hierarchy and joint range of motion of the preset digital human, so as to ensure that skeletal movement can drive mesh deformation, laying the foundation for subsequent skeletal retargeting and limb movement.
[0138] Here, in one example, the skeletal information includes the skeletal hierarchy (e.g., the spine is the parent bone, and the shoulder and hip joints are child bones), bone length (i.e., the three-dimensional spatial length of each bone, such as the length of the humerus and the length of the femur), joint type (e.g., rotational joints, gliding joints), and movement limits (e.g., the elbow flexion angle is 0-140°, and the hip rotation angle is -90° to 90°), thereby adapting to complex limb movements.
[0139] For example, in one example, a standard skeletal template can be preset based on human anatomy. An automatic skeletal binding tool can be used to align the standard skeletal template with the geometric center of the preset digital human using a point cloud registration algorithm. Then, a skeletal system can be bound to the preset digital human based on the standard skeletal template, thereby obtaining skeletal information including skeletal hierarchy, skeletal length, joint type, etc.
[0140] Step S902-2-2: Perform skinning processing based on the skeletal information of the preset digital human to obtain the weight values of the relationship between each vertex in the preset digital human and the skeleton in the preset digital human.
[0141] Here, skinning refers to assigning the degree of influence (or proportional weight, or weight value) of different skeletal movements to each vertex (also called mesh vertex) of the preset digital human's mesh model. In other words, the weight value represents the degree of influence of skeletal movements on the vertex (which can also be expressed as a proportion). In this way, the weight value obtained after skinning avoids problems such as clipping, stretching distortion, or stiff movements during skeletal movements, thereby ensuring that the preset digital human's limb movements are natural and realistic.
[0142] It should be noted that, in one example, the weight values range from 0 to 1. This allows for a visual representation of the influence of each bone on the mesh vertex. Furthermore, the sum of the weight values of all bones affecting the same vertex is 1 (ensuring that the vertex's motion state is unique and continuous). In practical applications, the closer the weight value is to 1, the stronger the bone's driving effect on the vertex's motion.
[0143] For example, in one example, based on the obtained skeleton information of the preset digital human, an initial weight is assigned to each vertex of the skeleton of the preset digital human using an automatic skinning tool. Subsequently, the initial weights are verified and optimized using a weight optimization tool, and finally the weight values of the relationship between each vertex of the preset digital human and the skeleton are obtained.
[0144] It should be noted that in practical applications, other skinning methods can be used for bone binding and skinning. This disclosure does not impose specific restrictions on the processing details of bone binding and skinning.
[0145] Thus, this disclosed solution provides a refined scheme for obtaining the animation sequence of each skeletal joint in the preset digital human based on the skeletal animation sequence. It can first eliminate the difference in bone length between the skeletal animation sequence and the preset digital human, and then perform bone relocation processing based on the skeletal binding relationship of the preset digital human and the weight values of the relationship between each vertex and the bone to obtain the animation sequence of the skeletal joint in the preset digital human. In this way, on the one hand, a skeletal system that conforms to the motion logic is constructed for the preset digital human through skeletal binding, realizing the basic support for limb movement; on the other hand, the mesh deformation transition is made natural through skinning processing, without obvious breakage or stiffness, realizing the naturalness and realism of the limb movement of the preset digital human.
[0146] This disclosure also provides a device for automatically binding a target 3D object to a digital human hand, such as... Figure 10 As shown, the device includes: The preprocessing unit 1001 is used to determine the hand skeleton of a preset digital human and a target 3D object to be bound to the hand skeleton; based on the geometric features of the target 3D object, determine the target gesture required for the hand skeleton to hold the target 3D object; determine the current pose information of the target 3D object and the current reference position of the hand skeleton; based on the current pose information of the target 3D object and the current reference position of the hand skeleton, determine the displacement information of the target 3D object moving to the reference position; and generate a hand closure simulation trajectory for the hand joints in the hand skeleton according to the target gesture. The simulation processing unit 1002 is used to perform collision detection on each hand joint in the hand skeleton based on the hand closure simulation trajectory after the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, so as to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
[0147] In a specific example of the disclosed solution, the preprocessing unit is specifically used for: The target 3D object is rendered from multiple angles to identify the target viewpoint required to display the target 3D object; Determine the object normal vector of the target 3D object from the target's perspective; wherein the object normal vector can characterize the current pose information of the target 3D object.
[0148] In a specific example of the disclosed solution, the preprocessing unit is specifically used for: Determine the line connecting the center of the palm to a specified finger joint in the hand bones; Based on the connection line, the current reference position of the hand bones is determined.
[0149] In a specific example of the disclosed solution, the preprocessing unit is specifically used for: The current pose information of the target 3D object is aligned with the current reference position of the hand bone in three-dimensional space to obtain an alignment vector representing the alignment effect. If the geometric center and centroid of the target 3D object both meet the visual requirements, the offset amount of the target 3D object toward the reference position is determined based on the alignment vector.
[0150] In a specific example of the scheme disclosed herein, the simulation processing unit is specifically used for: Based on the simulated hand closure trajectory, each hand joint in the hand skeleton is driven to perform a preset movement; During the preset movement of the hand joint, it is detected whether the joints at each level of the hand joint collide with the collision bodies bound to each level of the joint, so as to perform collision detection for the joints at each level of the hand joint. Once the collision detection is confirmed to be complete, the collision detection process ends.
[0151] In a specific example of the disclosed solution, a video processing unit is also included; wherein the video processing unit is configured to: Based on the relative positional relationship between each hand joint in the hand skeleton and the target 3D object, and the animation sequence of each bone joint in the preset digital human, rendering processing is performed to obtain a target video of limb movement for the preset digital human; wherein, in the target video, the hand skeleton of the preset digital human holds the target 3D object with the target gesture.
[0152] In a specific example of the scheme disclosed herein, the video processing unit is further configured to: Based on the limb movement video contained in the initial video data, determine the skeletal animation sequence of the 3D skeleton corresponding to the target body contained in the initial video data to perform limb movement. Based on the skeletal animation sequence, an animation sequence of each skeletal joint in the preset digital human is obtained; wherein, the hand skeleton of the preset digital human is obtained by decoupling from the animation sequence of each skeletal joint in the preset digital human.
[0153] In a specific example of the disclosed solution, the video processing unit is specifically used for: Extract key point information of the target body from each video frame of the initial video data; Based on the key point information of the target body in each video frame and the temporal information between video frames, a 3D skeleton is generated to perform limb movement skeletal animation sequence.
[0154] In a specific example of the disclosed solution, the video processing unit is specifically used for: Each video frame of the initial video data is input into the two-dimensional pose estimation model to obtain the key point information of the target body in each video frame.
[0155] In a specific example of the disclosed solution, the video processing unit is specifically used for: The key point information of the target object in each video frame is input into the pose regression model to obtain the coordinate information of the 3D skeleton corresponding to the key point information of the target object in each video frame. Based on the coordinate information of the 3D skeleton corresponding to the key point information of the target body in each video frame, and the temporal information between video frames, a skeletal animation sequence for limb movement of the 3D skeleton is obtained.
[0156] In a specific example of the disclosed solution, the video processing unit is specifically used for: Eliminate the difference between the bone length in the skeletal animation sequence and the bone length of the preset digital human; Based on the skeletal animation sequence, the binding relationship between bones in the preset digital human, and the weight values of the association between each vertex in the preset digital human and the bones in the preset digital human, the skeletal joints in the preset digital human after eliminating the bone length difference are subjected to bone relocation processing to obtain the animation sequence of the skeletal joints in the preset digital human.
[0157] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.
[0158] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0159] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 11 As shown, the electronic device includes a memory 1110 and a processor 1120. The memory 1110 stores a computer program that can run on the processor 1120. The number of memories 1110 and processors 1120 can be one or more. The memory 1110 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the method provided in the above-described method embodiments. The electronic device may also include a communication interface 1130 for communicating with external devices and performing data exchange and transmission.
[0160] If the memory 1110, processor 1120, and communication interface 1130 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0161] Optionally, in a specific implementation, if the memory 1110, processor 1120 and communication interface 1130 are integrated on a single chip, the memory 1110, processor 1120 and communication interface 1130 can communicate with each other through an internal interface.
[0162] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0163] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).
[0164] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, it can be non-transient storage media.
[0165] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0166] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0167] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0168] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0169] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A method for automatically binding a target 3D object to a digital human hand, characterized in that, The method includes: Determine the hand skeleton of the preset digital human, and the target 3D object to be bound to the hand skeleton; Based on the geometric features of the target 3D object, the target gesture required for the hand skeleton to hold the target 3D object is determined; Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones; Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object as it moves to the reference position; Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton; After the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, collision detection is performed on each hand joint in the hand skeleton based on the hand closure simulation trajectory to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
2. The method according to claim 1, wherein, Determining the current pose information of the target 3D object includes: The target 3D object is rendered from multiple angles to identify the target viewpoint required to display the target 3D object; Determine the object normal vector of the target 3D object from the target's perspective; wherein the object normal vector can characterize the current pose information of the target 3D object.
3. The method according to claim 2, wherein, Determining the current reference position of the hand bones includes: Determine the line connecting the center of the palm to a specified finger joint in the hand bones; Based on the connection line, the current reference position of the hand bones is determined.
4. The method according to claim 3, wherein, The step of determining the displacement information of the target 3D object to the reference position based on the current pose information of the target 3D object and the current reference position of the hand bones includes: The current pose information of the target 3D object is aligned with the current reference position of the hand bone in three-dimensional space to obtain an alignment vector representing the alignment effect. If the geometric center and centroid of the target 3D object both meet the visual requirements, the offset amount of the target 3D object toward the reference position is determined based on the alignment vector.
5. The method according to any one of claims 1-4, wherein, The collision detection of each hand joint in the hand skeleton based on the simulated hand closure trajectory includes: Based on the simulated hand closure trajectory, each hand joint in the hand skeleton is driven to perform a preset movement; During the preset movement of the hand joint, it is detected whether the joints at each level of the hand joint collide with the collision bodies bound to each level of the joint, so as to perform collision detection for the joints at each level of the hand joint. Once the collision detection is confirmed to be complete, the collision detection process ends.
6. The method according to any one of claims 1-5, further comprising: Based on the relative positional relationship between each hand joint in the hand skeleton and the target 3D object, and the animation sequence of each bone joint in the preset digital human, rendering processing is performed to obtain a target video of limb movement for the preset digital human; wherein, in the target video, the hand skeleton of the preset digital human holds the target 3D object with the target gesture.
7. The method according to claim 6, further comprising: Based on the limb movement video contained in the initial video data, determine the skeletal animation sequence of the 3D skeleton corresponding to the target body contained in the initial video data to perform limb movement. Based on the skeletal animation sequence, an animation sequence of each skeletal joint in the preset digital human is obtained; wherein, the hand skeleton of the preset digital human is obtained by decoupling from the animation sequence of each skeletal joint in the preset digital human.
8. The method according to claim 7, wherein, The process of determining the skeletal animation sequence for limb movement of the target body corresponding to the 3D skeleton contained in the initial video data based on the limb movement video includes: Extract key point information of the target body from each video frame of the initial video data; Based on the key point information of the target body in each video frame and the temporal information between video frames, a 3D skeleton is generated to perform limb movement skeletal animation sequence.
9. The method according to claim 8, wherein, The extraction of key point information of the target body in each video frame of the initial video data includes: Each video frame of the initial video data is input into the two-dimensional pose estimation model to obtain the key point information of the target body in each video frame.
10. The method according to claim 8, wherein, The process of generating a skeletal animation sequence for limb movement using key point information of the target body in each video frame and temporal information between video frames includes: The key point information of the target object in each video frame is input into the pose regression model to obtain the coordinate information of the 3D skeleton corresponding to the key point information of the target object in each video frame. Based on the coordinate information of the 3D skeleton corresponding to the key point information of the target body in each video frame, and the temporal information between video frames, a skeletal animation sequence for limb movement of the 3D skeleton is obtained.
11. The method according to claim 7, wherein, The step of obtaining the animation sequence of each skeletal joint in the preset digital human based on the skeletal animation sequence includes: Eliminate the difference between the bone length in the skeletal animation sequence and the bone length of the preset digital human; Based on the skeletal animation sequence, the binding relationship between bones in the preset digital human, and the weight values of the association between each vertex in the preset digital human and the bones in the preset digital human, the skeletal joints in the preset digital human after eliminating the bone length difference are subjected to bone relocation processing to obtain the animation sequence of the skeletal joints in the preset digital human.
12. A device for automatically binding a target 3D object to a digital human hand, comprising: A preprocessing unit is used to determine the hand skeleton of a preset digital human and the target 3D object to be bound to the hand skeleton; Based on the geometric features of the target 3D object, the target gesture required for the hand skeleton to hold the target 3D object is determined; Determine the current pose information of the target 3D object, and determine the current reference position of the hand bones; Based on the current pose information of the target 3D object and the current reference position of the hand bones, determine the displacement information of the target 3D object as it moves to the reference position; Based on the target gesture, generate a simulated hand closure trajectory for the hand joints in the hand skeleton; The simulation processing unit is used to perform collision detection on each hand joint in the hand skeleton based on the hand closure simulation trajectory after the target 3D object moves to the reference position where the hand skeleton is located based on the displacement information, so as to obtain the relative positional relationship between each hand joint in the hand skeleton and the target 3D object when the hand skeleton holds the target 3D object with the target gesture.
13. The apparatus according to claim 12, wherein, The preprocessing unit is specifically used for: The target 3D object is rendered from multiple angles to identify the target viewpoint required to display the target 3D object; Determine the object normal vector of the target 3D object from the target's perspective; wherein the object normal vector can characterize the current pose information of the target 3D object.
14. The apparatus according to claim 13, wherein, The preprocessing unit is specifically used for: Determine the line connecting the center of the palm to a specified finger joint in the hand bones; Based on the connection line, the current reference position of the hand bones is determined.
15. The apparatus according to any one of claims 12-14, wherein, The simulation processing unit is specifically used for: Based on the simulated hand closure trajectory, each hand joint in the hand skeleton is driven to perform a preset movement; During the preset movement of the hand joint, it is detected whether the joints at each level of the hand joint collide with the collision bodies bound to each level of the joint, so as to perform collision detection for the joints at each level of the hand joint. Once the collision detection is confirmed to be complete, the collision detection process ends.
16. The apparatus according to any one of claims 12-15, further comprising a video processing unit; wherein, The video processing unit is used for: Based on the relative positional relationship between each hand joint in the hand skeleton and the target 3D object, and the animation sequence of each bone joint in the preset digital human, rendering processing is performed to obtain a target video of limb movement for the preset digital human; wherein, in the target video, the hand skeleton of the preset digital human holds the target 3D object with the target gesture.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.