Method and device for generating whole-body posture of human-object interaction

Through the combination method of the dual-branch reciprocity diffusion model, contact predictor and contact guide interaction refiner, a coherent whole-body interaction posture in unknown scenarios is generated, which solves the problem of lack of whole-body coordination and poor generation effect of postures in the prior art.

CN120147513APending Publication Date: 2025-06-13Artificial Intelligence and Robotics Innovation Center of Hong Kong Institute of Innovation, Chinese Academy of Sciences +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161387.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the existing 3D human-object interaction generation methods, the posture lacks whole-body coordination and the generation effect is not good in unknown scenes.

Method used

The dual-branch reciprocity diffusion model is used to generate the initial interactive full-body posture of the human body, and is optimized through the contact predictor and the contact guide interactive refiner to ensure that the posture can also generate a coherent full-body interactive posture in unknown scenarios.

Benefits of technology

It realizes the generation of coherent whole-body interactive postures in unknown scenarios, solving the problem of lack of whole-body coordination and poor generation effect of postures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147513A_ABST
    Figure CN120147513A_ABST
Patent Text Reader

Abstract

The invention provides a human and object interaction whole body posture generation method and device, and the method comprises the steps: generating an initial interaction whole body posture of a human body based on the text data of an interaction action, the point cloud data of a target object, and a double-branch reciprocity diffusion model; inputting the initial interactive whole body posture and the point cloud data of the target object into a contact predictor, and predicting the contact position of the human body and the target object; and inputting the contact position and the initial interaction whole body posture into a contact guide interaction refinement device to obtain a target interaction whole body posture. Therefore, a coherent whole-body interaction posture can be generated in an unknown scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for generating a full-body posture for human-object interaction. Background Art

[0002] Generating realistic 3D human-object interactions from text descriptions is a research direction that has received much attention and has broad application potential in fields such as virtual reality, augmented reality, robotics, and animation. However, due to the lack of large-scale interaction data and the difficulty in ensuring physical rationality, especially in unknown scenarios (out-of-domain scenarios), generating high-quality 3D human-object interaction postures still faces many challenges.

[0003] Existing 3D human-object interaction generation mainly relies on taking text and the point cloud of an object as the input of a neural network, and directly generating the corresponding human-object contact postures by training a neural network model. Among them, the model usually includes two parts. One part is used to generate an initial contact posture, and the other part optimizes the posture generated by the first part to generate the final contact posture. Some generation models will additionally design a spatial attention module and a temporal attention module to improve temporal consistency and spatial consistency.

[0004] Existing 3D human-object interaction generation can be divided into two categories: those focusing on generating human body-object interactions or only on generating hand-object interactions, resulting in a lack of full-body coordination in the generated postures. At the same time, most existing solutions only consider generation within known scenarios, and the generation effect is not satisfactory when migrated to unknown scenarios. Summary of the Invention

[0005] The present invention provides a method and device for generating a full-body posture for human-object interaction, aiming to solve the defects in the prior art that the postures generated by 3D human-object interaction lack full-body coordination and the generation effect is not satisfactory when migrated to unknown scenarios, and to achieve the generation of coherent full-body interaction postures in unknown scenarios.

[0006] The present invention provides a method for generating a full-body posture for human-object interaction, including the following steps: Based on a text dataset of interaction actions, a point cloud dataset of a target object, and a dual-branch reciprocal diffusion model, generate an initial interactive full-body posture of the human body; Input the initial interactive full-body posture and the point cloud dataset of the target object into a contact predictor to predict the contact position between the human body and the target object; Input the contact position and the initial interactive full-body posture into a contact-guided interaction refiner to obtain a target interactive full-body posture.

[0007] A method for generating a full-body posture for human-object interaction provided by the present invention, wherein the text data set of the interaction actions is obtained according to the following method: Based on a large language model, semantic expansion is performed on the original text description of the interaction actions to obtain the text data set of the interaction actions.

[0008] A method for generating a full-body posture for human-object interaction provided by the present invention, wherein the point cloud data set of the target object is obtained according to the following method: Random geometric deformation operations are performed on the original point cloud data of the non-possible contact area of the target object to obtain the point cloud data set of the target object.

[0009] A method for generating a full-body posture for human-object interaction provided by the present invention, wherein generating an initial interactive full-body posture of a human body based on the text data set of the interaction actions, the point cloud data set of the target object, and a dual-branch reciprocal diffusion model includes: Based on the dual-branch structure module of the dual-branch reciprocal diffusion model, independent modeling is respectively performed on the text data set of the interaction actions and the point cloud data set of the target object to obtain the parameters of the interaction actions and the parameters of the target object; Based on the interaction module of the dual-branch reciprocal diffusion model, the parameters of the interaction actions and the parameters of the target object are jointly optimized, and based on the diffusion module of the dual-branch reciprocal diffusion model, diffusion processing is performed on the optimized parameters to obtain an initial interactive full-body posture of a human body.

[0010] A method for generating a full-body posture for human-object interaction provided by the present invention, wherein inputting the initial interactive full-body posture and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object includes: Inputting the initial interactive full-body posture and the point cloud data set of the target object into a contact predictor, and predicting the contact position between the human body and the target object based on the possible contact area of the target object and a pre-determined hand posture.

[0011] A method for generating a full-body posture for human-object interaction provided by the present invention, wherein inputting the contact position and the initial interactive full-body posture into a contact-guided interaction refiner to obtain a target interactive full-body posture includes: Inputting the contact position and the initial interactive full-body posture into a contact-guided interaction refiner, and adjusting the initial interactive full-body posture based on a contact distance loss and a contact normal plane loss to obtain a target interactive full-body posture.

[0012] The present invention also provides a full-body posture generation device for human-object interaction, including the following modules: A generation module, configured to generate an initial interactive full-body pose of a human body based on a text data set of interactive actions, a point cloud data set of a target object, and a two-branch reciprocal diffusion model; A prediction module, configured to input the initial interactive full-body pose and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; A refinement module, configured to input the contact position and the initial interactive full-body pose into a contact-guided interactive refinement device to obtain a target interactive full-body pose.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for generating a full-body pose of human-object interaction as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for generating a full-body pose of human-object interaction as described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for generating a full-body pose of human-object interaction as described in any one of the above is implemented.

[0016] The method and device for generating a full-body pose of human-object interaction provided by the present invention can generate an initial interactive full-body pose of a human body through a two-branch reciprocal diffusion model, can generate coherent full-body interactive actions, provide physical constraints according to a contact predictor, and optimize global rationality through a contact-guided interactive refinement device, so as to realize generating coherent full-body interactive poses in an unknown scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of the method for generating a full-body pose of human-object interaction provided by the present invention.

[0019] Figure 2 It is a schematic flowchart of an embodiment of the generation method provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the effect of an embodiment of the generation method provided by the present invention.

[0021] Figure 4 Schematic structural diagram of the full-body pose generation device for human-object interaction provided by the present invention.

[0022] Figure 5 Schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0024] Figure 1 Schematic flow diagram of the full-body pose generation method for human-object interaction provided by the present invention, as Figure 1 shown, the method includes the following steps: Step 100: Generate an initial interactive full-body pose of a human body based on a text data set of an interaction action, a point cloud data set of a target object, and a double-branch reciprocal diffusion model.

[0025] Step 101: Input the initial interactive full-body pose and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object.

[0026] Step 102: Input the contact position and the initial interactive full-body pose into a contact-guided interaction refiner to obtain a target interactive full-body pose.

[0027] Specifically, the target object in the embodiment of the present invention is an object that a human body needs to perform an interaction action on, which is preset.

[0028] First, the text data set of the interaction action and the point cloud data set of the target object can be input into the double-branch reciprocal diffusion model, so as to obtain the initial interactive full-body pose of the human body. Among them, the text data set of the interaction action may include texts describing action semantics, such as "grab a cup", "carry a box", etc.; the point cloud data set of the target object may include point cloud data representing the 3D shape of the target object, such as point cloud data of a cup, point cloud data of a box, etc.

[0029] The double-branch reciprocal diffusion model can respectively model the features of the human body and the object through two independent branches, and realize the collaborative optimization of the two through the reciprocal mechanism of the interaction module and the diffusion module. Generating the initial interactive full-body pose of the human body through the double-branch reciprocal diffusion model can avoid the problem of limb incoherence caused by local optimization.

[0030] After generating the initial interactive full-body pose of the human body, the initial interactive full-body pose and the point cloud dataset of the target object can be input into the contact predictor. By matching the hand pose in the initial interactive full-body pose with the geometric features of the object surface, the specific contact positions between the human body and the object can be predicted.

[0031] The contact positions can be represented by the coordinates of the contact points. For example, the contact position when the hand grasps a cup is the 3D coordinates of the contact point between the fingertips and the cup handle.

[0032] After obtaining the contact positions, the contact positions and the initial interactive full-body pose can be input into the contact-guided interaction refiner. The contact-guided interaction refiner is used to optimize the interaction pose between the human and the object. By introducing the physical constraints of the contact points, it ensures that the generated pose not only conforms to the semantic intention but also satisfies physical rationality, thereby obtaining a physically reasonable, natural, and smooth target interactive full-body pose.

[0033] The full-body pose generation method for human-object interaction provided by the present invention generates the initial interactive full-body pose of the human body through a dual-branch reciprocal diffusion model, which can generate coherent full-body interaction actions. According to the physical constraints provided by the contact predictor and the contact-guided interaction refiner to optimize the global rationality, it is thus possible to generate coherent full-body interaction poses even in unknown scenarios.

[0034] According to a full-body pose generation method for human-object interaction provided by the present invention, the text dataset of the interaction actions is obtained in the following manner: Based on a large language model, semantic expansion is performed on the original text description of the interaction actions to obtain the text dataset of the interaction actions.

[0035] Specifically, in the embodiments of the present invention, the text dataset of the interaction actions can be constructed in the following manner: First, collect the original text descriptions of the interaction actions (such as "grasp a cup", "carry a box"), and then use a large language model (LLM) (such as GPT-4o) to perform semantic expansion on these texts.

[0036] The LLM can generate various synonymous or near-synonymous expressions based on the context and action intention of the original description. For example, expand "grasp a cup" to "hold the cup handle", "pick up the water cup", "take the teacup", etc., while maintaining the consistency of the core semantics.

[0037] In addition, the model can further enrich the text diversity by rewriting the sentence pattern (such as changing "the box is lifted" to "lift the box") or introducing more complex action descriptions (such as "hold a heavy object with both hands" instead of "carry").

[0038] In some embodiments, to ensure the accuracy of the generated results, semantic similarity calculation (such as cosine similarity) and manual screening can be used to eliminate the extended text that deviates from the original intention.

[0039] Finally, these extended text descriptions can jointly constitute a text dataset for interaction actions, covering a wider range of semantic expression scenarios, providing diverse and semantically accurate input conditions for the subsequent pose generation model, thereby enhancing the model's understanding and generalization ability of open-domain text instructions.

[0040] According to a full-body pose generation method for human-object interaction provided by the present invention, the point cloud dataset of the target object is obtained in the following manner: Perform random geometric deformation operations on the original point cloud data of the non-possible contact area of the target object to obtain the point cloud dataset of the target object.

[0041] Specifically, in the embodiments of the present invention, the point cloud dataset of the target object can be constructed in the following manner: First, collect the original point cloud data of the target object, and then perform random geometric deformation operations on the non-possible contact positions of the target object. The non-possible contact positions of the target object are the areas that do not affect the contact function of the object. For example, if the target object is a cup, the cup mouth, cup bottom, inner wall of the cup, etc. of the cup are the non-possible contact positions of the target object.

[0042] Among them, the random geometric deformation operations can include operations such as rotation and stretching, so as to increase the diversity of the object while ensuring that the key contact areas remain unchanged.

[0043] Finally, the point cloud data obtained after these random geometric deformation operations can jointly constitute the point cloud dataset of the target object. These random geometric deformation operations help the model better learn the potential interaction relationship between humans and objects and enhance its adaptability to various forms.

[0044] According to a full-body pose generation method for human-object interaction provided by the present invention, based on the text dataset of interaction actions, the point cloud dataset of the target object, and the dual-branch reciprocal diffusion model, an initial interactive full-body pose of the human body is generated, including: Based on the dual-branch structure module of the dual-branch reciprocal diffusion model, independently model the text dataset of interaction actions and the point cloud dataset of the target object respectively to obtain the parameters of the interaction actions and the parameters of the target object; Based on the interaction module of the dual-branch reciprocal diffusion model, co-optimize the parameters of the interaction actions and the parameters of the target object, and based on the diffusion module of the dual-branch reciprocal diffusion model, perform diffusion processing on the optimized parameters to obtain the initial interactive full-body pose of the human body.

[0045] Specifically, an initial interaction pose is generated using a dual-branch reciprocal diffusion model, which can ensure the matching relationship between human body movements and the shape and position of the target object.

[0046] First, the text data set of the interaction action and the point cloud data set of the target object can be respectively input into the dual-branch reciprocal diffusion model.

[0047] In the dual-branch structure of the dual-branch reciprocal diffusion model, the pose parameters of the human and the object can be independently modeled, enabling the model to capture the characteristics of human body movements and the geometric and spatial features of the object respectively, thus ensuring the rationality and accuracy of the interaction process.

[0048] In this process, the modeling of the human body pose can rely on joint angles, bone structures, and movement trajectories, and the modeling of the target object can be based on the point cloud data to obtain its geometric shape, size, and possible interaction areas.

[0049] After the independent modeling is completed, the parameters of the interaction action and the parameters of the target object are obtained.

[0050] Then, according to the interaction module of the dual-branch reciprocal diffusion model, the parameters of the interaction action and the parameters of the target object can be co-optimized. The role of this module is to learn the potential interaction information between the human and the object, ensuring that the generated pose not only conforms to the laws of human movement but also maintains a reasonable contact relationship with the target object. By aligning and fusing the parameters of the interaction action and the parameters of the target object through the interaction module, and then performing diffusion processing on the optimized parameters according to the diffusion module, the human body pose can be adjusted according to the shape and position of the target object during generation, avoiding unreasonable interaction actions, such as the body penetrating the object or the contact points not matching.

[0051] Finally, based on the results of the diffusion module, a preliminary interaction pose is generated, which has preliminary interaction characteristics and provides a basis for subsequent contact position prediction, making the entire interaction process more in line with physical laws and actual application requirements.

[0052] According to a method for generating a full-body pose for human-object interaction provided by the present invention, the initial interaction full-body pose and the point cloud data set of the target object are input into a contact predictor to predict the contact position between the human body and the target object, including: Input the initial interaction full-body pose and the point cloud data set of the target object into the contact predictor, and predict the contact position between the human body and the target object based on the possible contact areas of the target object and the predetermined hand pose.

[0053] Specifically, using the contact predictor to determine the specific contact position between the human body and the target object can further optimize the interaction pose and improve the accuracy of the generation result.

[0054] First, input the initial interactive full-body pose generated by the double-branch reciprocal diffusion model and the point cloud dataset of the target object into the contact predictor, enabling it to comprehensively consider the human pose and the object's geometric shape, thereby predicting a reasonable contact area.

[0055] The main bases for the contact predictor to predict a reasonable contact area may include the possible contact area information of the target object and the pre-determined hand pose.

[0056] Among them, the possible contact area of the target object can be obtained by analyzing the object's geometric structure, surface features, and common human-computer interaction patterns. For example, if the target object is a chair, the possible contact areas of the target object are the seat surface, the backrest, etc.; if the target object is a cup, it can be determined that the possible contact areas of the target object are the cup handle, the cup wall, etc.

[0057] In addition, the hand pose of the human body plays a dominant role during the interaction process, and the manipulation of many objects depends on specific grasping or pressing actions. Therefore, during contact prediction, the contact predictor can, based on the possible contact area of the target object, combined with the pre-determined hand pose, predict the contact position between the human body and the target object.

[0058] For example, in the scenario of grasping an object, the contact predictor may analyze the distribution of the contact points between the fingers and the object surface to ensure the stability of the grasp, while in interactions such as pushing, pulling, or supporting, the system may pay more attention to the contact area of the palm to ensure that the interaction conforms to ergonomics and physical laws.

[0059] It can be understood that, in order to improve the prediction accuracy, the contact predictor can combine a deep learning model to identify the most likely contact patterns by learning a large amount of human-computer interaction data.

[0060] According to a method for generating a full-body pose for human-object interaction provided by the present invention, input the contact position and the initial interactive full-body pose into the contact-guided interaction refiner to obtain the target interactive full-body pose, including: Input the contact position and the initial interactive full-body pose into the contact-guided interaction refiner, and based on the contact distance loss and the contact normal plane loss, adjust the initial interactive full-body pose to obtain the target interactive full-body pose.

[0061] Specifically, according to the contact position obtained by the contact predictor and the initial interactive full-body pose generated by the double-branch reciprocal diffusion model, further optimize and refine the interactive pose to make the final target interactive pose more in line with physical logic and natural interaction laws.

[0062] First, the contact position predicted by the contact predictor and the initial interaction pose generated by the dual-branch reciprocal diffusion model are input into the contact-guided interaction refiner, enabling it to adjust and optimize the pose based on the contact information.

[0063] During this process, the contact-guided interaction refiner can adopt two main loss constraints: contact distance loss and contact normal plane loss, to ensure that the human pose not only correctly contacts the object, but also the interaction method conforms to real physical laws.

[0064] The role of the contact distance loss is to constrain the spatial position of the contact point, to ensure that the human hand or other interaction parts can accurately fit into the contact area of the target object. This loss function ensures that there are no non-physical hanging contacts or penetration problems during the interaction process by minimizing the distance between the predicted contact point and the target contact area. For example, in the scenario of grasping an object, the contact distance loss will prompt the finger joints to adjust to a reasonable grasping position, so that the fingers can closely fit the object surface, rather than floating in the air or having an unreasonable interpenetration with the object.

[0065] The contact normal plane loss further constrains the directionality of the contact point, to ensure that the interaction method between the human body and the object conforms to physical common sense in the contact direction. For example, in the case of the palm pressing the object surface, the contact direction of the palm should be close to the normal direction of the object surface to ensure reasonable force; when grasping a cylindrical object, the fingers should be distributed around the curved surface of the cylinder, rather than contacting the object in a non-ergonomic way. Through the optimization of this loss function, the system can adjust the rotation angle of the human joints to make the interaction method more natural and avoid unrealistic human poses.

[0066] During the refinement process, the contact-guided interaction refiner will comprehensively consider these two loss constraints and continuously adjust the human pose through iterative optimization until it finally reaches the optimal state. This optimization process not only corrects the subtle deviations in the initial interaction pose, but also further enhances the accuracy of the contact area between the human body and the object, making the finally generated target interaction pose more in line with physical logic in terms of spatial position, contact area, and contact direction.

[0067] The following further illustrates the full-body pose generation method for human-object interaction provided by the present invention through embodiments in specific application scenarios.

[0068] Figure 2 It is a schematic flowchart of the embodiment of the generation method provided by the present invention, as Figure 2As shown in the figure, the framework of this embodiment consists of three key components: a dual-branch reciprocal diffusion model, a contact-guided interaction refiner, and dynamic adaptation. The dual-branch reciprocal diffusion model generates combined whole-body interactions based on text descriptions and object point clouds. Subsequently, the contact-guided interaction refiner uses the predicted contact areas as guidance to adjust the interaction postures. This refiner also allows for additional inference-time guidance through the diffusion process to enhance the generated postures. To improve the model's generalization ability for previously unseen objects and various text descriptions, dynamic adaptation is added in this embodiment. It includes semantic adjustment and geometric deformation modules to enable the method of this embodiment to generalize and thus generate more robust and adaptable interactions.

[0069] The detailed steps of this embodiment will be given below: Step S0, in the data preprocessing stage, first use an LLM model (such as GPT-4o) to generate multiple synonymous descriptions of the original text to enhance the language diversity of the data and provide a richer semantic expression for the model. At the same time, for the non-possible contact positions of the object (i.e., the areas that do not affect the contact function of the object), perform random geometric deformation operations such as rotation and stretching to ensure that the key contact areas are not changed while increasing the diversity of the object. These deformation operations help the model better learn the potential interaction relationships between humans and objects and enhance its adaptability to various forms.

[0070] Step S1, the core of this stage is to generate an initial interaction posture through the dual-branch reciprocal diffusion model. First, input the data generated in step S0 into the model, and independently model the posture parameters of humans and objects in the dual-branch structure to accurately capture their respective characteristics. Subsequently, by introducing an interaction module, learn the potential interaction information between humans and objects and generate a preliminary human-object posture. The goal of this step is to obtain a posture with preliminary interaction characteristics to provide a basis for subsequent contact position prediction.

[0071] Step S2, in this step, input the initial posture generated in step S1 and the point cloud data of the object into the contact predictor. By using the known object contact information (such as the surface contact area) and hand postures, predict the specific contact positions of the object. The focus of this step is to clarify the contact points and contact modes between humans and objects to provide precise constraint conditions for further optimizing the interaction postures.

[0072] Step S3: Input the contact position obtained in Step S2 and the initial pose generated in Step S1 into the contact-guided interaction refiner for optimization. In the refiner, the interaction result is adjusted by the contact distance loss (constraining the distance of the contact point) and the contact normal plane loss (ensuring a reasonable contact direction and normal direction). The model further optimizes the poses of the human and the object in this step to make them more physically logical, while refining the accuracy and naturalness of the contact area, thus obtaining the final high-quality interaction result.

[0073] Figure 3 Schematic diagram of the effect of the embodiment of the generation method provided by the present invention, as Figure 3 shown, this embodiment shows the results of out-of-domain text-based and object-based tests, demonstrating the advantages of the method of this embodiment over existing alternatives. For the out-of-domain text test, this embodiment evaluates the handling of the unseen terms "deliver" and "elevate" by the model of this embodiment. The method of this embodiment successfully interprets and generates different poses for these terms. In the out-of-domain object test, the model of this embodiment performs well, generating reasonable grasping and action trends for the physical properties of each object and the corresponding actions, which well match the intention of the provided text. These results highlight the robustness of the method of this embodiment in out-of-domain text-based and object-based scenarios, always outperforming existing methods in accurate, context-sensitive motion and controlled object interaction.

[0074] Next, a full-body pose generation device for human-object interaction provided by the present invention will be described. The full-body pose generation device for human-object interaction described below can be mutually referred to corresponding to the full-body pose generation method for human-object interaction described above.

[0075] Figure 4 Schematic diagram of the structure of the full-body pose generation device for human-object interaction provided by the present invention, as Figure 4 shown, the device includes the following modules: A generation module 400, configured to generate an initial interactive full-body pose of a human body based on a text data set of interaction actions, a point cloud data set of a target object, and a double-branch reciprocal diffusion model; A prediction module 410, configured to input the initial interactive full-body pose and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; A refinement module 420, configured to input the contact position and the initial interactive full-body pose into a contact-guided interaction refiner to obtain a target interactive full-body pose.

[0076] According to a full-body pose generation device for human-object interaction provided by the present invention, the text data set of interaction actions is obtained in the following manner: Based on a large language model, semantic expansion is performed on the original text description of the interaction action to obtain a text dataset of the interaction action.

[0077] According to a full-body pose generation device for human-object interaction provided by the present invention, the point cloud dataset of the target object is obtained in the following manner: Random geometric deformation operations are performed on the original point cloud data of the non-possible contact area of the target object to obtain the point cloud dataset of the target object.

[0078] According to a full-body pose generation device for human-object interaction provided by the present invention, based on the text dataset of the interaction action, the point cloud dataset of the target object, and a double-branch reciprocal diffusion model, an initial interactive full-body pose of the human body is generated, including: Based on the double-branch structure module of the double-branch reciprocal diffusion model, independent modeling is respectively performed on the text dataset of the interaction action and the point cloud dataset of the target object to obtain the parameters of the interaction action and the parameters of the target object; Based on the interaction module of the double-branch reciprocal diffusion model, the parameters of the interaction action and the parameters of the target object are jointly optimized, and based on the diffusion module of the double-branch reciprocal diffusion model, diffusion processing is performed on the optimized parameters to obtain the initial interactive full-body pose of the human body.

[0079] According to a full-body pose generation device for human-object interaction provided by the present invention, the initial interactive full-body pose and the point cloud dataset of the target object are input into a contact predictor to predict the contact position between the human body and the target object, including: The initial interactive full-body pose and the point cloud dataset of the target object are input into the contact predictor, and based on the possible contact area of the target object and the pre-determined hand pose, the contact position between the human body and the target object is predicted.

[0080] According to a full-body pose generation device for human-object interaction provided by the present invention, the contact position and the initial interactive full-body pose are input into a contact-guided interaction refiner to obtain the target interactive full-body pose, including: The contact position and the initial interactive full-body pose are input into the contact-guided interaction refiner, and based on the contact distance loss and the contact normal plane loss, the initial interactive full-body pose is adjusted to obtain the target interactive full-body pose.

[0081] Figure 5 Schematic diagram of the structure of the electronic device provided by the present invention, as Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute the full-body pose generation method for human-object interaction provided by each of the above methods. The method includes: Generating an initial interactive full-body pose of the human body based on a text data set of interactive actions, a point cloud data set of the target object, and a dual-branch reciprocal diffusion model; Inputting the initial interactive full-body pose and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; Inputting the contact position and the initial interactive full-body pose into a contact-guided interaction refiner to obtain a target interactive full-body pose.

[0082] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0083] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the full-body pose generation method for human-object interaction provided by each of the above methods. The method includes: Generating an initial interactive full-body pose of the human body based on a text data set of interactive actions, a point cloud data set of the target object, and a dual-branch reciprocal diffusion model; Inputting the initial interactive full-body pose and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; Inputting the contact position and the initial interactive full-body pose into a contact-guided interaction refiner to obtain a target interactive full-body pose.

[0084] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a full-body pose generation method for human-object interaction provided by the above-mentioned various methods. The method includes: Generating an initial interactive full-body pose of a human body based on a text data set of interactive actions, a point cloud data set of a target object, and a double-branch reciprocal diffusion model; Inputting the initial interactive full-body pose and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; Inputting the contact position and the initial interactive full-body pose into a contact-guided interaction refiner to obtain a target interactive full-body pose.

[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating full-body postures for human-object interaction, characterized in that: include: Generate the initial interactive full-body posture of the human body based on the text dataset of interactive actions, the point cloud dataset of the target object and the two-branch reciprocal diffusion model; Inputting the initial interactive full-body posture and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; The contact position and the initial interaction full body posture are input into a contact-guided interaction refiner to obtain a target interaction full body posture.

2. The method for generating whole-body postures for human-object interaction according to claim 1, characterized in that: The text dataset of the interaction action is obtained in the following manner: Based on a large language model, semantic expansion is performed on the original text description of the interaction action to obtain a text dataset of the interaction action.

3. The method for generating whole-body postures for human-object interaction according to claim 1, characterized in that: The point cloud dataset of the target object is obtained in the following way: A random geometric deformation operation is performed on the original point cloud data of the non-possible contact area of ​​the target object to obtain a point cloud data set of the target object.

4. The method for generating whole-body postures for human-object interaction according to any one of claims 1 to 3, characterized in that: The text dataset based on the interactive action, the point cloud dataset of the target object and the double-branch reciprocal diffusion model generate the initial interactive full-body posture of the human body, including: Based on the dual-branch structure module of the dual-branch reciprocal diffusion model, the text dataset of the interactive action and the point cloud dataset of the target object are independently modeled to obtain parameters of the interactive action and parameters of the target object; Based on the interaction module of the double-branch reciprocal diffusion model, the parameters of the interactive action and the parameters of the target object are collaboratively optimized, and based on the diffusion module of the double-branch reciprocal diffusion model, the optimized parameters are diffused to obtain the initial interactive whole-body posture of the human body.

5. The method for generating whole-body postures for human-object interaction according to any one of claims 1 to 3, characterized in that: The step of inputting the initial interactive full-body posture and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object comprises: The initial interactive full-body posture and the point cloud data set of the target object are input into a contact predictor, and the contact position between the human body and the target object is predicted based on the possible contact area of ​​the target object and a predetermined hand posture.

6. The method for generating whole-body postures for human-object interaction according to any one of claims 1 to 3, characterized in that: The step of inputting the contact position and the initial interaction full body posture into a contact-guided interaction refiner to obtain a target interaction full body posture comprises: The contact position and the initial interactive full-body posture are input into a contact-guided interaction refiner, and the initial interactive full-body posture is adjusted based on a contact distance loss and a contact normal plane loss to obtain a target interactive full-body posture.

7. A whole-body posture generation device for human-object interaction, characterized in that: include: A generation module is used to generate the initial interactive full-body posture of the human body based on the text dataset of interactive actions, the point cloud dataset of the target object, and the two-branch reciprocal diffusion model; A prediction module, configured to input the initial interactive full-body posture and the point cloud data set of the target object into a contact predictor to predict the contact position between the human body and the target object; A refinement module is used to input the contact position and the initial interactive full-body posture into a contact-guided interaction refiner to obtain a target interactive full-body posture.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for generating a whole-body posture for human-object interaction as claimed in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a whole-body posture for human-object interaction as claimed in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a whole-body posture for human-object interaction as claimed in any one of claims 1 to 6 is implemented.