Invisible object 4D character interaction generation method based on text description
The 3D human-object interaction keyframe is restored through the two-stage method and dense 4D sequences are generated, which solves the problem of insufficient generalization ability of unknown object interaction synthesis in the prior art, and realizes natural and realistic 4D human-object interaction generation.
Patent Information
- Application Number
- CN202510565049.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The prior art has insufficient generalization capabilities when generating 4D human-object interaction sequences of unknown objects, and cannot achieve natural realistic dynamic interaction synthesis, especially in terms of multi-step interaction and object state migration.
The two-stage method is adopted, first recovering the 3D human-object interaction keyframes through the object position anchoring network, and then using the contact-perceptual diffusion model for timing interpolation to generate dense 4D human-object interaction sequences, reducing the dependence on large-scale 4D datasets.
The natural realistic 4D character-object interaction synthesis of unseen objects is realized, and the robust generalization ability of diverse object geometric forms is improved, ensuring time coherence and real contact dynamics.
Smart Images

Figure CN120491813A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for generating 4D character interactions of invisible objects based on text descriptions. Background Art
[0002] Human-environment interaction generation: Current research on human-environment interaction synthesis can be divided into two major directions: static object interaction and dynamic object interaction.
[0003] Static Object Interaction: Existing technologies based on regression models, diffusion models, and reinforcement learning methods can generate static scene actions such as sitting, lying down, and navigating confined spaces. However, they struggle with dynamic environments where objects move, deform, or change state (such as opening a door or rearranging furniture). Although recent research has begun to incorporate dynamic interaction, limitations remain in multi-step interaction, object state transfer, and real-time adaptability, hindering generalization capabilities for complex scenarios.
[0004] Dynamic Object Interaction: Early research used historical motion prediction to predict interactions. However, due to the scarcity of 4D datasets and physical plausibility constraints, high-quality 4D human-object interaction synthesis remains challenging. Recent research has combined physics and kinematics methods to improve the realism of full-body dynamic interactions, but the diversity of motions and the range of object interactions remain limited.
[0005] Zero-shot interaction generation: The main challenge of HOI synthesis lies in the scarcity of annotated datasets. Although existing datasets have laid a certain foundation, their scale is much smaller than general text-action datasets.
[0006] In summary, the main shortcomings of the existing technology are: since the existing 4D human-object interaction datasets are limited by the singleness of object categories and interaction patterns, the supervised training methods based on these datasets show poor generalization ability when facing unknown objects, and cannot achieve natural and realistic 4D human-object interaction sequence generation for unknown objects. Summary of the Invention
[0007] In view of this, the present invention provides a method for generating 4D characters interactively from invisible objects based on text descriptions, so as to at least solve the above technical problems.
[0008] According to a first aspect of an embodiment of the present invention, a method for generating 4D human interactions of invisible objects based on text descriptions is provided, comprising: stage one, 3D human-object interaction key frame recovery: obtaining a human motion sequence through a human motion model, and uniformly downsampling the human motion sequence to extract key frames of the human motion sequence; for each key frame, reconstructing a human mesh through an SMPL-X model and extracting vertex positions of the human mesh to form a human point cloud; an object position anchoring network using the human point cloud, object template point cloud and text prompts as input, predicting the object position, and generating sparse 3D human-object interaction key frames; stage two, 4D human-object interaction sequence generation: constructing a contact-aware diffusion model, using the sparse 3D human-object interaction key frames as input, and extracting conditional signals containing human posture and contact information from the sparse 3D human-object interaction key frames through a contact-aware encoder of the contact-aware diffusion model; based on the conditional signals, temporally interpolating the sparse 3D human-object interaction key frames through the contact-aware diffusion model to generate a temporally coherent dense 4D human-object interaction sequence.
[0009] Optionally, the object position anchoring network recovers the object position by inferring the spatial relationship between the human body and the object template and is trained on a hybrid dataset comprising Grab and Behave datasets.
[0010] Optionally, the contact perception encoder adopts a PointNet++ architecture to encode 3D human-object interaction key frames and extract contact perception features as the conditional signal.
[0011] Optionally, the contact-aware diffusion model also includes a contact-aware human-object interaction attention module, which dynamically aligns the contact-aware features with the latent variables of the contact-aware diffusion model through a cross-attention mechanism to ensure accurate integration of fine-grained spatial and contact information.
[0012] Optionally, the contact-aware diffusion model is pre-trained on the OMOMO dataset to learn basic action paradigms and object type priors to obtain robust human-object interaction spatial and temporal priors.
[0013] Optionally, the human motion model is an MDM model, and when extracting key frames, the key frames are selected by time averaging, wherein the selection of the number of key frames balances computational efficiency and motion fidelity.
[0014] According to a second aspect of an embodiment of the present invention, a system for generating 4D human interaction with invisible objects based on text description is provided, comprising: a key frame recovery module, configured to: obtain a human motion sequence through a human motion model, and uniformly downsample the human motion sequence to extract key frames of the human motion sequence; for each key frame, reconstruct a human mesh through an SMPL-X model and extract vertex positions of the human mesh to form a human point cloud; an object position anchoring network, which takes the human point cloud, object template point cloud and text prompts as input, predicts the object position, and generates sparse 3D human-object interaction key frames; a sequence generation module, configured to: construct a contact-aware diffusion model, which takes the sparse 3D human-object interaction key frames as input, and extracts conditional signals containing human posture and contact information from the sparse 3D human-object interaction key frames through a contact-aware encoder of the contact-aware diffusion model; based on the conditional signals, perform temporal interpolation on the sparse 3D human-object interaction key frames through the contact-aware diffusion model to generate a temporally coherent dense 4D human-object interaction sequence.
[0015] According to a third aspect of an embodiment of the present invention, there is provided an electronic device comprising a processor and a memory storing a program, wherein the program comprises instructions that, when executed by the processor, cause the processor to perform the steps of the method according to the first aspect.
[0016] According to a fourth aspect of an embodiment of the present invention, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method of the first aspect is implemented.
[0017] In summary, this invention proposes a new, universal framework for 4D human-object interaction synthesis. By decoupling spatial and temporal modeling, it enables natural and realistic 4D human-object interaction synthesis for unseen objects, effectively reducing the reliance on large-scale 4D human-object interaction datasets. During the temporal modeling phase, the proposed contact-aware diffusion model explicitly leverages interaction priors during 4D sequence generation, ensuring robust generalization across diverse object geometries. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0019] Figure 1 This is a flowchart of the steps of a method for interactively generating invisible objects and 4D characters based on text descriptions of the present invention.
[0020] Figure 2 For Figure 1 The corresponding architecture diagram of the invisible object 4D character interaction generation method based on text description.
[0021] Figure 3 Generate a result display diagram for 4D human-object interaction content.
[0022] Figure 4 A graph showing the results of object position anchoring network generation. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0024] See also Figure 1 、 Figure 2 The present invention provides a method for interactively generating invisible objects and 4D characters based on text description, comprising:
[0025] Phase 1: 3D human-object interaction keyframe recovery:
[0026] S11. Obtaining a human motion sequence through a human motion model, and uniformly downsampling the human motion sequence to extract key frames of the human motion sequence;
[0027] S12. For each key frame, reconstruct the human body mesh using the SMPL-X model and extract the vertex positions of the human body mesh to form a human body point cloud;
[0028] S13, object position anchoring network takes human point cloud, object template point cloud and text prompt as input, predicts object position and generates sparse 3D human-object interaction keyframes;
[0029] Phase 2: 4D human-object interaction sequence generation:
[0030] S21, constructing a contact-aware diffusion model, taking sparse 3D human-object interaction keyframes as input, and extracting conditional signals containing human posture and contact information from the sparse 3D human-object interaction keyframes through a contact-aware encoder of the contact-aware diffusion model;
[0031] S22. Based on the conditional signal, temporally interpolate the sparse 3D human-object interaction key frames through a contact-aware diffusion model to generate a temporally coherent dense 4D human-object interaction sequence.
[0032] Optionally, the object position anchoring network recovers the object position by inferring the spatial relationship between the human body and the object template and is trained on a hybrid dataset comprising Grab and Behave datasets.
[0033] Optionally, the contact perception encoder adopts a PointNet++ architecture to encode 3D human-object interaction key frames and extract contact perception features as the conditional signal.
[0034] Optionally, the contact-aware diffusion model also includes a contact-aware human-object interaction attention module, which dynamically aligns the contact-aware features with the latent variables of the contact-aware diffusion model through a cross-attention mechanism to ensure accurate integration of fine-grained spatial and contact information.
[0035] Optionally, the contact-aware diffusion model is pre-trained on the OMOMO dataset to learn basic action paradigms and object type priors to obtain robust human-object interaction spatial and temporal priors.
[0036] Optionally, the human motion model is an MDM model, and when extracting key frames, the key frames are selected by time averaging, wherein the selection of the number of key frames balances computational efficiency and motion fidelity.
[0037] In summary, the present invention proposes a new framework for 4D human-object generation, which utilizes a two-stage modeling method with spatiotemporal decoupling to achieve natural and realistic generation of 4D human-object interaction sequences for unknown objects. Specifically, in the first stage, 3D human-object interaction keyframes are reconstructed. For this purpose, an object position anchoring network is developed. Only human point clouds and object geometry templates are required to reconstruct 3D interaction keyframes, reducing dependence on 4D datasets. In the second stage, 4D human-object interaction sequences are generated. For this purpose, a contact-aware diffusion model is designed. The contact condition signals in the keyframes are extracted through a contact-aware encoder to achieve interpolation generation from sparse keyframes to dense time series.
[0038] Specifically, the solution of the present invention is further described according to the following examples:
[0039] To narrow the gap between datasets and real-world human-object interaction scenarios, this paper proposes a novel 4D human-object interaction sequence generation framework for unknown objects. This framework decomposes 4D human-object interaction sequence generation into two operational tasks:
[0040] (1) Reconstruct 3D human-object interaction keyframes for unknown objects;
[0041] (2) Interpolate the sparse 3D human-object interaction key frames into a temporally coherent dense 4D human-object interaction sequence.
[0042] For these two subtasks, the present invention develops a two-stage processing flow:
[0043] In the first stage, the object position anchoring network learns the human-object interaction pattern. It only needs to input the human point cloud and the object's geometric template to reconstruct the 3D human-object interaction keyframes. The network is trained on a 3D human-object interaction dataset, avoiding the need for large-scale 4D human-object interaction data.
[0044] The second stage employs a contact-aware diffusion model, extracting conditional signals containing human posture and contact information from keyframes through a contact-aware encoder, enabling temporal interpolation from keyframes to 4D sequences. By employing a spatiotemporal decoupling modeling strategy, this method significantly reduces reliance on 4D human-object interaction datasets, enabling the generation of 4D sequences of human-object interactions with unknown objects.
[0045] First, given a text and an object geometry template, our goal is to generate a natural and realistic human-object interaction sequence that conforms to the text description. The model architecture diagram of the method of the present invention is shown in the figure below. Figure 2 As shown, in the first stage, the present invention uses object geometry and human pose priors to recover keyframes of human-object interactions. In the second stage, the contact-aware diffusion model uses the human-object interaction keyframes and the encoded contact codes to generate 4D human-object interaction sequences. After training, the present invention can generalize to unseen objects based on the object geometry and related textual cues.
[0046] The goal of this paper is to synthesize 4D human-object interaction sequences based on text descriptions and unseen objects. This faces two main challenges:
[0047] (1) Generalize to unseen objects while maintaining spatial accuracy;
[0048] (2) Ensure temporal coherence and realistic contact dynamics.
[0049] To this end, we propose a two-stage process: 3D human-object interaction keyframe recovery, followed by 4D interpolation to maintain temporal coherence. In the first stage, we propose an object position anchoring network to learn human-object interaction patterns, enabling the recovery of 3D human-object interaction keyframes. In the second stage, we propose a contact-aware diffusion model to interpolate the sparse 3D human-object interaction keyframes into a temporally coherent 4D HOI sequence, and use a contact-aware encoder to encode the 3D human-object interaction keyframes into a conditional signal.
[0050] The present invention expresses human body motion as x h ∈R N×D , where N is the number of frames and D is the dimension of the human body posture. In each frame n, the human body posture x nContains global joint positions and local 6D continuous rotations. Using the SMPL-X model, the human body mesh is reconstructed from the pose and shape parameters. The object motion is represented by its global 3D position (center of mass) and rotation. Specifically, the human body and object motion are defined as:
[0051] x h =[j,q],x o =[o,r]
[0052] Phase 1, 3D human-object interaction keyframe recovery, includes two parts: human keyframe sampling and object position anchoring network design.
[0053] Starting from the text description p, the human motion x is obtained using the existing human motion model MDM (Human motion diffusion model) h First, key frames are extracted by uniformly downsampling the input human motion sequence x.
[0054] Specifically, given a motion sequence containing N frames, K = 5 key frames are selected by temporal averaging to preserve the key motion dynamics while minimizing redundancy. The choice of K balances computational efficiency and motion fidelity, which is verified in experiments. For each key frame, the human body mesh is reconstructed using the SMPL-X model (SMPL Extended) and the vertex positions V∈R are extracted. (K×M×3) , where M represents the number of mesh vertices. These vertices are considered as a point cloud of the human body. By operating on sparse keyframes rather than dense sequences, error propagation and computational overhead are reduced while capturing diverse interaction states.
[0055] The object position anchoring network recovers the object position by inferring the spatial relationship between the human body and the object template. It adopts the object position pop-up architecture. The network is trained on a hybrid dataset, combining the existing Grab dataset and Behave dataset, and enhanced with single-frame human-object interaction poses extracted from the existing 3DIR image dataset by the existing method CONTHO (Joint reconstruction of 3Dhuman and object via contact-based refinement transformer). This multi-source training strategy enhances the generalization ability of the object position anchoring network to unseen object shapes and interaction dynamics. By training on a point cloud that captures key topological information, rather than relying on coarse SMPL parameters and object poses, the network effectively captures fine-grained contact dynamics, which is critical for accurate and realistic 3D object position recovery. For each human keyframe V k ∈R (M×3), the network takes the human body point cloud, object template point cloud and text prompt p as input, predicts the position of the object, and thus forms a complete HOI frame
[0056] Phase 2: Contact-aware 4D human-object interaction interpolation. After establishing 3D human-object interaction keyframe recovery, we now interpolate these sparse keyframes into a temporally coherent motion sequence.
[0057] To this end, we design a contact-aware diffusion model that generates interaction sequences based on text descriptions, object geometry, and point cloud contact information, ensuring temporal consistency and geometric rationality. ContactDM follows a noise addition and removal framework to generate temporally coherent motion. The complete data representation in our model is:
[0058] τ=(x h , x o ),
[0059] It encapsulates human and object motion. This pose-based representation captures the 3D arrangement of key body joints (such as shoulders, elbows, and knees) as coordinates, providing a lightweight yet expressive representation for efficient training and inference. The model conditions its generation process on a set of signals c, including object geometry and textual descriptions.
[0060] To further enhance the model's ability to capture fine-grained human-object interactions, we introduce a contact-aware encoder and a contact-aware human-object interaction attention mechanism, which are key components of the contact-aware diffusion model. The contact-aware encoder efficiently processes sparse 3D human-object interaction keyframes to extract contact-aware features, while the contact-aware human-object interaction attention module dynamically aligns these features with the latent variables of the diffusion model through a cross-attention mechanism. This ensures the precise integration of fine-grained spatial and contact information, enabling the model to generate realistic and temporally coherent 4D human-object interaction sequences.
[0061] Contact sensing encoder:
[0062] Although diffusion models based on pose representation are lightweight and computationally efficient, they struggle to capture fine-grained details of 3D human-object contact regions. To address this limitation, we propose a contact-aware encoder to encode HOI point clouds and extract contact-aware features to enrich the representation with accurate spatial and interaction information, which is crucial for realistic synthesis.
[0063] Specifically, given a 3D human-object interaction keyframe This paper uses the PointNet++ architecture to encode 3D human-object interaction point clouds. Directly use PointNet++ to encode human and object point clouds and This can result in significant memory overhead.
[0064] In order to alleviate this problem, the present invention adopts an efficient sampling strategy. First, by selecting M o The farthest point to the object point cloud Downsample to obtain a sampling point cloud that retains the object's geometry Secondly, in order to accurately infer the contact relationship, we sample M h The nearest point to focus on the nearest body part, recorded as In the experiment, M o =500,M h = 1000. In order to distinguish the vertices of human body and objects, the unique hot encoding is introduced:
[0065]
[0066] in, and The one and zero vectors represent the human body and object points respectively. The final input point cloud is written as:
[0067]
[0068] The obtained point cloud is then processed by a point cloud encoder, which uses a multi-scale grouping strategy to extract hierarchical spatial features:
[0069] F i =PointEncoder(V k )∈R d ,
[0070] Among them, F i Denotes the encoded features of each frame, and d is the output feature dimension. The point cloud encoder aggregates features at multiple scales using local neighborhoods, ensuring geometrically and contact-aware representation learning.
[0071] Contact-aware human-object interaction attention mechanism:
[0072] The contact-aware human-object interaction attention module converts the encoded contact-aware features F i Connected with the conditional diffusion model. Unlike the tag connection that statically attaches the contact embedding to the input, the cross attention dynamically attaches F i Aligned with the latent variables of the diffusion model to achieve accurate and efficient feature integration.
[0073] Specifically, F i The projection is key K and value V, and the pose embedding E poseas the query Q. This enables the diffusion model to selectively focus on key contact areas during the generation process, ensuring that fine-grained spatial information guides the synthesis process. fused Calculated as:
[0074]
[0075] Among them, Q, K and V are E pose 、F i and F i Linear projection of .
[0076] Human-object interaction pre-training:
[0077] To enable the model to learn basic action patterns (e.g., lifting, picking up, and putting down) and object type priors, we pre-trained a conditional diffusion model and a contact-aware encoder on the OMOMO dataset. The pre-training process uses sampled real-world human-object meshes, object geometry, and text descriptions as conditional inputs. This ensures that the model acquires robust spatial and temporal priors for human-object interactions.
[0078] In addition, the present invention has been verified on multiple datasets, and has achieved advanced performance in the richness and authenticity of human motion and object trajectory generation, proving the effectiveness of the present invention. The quantitative results of the generation task of objects in the training set are shown in Table 1, and the quantitative results of the generation task of objects that are not seen in the training set are shown in Table 2. The qualitative results are shown in Figure 3 shown. Figure 4 This is the generated result of the object position anchoring network in stage 1. At the same time, as a design change, a stronger multi-view image diffusion model can be used as the base model.
[0079] Table 1 Quantitative results of generating existing objects in the training set
[0080]
[0081] Table 2 Quantitative results of generation on unseen objects in the training set
[0082]
[0083] In summary, this paper proposes a new, universal framework for 4D human-object interaction synthesis that decouples spatial and temporal modeling, enabling natural and realistic 4D human-object interaction synthesis for unseen objects. By decoupling spatial and temporal modeling, this paper effectively reduces the reliance on large-scale 4D human-object interaction datasets. During the temporal modeling phase, our proposed contact-aware diffusion model explicitly leverages interaction priors during 4D sequence generation, ensuring robust generalization to diverse object geometries.
[0084] As another example, an embodiment of the present invention further provides a system for generating 4D character interactions of invisible objects based on text descriptions, including:
[0085] Keyframe recovery module, used to:
[0086] Obtaining a human motion sequence through a human motion model and uniformly downsampling the human motion sequence to extract key frames of the human motion sequence;
[0087] For each key frame, the human body mesh is reconstructed using the SMPL-X model and the vertex positions of the human body mesh are extracted to form a human body point cloud;
[0088] The object position anchoring network takes human point cloud, object template point cloud and text prompt as input, predicts object position and generates sparse 3D human-object interaction keyframes;
[0089] Sequence generation module for:
[0090] A contact-aware diffusion model is constructed, which takes sparse 3D human-object interaction keyframes as input and extracts conditional signals containing human posture and contact information from the sparse 3D human-object interaction keyframes through a contact-aware encoder of the contact-aware diffusion model;
[0091] Based on the conditional signal, the sparse 3D human-object interaction key frames are temporally interpolated through a contact-aware diffusion model to generate a temporally coherent dense 4D human-object interaction sequence.
[0092] It should be understood that the invisible object 4D character interaction generation system based on text description in this embodiment is used to implement the corresponding methods in the aforementioned multiple method embodiments and has the beneficial effects of the corresponding method embodiments.
[0093] In summary, the proposed new universal framework for 4D human-object interaction synthesis achieves natural and realistic 4D human-object interaction synthesis for unseen objects by decoupling spatial and temporal modeling, effectively reducing the reliance on large-scale 4D human-object interaction datasets. During the temporal modeling phase, the proposed contact-aware diffusion model explicitly leverages interaction priors during 4D sequence generation, ensuring robust generalization across diverse object geometries.
[0094] As another example, the present invention also provides an electronic device, which will now be described as an electronic device that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0095] The electronic device may include: a processor, a communication interface, a memory, and a communication bus.
[0096] The processor, communication interface and memory communicate with each other through a communication bus. The communication interface is used to communicate with other electronic devices or servers.
[0097] The processor is used to execute programs, and specifically can execute the relevant steps in the above method embodiments.
[0098] Specifically, the program may include program codes including computer operation instructions.
[0099] The processor may be a CPU, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.
[0100] The memory is used to store programs and may include high-speed RAM memory or non-volatile memory, such as at least one disk storage.
[0101] When executed by a processor, the program is used to enable an electronic device to execute a method for interactively generating an invisible object 4D character based on text description of the present invention.
[0102] In addition, the specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiments, and will not be repeated here.
[0103] An exemplary embodiment of the present invention further provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods of the various embodiments of the present invention are implemented. The corresponding process descriptions in the aforementioned method embodiments can be referred to and will not be repeated here.
[0104] The method according to the embodiment of the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0105] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.
[0106] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0107] Finally, it should be noted that the above implementation methods are only used to illustrate the embodiments of the present invention, and are not limitations on the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims.
Claims
1. A method for interactively generating invisible objects and 4D characters based on text description, characterized in that: include: Phase 1: 3D human-object interaction keyframe recovery: Obtaining a human motion sequence through a human motion model and uniformly downsampling the human motion sequence to extract key frames of the human motion sequence; For each key frame, the human body mesh is reconstructed using the SMPL-X model and the vertex positions of the human body mesh are extracted to form a human body point cloud; The object position anchoring network takes human point cloud, object template point cloud and text prompt as input, predicts object position and generates sparse 3D human-object interaction keyframes; Phase 2: 4D human-object interaction sequence generation: A contact-aware diffusion model is constructed, which takes sparse 3D human-object interaction keyframes as input and extracts conditional signals containing human posture and contact information from the sparse 3D human-object interaction keyframes through a contact-aware encoder of the contact-aware diffusion model; Based on the conditional signal, the sparse 3D human-object interaction key frames are temporally interpolated through a contact-aware diffusion model to generate a temporally coherent dense 4D human-object interaction sequence.
2. The method according to claim 1, characterized in that The object position anchoring network recovers the object position by inferring the spatial relationship between the human body and the object template and is trained on a hybrid dataset including the Grab and Behave datasets.
3. The method according to claim 1, characterized in that The contact perception encoder adopts the PointNet++ architecture to encode 3D human-object interaction key frames and extract contact perception features as the conditional signal.
4. The method according to claim 3, characterized in that The contact-aware diffusion model also includes a contact-aware human-object interaction attention module, which dynamically aligns the contact-aware features with the latent variables of the contact-aware diffusion model through a cross-attention mechanism to ensure accurate integration of fine-grained spatial and contact information.
5. The method according to claim 4, characterized in that The contact-aware diffusion model is pre-trained on the OMOMO dataset to learn basic action paradigms and object type priors, obtaining robust human-object interaction spatial and temporal priors.
6. The method according to claim 1, characterized in that The human motion model is an MDM model, and when extracting key frames, the key frames are selected by time averaging, wherein the selection of the number of key frames balances computational efficiency and motion fidelity.
7. A system for interactively generating invisible objects and 4D characters based on text description, characterized in that: include: Keyframe recovery module, used to: Obtaining a human motion sequence through a human motion model and uniformly downsampling the human motion sequence to extract key frames of the human motion sequence; For each key frame, the human body mesh is reconstructed using the SMPL-X model and the vertex positions of the human body mesh are extracted to form a human body point cloud; The object position anchoring network takes human point cloud, object template point cloud and text prompt as input, predicts object position and generates sparse 3D human-object interaction keyframes; Sequence generation module for: A contact-aware diffusion model is constructed, which takes sparse 3D human-object interaction keyframes as input and extracts conditional signals containing human posture and contact information from the sparse 3D human-object interaction keyframes through a contact-aware encoder of the contact-aware diffusion model; Based on the conditional signal, the sparse 3D human-object interaction key frames are temporally interpolated through a contact-aware diffusion model to generate a temporally coherent dense 4D human-object interaction sequence.
8. An electronic device, characterized in that: include: processor; Memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the steps of the method according to any one of claims 1 to 6.
9. A computer storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
An augmented reality three-dimensional tracking registration method based on LINE-MOD template matching
CN109636854A
Video motion detection method and device based on key frame screening pixel blocks and medium
CN116168329A
Obstacle detection method, robot, equipment and storage medium
CN119888681A