Grounded human motion generation using open vocabulary scene and textual content

The method addresses the inefficiencies and biases of traditional human motion generation by using open-vocabulary scene and text context, enabling precise and efficient human motion generation in 3D scenes.

JP2025156125APending Publication Date: 2025-10-14FUJITSU LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025050964
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-10
Filing Date
2025-03-26
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional methods for generating human motion in 3D indoor scenes based on text descriptions are costly, time-consuming, and biased towards centering motion within the scene, lacking precise control and sufficient vocabulary flexibility.

Method used

A method using an open vocabulary scene and text context, involving a pre-trained visual language model and U-Net scene encoder to generate human motion by fusing scene and text features, with conditional motion generation to predict motion parameters for a parametric human model, incorporating open-vocabulary knowledge distillation and regularization losses for improved alignment.

Benefits of technology

Enables efficient, accurate, and timely human motion generation with improved alignment and control, addressing biases and mismatches in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025156125000001_ABST
    Figure 2025156125000001_ABST
Patent Text Reader

Abstract

To provide a method that generates a human movement using an open vocabulary scene and textual context.SOLUTION: A method comprises steps of: receiving input including a 3D point cloud of a scene including a target object and text including a natural language instruction related to the target object; applying a text tokenizer to the text to acquire tokenized text and applying a text encoder from a pre-trained visual language model to generate text features; applying a pre-trained U-Net scene encoder to the 3D point cloud to generate a first set of scene features, which are then down-sampled to acquire a second set of scene features; applying a conditional motion generator to a conditional latent acquired on the basis of fusion of the second scene features and the text features to predict motion parameters of a parametric human model toward the target object; and acquiring, on the basis of the motion parameters and the parametric human model, 3D human meshes for a plurality of motion frames.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 571,353, filed March 28, 2024, the entire contents of which are incorporated herein by reference.

[0002] The embodiments discussed in this disclosure relate to human motion generation using open vocabulary scenes and textual context. [Background technology]

[0003] Generating human motion in 3D indoor scenes based on text descriptions is challenging because motion generation requires joint modeling of the 3D scene, human motion, and natural language. Traditional methods often rely on generating 3D human motion that interacts with specified objects in a manner consistent with the given text description. However, generating diverse and semantically consistent human motion in a 3D scene can be costly and time-consuming in real-world scenarios. Additionally, traditional methods exhibit a bias toward generating motion that is centered within the scene.

[0004] The claimed subject matter in this disclosure is not limited to embodiments that solve any disadvantages or that operate only in such environments. Rather, this background is provided only to illustrate one example technology where some embodiments described in this disclosure may be practiced. Summary of the Invention

[0005] According to one aspect of an embodiment, there is provided a method for human motion generation using an open vocabulary scene and text context. The method may include a set of operations, which may include receiving an input comparing a 3D point cloud of a scene including a target object with text including natural language instructions associated with the target object. The set of operations may further include applying a text tokenizer to the text to obtain tokenized text and generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text. The set of operations may further include generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud and downsampling the first scene features to obtain second scene features. The set of operations may further include obtaining conditional latents based on a fusion of the second scene features and the text features and predicting a set of motion parameters for moving a parametric human model toward the target object over a specific time period by applying a conditional motion generator to the conditional latents. Furthermore, the set of operations may include obtaining a 3D human mesh for multiple motion frames based on a set of motion parameters and a parametric human body model.

[0006] The object and advantages of the embodiments will be realized and attained at least by the elements, features, and combinations particularly pointed out in the claims.

[0007] It is to be understood that both the foregoing general description and the following detailed description are provided by way of example and are explanatory only and are not restrictive of the invention as claimed. [Brief explanation of the drawings]

[0008] The exemplary embodiments will be described and explained with additional specificity and detail through the use of the accompanying drawings.

[0009] [Figure 1] FIG. 1 depicts an exemplary environment for human motion generation using open vocabulary scenes and textual context. [Figure 2] FIG. 1 is a block diagram illustrating an exemplary system for human motion generation using open vocabulary scenes and textual context. [Figure 3] FIG. 10 is a flowchart showing pre-training of the U-Net scene encoder. [Figure 4] FIG. 1 illustrates an example architecture diagram of a system for human motion generation using open vocabulary scenes and textual context. [Figure 5] FIG. 1 illustrates an example scenario of inference for a system for human motion generation using open vocabulary scenes and textual context. [Figure 6] FIG. 1 illustrates an example scenario of a shared open vocabulary visual language space using grounding. [Figure 7] FIG. 1 illustrates an example flowchart for human motion generation using open vocabulary scenes and text context.

[0010] All according to at least one embodiment described in this disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Some embodiments described herein relate to a method and system for human motion generation using open vocabulary scenes and text context. In this disclosure, the system may receive an input including a 3D point cloud of a scene with a target object and a natural language instruction associated with the target object. A text tokenizer may be applied to the text to obtain tokenized text. Then, text features may be generated by applying a text encoder from a pre-trained visual language model to the tokenized text. Additionally, first scene features may be generated by applying a pre-trained U-Net scene encoder to the 3D point cloud. These first scene features may be downsampled to obtain second scene features. A conditional latent may be obtained by fusing the second scene features and the text features. A conditional motion generator may be applied to the conditional latent to predict a set of motion parameters for the movement of a parametric human body model toward the target object over a specific time period. Furthermore, a 3D human mesh for multiple motion frames may be obtained based on the set of motion parameters and the parametric human body model.

[0012] Conventional methods for human motion generation involve populating a virtual 3D human into a 3D scene via textual control. Specifically, these methods use tokenized linguistic descriptions, vocabulary sizes, and RGB-colored 3D point clouds to model conditional probabilities and a set of human motion parameters, namely, global translation (t), global orientation (r), and body pose (θ). In addition, conventional methods utilize a differentiable SMPL-X body model to obtain a human mesh for each motion frame. However, generating motion through textual control presents several challenges. Users or developers may not have sufficient control over the generated motion, leading to a lack of precise control. Human motion generation may be assumed to start from a specific direction or location, resulting in coarse assumptions about location. Motion generation may also be biased toward the center of the scene. Pre-training using a closed vocabulary may lead to the prediction of a finite set of labels for each point in the 3D point cloud, which can be limiting. There may also be mismatches between text and image embeddings. Although a closed vocabulary may encompass a large data set, the closed vocabulary is often insufficient to meet the requirements, resulting in inadequate grounding.

[0013] The present disclosure may address these challenges through grounded human motion generation using open-vocabulary scene and text context. This approach may enable more efficient, accurate, and timely processing of datasets, leading to improved management and optimization of human motion generation. First, the system may be trained to minimize the distance between text embeddings and 3D point cloud scene feature embeddings. Second, it may provide a grounding framework for text- and scene-conditioned human motion generation. Third, the system may establish text-scene alignment in a visual language model space (e.g., CLIP space) by replacing pre-training of a closed-vocabulary scene encoder with open-vocabulary knowledge distillation. Additionally, the system may refine text-scene grounding by fine-tuning the scene encoder with two novel regularization losses that enhance recognition of target object category and size. Finally, the system may demonstrate substantially improved human motion alignment performance during sampling on the dataset for all teacher models.

[0014] Embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0015] 1 is a diagram illustrating an example environment for human motion generation using open vocabulary scenes and textual context, arranged in accordance with at least one embodiment described herein. Referring to FIG. 1, an environment 100 is shown. The environment 100 may include a system 102 hosting a pipeline of models 104 including a pre-trained visual language model 106, a pre-trained U-Net scene encoder 108, a downsampler 110, a fusion module 112, and a conditional motion generator 114. The environment 100 may further include a remote server 116 (which may store a dataset 118) and a communication network 122.

[0016] As used herein, the term "pre-trained" refers to a model that has been fine-tuned for a particular task or previously trained on a dataset before being used. In the context of a pre-trained visual language model 106 or a pre-trained U-Net scene encoder 108, the term may mean that the respective model has already been taught to recognize patterns and features in both visual and textual data through extensive training on a variety of open vocabulary datasets or 3D scene datasets.

[0017] The system 102 may include suitable logic, circuitry, and interfaces that may be configured to implement a model pipeline 104 for text and scene conditional human motion generation. Specifically, the system 102 may acquire inputs including a 3D point cloud 118B of a scene and text 118A including natural language instructions associated with a target object 120B in the scene. The system 102 may use the model pipeline 104 to generate a 3D human mesh 120A of a parametric human model for multiple motion frames of the scene based on the acquired inputs. Examples of the system 102 may include, but are not limited to, a computing device, a hardware-based annealer device, a digital annealer device, a quantum-based or quantum-inspired annealer device, a smartphone, a cellular telephone, a mobile phone, a gaming device, a mainframe machine, a server (or cluster of servers), a computer workstation, and / or a consumer electronics (CE) device.

[0018] The pre-trained visual language model 106 may be a neural network pre-trained for the task of semantic understanding of visual information in images and capable of assigning open-vocabulary text labels to the visual information. For example, the pre-trained visual language model 106 may be a Contrastive Language-Image Pre-Training (CLIP) model, an open-vocabulary image caption model, an open-vocabulary image segmentation model, or an open-vocabulary 3D scene understanding model. As used herein, the term "open vocabulary" refers to a model's ability to understand and process a wide range of words or terms not explicitly included in its training dataset. As an open-vocabulary semantic segmentation model, the pre-trained visual language model 106 may accurately assign semantic labels to each pixel in an image based on an arbitrary set of open-vocabulary text. As a CLIP model, the pre-trained visual language model 106 may be a multimodal visual language model capable of mapping image and text pairs to the same latent space. For example, the pre-trained visual language model 106 may use a vision transformer to encode the images of the scene and a text encoder to encode the text 118A into a common embedding space for comparison and retrieval.

[0019] In an exemplary embodiment, an open vocabulary image segmentation model may be designed to partition an image of a scene into meaningful regions based on an arbitrary text description. This method involves segmenting an image into semantically meaningful segments and classifying such segments with flexible, text-defined categories that may not have been seen during training. Similarly, an open vocabulary 3D scene understanding model may be designed to understand and interpret images without being restricted to a predefined set of object categories. Open vocabulary scene understanding models leverage large-scale visual language models (VLMs) and other multimodal foundational models to enable query and recognition of arbitrary object classes.

[0020] The pre-trained visual language model 106 may include a text encoder 106A and an image encoder 106B. The text encoder 106A may receive tokenized text as input. The text encoder 106A may convert the tokenized text into text embeddings. The tokenized text may be obtained by applying a tokenizer to the text received by the system 102. The text 118A may include a natural language instruction, such as "walk to the chair that is farthest from the end table." The end table may be, for example, a target object 120B.

[0021] Text tokenization may convert text (such as text 118A) into a sequence of tokens that can be processed by the pre-trained visual language model 106. The tokenized text may then be passed through a text encoder 106A, such as a Transformer model, which may process the tokenized text and generate embeddings for the tokenized text.

[0022] The embeddings from the Transformer model may be vector embeddings that can be projected into a common embedding space shared with image encoder 106B. The shape of the text embeddings is equal to the shape of the embeddings produced by image encoder 106B. During training, text encoder 106A may be configured to capture the semantics of the text and align the text features with the image features extracted (from the images) by image encoder 106B, enabling the pre-trained visual language model 106 to understand and generate text descriptions for images.

[0023] Image encoder 106B may receive image inputs, such as images (e.g., multi-view images) corresponding to 3D points or objects in a 3D point cloud of a scene. Image encoder 106B may process each received image through a convolutional neural network (e.g., ResNet), a Vision Transformer (ViT), or a suitable neural network-based encoder to generate an image feature vector. The image feature vector may be a structured representation that encapsulates the meaningful content and attributes of the image. The image feature vector may convert the visual information in the image into a form that can be understood and / or processed by other methods, enabling the extraction of semantic concepts, such as objects, scenes, contexts, or activities depicted in the image. Image encoder 106B may be jointly trained with text encoder 106A to map images and text to a shared latent space.

[0024] For text-scene conditional human motion generation, the system 102 may utilize the text encoder 106A during or after the fine-tuning stage (i.e., during the inference stage) of the pre-trained U-Net scene encoder 108. Similarly, the system 102 may utilize the image encoder 106B before the fine-tuning stage (i.e., in the pre-training stage of the pre-trained U-Net scene encoder 108).

[0025] The pre-trained U-Net scene encoder 108 may be applied to an acquired input, such as the 3D point cloud 118B, to generate first scene features. The first scene features may represent the 3D point cloud 118B as compact high-dimensional vectors that capture essential geometric and spatial features of the scene depicted in the 3D point cloud 118B. The first scene features may also include hierarchical information representing the structure of the scene and spatial relationships between 3D points or 3D objects in the scene for further analysis or processing.

[0026] The pre-trained U-Net scene encoder 108 may include an encoder and decoder pair that may be pre-trained on a dataset of 3D scene data and image pairs before being used in the model pipeline 104 for text- and scene-conditional human motion generation. In an exemplary embodiment, the pre-trained U-Net scene encoder 108 may be an encoder-decoder network that may be based on a Point Transformer. Specifically, the pre-trained U-Net scene encoder 108 may integrate the self-attention mechanism of the Point Transformer into the U-Net architecture to effectively process point cloud data (such as the 3D point cloud 118B). The encoder of the pre-trained U-Net scene encoder 408 (shown as encoder 108A in FIG. 4) may consist of a Point Transformer block that downsamples the point cloud data while capturing spatial relationships. The bottleneck layer of the pre-trained U-Net scene encoder 108 may also be a Point Transformer to extract high-level features. The decoder (shown in FIG. 4 as decoder 408B) may then upsample these features to the original resolution using additional point transformer blocks, using skip connections from the encoder to preserve spatial information. By way of example and not limitation, the pre-trained U-Net scene encoder 108 may include five encoder stages, each consisting of a transition down module and a varying number of point transformer blocks (2, 3, 4, 6, and 3, respectively). The decoder component may include five stages with transition up modules, each with two point transformer blocks. The output head on the decoder component may include a linear layer with ReLU activation and F units. Each point transformer block may incorporate a self-attention layer, a linear projection, and a residual skip connection.

[0027] According to one embodiment, the pre-trained U-Net scene encoder 108 may be fine-tuned for text and scene conditional human motion generation based on losses including two regularization losses associated with the category of the target object 120B and the size of the target object 120B.

[0028] The downsampler 110 may include logic, an interface, and / or code configured to perform downsampling of the first scene feature to obtain the second scene feature. For example, the downsampler 110 may randomly select a set of point feature vectors from the plurality of point feature vectors included in the first scene feature. Furthermore, the downsampler 110 may calculate a distance between each point feature vector in the set of point feature vectors and other point feature vectors among the plurality of point feature vectors. Around each point feature vector in the set of point feature vectors, the downsampler 110 may select a set of k-nearest neighbor vectors from the plurality of point feature vectors based on the distances. Furthermore, the downsampler 110 may apply an average pooling operation to the set of k-nearest neighbor vectors around each point feature vector in the set of point feature vectors to obtain a plurality of average-pooled vectors. The second scene feature may include a plurality of average-pooled vectors.

[0029] The fusion module 112 may include logic, interfaces, and / or code configured to obtain fused features that can be used to generate conditional latencies. The fusion module 112 may concatenate the second scene features with the text features to obtain concatenated features. Further, the fusion module 112 may apply a self-attention layer to the concatenated features to obtain the fused features.

[0030] The conditional motion generator 114 includes logic, interfaces, and / or code configured to obtain the 3D human mesh 120A for multiple motion frames. The conditional motion generator 114 may apply the generated conditional potential to predict a set of motion parameters for the movement of the parametric human model toward a target object (such as target object 120B) over a specific duration. Based on the set of motion parameters and the parametric human model, the conditional motion generator 114 may obtain the 3D human mesh 120A for multiple motion frames.

[0031] As used herein, the term "parametric human model" may refer to a computational model used to represent the human body with high realism. The model may use parameters to adjust the shape and pose of the body, include detailed anatomical structures of the face, hands, and body, and deform smoothly with movement. As an example, the parametric human model may be the Skinned Multi-Person Linear Model-eXtended (SMPL-X) or a variant thereof.

[0032] The remote server 116 may include logic, interfaces, and / or code configured to store the dataset 118, which includes text-3D data pairs (e.g., text 118A and 3D point cloud 118B). In at least one embodiment, the remote server 116 may be implemented as multiple distributed cloud-based resources through the use of several techniques known to those skilled in the art. In certain embodiments, the functionality of the remote server 116 may be incorporated, in whole or at least in part, into the system 102 without departing from the scope of the present disclosure.

[0033] The dataset 118 may be stored or cached on a device such as the remote server 116 or the system 102. The dataset 118 includes, on the remote server 116 or the system 102, text 118A in the form of a table or group of tables and a 3D point cloud 118B associated with a scene including a target object 120B. The 3D point cloud 118B may include a scene including the target object 120B and a 3D human mesh 120A. For example, the target object 120B may be any physical object in the 3D point cloud. The physical object may include at least one of a chair, a table, a blackboard, a television, etc. The dataset 118 may be hosted on multiple servers at the same location or separate locations. Operations of the dataset 118 may be performed using hardware including a processor, a microprocessor (e.g., for performing or controlling the execution of one or more operations), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).

[0034] The communication network 122 may include various communication media through which the system 102 may communicate with the remote server 116 or other devices. Examples of the communication network 122 may include, but are not limited to, the Internet, a cloud network, Wireless Fidelity (Wi-Fi), a Personal Area Network (PAN), a Local Area Network (LAN), a cellular network (such as a Long Term Evolution (or 4G) cellular network or a 5G cellular network), a satellite network (such as a network of low-earth orbit satellites), and / or a Metropolitan Area Network (MAN). Various devices in the environment 100 may connect to the communication network 122 using various wired and wireless communication protocols, including TCP / IP, UDP, HTTP, FTP, ZigBee, EDGE, IEEE 802.11, Li-Fi, IEEE 802.16, multi-hop communication, wireless access points (APs), device-to-device communication, cellular communication protocols, and Bluetooth.

[0035] During operation, the system 102 may receive a 3D point cloud 118B of a scene including a target object 120B and text 118A including natural language instructions associated with the target object 120B. According to one embodiment, the input may be retrieved from a dataset 118 stored on the remote server 116 or the system 102. According to another embodiment, the system 102 may receive the input via a user interface rendered on a user device (not shown). The user interface may include a text input for entering the natural language instructions and options for uploading the 3D point cloud 118B or performing a 3D scan of the scene to obtain the 3D point cloud 118B. The user interface may be part of a software application, such as animation software or robotics software.

[0036] Based on the input, the system 102 may implement a pipeline of models 104 for a conditional text and scene generation task. This task involves simultaneously leveraging both text and scene conditions and requires grounding between the two modalities. Specifically, the task objective is to identify a target object (e.g., target object 120B) among multiple instances of the same object class in a complex 3D scene (e.g., 3D point cloud 118B) guided by a text description of spatial relationships (e.g., text 118A), and subsequently generate a human movement (e.g., 3D human mesh 120A) that interacts with the target object 120B. This interaction may include, for example, movement toward the target object 120B.

[0037] In some cases, a task may be defined with the goal of placing a virtual 3D human's movements in a 3D scene via text control. Specifically, the model pipeline 104 uses conditional probabilities

number

number

number

number

number

number

[0038] Details of the implementation of model pipeline 104 and associated training / fine-tuning are described herein. System 102 may apply a text tokenizer to received text to obtain tokenized text. For example, a text tokenizer may be a function that can break down unstructured text, such as text 118A, into smaller units known as tokens. Tokens can be words, characters, subwords, or sentences, depending on the type of tokenization performed. Tokenization can be an important step in natural language processing tasks because it helps build context and meaning for system 102 by converting text into a form that can be easily processed and analyzed.

[0039] In another aspect, the system 102 can generate text features by applying the text encoder 106A of the pre-trained visual language model 106 to the tokenized text. The text encoder 106A may be, for example, a transformer-based text encoder. As another example, the text encoder 106A may be the text encoding component of an open vocabulary image segmentation model or a CLIP model. The text encoder 106A may process the tokenized text and convert the tokenized text into text embeddings. The text embeddings may include the semantic meaning of the text corresponding to the image of the scene including the target object 120B.

[0040] The system 102 may generate first scene features by applying a pre-trained U-Net scene encoder 108 to the 3D point cloud 118B. The generated first scene features may include multiple average-pooled vectors. In one embodiment, the pre-trained U-Net scene encoder 108 may be a Point Transformer-based neural network (both encoder and decoder blocks with residual skip connections) to calculate scene features for each 3D point of the 3D point cloud 118B. For example, the system 102 may supply position and color information of each 3D point of the 3D point cloud 118B to the pre-trained U-Net scene encoder 108 to generate first scene features, which may include a point feature vector for each 3D point of the 3D point cloud 118B.

[0041] Because the features extracted by the U-Net (i.e., the first scene features) may be generated from all N points in the 3D point cloud 118B, the output features (i.e., the first scene features) have a dimension of C×N. Considering all points for the fusion module 112 may not be feasible. Therefore, as described herein, a downsampler 110 may be used. The system 102 may use the downsampler 110 to downsample the first scene features to obtain the second scene features. Specifically, the downsampler 110 may be a logic block that can transform the first scene features by resampling them to a lower dimension. For example, the downsampler 110 may reduce the number of points from 32,768 in the first scene features to 2,048 in the second scene features by averaging features over k=16 nearest neighbors.

[0042] In an exemplary embodiment, downsampling may be performed using a k-nearest neighbor classifier. Downsampling may involve farthest point sampling and average pooling across the k-nearest neighbors. First, the system 102 may randomly select a set of point feature vectors from the plurality of feature vectors and calculate the distance between each point feature vector in the set of point feature vectors and other point feature vectors among the plurality of point feature vectors. Further, based on the calculated distances, the system 102 may select a set of k-nearest neighbor vectors from the plurality of point feature vectors surrounding each point feature vector in the set of feature vectors. The system 102 may obtain a plurality of average-pooled vectors by applying an average pooling operation to the set of k-nearest neighbor vectors surrounding each point feature vector in the set of point feature vectors.

[0043] The fusion module 112 may concatenate the second scene features with the text features to obtain concatenated features. Furthermore, the fusion module 112 may apply a self-attention layer to the concatenated features to obtain fused features. A conditional latent may be generated based on the fused features. For example, the resulting point features (i.e., the fused features) and associated scene coordinates may be passed through a dense ReLU and linear layer to obtain fused scene features. Finally, the fused scene and text features may be concatenated with a parametric human body model and transformed by a linear layer to generate a conditional latent. As used herein, the term “conditional latent” may refer to combined information of actions, interacting objects, and 3D scene context from two different modalities (i.e., the text 118A and the 3D point cloud 118B). In the context of a Conditional Variational Autoencoder (cVAE), the conditional latent (z) may be sampled from a distribution that can be conditioned on the input 3D point cloud 118B and the output segmentation, rather than being sampled from a simple Gaussian distribution. The conditional latents may enable the pre-trained U-Net scene encoder 108 to learn a complex and informative latent space that can capture the variability and uncertainty in the dataset 118.

[0044] By applying the model pipeline 104 to the conditional potentials, the system 102 may predict a sequence of motion parameters (such as 510 in FIG. 5) for the movement of the parametric human body model toward the target object 120B over a specific duration (e.g., 10 time steps). The sequence of motion parameters may include parameters associated with the global translation, global orientation, and body pose associated with the parametric body model. The sequence of motion parameters 120 may also be predicted over a specific duration (T) to determine multiple motion frames.

[0045] According to one embodiment, the system 102 may obtain a 3D human mesh 120A for multiple motion frames based on a set of motion parameters and a parametric human model. Specifically, a set of predicted motion parameters (over 1...T time steps) may be mapped to a parametric human model to generate multiple motion frames (over 1...T time steps). The mapping for each time step may result in a motion frame consisting of a parametric human model in a specific motion state (walking, sitting, or lying down). For example, the parametric human model may be sitting at time t=1 in one motion frame and standing at time t=2 in another motion frame, etc.

[0046] According to one embodiment, the pre-trained U-Net scene encoder 108 may be fine-tuned for text and scene conditional human motion generation based on losses including two regularization losses associated with the category of the target object 120B and the size of the target object 120B (as shown in Figures 4 and 5).

[0047] In one embodiment, the parametric human model may be SMPL eXpressive (SMPL-X). SMPL-X may be a unified body model that jointly models the human body, face, and hands. SMPL-X may use standard vertex-based linear blend skinning with learned corrective blend shapes, which may have 10,475 vertices and 54 nodes including the neck, jaw, eyeballs, and finger joints. SMPL-X is defined by the function M(θ, β, ψ), where θ represents the pose parameter, β represents the shape parameter, and ψ represents the facial expression parameter.

[0048] In another embodiment, a process for generating human movements based on text and scene conditions involves using a text description (e.g., text 120A) to identify a target object 118B from among multiple objects in a 3D point cloud 118B. Additionally, the process includes generating human movements that refer to or interact with the identified target object 120B. The generation of human movements is also influenced by the text description (e.g., text 118A).

[0049] Additionally, the process involves combining secondary scene features with text features using open vocabulary image segmentation. The process also incorporates two regularization losses for the category of the target object 120B and the size of the target object 120B. More details regarding human motion generation are provided, for example, in Figures 3-6.

[0050] FIG. 2 is a block diagram illustrating an example system for human motion generation using open vocabulary scenes and textual context, arranged in accordance with at least one embodiment described herein. FIG. 2 is described in conjunction with elements from FIG. 1. Referring to FIG. 2, a block diagram 200 of the system 102 is shown. The system 102 may include a processor 202, a memory 204, an I / O device 206, and a network interface 210. The I / O device 206 may include, for example, a display device 208. The memory 204 may store the pre-trained visual language model 106, the pre-trained U-Net scene encoder 108, the fusion module 112, and the conditional motion generator 114.

[0051] Processor 202 may include suitable logic, circuitry, and / or interfaces that can be configured to execute program instructions associated with different operations performed by system 102. Processor 202 may include any suitable special-purpose or general-purpose computer, computing entity, or processing device, including various computer hardware or software modules, and may be configured to execute instructions stored on any applicable computer-readable storage medium. For example, processor 202 may include a microprocessor, microcontroller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or any other digital or analog circuitry configured to interpret and / or execute program instructions and / or process data. While depicted as a single processor in FIG. 2 , processor 202 may include any number of processors that, individually or collectively, are configured to perform or direct the execution of any number of operations of system 102 as described in this disclosure. Additionally, one or more processors may reside on one or more different systems, such as different remote servers.

[0052] In some embodiments, processor 202 may interpret and / or execute program instructions and / or process data stored in memory 204. Processor 202 may execute the program instructions after they are loaded into memory 204. Some examples of processor 202 may be a Graphical Processing Unit (GPU), a Central Processing Unit (CPU), a Reduced Instruction Set Computer (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computer (CISC) processor, a co-processor, and / or combinations thereof.

[0053] The memory 204 may include suitable logic, circuitry, and / or interfaces that may be configured to store program instructions executable by the processor 202. In particular embodiments, the memory 204 may be configured to store information such as, but not limited to, a dataset including a 3D point cloud of a scene including a target object, text including natural language instructions associated with the target object, etc. The memory 204 may further store the pre-trained visual language model 106, the pre-trained U-Net scene encoder 108, the fusion module 112, and the conditional motion generator 114. In some respects, the pre-trained visual language model 106, the pre-trained U-Net scene encoder 108, the fusion module 112, and the conditional motion generator 114 may be located outside of the memory 204. The memory 204 may include a computer-readable storage medium for carrying or having computer-executable instructions or data structures stored thereon. Such a computer-readable storage medium may include any available medium that can be accessed by a general-purpose or special-purpose computer, such as the processor 202.

[0054] By way of example, and not limitation, such computer-readable storage media may include tangible or non-transitory computer-readable storage media including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory devices (e.g., solid-state memory devices), or any other storage medium that may be used to carry or store specific program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer. Combinations of the above may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause processor 202 to perform a particular operation or group of operations associated with system 102.

[0055] The I / O devices 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive user input. The I / O devices 206 may be further configured to provide output in response to the user input. The I / O devices 206 may include a variety of I / O devices that may be configured to communicate with other components, such as the processor 202 and the network interface 210. Examples of input devices may include, but are not limited to, a touchscreen, a keyboard, a mouse, a joystick, and / or a microphone. Examples of output devices may include, but are not limited to, a display device 208 and a speaker. The I / O devices 206 may be configured within the system 102 or external to the system 102.

[0056] The network interface 210 may communicate with the Internet, an intranet, and / or wireless networks, such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any of a number of communication standards, protocols, and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), VoIP, light fidelity (Li-Fi), or Wi-MAX.

[0057] In particular embodiments, the system 102 may include a model pipeline 104, a remote server 116, and a dataset 118. Modifications, additions, or omissions may be made to the system 102 without departing from the scope of the present disclosure. For example, in some embodiments, the system 102 may include any number of other components that may not be explicitly illustrated or described. The system 102, including the model pipeline 104, is described in detail in FIGS. 3, 4, 5, 6, and 7.

[0058] Figure 3 illustrates a flowchart for pre-training a U-Net scene encoder, according to one embodiment of the present disclosure. Figure 3 is described in conjunction with elements from Figures 1 and 2. Referring to Figure 3, an execution flow 300 is shown. The exemplary execution flow 300 may include a set of operations 302-320 that may be performed by one or more components of Figure 1, such as system 102. Operations may begin at 302 and may proceed to 304.

[0059] At 304, a 3D point cloud receiving operation may be performed. The U-Net scene encoder may be configured to receive a 3D point cloud (such as 3D point cloud 118B). In an example embodiment, the received 3D point cloud 118B of the scene may include target object 120B. The U-Net scene encoder may have the same network architecture as the pre-trained U-Net scene encoder 108 with randomly initialized weights. In some instances, the U-Net scene encoder may be referred to as an untrained version of the pre-trained U-Net scene encoder 108.

[0060] At 306, a 3D point selection operation may be performed. The system 102 may be configured to perform 3D point selection from the 3D point cloud in a first pass. The 3D points selected from the 3D point cloud may include position and color information for detection of the target object 120B. The color information may be in RGB format, for example.

[0061] A point feature vector extraction operation may be performed at 308. A U-Net scene encoder may be applied to the position and color information of the selected 3D points to extract the point feature vectors. In an exemplary embodiment, the U-Net scene encoder may be a Point Transformer-based encoder-decoder neural network.

[0062] At 310, an image may be acquired. The system 102 may acquire images corresponding to the selected 3D points. The acquired images may be single-view images or multi-view images of the same scene from views that include the selected 3D points of the 3D point cloud.

[0063] At 312, image feature vector extraction may be performed. An image encoder 106B of the pre-trained visual language model 106 may be applied to the captured image. The image encoder 106B may extract an image feature vector associated with the captured image. For example, the image encoder 106B of the pre-trained visual language model 106 may be an open vocabulary image segmentation model that may be applied to the captured image to extract the image feature vector.

[0064] The distance may be determined at 314. The system 102 may be configured to determine the distance between the point feature vector and the image feature vector.

[0065] At 316, the distance may be minimized. The system 102 may be configured to minimize the distance between the image feature vector and the point feature vector. At 316A, if the distance between the image feature vector and the point feature vector is greater than a threshold, the U-Net scene encoder (e.g., the pre-trained U-Net scene encoder 108) may re-execute steps 306 through 316 by incrementing the counter by 1 (i.e., i+1) and shifting the process to a second pass. The second pass process begins by selecting another 3D point from the 3D point cloud.

[0066] In one embodiment, the open vocabulary image segmentation model may be integrated as a teacher to train a U-Net scene encoder and obtain a pre-trained U-Net scene encoder 108. To minimize the distance between the image feature vectors and the text feature vectors, the system 102 may freeze the text encoder parameters of the open vocabulary image segmentation model.

[0067] In one embodiment, the distance between the image feature vector and the point feature vector may be minimized by maximizing the cosine similarity between the image feature vector and the point feature vector. The cosine similarity may be calculated based on the open vocabulary open scene loss, as given by Equation (1). The open vocabulary open scene loss may be used to train a U-Net scene encoder. For example, the open vocabulary open scene loss may be calculated as follows:

number

number

number

[0068] At 318, a pre-trained U-Net scene encoder 108 may be obtained. The U-Net scene encoder may be trained until the distance between the image feature vector and the point feature vector is minimized (e.g., as shown in equation (1)). When the distance between the image feature vector and the point feature vector is minimized, the pre-trained U-Net scene encoder 108 is obtained.

[0069] In one embodiment, the U-Net scene encoder may be pre-trained using a loss to achieve multimodal matching with the text encoder 106A of the pre-trained visual language model 106. For example, applying the U-Net scene encoder to 3D point location and color information generates point feature vectors (f 3d ) is extracted, and the image feature vector (f 2d ) may be extracted. The CLIP-based image encoder extracts the extracted image feature vectors (f 2d ) so that the system 102 can have a point feature vector (f 3d ) and image feature vector (f 2d ), the U-Net scene encoder uses that distance to generate a point feature vector (f 3d ) corresponds to the text 118A or image feature vector (f 2d ) may be trained to learn a shared embedding close to

[0070] 4 illustrates an example architecture diagram of a system for human movement generation using open vocabulary scenes and textual context, according to one embodiment of the present disclosure. FIG. 4 is described in conjunction with elements from FIG. 1, FIG. 2, and FIG. 3. The example architecture of system 102 illustrates the human movement generation shown in example environment 400. The human movement generation may be implemented by any suitable system, apparatus, or device, such as example system 102 of FIG. 1 or processor 202 of FIG. 2.

[0071] The system 102 may receive input including a 3D point cloud of a scene (S) 406 and text 118A. The 3D point cloud of the scene (S) 406 may include a target object 120B. The text 118A may include natural language instructions associated with the target object 120B. For example, the natural language instructions may refer to the use of natural language (as spoken or written by humans) to describe a task, guide an action, or provide constraints. The natural language instructions may be embedded in a model, such as the pre-trained visual language model 106, to enable the pre-trained visual language model 106 to accurately interpret and follow the instructions contained in the text 118A. Thus, the pre-trained visual language model 106 can generalize and perform a task based on the language instructions provided in the text 118A. Additionally, the system 102 may apply a text tokenizer to the text 118A to obtain tokenized text (L) 402.

[0072] The system 102 may generate text features by applying a text encoder 404 of the pre-trained visual language model 106 to the tokenized text (L) 402. The text features may be generated by passing the tokenized text (L) 402 through the text encoder 404 (e.g., a transformer-based encoder). The text encoder 404 may sequentially process the tokenized text (L) 402 and convert the tokenized text (L) 402 into text embeddings. The text embeddings may include the semantic meaning of the text corresponding to the 3D point cloud (S) 406 of the scene including the target object 120B. In one embodiment, the text encoder 404 of the pre-trained visual language model 106 may be an open-vocabulary semantic segmentation text encoder. For example, open-vocabulary semantic segmentation may involve labeling each point of the 3D point cloud with a semantic category of the target object 120B (or open-vocabulary text). Open vocabulary semantic segmentation may not be limited to a predefined set of classes of target objects 120B, but may be processed using a wide range of text descriptions that allow for identification and classification of target objects 120B within the 3D point cloud based on any input text (such as text 118A).

[0073] The system 102 may generate first scene features by applying a pre-trained U-Net scene encoder 408 to the 3D point cloud (S) 406. The generated first scene features may include multiple average-pooled vectors. The U-Net may be shaped like the letter U, with an encoder 408A, a decoder 408B, and a skip connection connecting the encoder 408A to the decoder 408B. Pre-training the encoder helps it learn useful features from a large amount of data and may then be fine-tuned for a specific task. The pre-trained U-Net scene encoder 408 may be based on Point Transformers and may be used to extract semantic information from the input 3D point cloud (S) 406 of the scene and transform the 3D point cloud (S) 406 of the scene into a latent representation vector.

[0074] In one embodiment, the system 102 may generate first scene features by feeding position and color information of each 3D point of the 3D point cloud to a pre-trained U-Net scene encoder 408. The first scene features may include a point feature vector for each 3D point of the 3D point cloud.

[0075] The system 102 may downsample the first scene feature using the downsampler 410. Furthermore, based on the downsampling performed by the downsampler 410 of the system 102, the system 102 may obtain a second scene feature. For downsampling of the first scene feature, the downsampler 410 may randomly select a set of point feature vectors from the plurality of point feature vectors included in the first scene feature. Furthermore, the downsampler 410 may calculate the distance between each point feature vector in the set of point feature vectors and other point feature vectors among the plurality of point feature vectors. Around each point feature vector in the set of point feature vectors, the downsampler 410 may select a set of k-nearest neighbor vectors from the plurality of point feature vectors based on the distance. For example, the k-nearest neighbor vectors may be the "k" closest data points to a query point in a vector space, which may be measured by a specified distance metric such as Euclidean distance. The query point may be a vector that may represent data points for which nearest neighbors can be determined. The query point may be used as input in a k-NN (k-nearest neighbor) search to identify and retrieve the most similar vector from the dataset 118 based on a defined similarity metric. The k-nearest neighbor may identify 3D points based on their proximity to the query point.

[0076] Additionally, the downsampler 102 of the system 102 may apply an average pooling operation to a set of k-nearest neighbor vectors around each point feature vector in the set of point feature vectors to obtain multiple average pooled vectors. In one embodiment, the multiple average pooled vectors may form a second scene feature. Each average pooled vector may replace the original set of k-nearest neighbor vectors, resulting in a downsampled (pooled) scene feature.

[0077] In one embodiment, downsampling may include point sampling and average pooling over k-nearest neighbors. The downsampler 110 may perform k-nearest neighbor classification. In k-nearest neighbor classification, feature vectors located nearby may be assumed to represent the same object. Therefore, downsampling may be performed by average pooling over k-NN points. For example, k-nearest neighbor classification may be a nonparametric instance-based learning method used in classification tasks. The k-nearest neighbor classification process may find the k-nearest neighbors to a query point and predict the class of the query point based on a majority vote of the classes in the neighborhood. The value of k may be a positive integer. Furthermore, the k-nearest neighbor classifier may assume that similar data points are located nearby.

[0078] In one embodiment, downsampling of the first scene features may be required when the pre-trained U-Net scene encoder 408 has features extracted from all N points of the 3D point cloud and the output scene features (i.e., the first scene features) are C×N dimensional, where N may be the number of points in the 3D point cloud 118B. Therefore, it may not be feasible to consider all points for the fusion module 412.

[0079] The system 102 may use a fusion module 412 to fuse the second scene features with the text features. The fusion module 102 of the system 102 may concatenate the second scene features with the text features to obtain concatenated features. Furthermore, the fusion module 412 may apply a self-attention layer to the concatenated features to obtain fused features. For example, the self-attention layer may compute single-head or multi-head self-attention of the input, such as the concatenated features. The self-attention layer may capture dependencies and relationships within the concatenated feature input. The self-attention mechanism may transform the concatenated feature input sequence into a three-vector: a query, a key, and a value. Furthermore, the self-attention layer may use the three vectors to determine the importance of each feature of the concatenated features in the sequence relative to others. Thus, self-attention enables the fusion module 412 to understand the context and assign appropriate weights to each feature of the concatenated features based on their relevance. Furthermore, a conditional latent 414 may be generated based on the fused features. In one embodiment, the fusion module 412 may pass the concatenated features and the position information of each 3D point in the 3D point cloud through a dense ReLU layer and a linear layer to obtain the fused features.

[0080] The system 102 may take as input the tokenized text (L) 402 and the 3D point cloud (S) 406 and provide as output the conditional latents 414. The conditional latents 414 may be c Alternatively, the output condition may be given by the following equation (2):

number

[0081] The system 102 may apply the conditional motion generator 416 to the conditional potentials 414 to predict a set of motion parameters for the movement of the parametric human model towards the target object 120B over a particular duration.

[0082] The conditional motion generator 416 may be a combination of a motion encoder and a motion decoder. The motion encoder of the conditional motion generator 416 may be represented by Encψ. Furthermore, the motion encoder of the conditional motion generator 416 may include a bidirectional gated recurrent unit (GRU) layer concatenated with the conditional latent 414, a residual block, and a linear output layer of Gaussian mean and covariance parameters. Furthermore, reparameterization may be applied to sample the conditional latent 414 (the sampled conditional latent may be represented by "Z"). The bidirectional GRU layer is a type of sequence processing model consisting of two gated recurrent units (GRUs). A bidirectional GRU processes an input sequence, such as the conditional latent 414, in two directions: one GRU processes the sequence from beginning to end (forward), and the other GRU processes the sequence from end to beginning (reverse). The outputs from both directions may then be concatenated to generate a final output. The GRU's bidirectional approach enables the conditional motion generator 416 to capture information from both the past and future of the input sequence, allowing the conditional motion generator 416 to understand the context of the entire sequence, such as natural language processing or speech recognition.

[0083] In one embodiment, the motion decoder of the conditional motion generator 416 may be represented by "Decφ". Furthermore, the motion decoder of the conditional motion generator 416 uses a linear layer to generate the conditional latent 414 (Z c) with a sampled conditional latent (Z) and processed by utilizing a sinusoidal positional embedding, a transform decoder, and a linear output layer. For example, the linear layer may be a connected layer or a dense layer, which may be a basic building block in a neural network where each input may be connected to each output by a weight. The linear layer performs a linear transform on the input data. Further details regarding the linear layer are omitted for brevity. Furthermore, the sinusoidal positional embedding may be used to encode the position of a token within a sequence by using sine and cosine functions. Furthermore, each dimension of the positional embedding corresponds to a sinusoidal wave with a different frequency. This ensures that the positional value falls between 0 and 1 regardless of the length of the sequence (such as a fused feature or a sequence of concatenated fused features and text features). The sinusoidal pattern allows the model to generalize to sequences of different lengths and recognize patterns across various positions in the data. Further details regarding the sinusoidal positional embedding are omitted for brevity.

[0084] In one embodiment, the conditional motion generator 416 may use the motion parameters 420 (i.e., initial parameters) and the conditional potentials 414 to provide reconstructed motion parameters 418. The reconstructed motion parameters 418 may be a predicted set of motion parameters for the motion of the parametric human body model.

[0085] The pre-trained U-Net scene encoder 408 may be fine-tuned for text and scene conditional human motion generation based on losses, including two regularization losses associated with the category of the target object 120B and the size of the target object 120B.

[0086] In one embodiment, the loss may be derived from a reconstruction loss between the true and predicted parametric human body model parameters and a regularization loss consisting of a KL (Kullback-Leibler) divergence loss. For example, two regularization losses associated with the category of the target object 120B and the size of the target object 120B may be given by the following equations (3), (4), and (5):

number

[0087] Figure 5 illustrates an exemplary scenario of inference for a system for human movement generation using open vocabulary scenes and textual context, according to one embodiment of the present disclosure. Figure 5 is described in conjunction with elements from Figures 1, 2, 3, and 4. With reference to Figure 5, an exemplary flow 500 is shown. The method illustrated in the exemplary environment 500 may be performed by any suitable system, apparatus, or device, such as the exemplary system 102 of Figure 1 or the processor 202 of Figure 2.

[0088] The system 102 may use a fusion module 502 to concatenate second scene features with text features to obtain concatenated features. The second scene features may be obtained by downsampling the first scene features. Furthermore, the text features may be generated based on applying a text encoder of a pre-trained visual language model to the tokenized text. Furthermore, the fusion module 502 may apply a self-attention layer to the concatenated features to obtain fused features. Furthermore, the conditional latent 504 may be generated based on the fused features. In one embodiment, the fusion module 502 may pass the concatenated features and position information of each 3D point in the 3D point cloud through a dense ReLU layer and a linear layer to obtain the fused features. The system 102 may take the tokenized text (L) 402 and the 3D point cloud (S) 406 as input and provide the conditional latent 504 as output. Additionally, the system 102 may apply a conditional motion generator 506 to the conditional latency 504 to predict a set of motion parameters for the movement of the parametric human model towards the target object 120B over a particular duration.

[0089] In one embodiment, the conditional motion generator 506 may use the motion parameters 420 and the conditional potential 504 to provide reconstructed motion parameters 508. Furthermore, the system 102 may obtain a 3D human mesh 512 for multiple motion frames 510 based on the set of motion parameters (i.e., the reconstructed motion parameters 508) and the parametric human model. The set of motion parameters and the parametric human model may be obtained for each time step from a set of time steps, i.e., t1 to t7, as shown at 510 in FIG. 5 .

[0090] In an exemplary embodiment, the system 102 may determine the target object 118B based on the text description (e.g., text 120A). Furthermore, the conditional motion generator 506 of the system 102 may generate multiple motion frames 510 near the target object 120B. Furthermore, the system 102 may calculate a distance between the multiple motion frames 510 and the determined target object 102B. Based on the determined distance, the system 102 may further obtain a 3D human mesh 512 that may be in the immediate vicinity of the target object 102B. For example, the distance may be calculated based on a random standard Gaussian latent, the conditional latent 504 at time step t, and an SMPL-X human mesh sampled from a signed distance function (SDF). The SDF may be evaluated based on a subset of S that corresponds to the target object 120B in the text 118A(L). The system 102 determines the shortest distance, and if the shortest distance is negative, the system 102 replaces the shortest distance with zero to ignore the penetration. Additionally, the system 102 may use the last movement frame t=T for walking, sitting, or lying down, and the first frame t=1 for standing up.

[0091] Figure 6 illustrates an exemplary scenario of a shared open vocabulary visual language space with grounding, according to one embodiment of the present disclosure. Figure 6 will be described in conjunction with elements of Figures 1, 2, 3, 4, and 5. Referring to Figure 5, an exemplary architecture 600 is shown. The method illustrated in exemplary architecture 600 may be performed by any suitable system, apparatus, or device, such as exemplary system 102 of Figure 1 or processor 202 of Figure 2. Architecture 600 may include a text input 602, a 3D point cloud 604 associated with a scene, an open vocabulary visual language space with grounding 606, a fusion 608 step, second scene features 610, and text features 612.

[0092] Architecture 600 represents an open-vocabulary grounding architecture that can be designed to enhance text-scene conditioned human motion generation. System 102 may establish text-scene relationships (groundings) prior to motion generation by leveraging the extensive grounding knowledge acquired by pre-trained visual language model 106 during pre-training.

[0093] The system 102 may pre-train a visual language model through pre-training an encoder 408A that distills knowledge from an open-vocabulary semantic image segmentation model on the dataset 118. Specifically, the system 102 creates correspondences between 3D scene points in the embedding space and text-aligned 2D viewpoint pixels and aligns the representation of the encoder 408A with the visual language model's text encoder 404. Furthermore, the system 102 may fuse text features 612 generated by the pre-trained visual language model 106 with second scene features 610 obtained by downsampling the first scene features (which may be generated by the pre-trained U-Net scene encoder 108). Upon feature fusion, the system 102 may fine-tune the pre-trained U-Net scene encoder 108 for text- and scene-conditioned human motion generation based on losses, including two regularization losses associated with the category of the target object 120B and the size of the target object 120B (and thus the grounding of the target object 120B).

[0094] In one embodiment, the architecture 600 employs a shared open vocabulary visual language space 606 for the generation of the generated text features 612 and the second scene features 610, and establishes an initial relationship (grounding) between the generated text features 612 and the generation of the second scene features 610. Furthermore, to normalize the relationship (grounding), the system 102 can classify and regress the corners of the bounding box of the target object 120B within the second scene features 610.

[0095] In one embodiment, the system 102 may identify the target object 120B according to the text input 602 and use grounding to generate human movements that more closely resemble the target object 120B in the open vocabulary visual language space 606. Furthermore, the human movements generated by the system 102 may have fine-grained orientations, such as those shown at 610 in FIG. 6, associated with the text input 602.

[0096] FIG. 7 illustrates a flowchart of an example of human movement generation using open vocabulary scenes and textual context, according to one embodiment of the present disclosure. FIG. 7 is described in conjunction with elements from FIGS. 1, 2, 3, 4, 5, and 6. Referring to FIG. 7, an example flow 700 is shown. The method illustrated in example flow 700 may be performed by any suitable system, apparatus, or device, such as example system 102 of FIG. 1 or processor 202 of FIG. 2. Although illustrated with discrete blocks, steps and operations associated with one or more of the blocks of flow 700 may be divided into additional blocks, combined into fewer blocks, or eliminated, depending on the particular implementation. Operations may start at 702 and proceed to 722.

[0097] At 704, input including a 3D point cloud of a scene and text may be received by the system 102. The 3D point cloud of the scene includes a target object and the text includes a natural language instruction associated with the target object.

[0098] A text tokenizer may be applied to the received text at 706. Application of the text tokenizer to the text may be performed to obtain tokenized text.

[0099] At 708, text features are generated. The text features may be generated by applying a pre-trained U-Net scene encoder to the 3D point cloud.

[0100] First scene features are generated at 710. The first scene features may be generated by applying a pre-trained U-Net scene encoder to the 3D point cloud.

[0101] At 712, the first scene feature is downsampled to obtain a second scene feature.

[0102] At 714, a conditional latent is obtained based on the fusion of the second scene feature and the text feature.

[0103] At 716, a set of motion parameters of the parametric body model over a particular duration is predicted by applying a conditional motion generator to the conditional latents.

[0104] At 718, a 3D human mesh is obtained for multiple motion frames based on the set of motion parameters and the parametric body model.

[0105] It should be noted that the user device having the display device 208 is provided merely as an example implementation of the system 102 of Figure 1 and should not be construed as limiting the scope of the present disclosure. The present disclosure may also be applicable to other modifications, deletions, or additions to the display device 208 without departing from the scope of the present disclosure.

[0106] The embodiments described herein may be used in many application fields, such as animated videos or movies, video games, robotics, augmented reality, virtual reality, and mixed reality, to utilize accurate global positioning systems, improved motion recovery systems for real-time applications where humanoid robot imitation learning enhances the robot's ability to accurately mimic human motion. Furthermore, the present disclosure may be used in text-to-motion generation, creating human motions based on text descriptions for various interactive applications. An advantage of the aforementioned applications may be the ability to generate diverse and contextually accurate human motions based on verbal prompts and scene information.

[0107] Various embodiments of the present disclosure may provide one or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause a system (such as system 102) to perform an operation. The operation may include receiving input comparing a 3D point cloud of a scene including a target object with text including natural language instructions associated with the target object. The operation may further include applying a text tokenizer to the text to obtain tokenized text. The operation may further include generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text. The operation may further include generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud. The operation may further include downsampling the first scene features to obtain second scene features. The operation may further include obtaining conditional latencies based on a fusion of the second scene features and the text features. The operations may further include predicting a set of motion parameters for movement of the parametric human model toward the target object over a particular time period by applying the conditional motion generator to the conditional potential. Further, the operations may include obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human model.

[0108] As indicated above, embodiments described in this disclosure may involve the use of a special-purpose or general-purpose computer (e.g., processor 202 of FIG. 2) that includes various computer hardware or software modules, as discussed in more detail below. Additionally, as indicated above, embodiments described in this disclosure may be implemented using a computer-readable medium that carries or has stored thereon computer-executable instructions or data structures (e.g., memory 204 or dataset 118 or a data input prepared based on dataset 118).

[0109] As used in this disclosure, the terms “module” and “component” may refer to a specific hardware implementation configured to perform the actions of the module or component and / or to a software object or software routine that may be stored on and / or executed by general-purpose hardware of the system 102 (e.g., a computer-readable medium, a processing device, etc.). In some embodiments, different components, modules, engines, and services described in this disclosure may be implemented as objects or processes that run on the system 102 (e.g., as separate threads). While some of the systems 102 and methods described in this disclosure are generally described as being implemented in software (stored on and / or executed by general-purpose hardware), specific hardware implementations or combinations of software and specific hardware implementations are possible and contemplated. In this description, a “computing entity” may be any system 102 defined above in this disclosure, or any module or combination of modules operating on the system 102.

[0110] Additionally, where a specific number of introduced claim provisions are intended, such intention will be expressly set forth in the claim; where such a provision is absent, such intention does not exist. For example, as an aid to understanding, the following appended claims may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim provisions. However, the use of such phrases should not be construed to suggest that the introduction of a claim provision with the indefinite article "a" or "an" limits any particular claim containing such introduced claim provision to embodiments containing only one such provision, even when the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"), and the same is true for any use of an indefinite article used to introduce a claim provision.

[0111] Furthermore, any disjunctive word or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both of the terms. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B" or "A and B."

[0112] All examples and conditional language set forth in this disclosure are intended for educational purposes to aid the reader in understanding the disclosure and the concepts the inventors have contributed to furthering the art, and should not be construed as being limited to the examples and conditions so specifically set forth. While embodiments of the present disclosure have been described in detail, various changes, substitutions, and alterations can be made thereto without departing from the spirit and scope of the present disclosure.

[0113] This disclosure discloses the following additional information. (Appendix 1) 1. A method executed by at least one processor, comprising: receiving an input, said input comprising: a 3D point cloud of a scene including a target object; text including natural language instructions associated with the target object; applying a text tokenizer to the text to obtain tokenized text; generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text; generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud; downsampling the first scene feature to obtain a second scene feature; obtaining a conditional latent based on a fusion of the second scene feature and the text feature; and predicting a set of motion parameters for a parametric human model moving towards a target object over a specific time period by applying a conditional motion generator to the conditional latents; obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human body model. (Appendix 2) 2. The method of claim 1, wherein the pre-trained visual language model is a Contrastive Language-Image Pre-Training (CLIP) model. (Appendix 3) 2. The method of claim 1, wherein the pre-trained U-Net scene encoder is a Point Transformer-based neural network. (Appendix 4) 2. The method of claim 1, further comprising: supplying position and color information of each 3D point of the 3D point cloud to the pre-trained U-Net scene encoder to generate the first scene features comprising a point feature vector for each 3D point of the 3D point cloud. (Appendix 5) selecting 3D points from the 3D point cloud; Extracting point feature vectors of the 3D points by applying a U-Net scene encoder to the position and color information of the 3D points; acquiring an image corresponding to the 3D point; extracting an image feature vector by applying an image encoder of the pre-trained visual language model to the image; 2. The method of claim 1, further comprising: obtaining the pre-trained U-Net scene encoder by pre-training the U-Net scene encoder until a distance between the image feature vector and the point feature vector is minimized. (Appendix 6) The downsampling step comprises: performing a random selection of a set of point feature vectors from a plurality of point feature vectors included in the first scene feature; calculating a distance between each point feature vector of the set of point feature vectors and another point feature vector of the plurality of point feature vectors; selecting a set of k-nearest neighbor vectors from the plurality of point feature vectors based on the distances around each point feature vector in the set of point feature vectors; applying an average pooling operation to the set of k-nearest neighbor vectors around each point feature vector of the set of point feature vectors to obtain a plurality of average pooled vectors; 2. The method of claim 1, wherein the second scene features include the plurality of average pooled vectors. (Appendix 7) 2. The method of claim 1, wherein the downsampling is performed using a k-nearest neighbor classifier. (Appendix 8) The fusion of the second scene feature and the text feature includes: concatenating the second scene features with the text features to obtain concatenated features; and and applying a self-attention layer to the concatenated features to obtain fused features, wherein the conditional latent is generated based on the fused features. (Appendix 9) 2. The method of claim 1, further comprising fine-tuning the pre-trained U-Net scene encoder for text and scene conditional human motion generation based on losses, including two regularization losses related to the target object category and the target object size. (Appendix 10) One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause the system to perform operations, including: receiving an input, said input comprising: a 3D point cloud of a scene including a target object; text including natural language instructions associated with the target object; applying a text tokenizer to the text to obtain tokenized text; generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text; generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud; downsampling the first scene feature to obtain a second scene feature; obtaining a conditional latent based on a fusion of the second scene feature and the text feature; and predicting a set of motion parameters for a parametric human model moving towards a target object over a specific time period by applying a conditional motion generator to the conditional latents; and obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human body model. (Appendix 11) 11. The one or more non-transitory computer-readable storage media of claim 10, wherein the pre-trained visual language model is a Contrastive Language-Image Pre-Training (CLIP) model. (Appendix 12) 11. The one or more non-transitory computer-readable storage media of claim 10, wherein the pre-trained U-Net scene encoder is a Point Transformer-based encoder-decoder neural network. (Appendix 13) 11. The one or more non-transitory computer-readable storage media of claim 10, wherein the operations further include supplying position and color information of each 3D point of the 3D point cloud to the pre-trained U-Net scene encoder to generate the tenth scene feature comprising a point feature vector for each 3D point of the 3D point cloud. (Appendix 14) The operation is selecting 3D points from the 3D point cloud; Extracting point feature vectors of the 3D points by applying a U-Net scene encoder to the position and color information of the 3D points; acquiring an image corresponding to the 3D point; extracting an image feature vector by applying an image encoder of the pre-trained visual language model to the image; obtaining the pre-trained U-Net scene encoder by pre-training the U-Net scene encoder until a distance between the image feature vector and the point feature vector is minimized. (Appendix 15) The downsampling is performing a random selection of a set of point feature vectors from a plurality of point feature vectors included in the first scene feature; calculating a distance between each point feature vector of the set of point feature vectors and another point feature vector of the plurality of point feature vectors; selecting a set of k-nearest neighbor vectors from the plurality of point feature vectors based on the distances around each point feature vector in the set of point feature vectors; applying an average pooling operation to the set of k-nearest neighbor vectors around each point feature vector of the set of point feature vectors to obtain a plurality of average pooled vectors; 11. The one or more non-transitory computer-readable storage media of claim 10, wherein the second scene features include the plurality of average pooled vectors. (Appendix 16) 11. The one or more non-transitory computer-readable storage media of claim 10, wherein the downsampling is performed using a k-nearest neighbor classifier. (Appendix 17) The fusion of the second scene feature and the text feature includes: concatenating the second scene features with the text features to obtain concatenated features; and 11. The one or more non-transitory computer-readable storage media of claim 10, wherein applying a self-attention layer to the concatenated features to obtain fused features, and wherein the conditional latent is generated based on the fused features. (Appendix 18) 11. The one or more non-transitory computer-readable storage media of claim 10, wherein the operations further include fine-tuning the pre-trained U-Net scene encoder for text and scene conditional human motion generation based on a loss, including a regularization loss related to the target object category and the target object size. (Appendix 19) 1. A system comprising: a memory for storing instructions; a processor coupled to the memory that executes the instructions to perform a process, the process comprising: receiving an input, said input comprising: a 3D point cloud of a scene including a target object; text including natural language instructions associated with the target object; applying a text tokenizer to the text to obtain tokenized text; generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text; generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud; downsampling the first scene feature to obtain a second scene feature; obtaining a conditional latent based on a fusion of the second scene feature and the text feature; and predicting a set of motion parameters for a parametric human model moving towards a target object over a specific time period by applying a conditional motion generator to the conditional latents; and obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human body model. (Appendix 20) The process comprises: selecting 3D points from the 3D point cloud; Extracting point feature vectors of the 3D points by applying a U-Net scene encoder to the position and color information of the 3D points; acquiring an image corresponding to the 3D point; extracting an image feature vector by applying an image encoder of the pre-trained visual language model to the image; 20. The system of claim 19, further comprising: obtaining the pre-trained U-Net scene encoder by pre-training the U-Net scene encoder until a distance between the image feature vector and the point feature vector is minimized.

Claims

1. 1. A method executed by at least one processor, comprising: receiving an input, said input comprising: a 3D point cloud of a scene including a target object; text including natural language instructions associated with the target object; applying a text tokenizer to the text to obtain tokenized text; generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text; generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud; downsampling the first scene features to obtain second scene features; obtaining a conditional latent based on a fusion of the second scene feature and the text feature; predicting a set of motion parameters for a parametric human model moving towards a target object over a specific time period by applying a conditional motion generator to the conditional latents; and obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human body model.

2. The method of claim 1 , wherein the pre-trained visual language model is a Contrastive Language-Image Pre-Training (CLIP) model.

3. The method of claim 1 , wherein the pre-trained U-Net scene encoder is a Point Transformer-based neural network.

4. 2. The method of claim 1, further comprising: providing position and color information of each 3D point of the 3D point cloud to the pre-trained U-Net scene encoder to generate the first scene features comprising a point feature vector for each 3D point of the 3D point cloud.

5. selecting 3D points from the 3D point cloud; Extracting point feature vectors of the 3D points by applying a U-Net scene encoder to the position and color information of the 3D points; acquiring an image corresponding to the 3D point; extracting an image feature vector by applying an image encoder of the pre-trained visual language model to the image; 2. The method of claim 1 , further comprising: obtaining the pre-trained U-Net scene encoder by pre-training the U-Net scene encoder until a distance between the image feature vector and the point feature vector is minimized.

6. The downsampling step comprises: performing a random selection of a set of point feature vectors from a plurality of point feature vectors included in the first scene feature; calculating a distance between each point feature vector of the set of point feature vectors and another point feature vector of the plurality of point feature vectors; selecting a set of k-nearest neighbor vectors from the plurality of point feature vectors based on the distances around each point feature vector in the set of point feature vectors; applying an average pooling operation to the set of k-nearest neighbor vectors around each point feature vector of the set of point feature vectors to obtain a plurality of average pooled vectors; The method of claim 1 , wherein the second scene features include the plurality of average-pooled vectors.

7. The method of claim 1 , wherein the downsampling is performed using a k-nearest neighbor classifier.

8. The fusion of the second scene feature and the text feature includes: concatenating the second scene features with the text features to obtain concatenated features; and applying a self-attention layer to the concatenated features to obtain fused features, wherein the conditional latent is generated based on the fused features.

9. 10. The method of claim 1, further comprising fine-tuning the pre-trained U-Net scene encoder for text- and scene-conditional human motion generation based on losses, including two regularization losses related to the target object category and the target object size.

10. One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause the system to perform operations, including: receiving an input, said input comprising: a 3D point cloud of a scene including a target object; text including natural language instructions associated with the target object; applying a text tokenizer to the text to obtain tokenized text; generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text; generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud; downsampling the first scene features to obtain second scene features; obtaining a conditional latent based on a fusion of the second scene feature and the text feature; predicting a set of motion parameters for a parametric human model moving towards a target object over a specific time period by applying a conditional motion generator to the conditional latents; and obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human body model.

11. 1. A system comprising: a memory for storing instructions; a processor coupled to the memory that executes the instructions to perform a process, the process comprising: receiving an input, said input comprising: a 3D point cloud of a scene including a target object; text including natural language instructions associated with the target object; applying a text tokenizer to the text to obtain tokenized text; generating text features by applying a text encoder of a pre-trained visual language model to the tokenized text; generating first scene features by applying a pre-trained U-Net scene encoder to the 3D point cloud; downsampling the first scene features to obtain second scene features; obtaining a conditional latent based on a fusion of the second scene feature and the text feature; predicting a set of motion parameters for a parametric human model moving towards a target object over a specific time period by applying a conditional motion generator to the conditional latents; and obtaining a 3D human mesh for a plurality of motion frames based on the set of motion parameters and the parametric human body model.