Natural Language Robot Motion Planner

By combining motion dynamics and semantic information, and using natural language descriptors and extended reality interfaces to train variational autoencoders, an interpretable robot motion planner is generated. This solves the problem that robots cannot naturally encode motion and improves the interpretability and collaborative efficiency of motion planning.

CN122299608APending Publication Date: 2026-06-30INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CORP
Filing Date
2025-11-11
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Robots lack the ability to naturally encode, link, or interpret continuous motion using language, which makes it impossible to create and adapt motion plans based on explicit user needs, affecting the interpretability and efficiency of motion planning.

Method used

By combining spatial information of motion dynamics and semantics, using natural language descriptors and extended reality interfaces, variational autoencoders are trained to generate interpretable robot motion planners, integrating linguistic cues and task semantics to generate interpretable motion trajectories.

Benefits of technology

It improves the interpretability and collaborative efficiency of robot motion planning, enhances the trust in human-robot collaboration and the smoothness of task execution, and reduces the cost of dataset collection and annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122299608A_ABST
    Figure CN122299608A_ABST
Patent Text Reader

Abstract

An apparatus includes: a memory configured to store: a first dataset including kinematics data representing multiple human-guided movements of a robot; a second dataset including linguistic descriptors of multiple human-guided movements of the robot; and a processor configured to generate a third dataset based on the first and second datasets, wherein the third dataset includes multiple motion primitives of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Despite advancements in artificial intelligence robotics and large language models (LLMs), robots still lack the ability to naturally encode, link, or interpret continuous motion using language. This limitation prevents robots from creating and adapting motion plans based on explicit user requirements, or generating context-based explanations of the reasons or manner in which motion is performed. Overcoming these challenges could significantly improve interpretability, simplify robot troubleshooting, and increase the efficiency of motion planning in demanding tasks involving humans and AI robots in various collaborative environments. Attached Figure Description

[0002] In the accompanying drawings, similar reference numerals are generally used throughout different views to refer to the same parts. The drawings are not necessarily drawn to scale; rather, the emphasis is usually on illustrating the exemplary principles of this disclosure. In the following description, various exemplary embodiments of this disclosure are described with reference to the following drawings, in which:

[0003] Figure 1 A language-driven, interpretable robot motion planner is described;

[0004] Figure 2 The generation of a corpus for a motion planner is described;

[0005] Figure 3 It depicts users cooperating using an extended reality interface;

[0006] Figure 4 The generation of motion primitives is described;

[0007] Figure 5 It depicts the trajectory generated based on language;

[0008] Figure 6 The motion dynamics components of the input to the variational autoencoder are described;

[0009] Figure 7 The motion planning process, hypothesis creation, and collision-free verification are described.

[0010] Figure 8 This is a flowchart of a motion planner; and

[0011] Figure 9 Depicting in Figure 8 A more detailed version of the creation of the AI-driven hypothetical configuration described in the text. Detailed Implementation

[0012] The following detailed description refers to the accompanying drawings, which illustrate exemplary details and embodiments in which aspects of this disclosure may be practiced by way of illustration.

[0013] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment or design described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0014] Throughout the accompanying drawings, it should be noted that, unless otherwise stated, similar reference numerals are used to depict the same or similar elements, features, and structures.

[0015] The phrases “at least one” and “one or more” can be understood to include a numerical quantity greater than or equal to one (e.g., one, two, three, four, [...] etc.). The phrase “at least one of…” relating to a group of elements can be used herein to mean at least one element from a group of elements. For example, the phrase “at least one of…” relating to a group of elements can be used herein to mean a selection: one of the listed elements, one of a plurality of listed elements, a plurality of individual listed elements, or a plurality of multiples of individual listed elements.

[0016] The terms “plural” and “multiple” in the description and claims clearly refer to a quantity greater than one. Therefore, any phrase explicitly invoking the preceding terms referring to a quantity of elements (e.g., “plural [elements]”, “multiple [elements]”) clearly refers to more than one of the elements. For example, the phrase “multiple” can be understood to include a numerical quantity greater than or equal to two (e.g., two, three, four, five, [...] etc.).

[0017] In the description and in the claims (if any), the phrases “group of…”, “set of…”, “heap of…”, “series of…”, “sequence of…”, “group of…”, etc., refer to a quantity equal to or greater than one, i.e., one or more. The terms “appropriate subset,” “reduced subset,” and “smaller subset” refer to a subset of the set that is not equal to the set, and illustratively refer to a subset of the set that contains fewer elements than the set.

[0018] As used herein, the term "data" can be understood to include information in any suitable analog or digital form, such as being provided as a file, a portion of a file, a collection of files, a signal or stream, a portion of a signal or stream, a collection of signals or streams, etc. Furthermore, the term "data" can also be used to mean a reference to information, for example, in the form of a pointer. However, the term "data" is not limited to the foregoing examples and can take various forms and represent any information as understood in the art.

[0019] For example, the terms “processor” or “controller” as used herein can be understood as any kind of technical entity that allows the processing of data. Data can be processed according to one or more specific functions performed by the processor or controller. Furthermore, a processor or controller as used herein can be understood as any kind of circuit, such as any kind of analog or digital circuit. Therefore, a processor or controller can be or includes analog circuits, digital circuits, mixed-signal circuits, logic circuits, processors, microprocessors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), integrated circuits, application-specific integrated circuits (ASICs), etc., or any combination thereof. Any other kind of implementation of the various functions described in further detail below can also be understood as a processor, controller, or logic circuit. It should be understood that any two (or more) processors, controllers, or logic circuits detailed herein can be implemented as a single entity with equivalent functionality, and conversely, any single processor, controller, or logic circuit detailed herein can be implemented as two (or more) separate entities with equivalent functionality.

[0020] As used herein, “memory” is understood to mean a computer-readable medium (e.g., a non-transitory computer-readable medium) in which data or information can be stored for retrieval. Therefore, references to “memory” included herein can be understood to refer to volatile or non-volatile memory, including random access memory (RAM), read-only memory (ROM), flash memory, solid-state memory, magnetic tape, hard disk drive, optical disk drive, 3D XPoint™, etc., or any combination thereof. Registers, shift registers, processor registers, data buffers, etc., are also included herein under the term memory. The term “software” refers to any type of executable instructions, including firmware.

[0021] As used herein, the term Extended Reality (XR) can generally be understood as an encompassing term that utilizes Augmented Reality (AR), Virtual Reality (VR), Mixed Reality (MR), or any combination thereof. In one example, XR can refer to the use of one or more sensors to detect human body movement (e.g., movement of one or more human limbs) and can be configured to generate sensor data that is used to control a robot to move in response to the movement of one or more human limbs. For example, a pick-up and place robot can be controlled with the aid of XR. In this way, a human moving their arm to simulate picking up an object, moving the object, and placing the object in an orientation different from where the object was picked up can be combined with XR to control a robot to perform similar picking, moving, and placing actions.

[0022] As used in this paper, Rapid Exploratory Random Tree (RRT) refers to one or more algorithms designed to search non-convex high-dimensional spaces by randomly constructing space-filling trees. In this way, the tree can be generated incrementally or progressively, such as by using samples randomly drawn from the search space. RRTs can inherently tend to evolve towards large unsearched regions of the problem. RRTs can be particularly useful for problems with obstacles and differential constraints (e.g., motion dynamics problems), and therefore have many applications in motion planning for autonomous robots. Essentially, many RRTs generate open-loop trajectories for nonlinear systems with state constraints. RRTs can be particularly useful in computing approximate control policies to control high-dimensional nonlinear systems with both state and motion constraints.

[0023] Traditional motion planners for robotic applications ignore or fail to consider task semantics (e.g., adjective cues and non-geometric context) when searching for continuous motion trajectories in variable, semi-structured tasks. Traditional motion planners are geometrically constrained and neglect valuable semantic subspaces. That is, current motion planners do not use semantic cues during exploratory search of the space, which presents an opportunity to improve performance through hints to generate enhanced downstream capabilities.

[0024] The core problem is found during motion hypothesis generation, where only robot joint states are considered, ignoring or disregarding the semantic context of the task. Without using the semantic dimension of the task as guidance, solutions are limited to geometric insights, lacking semantic trajectories (e.g., language-driven trajectories) that help interpret actions to humans. Therefore, without a semantic subspace (e.g., language embeddings) that connects the robot's movement to the meaning of a specific task, motion planning algorithms (whether sampling, optimization, or hybrid) often get stuck in disconnected non-convex regions. This limitation impacts the quality of the solution and can hinder rapid implementation of motion planning. By creating a space that combines both dynamical movement and semantic information, geometrically distant sub-trajectories can be linked to semantically meaningful ones, leading to rapidly interpretable solutions for motion planning. Integrating task semantics during motion planning through linguistic cues (e.g., nouns and verbs) and / or descriptive contextual cues (e.g., adverbs and adjectives) can address significant limitations and provide high-value opportunities.

[0025] Furthermore, it is noteworthy that the robot's joint trajectory planning is uninterpretable and lacks a link between semantic pre-state and semantic post-state. That is, interpretable motion is crucial for effective human-robot collaboration, especially in applications involving dialogue interfaces. In such scenarios, dialogue that interprets the type or attributes of motion, such as those conceptualized by humans, can be expected or even necessary. Extending this point, ensuring that human-robot collaboration includes motion demonstrated by humans or synthesized by the robot, and that motion can be described and exchanged in natural language by all intelligent agents, creates a motor language capability that facilitates productive and trustworthy collaboration between humans, AI robots, and other AI. Moreover, motor language improves various AI tasks through cues. The integration of motor language cueing mechanisms can significantly enhance AI task performance. Therefore, linguistic cues derived from semantic motion trajectories associated with motion dynamics planning can assist other control and perception tasks by providing cues.

[0026] Regarding human-to-robot programming, consider tasks such as describing how to assemble CPU and memory modules onto a motherboard, through demonstration, teleoperation, and / or language. When teleoperating a robot, humans describe each movement in natural language, which adds semantic cues in the form of nouns, adjectives, and adverbs. For example, when reaching and removing a memory module from a memory module tray, a human verbal description might be, "Move quickly over the memory tray, then carefully grasp the module around the corner, lifting it vertically and slowly until it is a few centimeters off the tray." This example highlights how careful manipulation requiring delicate handling and gentle movements can be combined with rapid transport operations to efficiently perform the task at hand.

[0027] In the second example use case, consider an application in the field of service robotics where a human teaches a robot how to open a drawer, grasp a cup, fill a cup with water, and transport the cup to a tray. The task description provided along with the robot's XR teleoperation could be: "Reach for the drawer handle and make a cage-like grasp; and make an compliant pull to accommodate the hinged movement of the door. Reach for the glass through the side and grasp it. Gently slide the glass out and move it to an upright position under the faucet." After the faucet is running and the glass is full, the description continues as follows: "Transport the glass upright to the tray, avoiding violent movements. Gently place the glass on the tray." This use case demonstrates the flexibility of teaching long-duration tasks with implicit action segmentation and explicit task constraints.

[0028] When generating language-driven, interpretable robot motion planners, the following must be considered: (1) how to economically create online robot motion planning that benefits from language-driven semantics using text descriptors and human demonstrations; (2) how to use language cues to help the robot's motion planner generate bimodal motion dynamics plans that consider both task-related and embodied physical and semantic aspects; and (3) how to clearly describe these plans in detail (e.g., stepwise) or as a whole (e.g., holistic) using human natural language, and how to interpret them for other downstream AI tasks using joint trajectories and language (e.g., using text or speech).

[0029] With the increasing prominence and importance of human-robot applications, the core problem of fully generating and utilizing motion planning in motion dynamics, leveraging compact and efficient language descriptors during hypothesis generation, which is enhanced by probabilistic generative AI. In such tasks, language plays a crucial role in both input and output, as well as in describing motion (e.g., gentle, slow, robust, responsive). Using the principles disclosed herein, motion-to-sentence and sentence-to-motion combinations of AI can be enhanced by employing natural human-robot interfaces. In this way, the proposed language-driven motion planning and training methods can bridge the AI ​​gap between small language embedding models and robot motion, thereby facilitating intuitive human-robot collaboration.

[0030] One benefit of the framework disclosed herein is the guarantee of economical motion. In this way, motion hypotheses contained in a customized latent space formed according to the task specification are collected via multimodal human demonstrations (e.g., using natural language, using extended reality controllers). This can be used for data production of variational autoencoders (VAEs), which are trained to synthesize data via probabilistic generative models, thereby avoiding costly dataset collection and annotation (e.g., datasets entirely generated by humans and / or manually labeled). The framework disclosed herein enables the training of context-sensitive, scalable, and economically generated motion planning data via low-dimensional probabilistic motion primitives (MPs). As will be described in more detail, these probabilistic MPs can be created based on human demonstrations, such as through the use of XR devices.

[0031] Improved motion dynamics and semantic exploration performance can be achieved using (compact) probabilistic neural networks (such as VAEs), where joint semantic and motion dynamics batch inference can be obtained during inference through the parallelization of hypothesis generation. This can be achieved by using a single encoding process and multiple variational decoding processes.

[0032] A robot's ability to rapidly and coherently describe motion in natural language provides improved trust in the context of human-robot interactions. This can be observable, for example, by demonstrating fluent and predictable psychomotor skills in responding to different scenarios, manifesting as thought processes in concrete actions. That is, motion planning with a well-defined linguistic framework can link each joint step in the motion plan, thus providing a detailed or generalized description of each of multiple actions.

[0033] Attention has now turned to the problem of how to create online effective kinematic and dynamic (e.g., motion dynamics) robot motion planning that integrates geometric and semantic information to produce collision-free and coherent joint trajectories, accompanied by natural language descriptions that are interpretable at individual steps and as the entire trajectory.

[0034] Figure 1 A language-driven, interpretable robot motion planner is described. Figure 1 The top section depicts the offline phase, which includes: a section for determining language semantic embeddings and thus labeled "Language Semantic Embedding Module"; and a section for determining human descriptive demonstrations, labeled "Human Descriptive Demonstration" section. Figure 1 The bottom section depicts the online phase, which is labeled "Language-Driven Robotic Motion Planner".

[0035] Starting in the offline phase, language embeddings are created from robot manuals and application documentation using unsupervised learning, thus converting sentences into feature vectors. Human-robot task co-execution (such as via an XR interface) can be used to generate motion primitives, thereby creating a large, annotated dataset. These datasets can be used to train a variational autoencoder to map the robot's state to the next step, while being biased by language descriptors and random seeds, thus linking the semantic model and the motion dynamics model. Starting in the offline phase, the process begins with the creation of language embeddings (such as from robot manuals and / or application corpora 102). These manuals and process descriptions can be referred to as, or understood as, a corpus (a1). This corpus can be processed using an unsupervised learning model (a2) (referred to herein as an unsupervised learning sentence encoder 104) to convert short descriptive sentences S into features in the semantic space R. k eigenvectors within (Depicted as a corpus-semantic statement to vector 106). This can be performed on any single robot, any group of robots, one or more robots of any category, or other robots. In some cases, the corpus may refer to the robot manual for any particular robot and / or any other documents related to a particular robot, but may be applicable to other similar robots. The specific model may be a statement encoder, which may be an AI that encodes text into high-dimensional vectors. These vectors can then be used for text classification, semantic similarity, clustering, and other natural language tasks. The statement encoder can be selectively and / or trained to receive text containing statements, phrases, or short paragraphs, and to convert the text containing statements, phrases, or short paragraphs into multi-dimensional vector outputs. Any statement encoder or any other model capable of performing the tasks disclosed herein can be used in this regard. Specific statement encoders are not discussed at this stage. Rather, those skilled in the art will recognize how to determine a suitable statement encoder for the tasks disclosed herein.

[0036] Now we move on to the next part of the offline phase, publicly presented as a “human descriptive demonstration,” where humans can perform multiple human or robotic tasks collaboratively (labeled as a small dataset: ηXR-based human motion narrative demonstration S112), such as controlling a robot to perform specific tasks (e.g., picking up objects, moving objects, connecting objects, separating objects, releasing objects, etc.). For example, XR can be used for these collaborative tasks, such that one or more sensors are attached to or otherwise configured to capture or detect the human operator's movements. In this way, robot movement can be controlled in a more human-like manner. For example, a human can control a robotic arm to pick up an object in a first position and place it in a second position. Humans can achieve this using smooth, fluid motion, and because the machine is controlled by a human, the robot can also perform such picking and moving tasks with similar fluidity and smoothness, or, in any case, with greater fluidity than could be achieved by programming movement at the joint level. Unless otherwise stated, while it may be difficult to program a robot to perform tasks in a fluid, human-like manner, it is capable of doing so when controlled based on the movement of human limbs. Humans can perform collaborative tasks multiple times, thereby generating a small dataset of human-controlled robot movements.

[0037] Furthermore, humans can use one or more natural language descriptors to narrate the cooperative task. For example, this could include “carefully pick up the object,” “move the object slowly and smoothly,” or “gently put down the object,” or any other number or type of task descriptors. The narration can include one task descriptor for a single task, or multiple task descriptors for a single task. Task descriptors can be single words, a few words, phrases, statements, or multiple statements. That is, human-robot cooperative task execution (b1) as described above might be accompanied by a brief narration, which can provide clues to compose a generative motion primitive (b2) capable of synthesizing a large number of trajectory samples κ (labeled as variations in κ per demonstrated bendable motion primitive sample 114) by employing probabilistic motion primitives to simulate 6D offsets from the start and end points of the robot's end effector.

[0038] Based on these motion primitives, a large dataset (b3) with relevant linguistic annotations can be generated, i.e., η·κ samples 116 for each motion segment. Combining this data and linguistic embedding, a VAE can be trained to map from the instantaneous dynamic state of robot T to the next discrete step T′ biased by an ε-linguistic descriptor (b4) and a Gaussian distributed random seed σ (described as unsupervised learning decoupled VAE 118). This generative model The input is mapped to a distribution that fuses the linguistic-semantic manifold with a robot-specific trained motion dynamics model. The key is: within the semantic motion dynamics space, (via σ) the constraint triples [T] on feasible paths linking... O ,T f ,T i The ability to sample different robot task configurations within [ ] |{T′}|.

[0039] The AI ​​hypothesis generator (described as a joint language-driven semantic & motion dynamics hypothesis generator 120) is a technology enabler. In this way, the role of the language-driven hypothesis generator is to generate potential next steps in motion planning in batches, as shown in the online phase. At runtime, the input is the first or current 6D pose T of the end effector. O (c1), and the target or final pose T f (c2). At each step T along the trajectory i At this point, multiple checks may be needed to ensure collision-free movement (e.g., movement where the robot does not collide with anything in its environment). Regarding the environment and the robot itself (c3), computationally, note that this allows for the simultaneous generation of multiple output hypotheses in batch form, which requires a single encoding process for multiple variational decoding generation (b5). This results in at each step (T i ,ε iThe language embedding at the point () uses interpolation and generalization capabilities to describe motion. This can be important, at least because it allows for the creation of motion plans and their description at each step. In this way, both the user and the AI ​​agent can derive more explicit semantic information from the trajectory. These added cues inform the secondary task being solved, the motion style, the overall process logic, and hidden constraints, thus revealing a semantically endowed motion plan with superior control and interface outcomes. This capability transforms human-robot collaboration, enabling large-scale mutual knowledge transfer between humans and AI.

[0040] Now we turn our attention to how to ensure that robot motion is interpretable by humans, specifically: demonstrating generalization ability while maintaining task-specific motion signatures, thereby facilitating trusting and intuitive movement through cost-effective and reliable human-robot collaboration.

[0041] During the online phase (see Figure 1 In the “online phase” of motion planning, the first pose 132 and the target pose 134 (e.g., the first 6D pose and the target 6D pose) are inputs, and collision-free movement is ensured at each step. As illustrated herein, multiple hypotheses 136 can be generated in a single encoding process, providing a language embedding to describe the motion at each step. This enhances human-robot collaboration by interpreting actions, providing more information about the task, motion style, process logic, constraints, and meaningful paths. Unlike existing LLMs and motion planners that operate at high symbolic levels, the principles and methods disclosed herein work at successive sub-symbolic levels, bridging the gap between motion and language with minimal human demonstration. This results in shallower, more energy-efficient, and more timely models compared to LLMs, allowing for both geometric and dynamic feasibility in online and onboard motion planning.

[0042] In the offline phase, unsupervised learning is used to create language embeddings from robot manuals and application documentation, transforming short descriptive statements (from a string of sentences to short paragraphs) into k-dimensional feature vectors within the domain semantic space. Human-robot task co-execution (immersive operation) via an XR (teleoperation) interface (accompanied by short speech-to-text narration) enables the compilation of motion primitives for synthetic feasible trajectories; i.e., language plus motion as learning samples. This process economically creates a large-scale dataset with rich language annotations for specific robot embodiments and kinematics. The dataset is used for unsupervised training with a variational autoencoder (VAE), mapping the robot's dynamic states to the variational next step, biased by language descriptors and randomly generated seeds. This model training links the linguistic-semantic manifold with the robot's specific (kinematic and dynamic) model, allowing for the generation of samples for different task configurations within the semantic-motion dynamics space. In the online phase: i) language cues, such as "smoothly pick up the tray," ii) the initial 6D pose of the end effector, and iii) the target pose of the end effector are the three inputs. At each planning / search step along the trajectory, check (for example, see...) Figure 4 Ensure collision-free movement. Note that the above method allows for the simultaneous generation of multiple output hypotheses; therefore, a single encoding process is required for multiple decoding generation. The resulting plan provides language embeddings at each step, thus providing a fine-grained description of motion with interpolation and abstract generalization capabilities. This allows the robot to interpret its actions at each step via natural language, enhancing human-robot collaboration and providing more information about secondary tasks, motion style, process logic, hidden constraints, and meaningful paths that convey the task and robot intent.

[0043] Figure 2 The generation of language embeddings from robot manuals and application documents (collectively referred to as corpus W) is described above for robot manuals and / or application corpus 102. First, various documents, such as robot manuals and application corpora (e.g., documents related to one or more tasks to be performed), can be selected 202. The selection of documents used in this way is flexible to some extent in both quantity and subject matter, and the selected documents can be applied to a variety of use cases and robot types. This data is encoded 204 such that ∈ ∈ R k For example, the information gathered from these texts allows the underlying processor model to create combinations of deep neural network (DNN) structures, such as, but not limited to, Skipgram embeddings 206. This means that the corpus W contains the vocabulary H. k Statements within It is possible to use L1, L2, cosine, and other metrics to be semantically related. This means that for f(S) i )→εi ∈R k and f(S) i )→ε i ∈R k ,exist In the case of S, although i and S j They have different expressions, but S i and S j Similar robotic movements—actions (changing the sequence of words or using synonyms)—are described. This can be particularly important for constructing a semantic space, where locality implies the meaning of sentence embeddings. In such a space R... k In this context, traversing point clusters defines the gradual semantic changes. In practice, this can be trained once using an unsupervised DNN and subsequently applied to various robot and task types contained in the corpus. Note that this embodiment-independent process... Figure 1 The topmost part corresponds to this. It is also worth noting that such embeddings can be created using Retrieval Automatic Generation (RAG) and Large Language Models (LLM). However, the ability to achieve this in small-scale computational developments on a board would require smaller language embedding models.

[0044] Regarding the task collaboration described above, one or more human-robot task collaborations can be performed, such as by using one or more XR teleoperation interfaces. Figure 3 The diagram depicts a user performing cooperative actions using an XR teleoperation interface, where the user controls the robot to move in a manner corresponding to the user's human movements. As illustrated in the figure, these cooperative actions are accompanied by brief narrations (e.g., from the user) that provide necessary linguistic cues for the task. That is, the human issues movement commands via controllers and also uses voice. The human can include style and / or intent, such as slow, cautious, smooth, etc. In this way, the human user acts as a bridge between linguistic embedding and the robot's combined kinematic dynamics in a simple and natural teleoperation demonstration.

[0045] While human-controlled motion has value as a linguistic descriptor, generating sufficient experimental data by humans may be impractical or undesirable. That is, due to the cost of human annotation when describing such a set of action segments (the entire task is a sequence of multiple segments), this subprocess is ideally executed on a small scale, such as 2-8 demonstration trials per application domain, although fewer or more trials could be used. If an appropriate number of trials are performed, it can be expected that the raw data from these trials can be transformed into a flexible probabilistic model to generate large training datasets on demand in real-time during VAE training.

[0046] Each narrative demonstration will be influenced by a certain amount of human variability, creating unique motion vectors, even when the task used for the demonstration is the same as in other demonstrations. In other words, a human repeating the same pick-and-place operations will inevitably change the operation in terms of position or velocity each time. Of course, users can try teleoperating the robot to understand the embodied limitations of the physical system. This practice can help achieve a more consistent mapping of adjectives in the narrative. For example, when a user commands the robot to move along a trajectory described as "fast," the robot can perform movement at 60% of its maximum achievable speed at its fastest point. Higher attributes (such as "maximum speed" or "maximum / minimum acceleration") can correspond to 10%–90% of the robot's capabilities, thus providing a tolerance margin for trajectory generation. The lexical scale can be optionally constrained based on any of the following: industry, use case, human proximity, or target object, without affecting the overall algorithm and interface.

[0047] Generative motion primitives can simulate novel 6D end-effector poses (scene offsets and velocities) from the start and end points of a robot's end-effector using probabilistic motion primitives, synthesizing numerous trajectory samples κ. The challenge lies in generating trajectories adapted to different start and end positions while capturing the narrative of the demonstration. This can be achieved by dividing the problem into two parts. First, generating safe motion dynamic trajectories. Second, generating a natural language description of the generated trajectories.

[0048] Given a set of demonstrations, the motion planner can generate three distinct MPs, representing the average, upper bound, and lower bound of the demonstration trajectory, as shown in... Figure 4 This is depicted in Figures 402-408, which illustrate the demonstration and then divide the upper and lower boundaries of the movement of the demonstration by dimension. For example, Figure 402 depicts the average movement and its upper and lower boundaries in the first dimension; Figure 404 depicts the average movement and its upper and lower boundaries in the second dimension; Figure 406 depicts the average movement and its upper and lower boundaries in the third dimension; and Figure 408 depicts the average movement and its upper and lower boundaries in the fourth dimension. These can be based on multiple motion primitives as depicted in Figure 410. At runtime, the motion planner can solve an optimization problem to generate a controller that guarantees adherence to the average trajectory while always remaining within the boundaries. Depending on the parameters of the model, the obtained probability MP can be adjusted in different ways, for example, to closely follow the reference at the cost of increased control effort, or vice versa. Therefore, by exploring the design space of the probability MP, various trajectories that follow the safety boundaries can be generated from the demonstration. To do this effectively, different methods can be used, such as Monte Carlo search or Bayesian optimization.

[0049] For the second offline component, an LLM (Local Language Model) can be used to describe the trajectories. To achieve this efficiently, a scene graph is constructed that captures the key contextual information needed to describe the robot's movement within the scene. Assuming the robot moves at the same rate whether an action is labeled as cautious movement or slow movement, the relevant reviewer can distinguish between the two cases. The latter is useful when the robot is moving in free space (i.e., without obstacles nearby), while the former makes sense when the robot is navigating in cluttered space near fragile objects. Therefore, for demonstration of each label, a frame sequence is attached, in which the robot's state relative to scene elements in different views is abstracted. Subsequently, a new random scene is synthesized, and the robot is moved from a random start condition to a target condition using a DMP (Discrete Language Model). The frame sequence can be obtained from the generated scene as described above, and can then be input into the LLM along with offline examples. The LLM can then be asked to generate a narrative based on the examples. This process is repeated to create a dataset consisting of pairs of robot trajectories and corresponding natural language descriptors. This process allows for the economical and rapid creation of large datasets with relevant language annotations, resulting in η·κ samples for each motion segment. Figure 5 The trajectory generated from language as described above is depicted. In this example, the robot is instructed to perform a picking action slowly, and the motion planner generates λ samples based on a single motion primitive (L-PMP refers to language-based trajectories based on the main motion primitive).

[0050] Unsupervised generative AI training can then be performed using a motion dynamics hypothesis generator (such as by using a VAE). Figure 5 The training and inference of a VAE used for motion planning are described. This training typically involves creating motion dynamics hypotheses, such as by using... Figure 3 and Figure 4 The data generated during the depiction process.

[0051] Figure 6 The kinematic components of the input to the variational autoencoder were described, and Figure 7 The motion planning process, hypothesis creation, and collision-free verification are described. Figure 7 Eight steps for processing language prompt (0)601 are also shown. Go to Figure 6 This demonstrates a rigid transformation with six degrees of freedom (e.g., 6DOF, 6D) from the first frame (frame A) to the final frame (frame E). First, in step 602 (also understood as frame A), the first 6D end effector pose and its associated joint angles n>=6 are provided by the dataset. Second, in step 604, the final 6D pose of the end effector is also provided. Third, in step 606, at each step of the planning, the transformation... This is used to calculate the first [function] based on the motion primitives by numerical differentiation or by minimizing constraints and ensuring numerical stability. Second derivative and second Fourth, in step 608, these motion dynamics elements are combined with user-defined language embeddings ∈ to define the input points in the semantic motion dynamics space.

[0052] Go to Figure 7 Because VAEs use stochastic backpropagation (e.g., using reparameterization 704) for training 702, once training is complete, multiple decoding processes can be applied to a single input point, thus generating multiple hypotheses for the RRT, like the motion planner 708. One advantage of VAEs is their use of linguistic semantic embeddings during the search phase, which modulate the motion dynamics trajectory by shortening or lengthening the motion dynamics distance. This results in non-obvious geometric paths that preserve the semantic signature of the motion, as illustrated in 710 by an extended hyperellipsoid. This extended hyperellipsoid is provided as an illustrative concept, and it is explicitly stated that the shape can vary depending on the situation. Finally, as in... Figure 8 As shown, motion planning is a sequence of motion dynamic configurations with collision checks. This includes language embedding at both the step and trajectory levels, which can be very valuable for various downstream tasks.

[0053] Now will be used Figure 8 Explain in more detail Figure 1 During the online phase, Figure 8 A motion planner is described, which is itself a module implementing a sample-based approach. While several options are available, RRT may be particularly suitable. This motion planner is capable of generalizing the generated trajectory in an embedded semantic-linguistic space throughout the trajectory, and / or the motion planner can reduce verbosity by progressively interpreting the generated trajectory by selecting curvature inflection points in the feature space. Using a hypothesis generator (e.g., unsupervised probabilistic learning in the VAE described above), the motion planner generates hypotheses for the next state based on user prompts. The motion planner checks (e.g., evaluates) these hypotheses regarding velocity, acceleration, joint constraints, and blank space in the scene, and creates possible state-space paths. These checks and progressive tree graph construction inherit the benefits of RRT regarding convergence and on-the-fly solutions (e.g., algorithms or solutions that can return valid results even if interrupted before completion, such that the longer the algorithm is allowed to run, the higher the quality of the results). Finally, the motion planner can represent the trajectory solution as a generalized language descriptor that can be decoded into various human languages ​​as needed. Given the exploratory nature of the enhanced RRT, users do not need to define coordinates in task space or joint space at any point. While it is possible to check... Figure 8Deep neural networks (DNNs) are used to add clues about the occupant, but the motion planner is validating... Figure 8 The assumptions created in the model still require collision testing (818). Therefore, margin distance is considered a language-derived clue and an additional feature of this implementation.

[0054] Figure 8 A stepwise sampling-based method for selecting the motion planner is described. The motion planner receives data from robot T. O The initial position / pose (e.g., the initial or current 6D pose of the end effector) and motion language cue S, where S can be a statement or short paragraph defining any of the style, goal, or attribute of the desired motion planner 802. The motion planner processes the initial position T. O The motion planner receives motion language cues S (as described in detail above) to generate motion vector 804. Motion vector 804 is then added to the motion tree as initial configuration 806. The motion planner also receives the target position T. f Target location T f This can be understood as the target position / pose or the final position / pose 810. The motion planner calculates the distance from the leaf to the target position T. f The kinematic distance is 808.

[0055] Regarding trajectory planning, the device performs multiple collision checks, where, in each step T iAt this point, the motion planner performs multiple checks along the trajectory to ensure collision-free movement with respect to the environment (e.g., the robot does not collide with obstacles or external objects) and the robot itself (e.g., the machine does not move in a way that causes it to collide with itself). If the target is reached 812 without collision, the motion planner can then apply a language-based motion abstraction model to generate a brief natural language description 817 of the movement derived from step 812, and motion planning is complete 815. Assuming the target has not been reached, the motion planner can evaluate whether a timer has expired 814 (e.g., the motion planner can have a maximum time allocation for the selected trajectory to avoid undesirable delays). If time has elapsed, the process may then fail 816. However, assuming the timer has not expired, the motion planner can then implement the model to generate an AI-driven hypothetical configuration 818. After generating the hypothetical configuration at 818, the motion planner considers whether the motion hypothetical configuration is kinematically feasible (e.g., corresponding to available joint movements, acceptable velocities and / or accelerations, etc.) and within blank space (e.g., collision-free) 820. If this is true, the motion planner then adds this hypothesis configuration to the RRT tree and stores the corresponding dynamic embedding at 822. Afterward, the hypothesis failure counter is reset to 0 at 810, and the motion dynamics distance is calculated at 808. However, if the motion planner determines at 820 that the hypothesis configuration is not dynamically feasible or is not in the blank space, the motion planner then increments the hypothesis failure counter at 824. The hypothesis failure counter can have an allowable maximum value, and the motion planner can determine whether the maximum value has been reached at 826. If the maximum value has been reached, the motion planner can then lock ∈′ i And accordingly, a hypothesis configuration is created 828. If a request to exit or a timeout occurs 830, the process then fails. Otherwise, the motion planner determines whether the new hypothesis corresponds to a space that is feasible in motion dynamics 820, and the process continues as described above. If the hypothesis counter does not reach its threshold at 826, the process then returns to the generation of the AI-driven hypothesis configuration at 818.

[0056] Note that the motion planner is capable of batch hypothesis generation, whereby the motion planner generates multiple output hypotheses simultaneously in batches, which may require only a single encoding process for multiple variational decodings.

[0057] The resulting plan is generated at each step (T) i ,ε iThe system provides language embeddings, enabling the description of motion with interpolation and generalization capabilities. This allows robots to create motion plans and describe their actions at each step, providing users and other AI agents with more information about any of the secondary tasks, motion styles, overall process logic, hidden constraints, and optimal paths. This capability can significantly enhance human-robot collaboration and mutual knowledge transfer.

[0058] Figure 9 Depicting as in Figure 8 This is a more detailed version of the AI-driven hypothesis configuration creation 818 depicted in the diagram. Creating interpretable robot motion planning involves language (semantics), kinematics (geometry), and dynamics (mass, force & acceleration) to describe meaningful actions beyond just joint trajectories and occupied space. This robot motion planning can be created by generating and selecting partial motion segments 902 (e.g., used as the next configuration of the hypothesis in a sampling search) in a semantic motion dynamics feature space using probabilistic generative artificial intelligence (AI) (e.g., a probabilistic encoder 904), while enhancing deterministic fast exploratory tree (RRT) validation. The probabilistic encoder 904 determines the mean (μ) 906, perturbation factor (∈) 908, and standard deviation (σ) 910 combined according to z = μ + σ ∠ 912, where z 912 is a sample of a probability distribution. This can be understood as a reparameterization trick of VAR, which allows differentiable sampling in a stochastic model. The probabilistic decoder 914 generates a modified hypothesis x′ 916 based on z. This enables collision-free, dynamically achievable robotic motion, accompanied by explicit natural language descriptions. This enhances convergence, reveals asymptotic guarantees, and improves trust between humans and robots during task collaboration and knowledge transfer. Dimensionality can vary within a variable molecular space.

[0059] As the abstract generalization disclosed in this article refers to shortening along the path ∈′0,...∈′ m The process of collecting a set of text embeddings to create a representation of the sequence ∈′0,...∈′ m The most important information block in the summary ∈ ∑ Existing summarization methods can be categorized into two types: extractive and abstractive. Abstractive summarizers generate novel text fragments that convey the most prominent concepts prevalent in the source.

[0060] Additionally, in Figure 8 The process described can be optimized by extending it with a node reconnection feature based on a cost function. In one approach, the cost function for each node can be a convex combination of the distance to the tree and the similarity of the trajectory with respect to the command. This similarity can be determined by evaluating the decoder components. By reconnecting nodes according to this cost function, the solution asymptotically converges to the most similar trajectory.

[0061] In some cases, it may be expected that robots will actually teach humans various physical tasks. In such cases, teaching can be aided by examples and / or narration. For this, the roles of teacher and student are reversed. Assuming the robot has already learned to use appropriate movement styles to perform tasks, it can then demonstrate to the human how the movements must be performed to meet specific task constraints, while also describing the movement being performed in natural language at each step. Continuing with the motherboard assembly example above, the robot can perform the actions while describing them: “Gently grasp the memory module from the top corner,” “Carefully lift the memory module until it leaves the tray,” “Quickly transport the memory module over the target memory slot,” “Precisely align the memory module with the slot and gently insert it,” “Press firmly from the top corner to latch the memory module.”

[0062] Furthermore, since the initial, final, and intermediate conditions of the task are unknowable to the robot, human motion can be translated into a linguistic description, and natural language feedback can be provided. This can be achieved by using existing human motion capture vision systems to detect human trajectories, obtain a description of the human trajectory, and calculate the distance from the executed trajectory to the desired trajectory. Therefore, natural language feedback can be communicated, such as "lift the memory higher above the tray" or "align the memory with the slot, then slowly insert the memory into the slot."

[0063] Additional aspects will be revealed through examples:

[0064] In Example 1, an apparatus includes: a memory configured to store: a first dataset including kinematic data representing multiple human-guided movements of a robot, and a second dataset including linguistic descriptors of multiple human-guided movements of the robot; and a processor configured to generate a third dataset based on the first and second datasets, wherein the third dataset includes multiple motion primitives of the robot.

[0065] In Example 2, the apparatus of Example 1, wherein generating a plurality of motion primitives of the robot includes: generating a plurality of end effector poses based on the starting point of each of a plurality of human-guided movements of the robot and the corresponding ending point of each of the plurality of human-guided movements of the robot, wherein each end effector pose of the plurality of end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

[0066] In Example 3, the apparatus of Example 2, wherein the three-dimensional end effector orientation includes: end effector roll, pitch and yaw; end effector rotation matrix; or end effector quaternion.

[0067] In Example 4, the apparatus of either Example 1 or Example 3 further includes: a fourth dataset comprising multiple multidimensional feature vectors of the robot, wherein the processor is further configured to: generate a fifth dataset based on the third and fourth datasets, wherein the fifth dataset comprises multiple motion trajectories of the robot.

[0068] In Example 5, the apparatus of Example 4 is used, wherein generating multiple motion trajectories includes: the processor generating an upper boundary, a lower boundary, and an average value of movement corresponding to a first dataset; and generating motion trajectories that follow a three-dimensional space around the average value while remaining between the upper and lower boundaries.

[0069] In Example 6, the apparatus of Example 4 or Example 5 further includes: a processor executing a model to generate a language description of each of the plurality of motion trajectories, and labeling each of the plurality of motion trajectories with the corresponding language description.

[0070] In Example 7, the apparatus of any one of Example 4 or Example 6 further includes: a probabilistic model, wherein the processor is further configured to execute the probabilistic model to generate a plurality of motion dynamics hypotheses based on the fifth data.

[0071] In Example 8, the apparatus of Example 7 is provided, wherein each of the plurality of kinematics assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

[0072] In Example 9, the apparatus of Example 7 or Example 8, wherein each of the plurality of kinematic assumptions includes data corresponding to the constraints of joints in the plurality of joints of the robot.

[0073] In Example 10, the apparatus of any one of Examples 7 to 9, wherein each of the plurality of kinematics hypotheses includes: data corresponding to a language descriptor among the plurality of language descriptors.

[0074] In Example 11, the apparatus of any one of Examples 7 to 10, wherein the probabilistic model is a variational autoencoder configured to perform unsupervised learning on the fifth data.

[0075] In Example 12, the apparatus of any one of Examples 7 to 11, wherein the language descriptor is a first language descriptor, further includes: first multidimensional pose data representing an initial pose of the robot; second multidimensional pose data representing a target pose of the robot; and a second language descriptor, wherein the processor is further configured to: select a kinematic hypothesis from a plurality of kinematic hypotheses based on the first multidimensional pose data, the second multidimensional pose data, and the second language descriptor.

[0076] In Example 13, the apparatus of Example 12 is used, wherein the second language descriptor corresponds to the embedding of feature vectors in a plurality of multidimensional feature vectors of the robot.

[0077] In Example 14, the apparatus of Example 12 or Example 13, wherein the kinematics assumption selected from a plurality of kinematics assumptions includes: based on any one of the following: velocity limits of robot joints, acceleration limits of robot joints, range of motion of robot joints, or the absence of a blank space for performing kinematics assumptions, the kinematics assumption is selected.

[0078] In Example 15, the apparatus of any one of Examples 2 to 14, wherein generating the motion primitives of the robot includes: generating a velocity or acceleration of movement based on a timestamp indicating the timing of a first position and the timing of a second position after the first position.

[0079] In Example 16, the apparatus of Example 14 or Example 15, wherein the processor is further configured to use the model to generate a natural language description of motion corresponding to a selected kinematics hypothesis.

[0080] In Example 17, the apparatus of Example 16 is used, wherein the processor is configured to cause the robot to generate audible signals with natural language descriptions and simultaneously perform motions corresponding to selected motion dynamics assumptions.

[0081] In Example 18, the device of either Example 1 or Example 17 is configured as part of a robot or a server.

[0082] In Example 19, an apparatus includes: a memory configured to store: a first dataset including kinematic data representing multiple human-guided movements of a robot, and a second dataset including linguistic descriptors of multiple human-guided movements of the robot; and a processor configured to generate a third dataset based on the first and second datasets, wherein the third dataset includes multiple motion primitives of the robot.

[0083] In Example 20, the apparatus of Example 19, wherein generating motion primitives of the robot includes: generating a plurality of end effector poses based on the starting point of each of a plurality of human-guided movements of the robot and the corresponding ending point of each of the plurality of human-guided movements of the robot, wherein each end effector pose of the plurality of end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

[0084] In Example 21, the apparatus of Example 20, wherein the three-dimensional end effector orientation includes: end effector roll, pitch and yaw; end effector rotation matrix; or end effector quaternion.

[0085] In Example 22, the apparatus of any one of Example 19 or Example 21 further includes: a fourth dataset comprising multiple multidimensional feature vectors of the robot, wherein the processor is further configured to: generate a fifth dataset based on the third and fourth datasets, wherein the fifth dataset comprises multiple motion trajectories of the robot.

[0086] In Example 23, the apparatus of Example 22, wherein generating multiple motion trajectories includes: the processor generating an upper boundary, a lower boundary, and an average value of movement corresponding to a first dataset; and generating motion trajectories that follow a three-dimensional space around the average value while remaining between the upper and lower boundaries.

[0087] In Example 24, the apparatus of Example 22 or Example 23 further includes: a processor executing a model to generate a language description of each of the plurality of motion trajectories, and labeling each of the plurality of motion trajectories with the corresponding language description.

[0088] In Example 25, the apparatus of any one of Example 22 or Example 24 further includes: a probabilistic model, and wherein the processor is further configured to execute the probabilistic model to generate a plurality of motion dynamics hypotheses based on the fifth data.

[0089] In Example 26, the apparatus of Example 25, wherein each of the plurality of kinematics assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

[0090] In Example 27, the apparatus of Example 25 or Example 26, wherein each of the plurality of kinematic assumptions includes data corresponding to the constraints of joints in the plurality of joints of the robot.

[0091] In Example 28, the apparatus of any one of Examples 25 to 27, wherein each of the plurality of kinematics assumptions includes: data corresponding to a language descriptor among the plurality of language descriptors.

[0092] In Example 29, the apparatus of any of Examples 25 to 28, wherein the probabilistic model is a variational autoencoder configured to perform unsupervised learning on the fifth data.

[0093] In Example 30, an apparatus includes: a memory configured to store: first multidimensional pose data representing an initial pose of a robot, second multidimensional pose data representing a target pose of the robot, a language descriptor for the desired motion of the robot, and a plurality of motion dynamics hypotheses for the movement of the robot; and a processor configured to select a motion dynamics hypothesis from the plurality of motion dynamics hypotheses based on the first multidimensional pose data, the second multidimensional pose data, and the language descriptor.

[0094] In Example 31, the apparatus of Example 30 is used, wherein the language descriptor corresponds to the embedding of feature vectors in a plurality of multidimensional feature vectors of the robot.

[0095] In Example 32, the apparatus of Example 30 or Example 31, wherein the kinematics assumption selected from a plurality of kinematics assumptions includes: a velocity limit of the robot's joints, an acceleration limit of the robot's joints, a range of motion of the robot's joints, or the absence of a blank space for performing the kinematics assumption, the kinematics assumption is selected.

[0096] In Example 33, the apparatus of any one of Examples 31 to 32, wherein generating the motion primitives of the robot includes: generating a velocity or acceleration of movement based on a timestamp indicating the timing of a first position and the timing of a second position after the first position.

[0097] In Example 34, the apparatus of Example 32 or Example 33, wherein the processor is further configured to: use the model to generate a natural language description of motion corresponding to the selected kinematics assumptions.

[0098] In Example 35, the apparatus of Example 34 is provided, wherein the processor is configured to cause the robot to generate audible signals with natural language descriptions and simultaneously perform motions corresponding to selected motion dynamics assumptions.

[0099] In Example 36, a non-transitory computer-readable medium includes instructions that, when executed by a processor, cause the processor to: generate a third dataset based on a first dataset and a second dataset, the first dataset including kinematics data representing multiple human-guided movements of a robot, the second dataset including language descriptors of multiple human-guided movements of the robot, wherein the third dataset includes multiple motion primitives of the robot.

[0100] In Example 37, the non-transitory computer-readable medium of Example 36, wherein the instructions are configured to cause a processor to generate a plurality of motion primitives of a robot, including: the instructions being configured to cause the processor to generate a plurality of end effector poses based on the starting point of each of a plurality of human-guided movements of the robot and the corresponding ending point of each of the plurality of human-guided movements of the robot, wherein each of the plurality of end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

[0101] In Example 38, the non-transitory computer-readable medium of Example 37, wherein the three-dimensional end effector orientation includes: end effector roll, pitch, and yaw; the rotation matrix of the end effector; or the quaternion of the end effector.

[0102] In Example 39, the non-transitory computer-readable medium of any of Examples 36 or 38 further includes: a fourth dataset comprising multiple multidimensional feature vectors of the robot, wherein the command is further configured to cause the processor to: generate a fifth dataset based on the third and fourth datasets, wherein the fifth dataset comprises multiple motion trajectories of the robot.

[0103] In Example 40, the non-transitory computer-readable medium of Example 39, wherein the instructions are configured to cause the processor to generate multiple motion trajectories, including: the instructions are configured to cause the processor to generate an upper boundary, a lower boundary, and an average value of movement corresponding to a first dataset; and to generate motion trajectories that follow a three-dimensional space around the average value while remaining between the upper and lower boundaries.

[0104] In Example 41, the non-transitory computer-readable medium of Example 39 or Example 40, wherein the instructions are further configured to cause the processor to: execute the model to generate a language description of each of the plurality of motion trajectories, and to label each of the plurality of motion trajectories with the corresponding language description.

[0105] In Example 42, the non-transitory computer-readable medium of any of Example 39 or Example 41 further includes: a probabilistic model, wherein the instructions are further configured to cause the processor to: execute the probabilistic model to generate a plurality of motion dynamics hypotheses based on the fifth data.

[0106] In Example 43, the non-transitory computer-readable medium of Example 42, wherein each of the plurality of kinematics assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

[0107] In Example 44, the non-transitory computer-readable medium of Example 42 or Example 43, wherein each of the plurality of kinematics assumptions includes data corresponding to the constraints of joints in the plurality of joints of the robot.

[0108] In Example 45, a non-transitory computer-readable medium of any of Examples 42 to 44, wherein each of the plurality of motion dynamics hypotheses includes data corresponding to a language descriptor among the plurality of language descriptors.

[0109] In Example 46, a non-transitory computer-readable medium of any of Examples 42 through 45, wherein the probabilistic model is a variational autoencoder configured to perform unsupervised learning on the fifth data.

[0110] In Example 47, the non-transitory computer-readable medium of any of Examples 42 to 46, wherein the language descriptor is a first language descriptor, further includes: first multidimensional pose data representing an initial pose of the robot; second multidimensional pose data representing a target pose of the robot; and a second language descriptor, wherein instructions are further configured to cause the processor to: select a kinematic hypothesis from a plurality of kinematic hypotheses based on the first multidimensional pose data, the second multidimensional pose data, and the second language descriptor.

[0111] In Example 48, the non-transitory computer-readable medium of Example 47, wherein the second language descriptor corresponds to the embedding of feature vectors in multiple multidimensional feature vectors of the robot.

[0112] In Example 49, the non-transitory computer-readable medium of Example 47 or Example 48, wherein instructions are configured to cause a processor to select a kinematics hypothesis from a plurality of kinematics hypotheses, including: instructions being configured to cause the processor to select a kinematics hypothesis based on any one of the following: a velocity limit of a robot joint, an acceleration limit of a robot joint, a range of motion of a robot joint, or the absence of a blank space for executing a kinematics hypothesis.

[0113] In Example 50, a non-transitory computer-readable medium of any of Examples 37 to 49, wherein instructions are configured to cause a processor to generate a plurality of motion primitives of a robot, including: instructions being configured to cause the processor to generate a velocity or acceleration of movement based on a timestamp indicating timing of a first position and timing of a second position following the first position.

[0114] In Example 51, the non-transitory computer-readable medium of Example 49 or Example 50, wherein the instructions are further configured to cause the processor to: use the model to generate a natural language description of motion corresponding to selected kinematics assumptions.

[0115] In Example 52, the non-transitory computer-readable medium of Example 51, wherein the instructions are further configured to cause the processor to: cause the robot to generate audible signals with natural language descriptions, and simultaneously perform motions corresponding to selected motion dynamics assumptions.

[0116] In Example 53, a non-transitory computer-readable medium includes instructions that, when executed by a processor, are configured to cause the processor to: generate a third dataset based on a first dataset and a second dataset, the first dataset including kinematics data representing multiple human-guided movements of a robot, the second dataset including language descriptors of multiple human-guided movements of the robot, wherein the third dataset includes multiple motion primitives of the robot.

[0117] In Example 54, the non-transitory computer-readable medium of Example 53, wherein the instructions are configured to cause the processor to generate a plurality of motion primitives of a robot, including: the instructions are configured to cause the processor to generate a plurality of end effector poses based on the starting point of each of a plurality of human-guided movements of the robot and the corresponding ending point of each of the plurality of human-guided movements of the robot, wherein each of the plurality of end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

[0118] In Example 55, the non-transitory computer-readable medium of Example 54, wherein the three-dimensional end effector orientation includes: end effector roll, pitch, and yaw; the rotation matrix of the end effector; or the quaternion of the end effector.

[0119] In Example 56, the non-transitory computer-readable medium of any of Examples 53 or 55 further includes: a fourth dataset comprising multiple multidimensional feature vectors of the robot, wherein execution is further configured to cause the processor to: generate a fifth dataset based on the third and fourth datasets, wherein the fifth dataset comprises multiple motion trajectories of the robot.

[0120] In Example 57, the non-transitory computer-readable medium of Example 56, wherein the instructions are configured to cause the processor to generate multiple motion trajectories, including: the instructions are configured to cause the processor to generate an upper boundary, a lower boundary, and an average value of movement corresponding to a first dataset; and to generate motion trajectories that follow a three-dimensional space around the average value while remaining between the upper and lower boundaries.

[0121] In Example 58, or the non-transitory computer-readable medium of Example 56 or Example 57, the instructions are further configured to cause the processor to: execute the model to generate a language description of each of the plurality of motion trajectories, and to label each of the plurality of motion trajectories with the corresponding language description.

[0122] In Example 59, a non-transitory computer-readable medium of either Example 56 or Example 58, wherein the instructions are further configured to cause the processor to: execute a probabilistic model to generate multiple motion dynamics hypotheses based on the fifth data.

[0123] In Example 60, the non-transitory computer-readable medium of Example 59, each of the plurality of kinematics assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

[0124] In Example 61, the non-transitory computer-readable medium of Example 59 or Example 60, wherein each of the plurality of kinematics assumptions includes data corresponding to the constraints of joints in a plurality of joints of the robot.

[0125] In Example 62, a non-transitory computer-readable medium of any of Examples 59 to 61, wherein each of the plurality of motion dynamics hypotheses includes data corresponding to a language descriptor among the plurality of language descriptors.

[0126] In Example 63, the non-transitory computer-readable medium of any of Examples 59 to 62, wherein the probabilistic model is a variational autoencoder configured to perform unsupervised learning on the fifth data.

[0127] In Example 64, a non-transitory computer-readable medium includes instructions that, when executed by a processor, cause the processor to: select a kinematics hypothesis from a plurality of kinematics hypotheses for the movement of the robot based on first multidimensional pose data representing an initial pose of the robot, second multidimensional pose data representing a target pose of the robot, and a language descriptor of the robot's desired motion.

[0128] In Example 65, the non-transitory computer-readable medium of Example 64, wherein the language descriptor corresponds to the embedding of feature vectors in a plurality of multidimensional feature vectors of the robot.

[0129] In Example 66, or the non-transitory computer-readable medium of Example 64 or Example 65, instructions are configured to cause a processor to select a kinematics hypothesis from a plurality of kinematics hypotheses, including: instructions being configured to cause the processor to select a kinematics hypothesis based on any one of the following: a velocity limit of a robot joint, an acceleration limit of a robot joint, a range of motion of a robot joint, or the absence of a blank space for executing a kinematics hypothesis.

[0130] In Example 67, the non-transitory computer-readable medium of any of Examples 65 to 66, wherein the instructions are configured to cause a processor to generate a plurality of motion primitives of a robot, including: the instructions being configured to cause the processor to generate a velocity or acceleration of movement based on a timestamp indicating a timing of a first position and a timing of a second position after the first position.

[0131] In Example 68, the non-transitory computer-readable medium of Example 66 or Example 67, wherein the instructions are further configured to cause the processor to: use the model to generate a natural language description of motion corresponding to the selected kinematics assumptions.

[0132] In Example 69, the non-transitory computer-readable medium of Example 68, wherein the instructions are further configured to cause the processor to: cause the robot to generate audible signals with natural language descriptions, and simultaneously perform motions corresponding to selected motion dynamics assumptions.

[0133] In Example 70, a method includes: generating a third dataset based on a first dataset and a second dataset, the first dataset including kinematics data representing multiple human-guided movements of a robot, the second dataset including linguistic descriptors of multiple human-guided movements of the robot, wherein the third dataset includes multiple motion primitives of the robot.

[0134] In Example 71, the method of Example 70, wherein generating multiple motion primitives of the robot includes: generating multiple end effector poses based on the starting point of each of the multiple human-guided movements of the robot and the corresponding ending point of each of the multiple human-guided movements of the robot, wherein each end effector pose of the multiple end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

[0135] In Example 72, the method of Example 71, wherein the three-dimensional end effector orientation includes: end effector roll, pitch and yaw; end effector rotation matrix; or end effector quaternion.

[0136] In Example 73, the method of any one of Example 70 or Example 72 further includes: a fourth dataset comprising multiple multidimensional feature vectors of the robot, and also includes: generating a fifth dataset based on the third and fourth datasets, wherein the fifth dataset comprises multiple motion trajectories of the robot.

[0137] In Example 74, the method of Example 73, wherein generating multiple motion trajectories includes: generating an upper boundary, a lower boundary, and a mean value of the movement corresponding to a first dataset; and generating a motion trajectory that follows a three-dimensional space around the mean value while remaining between the upper and lower boundaries.

[0138] In Example 75, the method of Example 73 or Example 74 further includes: executing the model to generate a language description of each of the plurality of motion trajectories, and labeling each of the plurality of motion trajectories with the corresponding language description.

[0139] In Example 76, the method of any of Example 73 or Example 75 further includes: performing a probabilistic model to generate multiple motion dynamics hypotheses based on the fifth data.

[0140] In Example 77, the method of Example 76, wherein each of the plurality of kinematic assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

[0141] In Example 78, the method of Example 76 or Example 77, each of the plurality of kinematic assumptions includes: data corresponding to the constraints of the joints in the plurality of joints of the robot.

[0142] In Example 79, the method of any of Examples 76 to 78, wherein each of the plurality of motion dynamics assumptions includes: data corresponding to a language descriptor among the plurality of language descriptors.

[0143] In Example 80, the method of any of Examples 76 to 79, wherein the probabilistic model is a variational autoencoder configured to perform unsupervised learning on the fifth data.

[0144] In Example 81, the method of any one of Examples 76 to 80, wherein the language descriptor is a first language descriptor, further includes: selecting a kinematic hypothesis from a plurality of kinematic hypotheses based on a first multidimensional pose data representing the robot’s initial pose, a second multidimensional pose data representing the robot’s target pose, and a second language descriptor.

[0145] In Example 82, the method of Example 81 is used, wherein the second language descriptor corresponds to the embedding of feature vectors in multiple multidimensional feature vectors of the robot.

[0146] In Example 83, the method of Example 81 or Example 82, wherein selecting a kinematics hypothesis from a plurality of kinematics hypotheses includes: making the processor select a kinematics hypothesis based on any one of the following: velocity limits of robot joints, acceleration limits of robot joints, range of motion of robot joints, or the absence of a blank space for executing kinematics hypotheses.

[0147] In Example 84, the method of any of Examples 71 to 83, wherein generating multiple motion primitives of the robot includes: generating a velocity or acceleration of movement based on a timestamp indicating the timing of a first position and the timing of a second position after the first position.

[0148] In Example 85, the method of Example 83 or Example 84 further includes: using the model to generate a natural language description of the motion corresponding to the selected motion dynamics assumptions.

[0149] In Example 86, the method of Example 85 further includes: enabling the robot to generate audible signals with natural language descriptions, and simultaneously performing motions corresponding to selected motion dynamics assumptions.

[0150] In Example 87, a method includes: generating a third dataset based on a first dataset and a second dataset, the first dataset including kinematics data representing multiple human-guided movements of a robot, the second dataset including linguistic descriptors of multiple human-guided movements of the robot, wherein the third dataset includes multiple motion primitives of the robot.

[0151] In Example 88, the method of Example 87, wherein generating multiple motion primitives of the robot includes: generating multiple end effector poses based on the starting point of each of the multiple human-guided movements of the robot and the corresponding ending point of each of the multiple human-guided movements of the robot, wherein each end effector pose of the multiple end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

[0152] In Example 89, the method of Example 88, wherein the three-dimensional end effector orientation includes: end effector roll, pitch and yaw; end effector rotation matrix; or end effector quaternion.

[0153] In Example 90, the method of any one of Example 87 or Example 89 further includes: a fourth dataset comprising multiple multidimensional feature vectors of the robot, and also includes: generating a fifth dataset based on the third and fourth datasets, wherein the fifth dataset comprises multiple motion trajectories of the robot.

[0154] In Example 91, the method of Example 90, wherein generating multiple motion trajectories includes: generating an upper boundary, a lower boundary, and an average value of the movement corresponding to a first dataset; and generating a motion trajectory that follows a three-dimensional space around the average value while remaining between the upper and lower boundaries.

[0155] In Example 92, the method of Example 90 or Example 91 is used, wherein a model is executed to generate a language description of each of a plurality of motion trajectories, and each of the plurality of motion trajectories is labeled with a corresponding language description.

[0156] In Example 93, the method of any of Example 90 or Example 92 further includes: performing a probabilistic model to generate multiple motion dynamics hypotheses based on the fifth data.

[0157] In Example 94, the method of Example 93, wherein each of the plurality of kinematic assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

[0158] In Example 95, the method of Example 93 or Example 94, each of the plurality of kinematic assumptions includes: data corresponding to the constraints of joints in the plurality of joints of the robot.

[0159] In Example 96, according to the method of any one of Examples 93 to 95, each of the plurality of motion dynamics hypotheses includes: data corresponding to the language descriptors in the plurality of language descriptors.

[0160] In Example 97, the method of any of Examples 93 to 96, wherein the probabilistic model is a variational autoencoder configured to perform unsupervised learning on the fifth data.

[0161] In Example 98, a method includes: selecting a kinematics hypothesis from a plurality of kinematics hypotheses for the movement of the robot based on a first multidimensional pose data representing an initial pose of the robot, a second multidimensional pose data representing a target pose of the robot, and a language descriptor of the robot’s desired motion.

[0162] In Example 99, the method of Example 98 is used, where the language descriptor corresponds to the embedding of feature vectors in multiple multidimensional feature vectors of the robot.

[0163] In Example 100, the method of Example 98 or Example 99, wherein the kinematics assumption selected from a plurality of kinematics assumptions includes: based on the velocity limit of the robot's joints, the acceleration limit of the robot's joints, the range of motion of the robot's joints, or the absence of a blank space for performing the kinematics assumption, the kinematics assumption is selected.

[0164] In Example 101, the method of any one of Examples 99 to 100, wherein generating multiple motion primitives of the robot includes: generating a velocity or acceleration of movement based on a timestamp indicating the timing of a first position and the timing of a second position after the first position.

[0165] In Example 102, the method of Example 100 or Example 101 is used, wherein a model is used to generate a natural language description of motion corresponding to a selected motion dynamics hypothesis.

[0166] In Example 103, the method of Example 102 is used, wherein the robot generates an audible signal with a natural language description and simultaneously performs motions corresponding to selected motion dynamics assumptions.

[0167] While the descriptions and diagrams above depict components as discrete elements, those skilled in the art will recognize the various possibilities of combining or integrating discrete elements into a single component. These possibilities can include combining two or more circuits to form a single circuit, mounting two or more circuits on a common chip or chassis to form an integrated component, executing discrete software components on a common processor core, and so on. Conversely, those skilled in the art will recognize the possibility of separating a single component into two or more discrete components, such as dividing a single circuit into two or more discrete circuits, separating a chip or chassis into discrete components initially provided on it, separating a software component into two or more segments and executing each segment on a separate processor core, and so on.

[0168] It is recognized that the implementation of the methods detailed herein is demonstrative in nature and is therefore understood to be capable of being implemented in the corresponding device. Similarly, it is recognized that the implementation of the devices detailed herein is understood to be capable of being implemented as the corresponding methods. Therefore, it is understood that the device corresponding to the methods detailed herein may include one or more components configured to perform each aspect of the relevant methods.

[0169] All abbreviations defined in the above description additionally apply to all claims included herein.

Claims

1. An apparatus comprising: Memory, the memory being configured to store: The first dataset includes motion dynamics data representing multiple human-guided movements of the robot, and The second dataset includes the language descriptors of the robot's multiple human-guided movements; as well as A processor configured to generate a third dataset based on the first dataset and the second dataset, wherein the third dataset includes multiple motion primitives of the robot.

2. The apparatus of claim 1, wherein, Generating the plurality of motion primitives of the robot includes: generating a plurality of end effector poses based on the starting point and the corresponding ending point of each of the plurality of human-guided movements of the robot. Each of the plurality of end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

3. The apparatus of claim 2, wherein, The three-dimensional end effector orientation includes: end effector roll, pitch, and yaw; the end effector rotation matrix; or the end effector quaternion.

4. The apparatus according to any one of claims 1 or 3, further comprising: The fourth dataset includes multiple multidimensional feature vectors of the robot. The processor is further configured to generate a fifth dataset based on the third and fourth datasets, wherein the fifth dataset includes multiple motion trajectories of the robot, and Preferably, generating the plurality of motion trajectories includes: the processor generating an upper boundary, a lower boundary, and an average value of the movement corresponding to the first dataset; and generating a motion trajectory that follows a three-dimensional space around the average value while remaining between the upper boundary and the lower boundary.

5. The apparatus according to claim 4, further comprising: The processor executes a model to generate a language description for each of the plurality of motion trajectories, and labels each of the plurality of motion trajectories with the corresponding language description.

6. The apparatus according to any one of claims 4 or 5, further comprising: A probabilistic model, wherein the processor is further configured to execute the probabilistic model to generate multiple kinematic hypotheses based on the fifth data, and Preferably, each of the plurality of kinematics assumptions includes: velocity and / or acceleration data corresponding to the movement of each of the plurality of joints of the robot.

7. The apparatus of claim 6, wherein, Each of the plurality of kinematics assumptions includes data corresponding to the constraints of joints in the plurality of joints of the robot.

8. The apparatus of claim 6 or 7, wherein, Each of the plurality of motion dynamics hypotheses includes: data corresponding to a language descriptor among the plurality of language descriptors.

9. The apparatus of any one of claims 6-8, wherein, The probabilistic model is a variational autoencoder, which is configured to perform unsupervised learning on the fifth data.

10. An apparatus comprising: Memory, the memory being configured to store: The first multidimensional pose data represents the robot's initial pose. The second multidimensional pose data represents the target pose of the robot. A language descriptor for the robot's desired motion, and Multiple kinematic assumptions for the movement of the robot; as well as A processor configured to: select a kinematic hypothesis from the plurality of kinematic hypotheses based on the first multidimensional pose data, the second multidimensional pose data, and the language descriptor.

11. The apparatus of claim 10, wherein, The language descriptor corresponds to the embedding of feature vectors in the robot's multiple multidimensional feature vectors.

12. The apparatus of claim 10 or 11, wherein, The kinematics hypothesis selected from the plurality of kinematics hypotheses includes: selecting the kinematics hypothesis based on any one of the following: velocity limits of the robot's joints, acceleration limits of the robot's joints, range of motion of the robot's joints, or the absence of a blank space for executing the kinematics hypothesis. Preferably, it further includes: determining the speed or acceleration of movement based on a timestamp, the timestamp indicating the timing of a first position and the timing of a second position after the first position.

13. The apparatus of claim 11 or 12, wherein, The processor is further configured to: use the model to generate a natural language description of the motion corresponding to the selected motion dynamics assumptions, and Preferably, the processor is configured to cause the robot to generate an audible signal describing the natural language, and simultaneously perform the motion corresponding to the selected motion dynamics hypothesis.

14. The apparatus of any of claims 10 to 13, wherein, The device is configured as a robot.

15. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to: Based on the first and second datasets, a third dataset is generated. The first dataset includes kinematics data representing multiple human-guided movements of the robot, and the second dataset includes linguistic descriptors of the multiple human-guided movements of the robot. wherein, The third dataset includes multiple motion primitives of the robot.

16. The non-transitory computer-readable medium of claim 15, wherein, The instructions are configured to cause the processor to generate the plurality of motion primitives of the robot, including: the instructions are configured to cause the processor to generate a plurality of end effector poses based on the starting point of each of the plurality of human-guided movements of the robot and the corresponding ending point of each of the plurality of human-guided movements of the robot. Each of the plurality of end effector poses includes a three-dimensional end effector position and a three-dimensional end effector orientation.

17. The non-transitory computer-readable medium of claim 15, wherein, The three-dimensional end effector orientation includes: end effector roll, pitch, and yaw; the end effector rotation matrix; or the end effector quaternion.

18. The non-transitory computer-readable medium of claim 15, further comprising: The fourth dataset includes multiple multidimensional feature vectors of the robot. The instructions are further configured to cause the processor to generate a fifth dataset based on the third dataset and the fourth dataset, wherein the fifth dataset includes multiple motion trajectories of the robot.

19. The non-transitory computer-readable medium of claim 15, wherein, The instructions are configured to cause the processor to generate the plurality of motion trajectories, including: the instructions are configured to cause the processor to generate an upper boundary, a lower boundary, and an average value of movement corresponding to the first dataset; and to generate a motion trajectory that follows a three-dimensional space around the average value while remaining between the upper boundary and the lower boundary.

20. The non-transitory computer-readable medium of claim 19, wherein, The instructions are also configured to cause the processor to: execute a model to generate a language description for each of the plurality of motion trajectories, and to label each of the plurality of motion trajectories with the corresponding language description.