Method, device and product for determining motion of robot arm
IRASim addresses the limitations of existing robot simulation by using generative models and diffusion transformers to create realistic robot arm motion simulations, achieving scalable and accurate motion prediction.
Patent Information
- Application Number
- PCT/CN2024/097571
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-11
AI Technical Summary
Existing robot simulation methods are limited by the cost and safety issues of real robots, lack visual realism, and are not scalable, with machine learning models only capable of simulating simple robot arm actions.
The development of Interactive Real-Robot Action Simulators (IRASim) that leverage generative models to generate highly realistic videos of robot arm motions using diffusion transformers and attention mechanisms, ensuring consistency with initial frames and adherence to action trajectories.
IRASim effectively simulates robot arm motions in a scalable and accurate manner, generating long-horizon, high-resolution videos that closely resemble ground-truth videos, outperforming baseline methods and demonstrating improved human evaluation results.
Smart Images

Figure CN2024097571_11122025_PF_FP_ABST
Abstract
Description
METHOD, DEVICE AND PRODUCT FOR DETERMINING MOTION OF ROBOT ARMFIELD
[0001] The present disclosure generally relates to robot simulation, and more specifically, to methods, devices, and computer program products for determining motions of robot arms.BACKGROUND
[0002] Scalable robot learning in the real world is limited by the cost and safety issues of real robots, and rolling out robot trajectories in the real world may be time-consuming and labor-intensive. Efforts have been made to create powerful physical simulators, while they are still not visually realistic enough. Also, the physical simulators are not scalable because it takes efforts to build new environments in simulation. Machine learning models have witnessed remarkable progress in recent years, and have been applied in the robot simulation field. By now, the machine learning models can only simulate simple actions of the robot arm. At this point, it is desired to improve performance of the machine learning models, such that the robot arm motion may be simulated in a more effective and accurate way.SUMMARY
[0003] In a first aspect of the present disclosure, there is provided a method for determining a motion of a robot arm. In the method, a state image that specifies an initial state of the motion of the robot arm is obtained. An action sequence is obtained, the action sequence specifying a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively. An image sequence that describes the motion of the robot arm is determined based on the state image and action sequence.
[0004] In a second aspect of the present disclosure, there is provided an electronic device. The electronic device comprises: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method according to the first aspect of the present disclosure.
[0005] In a third aspect of the present disclosure, there is provided a computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method according to the first aspect of the present disclosure.
[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0007] BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0008] Through the more detailed description of some implementations of the present disclosure in the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more apparent, wherein the same reference generally refers to the same components in the implementations of the present disclosure.
[0009] Fig. 1 illustrates an example environment for robot simulation;
[0010] Fig. 2 illustrates an example diagram for determining a motion of a robot arm according to implementations of the present disclosure;
[0011] Fig. 3 illustrates an example diagram of a model for determining a motion of a robot according to implementations of the present disclosure;
[0012] Fig. 4 illustrates an example diagram of an attention block based on a video-level condition according to implementations of the present disclosure;
[0013] Fig. 5 illustrates an example diagram of an attention block a frame-level condition according to implementations of the present disclosure;
[0014] Fig. 6 illustrates an example diagram of an action simulation result of a robot arm according to implementations of the present disclosure;
[0015] Fig. 7 illustrates an example flowchart of a method for determining a motion of a robot arm according to implementations of the present disclosure; and
[0016] Fig. 8 illustrates a block diagram of a computing device in which various implementations of the present disclosure can be implemented.DETAILED DESCRIPTION
[0017] Principle of the present disclosure will now be described with reference to some implementations. It is to be understood that these implementations are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.
[0018] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0019] References in the present disclosure to “one implementation, ” “an implementation, ” “an example implementation, ” and the like indicate that the implementation described may include a particular feature, structure, or characteristic, but it is not necessary that every implementation includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same implementation. Further, when a particular feature, structure, or characteristic is described in connection with an example implementation, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other implementations whether or not explicitly described.
[0020] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example implementations. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0021] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of example implementations. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0022] Principle of the present disclosure will now be described with reference to some implementations. It is to be understood that these implementations are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below. In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.
[0023] It may be understood that data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with requirements of corresponding laws and regulations and relevant rules.
[0024] It may be understood that, before using the technical solutions disclosed in various implementation of the present disclosure, the user should be informed of the type, scope of use, and use scenario of the personal information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.
[0025] For example, in response to receiving an active request from the user, prompt information is sent to the user to explicitly inform the user that the requested operation will need to acquire and use the user’s personal information. Therefore, the user may independently choose, according to the prompt information, whether to provide the personal information to software or hardware such as electronic devices, applications, servers, or storage media that perform operations of the technical solutions of the present disclosure.
[0026] As an optional but non-limiting implementation, in response to receiving an active request from the user, the way of sending prompt information to the user, for example, may include a pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose “agree” or “disagree” to provide the personal information to the electronic device.
[0027] It may be understood that the above process of notifying and obtaining the user authorization is only illustrative and does not limit the implementation of the present disclosure. Other methods that satisfy relevant laws and regulations are also applicable to the implementation of the present disclosure.
[0028] Referring to Fig. 1 for a brief introduction of the robot simulation environment, here Fig. 1 illustrates an example environment 100 for robot simulation. In Fig. 1, a robot 112 is provided and arm (s) in the robot 112 may be controlled to desired position (s) for implementing a task, for example, moving an object from one position to another position, and the like. Machine learning models have been applied in the robot simulation field. Here, the motion of the robot may include motion (s) of the robot arm (s) of the robot. When the robot have several robot arms, the motion may include motions of respective robot arms. As shown in Fig. 1, models may predict future observations (for example, an initial image 110 of the robot) based on the current observation and actions. Especially, text-to-video models may generate videos (for example, a video 120) for robots, and recurrent state space models may learn a latent representation of states through modeling a world model for reinforcement learning. Generative models may leverage videos, texts, and actions to generate photorealistic driving scenes.
[0029] Rolling out policies in the real world are essential in scaling up robot learning. Firstly, it is necessary for model evaluation. Secondly, as real-robot data are scarce for the reason that data collection often requires costly human demonstrations, an alternative is to roll out a policy to collect data. Finally, real-robot reinforcement learning requires rolling out robots in the real world to collect trajectories. Further, accurate levels for the models are not satisfied, especially, these models can only provide videos of robot motions from texts and control the robot arms to move in 2D space with language instructions, instead of controlling the robot motion in an accurate way. For example, although the robot may implement a task which is defined by texts, the action trajectory of the robot arm cannot be controlled.
[0030] In view of the above, the present disclosure aims to build a real-robot action simulator as an efficient and scalable alternative for real-world policy rollout. Specifically, Interactive Real-Robot Action Simulators (IRASim) is provided for managing the robot arm motion, which leverages the power of generative models to generate extremely realistic videos of a robot arm that executes a given action trajectory, starting from an initial given frame. Further, effectiveness of the proposed solution is tested based on several real-robot datasets and perform extensive experiments. Results show that the proposed solution outperforms all the baseline methods and is more preferable in human evaluations.
[0031] Referring to Fig. 2 for a brief description of the present disclosure, here Fig. 2 illustrates an example diagram 200 for determining a motion of a robot arm according to implementations of the present disclosure. As shown in Fig. 2, a state image 210 and an action sequence 220 may be obtained. Here, the state image 210 may specify an initial state of the motion of the robot arm (s) , and the action sequence 220 may specify a plurality of actions of the robot arm (s) at a plurality of time points of the motion of the robot arm (s) , respectively (for example, N-1 actions for N-1 time points. Further, an image sequence 230 may be determined for describing the motion of the robot arm (s) based on the state image 210 and action sequence 220.
[0032] Usually, the robot includes more than one robot arms which are connected by joins between two arms. At this point, the state image 210 may be a picture of all the robot arms or a portion of the robot arms. When multiple robot arms are included, the action sequence 220 may include actions of the multiple robot arms at each time point. Alternatively, the action sequence 220 may include the action of a robot arm which is deployed at the end of the multiple robot arms.
[0033] With these implementations of the present disclosure, the motion of the robot arm may be simulated in an accurate way, and images in the image sequence 230 may correspond to actions in the action sequence 220. In other words, , image 1#in the image sequence 230 may corresponds to action 1#in the action sequence 220, image 2#in the image sequence 230 may corresponds to action 2#in the action sequence 220, and the like. Therefore, the action trajectory of the robot arm may be determined in an accurate way.
[0034] In the context of the present disclosure, the trajectory-to-video generation task aims at predicting the video of a robot that executes a trajectory given the initial frame I1 (i.e., the state image 210) and the action trajectory a1: N-1 (i.e., the action sequence 220) : I2: N=f (I1, a1: N-1) (1)
[0035] In Formula (1) , N denotes the number of images in the video, ai denotes the action at the ith timestep. The present disclosure focuses on predicting videos for robot arms, and a typical action space for robot arms may contain 7 degrees of freedom (DoFs) , i.e., 3 DoFs for describing translation in the 3D space, 3 DoFs for 3D rotation, and 1 DoF for a state of the robot arm (for example, with respect to a gripper at the end of the robot arm, “1” may denote a close state and “0” may denote an open state. It is to be understood that more or less DoFs may be defined for the actions. For example, more DoFs may be defined for the state of the robot arm with a complex tool which have more states.
[0036] In implementations of the present disclosure, diffusion models may be adopted for building the prediction models. In order to determine the image sequence, a noise image sequence associated with the image sequence may be obtained, the noise image sequence comprising the state image and a plurality of noise images corresponding to the plurality of actions, respectively. Then, the image sequence may be determined based on the noise image sequence. Specifically, Diffusion models may be adapted in determining the image sequence.
[0037] Before delving into the proposed solution, the following paragraphs briefly review preliminaries of diffusion models. Diffusion models typically includes a forward process and a reverse process. The forward process gradually adds Gaussian noises to data x0 over Ttimesteps. It can be formulated as where xt donates the diffused data at the tth diffusion timestep and is a constant defined by a variance schedule. The reverse process starts from and gradually remove the noise to recover x0. It may be mathematically expressed as where μθ () and Σθ () denote the mean and covariance functions, respectively, and can be parameterized via a neural network.
[0038] In the training phase, a timestep is defined as t∈ [1, T] and may be obtained via the reparameterization, where The present disclosure leverages the simplified training objective to train a noise prediction model ∈θ as below:
[0039] In the inference phase, the present disclosure generates x0 by first sampling xT from and iteratively compute the following Formula (3) until t=0:
[0040] For conditional diffusion processes, the noise prediction model ∈θ may be parameterized as ∈θ (xt, t, c) where c is the condition that controls the generation process. The present disclosure uses superscript and subscript to indicate the timestep of a frame in the input video and the diffusion timestep, respectively. With these implementations of the present disclosure, the noise model may be built based on the diffusion model, and then the image sequence may be determined based on the noise image sequence in an effective and accurate way.
[0041] In implementations of the present disclosure, in order to determine the image sequence based on the noise image sequence, a noise sequence representation of the noise image sequence may be generated; a sequence representation corresponding to the image sequence may be obtained by a noise model based on the noise sequence representation; and then the image sequence may be determined based on the sequence representation. Specifically, the noise image sequence may include the state image at position 1#and (N-1) noise images at positions 2#to N#. Then, the noise image sequence may be generated, for example, based on images in the noise image sequence, the action sequence, and the diffusion time step. Further, the reverse process of the diffusion may be implemented to remove the noise and recover the images in the image sequence. With these implementations of the present disclosure, the powerfully processing ability of the diffusion model may be used for the robot motion simulation.
[0042] The present disclosure adopts Diffusion Transformers as the backbone, and aims to address three key aspects: 1) consistency with the initial frame, 2) adherence to the given action trajectory, and 3) computation efficiency. In the following, the present disclosure describes details and discuss design choices to achieve the aforementioned objectives.
[0043] In implementations of the present disclosure, in order to generate the sequence representation, a latent sequence representation of the noise image sequence may be generated by mapping the noise image sequence from an image space to a latent space; and the noise sequence representation may be determined based on latent sequence representation. Directly diffusing the entire video in the pixel space is time-consuming and requires substantial computation to generate long videos with high resolutions. The present disclosure performs the diffusion process in a low-dimension latent space z instead of the pixel space for improving the computation efficiency. For example, a pre-trained variational autoencoder (VAE) may be leveraged to compress each frame Ii in the noise video to a latent representation with the VAE encoder zi=Enc (Ii) , where i ∈ {1, 2, …, N} . The latent representation can be decoded back to the pixel space with the VAE decoder Ii=Dec (zi) . With these implementations of the present disclosure, the computation complexity level may be reduced and thus the computation efficiency may be increased.
[0044] Referring to Fig. 3 for more details about the noise model, Fig. 3 illustrates an example diagram 300 of a model for determining a motion of a robot arm according to implementations of the present disclosure. As shown in Fig. 3, the present disclosure provides a conditional diffusion model operating in the latent space z of the VAE. The condition c includes the latent representation of the initial frame of a video z1=Enc (I1) and an action trajectory a1: N-1. The diffusion target is to work out the latent representations of the subsequent N-1 frames of the video in which the robot executes the action trajectory, i.e. x=z2: N.
[0045] As shown in Fig. 3, the first frame 310 and the noise frames 312 may be divided into p patches (for example, patch 360, 362, and the like) , and at a block 320, tokens related to the first frame 310 and the noise frames 312 may be determined. The action trajectory 314 may be inputted into the MLP (Multilayer Perceptron) block 322 to obtain tokens related to the action sequence, and further the diffusion timestep 316 (donated as t) is also considered in the embedding block 324. At this point, tokenization is implemented and representations about the noise sequence, the conditions (including the initial image and the action sequence) , and the timestep are obtained. Specifically, each latent representation z1=Enc (I1) contains P tokens of D dimensions, where P denotes the number of patches per frame. By sequencing the latent representations of all frames by timestep order, the video is tokenized to N*P tokens.
[0046] In implementations of the present disclosure, in order to obtain the sequence representation, scale and shift parameters associated with the noise sequence representation may be determined by performing an attention operation on the latent sequence representation, the attention operations comprising at least any of: a spatial attention operation and a temporal attention operation. Still referring to Fig. 3, at an attention block 330, spatial and temporal positional embeddings are added to the tokens to allow awareness of patch positions within frames and timesteps in the video, respectively. The VAE is frozen throughout the training process. With these implementations of the present disclosure, the spatial attention may be implemented at a spatial attention block 332, and the temporal attention may be implemented at a temporal attention block 334. At a block 340, the scale and shift parameters may be used for determining the representation and then the noise prediction 350 may be determined. Therefore, the scale and shift parameters may be determined in a more effective and accurate way, and thus the noise prediction may be determined accurately.
[0047] Generally, standard transformer blocks may apply Multi-Head Self-Attention (MHA) to all tokens in the input token sequence, resulting in quadratic computation cost. Different from the standard transformer, the present disclosure leverages the memory-efficient spatial-temporal attention mechanism in the transformer block to reduce the computation cost. Specifically, each block includes a spatial attention block and a temporal attention block. In the spatial attention block, MHA is confined to tokens within a frame to model intra-frame interaction. In the temporal attention block, MHA is confined to tokens at an identical patch position across all the frames to model inter-frame interaction.
[0048] For a sequence of N*P tokens, the spatial attention operates on the 1*P tokens within each frame; temporal attention operates on the N*1 tokens across the N timesteps. Compared to attending over all the N*P tokens at a time, the spatial-temporal attention greatly decreases the computation cost which makes generating long and high-resolution videos feasible.
[0049] The present disclosure provides various ways for leveraging conditions to predict the noise. Generally, the conditions may include the initial frame condition and the action trajectory condition. The initial frame condition is achieved by treating the initial frame as the ground-truth portion in the input video sequence. That is, during training, the present disclosure only adds noise to the tokens corresponding to the frames 2#to N#, while keeping those of the initial frame z1 intact as it does not need to be predicted. Here, the diffusion loss is only computed upon the frames 2#to N#, i.e., z2: N. This condition approach ensures consistency with the initial frame by enabling the predicted frames to interact with it via attention mechanism.
[0050] For the action trajectory condition, a naive approach to impose the trajectory condition is to encode the trajectory as one embedding and append it to the input token sequence as an in-context condition. However, considering Diffusion Transformers demonstrate that adaptive normalization performs better than in-context condition, the present disclosure adopts this design to achieve trajectory condition.
[0051] In implementations of the present disclosure, two types of conditions are provided: a video-level condition and a frame-level condition. In order to determine the scale and shift parameters, the scale and shift parameters may be determined by performing the spatial attention operation on the latent sequence representation under a constraint of the state image and the action sequence (which is called as the video-level condition) . In implementations of the present disclosure, in order to determine the scale and shift parameters, with respect to a target image of a plurality of images in the image sequence, the scale and shift parameters may be determined and for a target position for a patch in a plurality of patches of the target image by performing the temporal attention operation on a plurality of patches corresponding to the target position in the plurality of images (which is called as the frame-level condition) .
[0052] Referring to Fig. 4 for more details, here Fig. 4 illustrates an example diagram 400 of an attention block based on a video-level condition according to implementations of the present disclosure. As shown in Fig. 4, the present disclosure uses a linear layer to encode the entire trajectory into a single embedding for condition. The embedding is then added to the embedding of the diffusion timestep t for generating the scale parameters γ and α and the shift parameters β for each spatial and temporal attention block. These parameters control the video generation via shifting the distribution of the token embeddings in the transformer block.
[0053] In the video-level condition, it obtains the conditioning embedding 412 (represented by cST) by adding the diffusion timestep embedding to the trajectory embedding. cST is used to control the input tokens 410 and regress the scale parameters γ and α, as well as the shift parameters β. Specifically, the computation of the spatial block is as follows: x=x+ (1+α1) ×MHA (γ1×LayerNorm (x) +β1) (4) x=x+ (1+α2) ×FFN (γ2×LayerNorm (x) +β2) (5)
[0054] In Formulas (4) and (5) , x is represented by a shape of (N, P, D) and denotes the token embeddings. Further, x is reshaped as (P, N, D) before entering the temporal block 334, in other words, positions for P, N are swapped. The computation of the temporal block is: x=x+ (1+α3) ×MHA (γ3×Layernorm (x) +β3) (6) x=x+ (1+α4) ×FFN (γ4×layerNorm (x) +β4) (7)
[0055] Specifically, blocks 420, 423, 430, and 433 may perform the scale &shift operations, blocks 421 and 431 may perform the multi-head attention operation, blocks 422, 425, 432 and 435 may perform the scale operations, and blocks 424 and 434 may perform the feed forward operations. Here, layer normalization is performed before scaling and shifting. With these implementations of the present disclosure, the video-level condition may ensure that the noise predictions may be controlled by the action sequence in a more effective and accurate way.
[0056] In implementations of the present disclosure, in order to determine the scale and shift parameters, the scale and shift parameters may be determined by performing the spatial attention operation on a portion of the latent sequence corresponding to a target action in the plurality of actions under a constraint of the state image and the target action.
[0057] Regarding the frame-level condition, the trajectory in the trajectory-to-video task is a finer description. Each action in the trajectory defines how the robot should move in each frame. At this point, each generated frame must be match with its corresponding action in the trajectory. To achieve this precise frame-level alignment, the generation of each frame is based on its corresponding action. Instead of encoding the entire action trajectory into a single embedding, the present disclosure uses a linear layer to encode each action into an individual embedding. The diffusion timestep embedding is added to each action embedding to generate the scale and shift parameters for each individual frame in the spatial block. The scale and shift parameters of the temporal block for all frames share the same conditioning embedding which is derived similarly as in video-level condition.
[0058] Referring to Fig. 5 for more details, Fig. 5 illustrates an example diagram 500 of an attention block based on a frame-level condition according to implementations of the present disclosure. As shown in Fig. 5, spatial attention blocks and temporal attention blocks are conditioned differently. The derivation of the conditioning embedding for temporal attention blocks cT is the same as in video-level condition, where the diffusion timestep embedding is added to the trajectory embedding. Different frames are conditioned differently in spatial attention blocks. The conditioning embedding of spatial attention blocks for the ith frame is donated as To derive the ith action in the trajectory is first encoded to an embedding through a linear layer. The diffusion timestep embedding is then added to the encoded embedding to obtain Here, (represented by embedding 512) and cT (represented by embedding 514) are used to regress the corresponding scale parameters γ and α, as well as the shift parameters β. While the computation of the temporal blocks is the same as the video-level condition, the computation of spatial blocks is different:
[0059] In the above formulas, denote the scale and shift parameters for the ith frame, and they are regressed from Specifically, blocks 520, 523, 530, and 533 may perform the scale &shift operations, blocks 521 and 531 may perform the multi-head attention operation, blocks 522, 525, 532 and 535 may perform the scale operations, and blocks 524 and 534 may perform the feed forward operations. With these implementations of the present disclosure, different frames are conditioned differently in spatial attention blocks, and thus the video-level condition may ensure that the noise predictions may be controlled by the action sequence in a more effective and accurate way.
[0060] In implementations of the present disclosure, in order to generate the sequence representation corresponding to the image sequence, a noise prediction associated with the noise sequence representation may be determined by the noise model based on the scale and shift parameters; and then the image sequence may be determined based on the noise prediction and the noise image sequence.
[0061] The output layer contains a linear layer which outputs the noise prediction is used to compute the L2 loss with the ground-truth noise during training (Formula (2) ) . Note that IRASim only predicts the mean of the noise but not the covariance. During inference, the present disclosure samples xT from and gradually denoise it via Formula (3) to obtain the predicted latent representation of the frames 2#to N#: The predicted video frames can be decoded with the VAE decoder
[0062] In implementations of the present disclosure, in order to obtain the noise model, a reference image sequence and a reference action sequence may be obtained; a reference noise image sequence may be obtained by adding noise into the reference image sequence to generating a reference noise image sequence; and then the noise model may be updated based on the reference image sequence, the reference action sequence, and reference image sequence. Specifically, reference data may be obtained for training the noise model as described in Fig. 3. For example, the noise model may be built based on the network structure of Fig. 3. Then, the reference image sequence and the reference action sequence may be ground-truth data. Then, noise may be gradually added in to images 2#to N#in the reference image sequence, and then the noise model may be trained iteratively based on the training data. With these implementations of the present disclosure, the noise model may learn knowledge among the image sequence, the noise image sequence and the action sequence. Therefore, the trained noise may predict accurate noise based on the initial image, the noise images, and the action sequence.
[0063] Fig. 6 illustrates an example diagram 600 of an action simulation result of a robot according to implementations of the present disclosure. As shown in Fig. 6, a state image 610 and an action sequence 620 may be inputted into a noise model 630, and then an image sequence 640 starting from the state image 610 may be determined. Further, extensive experiments are performed on IRASim benchmark which includes three challenging real-robot datasets to verify the performance of IRASim. Statistic of the experiments shows that, the proposed solution is effective on solving the trajectory-to-video task on various datasets with different action spaces. With these implementations of the present disclosure, the proposed solution generates videos of a robot that executes an action trajectory given the initial frame. Results show that the proposed solution is able to generate long-horizon and high-resolution videos that are almost visually indistinguishable from ground-truth videos.
[0064] The above paragraphs have described details for determining a motion of a robot arm. According to implementations of the present disclosure, a method is provided for determining a motion of a robot arm. Reference will be made to Fig. 7 for more details about the method, where Fig. 7 illustrates an example flowchart of a method 700 for determining a motion of a robot arm according to implementations of the present disclosure. At block 710, obtaining a state image that specifies an initial state of the motion of the robot arm. At block 720, obtaining an action sequence that specifies a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively. At block 730, determining an image sequence that describes the motion of the robot arm based on the state image and action sequence.
[0065] In implementations of the present disclosure, determining the image sequence comprises: obtaining a noise image sequence associated with the image sequence, the noise image sequence comprising the state image and a plurality of noise images corresponding to the plurality of actions, respectively; and determining the image sequence based on the noise image sequence.
[0066] In implementations of the present disclosure, determining the image sequence based on the noise image sequence comprises: generating a noise sequence representation of the noise image sequence; obtaining a sequence representation corresponding to the image sequence by a noise model based on the noise sequence representation; and determining the image sequence based on the sequence representation.
[0067] In implementations of the present disclosure, generating the sequence representation comprises: generating a latent sequence representation of the noise image sequence by mapping the noise image sequence from an image space to a latent space; and determining the noise sequence representation based on latent sequence representation.
[0068] In implementations of the present disclosure, obtaining the sequence representation comprises: determining scale and shift parameters associated with the noise sequence representation by performing an attention operation on the latent sequence representation, the attention operations comprising at least any of: a spatial attention operation and a temporal attention operation.
[0069] In implementations of the present disclosure, determining the scale and shift parameters comprises: determining the scale and shift parameters by performing the spatial attention operation on the latent sequence representation under a constraint of the state image and the action sequence.
[0070] In implementations of the present disclosure, determining the scale and shift parameters comprises: determining the scale and shift parameters by performing the spatial attention operation on a portion of the latent sequence corresponding to a target action in the plurality of actions under a constraint of the state image and the target action.
[0071] In implementations of the present disclosure, determining the scale and shift parameters comprises: with respect to a target image of a plurality of images in the image sequence, for a target position for a patch in a plurality of patches of the target image, determining the scale and shift parameters by performing the temporal attention operation on a plurality of patches corresponding to the target position in the plurality of images.
[0072] In implementations of the present disclosure, generating the sequence representation corresponding to the image sequence comprises: determining a noise prediction associated with the noise sequence representation by the noise model based on the scale and shift parameters; and determining the image sequence based on the noise prediction and the noise image sequence.
[0073] In implementations of the present disclosure, the noise model is obtained by: obtaining a reference image sequence and a reference action sequence; a reference noise image sequence by adding noise into the reference image sequence to generating a reference noise image sequence; and updating the noise model based on the reference image sequence, the reference action sequence, and reference image sequence.
[0074] According to implementations of the present disclosure, an apparatus is provided for determining a motion of a robot arm. The apparatus comprises: an image obtaining module, being configured for obtaining a state image that specifies an initial state of the motion of the robot arm; an action obtaining module, being configured for obtaining an action sequence that specifies a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively; and a determining obtaining module, being configured for determining an image sequence that describes the motion of the robot arm based on the state image and action sequence. The apparatus further comprise other modules being configured for implementing other steps in the above method.
[0075] According to implementations of the present disclosure, an electronic device is provided for implementing the method 700. The electronic device comprises: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method for determining a motion of a robot arm. The method comprises: obtaining a state image that specifies an initial state of the motion of the robot arm; obtaining an action sequence that specifies a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively; and determining an image sequence that describes the motion of the robot arm based on the state image and action sequence.
[0076] In implementations of the present disclosure, determining the image sequence comprises: obtaining a noise image sequence associated with the image sequence, the noise image sequence comprising the state image and a plurality of noise images corresponding to the plurality of actions, respectively; and determining the image sequence based on the noise image sequence.
[0077] In implementations of the present disclosure, determining the image sequence based on the noise image sequence comprises: generating a noise sequence representation of the noise image sequence; obtaining a sequence representation corresponding to the image sequence by a noise model based on the noise sequence representation; and determining the image sequence based on the sequence representation.
[0078] In implementations of the present disclosure, generating the sequence representation comprises: generating a latent sequence representation of the noise image sequence by mapping the noise image sequence from an image space to a latent space; and determining the noise sequence representation based on latent sequence representation.
[0079] In implementations of the present disclosure, obtaining the sequence representation comprises: determining scale and shift parameters associated with the noise sequence representation by performing an attention operation on the latent sequence representation, the attention operations comprising at least any of: a spatial attention operation and a temporal attention operation.
[0080] In implementations of the present disclosure, determining the scale and shift parameters comprises: determining the scale and shift parameters by performing the spatial attention operation on the latent sequence representation under a constraint of the state image and the action sequence.
[0081] In implementations of the present disclosure, determining the scale and shift parameters comprises: determining the scale and shift parameters by performing the spatial attention operation on a portion of the latent sequence corresponding to a target action in the plurality of actions under a constraint of the state image and the target action.
[0082] In implementations of the present disclosure, determining the scale and shift parameters comprises: with respect to a target image of a plurality of images in the image sequence, for a target position for a patch in a plurality of patches of the target image, determining the scale and shift parameters by performing the temporal attention operation on a plurality of patches corresponding to the target position in the plurality of images.
[0083] In implementations of the present disclosure, generating the sequence representation corresponding to the image sequence comprises: determining a noise prediction associated with the noise sequence representation by the noise model based on the scale and shift parameters; and determining the image sequence based on the noise prediction and the noise image sequence.
[0084] In implementations of the present disclosure, the noise model is obtained by: obtaining a reference image sequence and a reference action sequence; a reference noise image sequence by adding noise into the reference image sequence to generating a reference noise image sequence; and updating the noise model based on the reference image sequence, the reference action sequence, and reference image sequence.
[0085] According to implementations of the present disclosure, a computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform the method 700.
[0086] Fig. 8 illustrates a block diagram of a computing device 800 in which various implementations of the present disclosure can be implemented. It would be appreciated that the computing device 800 shown in Fig. 8 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the present disclosure in any manner. The computing device 800 may be used to implement the above method in implementations of the present disclosure. As shown in Fig. 8, the computing device 800 may be a general-purpose computing device. The computing device 800 may at least comprise one or more processors or processing units 810, a memory 820, a storage unit 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.
[0087] The processing unit 810 may be a physical or virtual processor and can implement various processes based on programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the computing device 800. The processing unit 810 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller, or a microcontroller.
[0088] The computing device 800 typically includes various computer storage medium. Such medium can be any medium accessible by the computing device 800, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 820 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM) ) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 830 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk, or another other media, which can be used for storing information and / or data and can be accessed in the computing device 800.
[0089] The computing device 800 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in Fig. 8, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0090] The communication unit 840 communicates with a further computing device via the communication medium. In addition, the functions of the components in the computing device 800 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 800 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.
[0091] The input device 850 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 860 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 840, the computing device 800 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the computing device 800, or any devices (such as a network card, a modem, and the like) enabling the computing device 800 to communicate with one or more other computing devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .
[0092] In some implementations, instead of being integrated in a single device, some, or all components of the computing device 800 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some implementations, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various implementations, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.
[0093] The functionalities described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs) , Application-specific Integrated Circuits (ASICs) , Application-specific Standard Products (ASSPs) , System-on-a-chip systems (SOCs) , Complex Programmable Logic Devices (CPLDs) , and the like.
[0094] Program code for carrying out the methods of the subject matter described herein may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely or partly on a machine, executed as a stand-alone software package partly on the machine, partly on a remote machine, or entirely on the remote machine or server.
[0095] In the context of this disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0096] Further, while operations are illustrated in a particular order, this should not be understood as requiring that such operations are performed in the particular order shown or in sequential order, or that all illustrated operations are performed to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular implementations. Certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Rather, various features described in a single implementation may also be implemented in multiple implementations separately or in any suitable sub-combination.
[0097] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
[0098] From the foregoing, it will be appreciated that specific implementations of the presently disclosed technology have been described herein for purposes of illustration, but that various modifications may be made without deviating from the scope of the disclosure. Accordingly, the presently disclosed technology is not limited except as by the appended claims.
[0099] Implementations of the subject matter and the functional operations described in the present disclosure can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing unit” or “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0100] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document) , in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code) . A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0101] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0102] It is intended that the specification, together with the drawings, be considered exemplary only, where exemplary means an example. As used herein, the use of “or” is intended to include “and / or” , unless the context clearly indicates otherwise.
[0103] While the present disclosure contains many specifics, these should not be construed as limitations on the scope of any disclosure or of what may be claimed, but rather as descriptions of features that may be specific to particular implementations of particular disclosures. Certain features that are described in the present disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0104] Similarly, while operations are illustrated in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the implementations described in the present disclosure should not be understood as requiring such separation in all implementations. Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in the present disclosure.
Claims
1.A method for determining a motion of a robot arm, comprising:obtaining a state image that specifies an initial state of the motion of the robot arm;obtaining an action sequence that specifies a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively; anddetermining an image sequence that describes the motion of the robot arm based on the state image and action sequence.2.The method according to claim 1, wherein determining the image sequence comprises:obtaining a noise image sequence associated with the image sequence, the noise image sequence comprising the state image and a plurality of noise images corresponding to the plurality of actions, respectively; anddetermining the image sequence based on the noise image sequence.3.The method according to claim 2, wherein determining the image sequence based on the noise image sequence comprises:generating a noise sequence representation of the noise image sequence;obtaining a sequence representation corresponding to the image sequence by a noise model based on the noise sequence representation; anddetermining the image sequence based on the sequence representation.4.The method according to claim 3, wherein generating the sequence representation comprises:generating a latent sequence representation of the noise image sequence by mapping the noise image sequence from an image space to a latent space; anddetermining the noise sequence representation based on latent sequence representation.5.The method according to claim 4, wherein obtaining the sequence representation comprises: determining scale and shift parameters associated with the noise sequence representation by performing an attention operation on the latent sequence representation, the attention operations comprising at least any of: a spatial attention operation and a temporal attention operation.6.The method according to claim 5, wherein determining the scale and shift parameters comprises: determining the scale and shift parameters by performing the spatial attention operation on the latent sequence representation under a constraint of the state image and the action sequence.7.The method according to claim 5, wherein determining the scale and shift parameters comprises: determining the scale and shift parameters by performing the spatial attention operation on a portion of the latent sequence corresponding to a target action in the plurality of actions under a constraint of the state image and the target action.8.The method according to claim 5, wherein determining the scale and shift parameters comprises: with respect to a target image of a plurality of images in the image sequence,for a target position for a patch in a plurality of patches of the target image, determining the scale and shift parameters by performing the temporal attention operation on a plurality of patches corresponding to the target position in the plurality of images.9.The method according to claim 5, wherein generating the sequence representation corresponding to the image sequence comprises:determining a noise prediction associated with the noise sequence representation by the noise model based on the scale and shift parameters; anddetermining the image sequence based on the noise prediction and the noise image sequence.10.The method according to claim 3, wherein the noise model is obtained by:obtaining a reference image sequence and a reference action sequence;a reference noise image sequence by adding noise into the reference image sequence to generating a reference noise image sequence; andupdating the noise model based on the reference image sequence, the reference action sequence, and reference image sequence.11.An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method for determining a motion of a robot arm, comprising:obtaining a state image that specifies an initial state of the motion of the robot arm;obtaining an action sequence that specifies a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively; anddetermining an image sequence that describes the motion of the robot arm based on the state image and action sequence.12.The method according to claim 11, wherein determining the image sequence comprises:obtaining a noise image sequence associated with the image sequence, the noise image sequence comprising the state image and a plurality of noise images corresponding to the plurality of actions, respectively; anddetermining the image sequence based on the noise image sequence.13.The method according to claim 12, wherein determining the image sequence based on the noise image sequence comprises:generating a noise sequence representation of the noise image sequence;obtaining a sequence representation corresponding to the image sequence by a noise model based on the noise sequence representation; anddetermining the image sequence based on the sequence representation.14.The method according to claim 13, wherein generating the sequence representation comprises:generating a latent sequence representation of the noise image sequence by mapping the noise image sequence from an image space to a latent space; anddetermining the noise sequence representation based on latent sequence representation.15.The method according to claim 14, wherein obtaining the sequence representation comprises: determining scale and shift parameters associated with the noise sequence representation by performing an attention operation on the latent sequence representation, the attention operations comprising at least any of: a spatial attention operation and a temporal attention operation.16.The method according to claim 15, wherein determining the scale and shift parameters comprises any of:determining the scale and shift parameters by performing the spatial attention operation on the latent sequence representation under a constraint of the state image and the action sequence; anddetermining the scale and shift parameters by performing the spatial attention operation on a portion of the latent sequence corresponding to a target action in the plurality of actions under a constraint of the state image and the target action.17.The method according to claim 15, wherein determining the scale and shift parameters comprises: with respect to a target image of a plurality of images in the image sequence,for a target position for a patch in a plurality of patches of the target image, determining the scale and shift parameters by performing the temporal attention operation on a plurality of patches corresponding to the target position in the plurality of images.18.The method according to claim 15, wherein generating the sequence representation corresponding to the image sequence comprises:determining a noise prediction associated with the noise sequence representation by the noise model based on the scale and shift parameters; anddetermining the image sequence based on the noise prediction and the noise image sequence.19.The method according to claim 3, wherein the noise model is obtained by:obtaining a reference image sequence and a reference action sequence;a reference noise image sequence by adding noise into the reference image sequence to generating a reference noise image sequence; andupdating the noise model based on the reference image sequence, the reference action sequence, and reference image sequence.20.A non-transitory computer program product, the non-transitory computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method for determining a motion of a robot arm, comprising:obtaining a state image that specifies an initial state of the motion of the robot arm;obtaining an action sequence that specifies a plurality of actions of the robot arm at a plurality of time points of the motion of the robot arm, respectively;determining an image sequence that describes the motion of the robot arm based on the state image and action sequence.
Citation Information
Patent Citations
Simulation of the machining of a workpiece
CN103093036A
Multitask strategy learning method based on diffusion model
CN117474075A
Authoring system and method, and storage medium
CN1392824A
Artificial intelligence system for learning robotic control policies
US10792810B1
Method and device for correcting motion of robotic arm
WO2019114339A1