Method, device, apparatus and medium for operating a robot arm

Through the motion model training method that combines pre-training and fine-tuning, the problems of low accuracy and efficiency of robot arm motion control are solved, efficient motion execution in complex environments is achieved, and the operational capabilities of the robot arm are improved.

CN117086881BActive Publication Date: 2025-09-12BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311286388.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-09-12
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

The existing technology has low accuracy and efficiency in motion control of robotic arms, especially when visual data and language data are combined, making it difficult to effectively perform complex tasks.

Method used

The action model is trained by combining pre-training and fine-tuning. Pre-training is performed using reference data including human actions and language descriptions. The action model is then fine-tuned using robot arm data, including language encoder, state encoder, image encoder and decoder, to improve the accuracy of action prediction through machine learning technology.

Benefits of technology

It improves the accuracy and efficiency of the robot arm's motion control in complex environments, enables it to stably perform tasks under multiple interference factors, and enhances the robot arm's operational capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117086881B_ABST
    Figure CN117086881B_ABST
Patent Text Reader

Abstract

Methods, devices, equipment, and media for operating a robotic arm are provided. In one method, a language description for specifying a goal to be achieved by a robotic arm is received; and a current state of the robotic arm is obtained. According to an action model, an action to be performed by the robotic arm is determined based on the language description and the current state, wherein the action model is pre-trained using reference data including relevant data of a human arm. Utilizing the exemplary implementation of the present disclosure, the problem of insufficient training data for the robotic arm can be alleviated. Furthermore, the pre-trained action model can grasp the basic knowledge about the correlation between the language description and the human action, and can obtain a more accurate action model, thereby obtaining the action of the robotic arm that matches the language description in a more efficient manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary implementations of the present disclosure relate generally to robotic control, and more particularly to methods, apparatuses, devices, and computer-readable storage media for operating a robotic arm. Background Art

[0002] In recent years, robotics has rapidly developed and is widely used in a variety of technological fields. For example, on factory production lines, robotic arms can perform tasks such as processing, grasping, sorting, and packaging. Furthermore, machine learning has also been widely used in a variety of application scenarios. Therefore, it is desirable to combine robotics and machine learning to control robotic operations in a simpler and more efficient manner. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for operating a robotic arm is provided. In the method, a verbal description specifying a goal to be achieved by the robotic arm is received; a current state of the robotic arm is obtained; and an action to be performed by the robotic arm is determined based on the verbal description and the current state according to an action model.

[0004] In a second aspect of the present disclosure, a device for operating a robotic arm is provided. The device includes: a receiving module configured to receive a verbal description specifying a goal to be achieved by the robotic arm; an acquiring module configured to acquire a current state of the robotic arm; and a determining module configured to determine, according to an action model, an action to be performed by the robotic arm based on the verbal description and the current state.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.

[0007] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0009] Figure 1 A block diagram illustrating an application environment using a robotic arm according to an exemplary implementation of the present disclosure is shown;

[0010] Figure 2 A block diagram for operating a robotic arm according to some implementations of the present disclosure is shown;

[0011] Figure 3 A block diagram illustrating input / output of an action model according to some implementations of the present disclosure;

[0012] Figure 4 A block diagram illustrating the structure of an action model according to some implementations of the present disclosure;

[0013] Figures 5A to 5E A block diagram illustrating the structures of various encoders and decoders in an action model according to some implementations of the present disclosure;

[0014] Figure 6 A block diagram illustrating a process for performing pre-training according to some implementations of the present disclosure is shown;

[0015] Figure 7 A block diagram illustrating a correspondence between various data in a process of performing pre-training according to some implementations of the present disclosure;

[0016] Figure 8 A block diagram illustrating a process for performing fine-tuning according to some implementations of the present disclosure is shown;

[0017] Figure 9 A block diagram illustrating a correspondence between various data in a process of performing fine-tuning according to some implementations of the present disclosure;

[0018] Figure 10 A block diagram illustrating the inference phase according to some implementations of the present disclosure is shown;

[0019] Figure 11 A block diagram illustrating a comparison between results obtained using an action model according to some implementations of the present disclosure and existing technical solutions;

[0020] Figure 12 A block diagram illustrating a comparison between results obtained by using different methods to train an action model according to some implementations of the present disclosure;

[0021] Figure 13A flow chart illustrating a method for operating a robotic arm according to some implementations of the present disclosure is shown;

[0022] Figure 14 A block diagram illustrating an apparatus for operating a robotic arm according to some implementations of the present disclosure; and

[0023] Figure 15 A block diagram is shown of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION

[0024] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0025] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.

[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0027] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0028] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0029] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0030] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0031] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.

[0032] Sample Environment

[0033] In recent years, robotics and machine learning technologies have been widely used in many application scenarios. Figure 1 Describe an application environment according to an example implementation of the present disclosure, Figure 1 A block diagram 100 is shown of an application environment 110 using a robotic arm according to an exemplary implementation of the present disclosure. Figure 1 As shown, in the application environment 110, a robotic arm 120 can be used to manipulate various objects in the application environment 110. Here, the objects can be various fruits and / or vegetables. For example, the robotic arm can be used to grab the objects on the table and place them on a plate; for another example, the robotic arm can be used to place the objects on the plate to a specified location on the table, and so on.

[0034] Machine learning models based on visual and language data have been developed for controlling robotic motion. However, the accuracy and efficiency of existing solutions are unsatisfactory. Therefore, it is desirable to combine robotics and machine learning techniques to control robotic operations in a simpler and more efficient manner.

[0035] Overview of the operation process

[0036] In order to at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for operating a robot arm is proposed. Figure 2Describes an overview of an exemplary implementation according to the present disclosure, Figure 2 A block diagram 200 for operating a robotic arm according to some implementations of the present disclosure is shown. Figure 2 As shown, the motion of the robotic arm 220 may be controlled based on the motion model 240. For example, a user 212 may specify a language description 210 of a goal to be achieved by the robotic arm.

[0037] An example implementation according to the present disclosure will be described below in an English language environment. Alternatively and / or additionally, the technical solution according to an example implementation according to the present disclosure can be executed in other language environments. For example, the robot can be controlled in Chinese, English, Japanese, Chinese, and other environments. For example, the robot can be controlled in application environments of different languages ​​based on the multilingual capabilities provided by machine learning technology. For ease of description, the process of controlling the robot will be described below using only the example of picking up and placing an object. Alternatively and / or additionally, the robot arm can perform other actions, for example, the robot arm can be used to process parts to predetermined sizes, package various items, and so on.

[0038] Furthermore, the current state 230 of the robot arm 220 can be obtained. It should be understood that the current state 230 may include various aspects of data, such as image data of the robot arm, posture data of the robot arm, and the state of a tool (e.g., a fixture, a tool, etc.) secured to the end of the robot arm. The language description 210 and the current state 230 can be input into an action model 240 (e.g., a GR-1 model). The action model 240 then determines an action 250 to be performed by the robot arm 220 based on the language description 210 and the current state 230. The action 250 may represent the difference between the current and next postures of the robot arm, as well as the difference between the current and next states of the tool.

[0039] It should be understood that the action model 240 herein can be obtained based on a pre-training process and a fine-tuning process. Specifically, the action model 240 is pre-trained using reference data including relevant data of a person's arm, that is, pre-trained using reference data that does not include relevant data of a robot arm. For example, the pre-training process can be performed using relevant data including language data and person's actions. This can alleviate the problem of insufficient training data for the robot arm. Furthermore, the pre-trained action model can master the basic knowledge about the correlation between language descriptions and person's actions. Furthermore, in the fine-tuning process, the relevant data of the robot arm can be used to further train the action model 240. In this way, the action of the robot arm that matches the language description can be obtained in a more effective manner.

[0040] Detailed process of operation

[0041] An overview of an example implementation according to the present disclosure has been described, and below, more details of an example implementation according to the present disclosure will be described. Figure 3 A block diagram 300 illustrating the input / output of an action model according to some implementations of the present disclosure is shown. Figure 3 As shown, inputs and learnable tokens of the action model 240 are shown below the action model 240 .

[0042] Specifically, legend 310 represents language-related data, such as a language description for specifying a goal to be achieved by the robot arm; legend 312 represents image-related data in the current state, such as image data captured by an image capture device on or near the robot arm; and legend 314 represents other state-related data in the current state, such as the posture of the robot arm and the state of the tool, etc. Regarding learnable tags, legend 320 may represent an action tag, such as the action tag that can be determined based on the action query; and legend 322 may represent an image tag, such as the action tag that can be determined based on the imager.

[0043] According to an exemplary implementation of the present disclosure, output data of the action model 240 is shown above the action model 240. For example, legend 330 may represent the relevant action output by the action model 240, and legend 330 may represent the image prediction output by the action model 240 of the scene in which the robot arm performs the relevant action, that is, the predicted corresponding scene when the robot arm performs the action.

[0044] In the following, the structure of the action model 240 is first described, Figure 4 A block diagram 400 is shown showing the structure of an action model according to some implementations of the present disclosure. Figure 4 As shown, the action model 240 may include a language encoder 410, a state encoder 420, an image encoder 430, an action decoder 440, and an image decoder 450. It should be understood that Figure 4 This is merely illustrative, and the motion model 240 may further include other network structures, and some network structures may be shared between encoders and decoders.

[0045] It should be understood that the encoders and decoders described above can be used to process data during the training and inference processes of the motion model 240. For example, during the training phase, the encoders and decoders can process data in a training dataset (also referred to as reference data), and during the inference phase, they can process currently acquired data to be processed.

[0046] See also Figures 5A to 5EMore details of an example implementation according to the present disclosure are described. Figure 5A A block diagram 500A shows the structure of the speech encoder 410 in the action model according to some implementations of the present disclosure. Figure 5A As shown, the language encoder 410 may include a text encoder 510, a multilayer perceptron 512 (Multilayer Perceptron, abbreviated MLP), and corresponding features (e.g., embedding) 514. It should be understood that the text encoder 510 here can be implemented based on the contrastive language image pre-training (Contrastive Language Image Pre-training, abbreviated CLIP) technology. It should be understood that Figure 5A This is merely illustrative, and alternatively and / or additionally, the language encoder 410 may include more, fewer, and / or different components to extract features 514 from the input language description.

[0047] Figure 5B A block diagram 500B illustrates the structure of the state encoder 420 in the action model according to some implementations of the present disclosure. Figure 5B As shown, the state encoder 420 may include multiple MLPs 520, 522, and 524. MLPs 520 and 522 may receive the posture of the robot arm and the state of the tool, respectively. According to an exemplary implementation of the present disclosure, the posture of the robot arm may be described using a vector comprising six degrees of freedom, for example, [pos1, pos2, pos3, rot1, rot2, rot3]. The first three dimensions of the vector represent the position of the robot arm, and the last three dimensions represent the orientation of the robot arm. For different types of tools, the state of the tool may be represented in different formats. Assuming that the tool is a fixture, the state may include an open state represented by 0 and a closed state represented by 1. Assuming that the tool is a drill, the state may include, for example, the speed and model of the tool, and so on.

[0048] MLP 524 can receive the outputs of MLP 520 and 522 and generate relevant features 526 of the state of the robot arm. It should be understood that Figure 5B This is merely illustrative, and alternatively and / or additionally, the state encoder 420 may include more, fewer, and / or different parts to extract the features 526 from the input state data.

[0049] Figure 5C A block diagram 500C shows the structure of the image encoder 430 in the motion model according to some implementations of the present disclosure. Figure 5CAs shown, the image encoder 430 may include a masked autoencoder 530 (MAE) and a perceptron resampler 532. The MAE 530 may receive one or more images and output corresponding features. The perceptron resampler 532 can integrate the features This generates a smaller number of features Furthermore, features can be connected and Then the connected features are used as the features of the image. It should be understood that Figure 5C This is merely illustrative, and alternatively and / or additionally, the image encoder 430 may include more, fewer, and / or different parts to extract the features 534 from the input state data.

[0050] Specifically, the image frame data can be encoded by the pre-trained MAE. is used as the global representation of the image. Corresponding to the slice mark is used as a local representation, which is further processed by the perceptron resampler to reduce the number of labels. During pre-training, there can be only one image at each time point; during fine-tuning, the image data can include images captured from both static capture devices (e.g., mounted at a fixed position near a robot arm) and dynamic capture devices (e.g., mounted at the end of a robot arm). The number of images can be adjusted flexibly.

[0051] Figure 5D A block diagram 500D shows the structure of the action decoder 440 in the action model according to some implementations of the present disclosure. Figure 5D As shown, the motion decoder 440 may include a plurality of MLPs 540, 542, and 544. The MLP 540 may receive motion-related features, and the MLPs 542 and 544 may respectively output features a related to the posture of the robot arm. arm and tool status-related characteristics a tool It should be understood that Figure 5D This is merely illustrative; alternatively and / or additionally, the action decoder 430 may include more, fewer, and / or different parts to map the input feature data to corresponding actions.

[0052] Figure 5E A block diagram 500E shows the structure of the motion picture decoder 450 in the motion model according to some implementations of the present disclosure. Figure 5EThe image decoder 450 shown may include a visual decoder 550, which may receive features related to each image and corresponding mask marks (e.g., 560) to output image blocks 552, ..., 554, ..., 556 corresponding to each mask, and then generate a corresponding image 558. It should be understood that Figure 5E This is merely illustrative, and alternatively and / or additionally, the image decoder 450 may include more, fewer, and / or different parts to map the input feature data to the corresponding image.

[0053] Having described the structures of the various encoders and decoders in the action model 240, the process for training the action model 240 will be described in detail below. It should be understood that the training process may include a pre-training process and a fine-tuning process. Specifically, the action model may be pre-trained using a reference character video including reference character actions and a reference language description describing the character video to obtain a pre-trained action model. Using the example implementation of the present disclosure, the pre-training process can enable the action model 240 to grasp the basic knowledge about the relationship between language representations and actions, thereby improving the accuracy of the subsequent fine-tuning process.

[0054] First, we provide an introduction to the notations that may be used in the training and inference process. We can use language-conditioned video prediction as a pre-training task for video generation. Specifically, we can pre-train the action model (including parameters π) to predict the action of a given video with a language description l and a sequence of video frames o at time points th to t. t-h:t In the case of , the corresponding video frame at the future time point t+f can be predicted. As shown in the following formula 1:

[0055] π(l,o t-h:t )→o t+f Formula 1

[0056] In the above formula, π represents the parameters of the motion model 240, l represents the language description, o represents the image of the video frame, t represents the current time point, and o t-h:t Represents a sequence of video frames within the previous time range before the current time point, o t+f An image representing the corresponding video frame at the future time point t+f.

[0057] During the pre-training process, the action model 240 can be pre-trained using the video and its related language description collected from the character video dataset. At this time, each training data in the dataset can be represented in the following format:

[0058] v={l, o1, o2, ..., o T} Formula 2

[0059] In the above formula, ν represents the relevant training data (also called reference data) of a video in the dataset, l represents the language description of the video, o1, o2, ..., o T Denotes each video frame in the video, and T denotes the number of video frames.

[0060] See also Figure 6 Describes more details about pre-training. Figure 6 A block diagram 600 illustrating a process for performing pre-training according to some implementations of the present disclosure is shown. Figure 6 As shown, the person video dataset 640 used for the pre-training process 650 may include multiple reference person videos and corresponding reference language descriptions. For example, reference data 610 may include reference person videos 614 and reference language descriptions 612; reference data 620 may include reference person videos 624 and reference language descriptions 622; ..., reference data 630 may include reference person videos 634 and reference language descriptions 632.

[0061] The various reference character videos may relate to the same or different purposes. For example, the character in reference character video 614 is holding an onion in his left hand, the character in reference character video 624 is adjusting a plant in his hand, ..., the character in reference character video 634 is wiping a stair railing with a sponge, etc. According to an exemplary implementation of the present disclosure, corresponding data may be extracted from the various reference data and input into the motion model 240 to obtain corresponding prediction values.

[0062] See also Figure 7 More details on determining the loss function and updating the motion model 240 are described in detail. Figure 7 A block diagram 700 is shown showing the correspondence between various data in the process of performing pre-training according to some implementations of the present disclosure. Figure 7 In the example of the reference data 610, the pre-training process is described. Specifically, during the pre-training of the motion model 240, a first set of reference frames 710 and a second set of reference frames 720 subsequent to the first set of reference frames 710 can be extracted from the reference character video 614.

[0063] It should be understood that the first group of reference frames 710 includes different numbers of video frames. For example, the first group of reference frames 710 may include the th to t video frames in the video (here, h may represent a pre-specified positive integer). The second group of reference frames 720 may include one or more video frames at different time points, for example, a video frame at time point t+f (here, f may represent a pre-specified positive integer). For ease of description, the following description will only take the example of the second group of reference frames 720 including only one video frame. In the case where the second group of reference frames 720 includes multiple video frames, each video frame may be determined in a similar manner.

[0064] The motion model 240 may be used to determine a prediction 712 of the second set of reference frames 720 based on the reference language description 612 (corresponding to the example 310) and the first set of reference frames 710 (corresponding to the example 312). Further, a loss 730 (for ease of description, this loss may be referred to as a first loss, for example, represented as L) between the prediction 712 of the second set of reference frames and the second set of reference frames 720 (i.e., the true value, corresponding to the example 332) may be used. video1 ), update the action model 240.

[0065] It should be understood that due to the limited amount of training data on robotic arms and the limited types of robotic arm movements, directly using the robotic arm training dataset cannot effectively extract the correlation between language descriptions and robotic arm movements. Using the exemplary implementation of the present disclosure, by using a training dataset that includes human video data, the pre-training process can enable the motion model 240 to understand relevant knowledge about rich movements, thereby further helping to improve the accuracy of the motion model 240 during the subsequent fine-tuning process.

[0066] Furthermore, a reference robot arm video including reference robot arm motions and a reference motion language description describing the robot arm video can be used to fine-tune the pre-trained motion model, thereby obtaining a fine-tuned motion model. Using the exemplary implementations of the present disclosure, the parameters of the motion model 240 can be further optimized using data related to the robot arm based on the pre-training process, thereby improving the accuracy of the motion model 240.

[0067] Figure 8 A block diagram 800 illustrating a process for performing fine-tuning according to some implementations of the present disclosure is shown. Figure 8As shown, the robot video dataset 840 used for fine-tuning process 850 may include multiple reference robot videos and corresponding reference action language descriptions. Furthermore, more relevant information about the robot arm may be included, such as the robot arm's state and the action to be performed, etc. For example, reference data 810 may include reference robot videos 814, reference action language descriptions 812, states 816, and actions 818; reference data 820 may include reference human videos 824, reference language descriptions 822, states 826, and actions 828; ..., reference data 830 may include reference human videos 834, reference language descriptions 832, states 836, and actions 838.

[0068] According to an example implementation of the present disclosure, the state herein may include at least one of the following: a reference pose of a reference robot arm and a reference state of a reference tool of the reference robot arm. For example, the state may be represented based on the vector format described above. According to an example implementation of the present disclosure, the reference action may involve at least one of the following: a change in the reference pose and a change in the reference state. For the pose, this may involve a change in the six degrees of freedom described above; for the tool state, this may include a change in the on / off state of the gripper, for example, from on to off, etc.

[0069] Each video may involve the same or different purposes. For example, the robot arm in reference robot video 814 picks up broccoli, the robot arm in reference robot video 824 places an item on a plate, ..., the robot arm in reference robot video 834 places a green pepper on a plate, etc. According to an exemplary implementation of the present disclosure, corresponding data can be extracted from each video and input into the motion model 240 to obtain corresponding fine-tuning.

[0070] According to an exemplary implementation of the present disclosure, the loss of the video portion during fine-tuning is determined in the same manner as Figure 7 The method shown is similar and will not be described in detail. Specifically, the third group of reference frames and the fourth group of reference frames after the third group of reference frames can be extracted from the reference robot arm video. The pre-trained action model can be used to determine the prediction of the fourth group of reference frames (i.e., the true value) based on the reference action language description and the third group of reference frames. Further, the loss between the prediction of the fourth group of reference frames and the fourth group of reference frames (for the sake of convenience, the loss can be called the second loss, for example, represented as L) can be used. video2 ), update the action model.

[0071] Furthermore, the states and actions in the reference data can be used to determine the corresponding loss function during the fine-tuning process. In the following, the relevant process will be described using the reference data 810 as an example. Specifically, the reference current state (e.g., Figure 8 816 in the state) and the reference action (e.g., Figure 8 Further, a pre-trained action model can be used to determine a prediction of a reference action based on the reference current state, the reference action language description, and the third set of reference frames, and the action model can be updated based on a third loss between the prediction of the reference action and the reference action.

[0072] Figure 9 A block diagram 900 shows the correspondence between various data in the process of performing fine-tuning according to some implementations of the present disclosure. Figure 9 As shown, the reference action language description 812 (corresponding to the example 310), the third set of reference frames 910 (corresponding to the example 312) extracted from the reference robot video 814, and the state 816 (corresponding to the example 314) can be used to determine the prediction 920 of the action. Further, the corresponding loss 930 can be determined based on the difference between the true action 818 and the prediction 920, and then the loss 930 can be used to update the action model 240. Here, the loss 930 can involve multiple aspects, for example, the loss L related to the posture of the robot arm arm , and the loss L associated with the tool tool .

[0073] According to an exemplary implementation of the present disclosure, the first number of the first set of reference frames may be equal to the third number of the third set of reference frames, and the second number of the second set of reference frames may be equal to the fourth number of the fourth set of reference frames. That is, during the pre-training process and the fine-tuning process, the formats of the corresponding video frame groups are the same. In this way, the motion model 240 can be trained in a unified manner, thereby improving the performance of the motion model 240.

[0074] According to an example implementation of the present disclosure, the fine-tuning process involves a multi-task model, and the fine-tuning process can be continued for the pre-trained motion model 240 described above (that is, the motion model 240 in the initial stage of the fine-tuning process and the pre-trained motion model 240 share the same model parameters). Specifically, the fine-tuning process can be performed based on the following formula:

[0075] π(l,o t-h:t , s t-h:t )→a t , o t+f Formula 3 In the above formula, s t-h:trepresents the state of the robot arm in the previous time range before the current time point, a t Indicates the action that the robot arm will perform, and the meanings of other symbols are the same as those in the formula described above. Specifically, a training dataset including N reference data (i.e., trajectories) for M different tasks can be accessed. Each trajectory can include a language representation, a sequence of video frames, states, and actions:

[0076] τ={l,o1,s1,a1,o2,s2,a2,...,o T , s T , a T} Formula 4

[0077] According to an example implementation of the present disclosure, the feature dimensions of each modality can be reduced by a linear layer before being input to the transformer. For action prediction, the posture of the robot arm and the state of the tool can be predicted separately. For simplicity, the action is referred to as [ACT]. For image prediction, the future frame can be predicted. For simplicity, the image is referred to as [OBS]. During pre-training, the various labels can be arranged in the following order:

[0078] (l,o t-h ,[OBS],l,o t-h+1 ,[OBS],...,l,o t ,[OBS],) Formula 5 During fine-tuning, you can arrange the markers in the following order:

[0079] (l,s t-h , o t-h ,[OBS],[ACT],l,s t-h+1 ,...,l,s t , o t , [OBS], [ACT]) Formula 6

[0080] It should be understood that language tags are repeated in each time step to avoid being overwhelmed by other modalities. In order to take into account temporal information, temporal features can be added to the tags. Within one time step, all devices can share the same temporal features. Since tags of different modalities are encoded in different ways, there is no need to add embeddings to disambiguate the modalities. A causal attention mechanism can be adopted. That is, during pre-training, all tags including [OBS] tags can only process all language and image tags, but not past [OBS] tags. During fine-tuning, all tags (including [ACT] and [OBS] tags) can only process relevant tags for language, image, and state, but not past [ACT] or [OBS] tags.

[0081] The output from the [ACT] tag is passed through a linear layer in order to predict the actions of the robot arm and tool (see above). Figure 5D Specifically, the decoder can be implemented using a self-attention module and an MLP module. The image decoder performs operations on the output corresponding to [OBS] and the mask label (as described above). Figure 5D Description). Each mask token is a shared learnable vector corresponding to a positional encoding. The output corresponding to the mask token reconstructs the predicted future image patch.

[0082] According to an example implementation of the present disclosure, the publicly available Ego4D dataset (or other datasets) can be used to perform the pre-training process. During the fine-tuning process, videos of the robot dataset can be sampled and end-to-end optimization can be performed using the causal behavior cloning loss and the video prediction loss. Specifically, the loss function is as follows:

[0083] L=L arm +λ1L tool +λ2L video2 Formula 7

[0084] Specifically, for image prediction, the image can be predicted in the subsequent f=3 steps and supervised using MSE (mean square error) loss. For the posture of the robot arm, the Smooth-L1 loss can be used to learn the movement of the robot arm. For the state of the tool, the binary cross entropy (BCE) loss can be used. In the above formula, λ1 and λ2 represent predetermined weight coefficients, which can be set to 0.01, 0.1 or other values, respectively.

[0085] According to an exemplary implementation of the present disclosure, the action model 240 may be pre-trained and may be used directly during the inference phase to perform a desired task. For example, at the positions shown in the diagrams 310, 312, and 314, a language description, image data of the current state, and data of the robot arm may be input into the action model 240, and the corresponding action may be obtained at the output position shown in the diagram 330 of the action model 240. Alternatively and / or additionally, a corresponding image prediction may be obtained at the output position shown in the diagram 332 of the action model 240.

[0086] By using the exemplary implementation of the present disclosure, the knowledge in the action model 240 can be fully utilized to predict the action and corresponding image when the robot arm achieves a certain goal. Figure 10 Describes the inference phase in more detail. Figure 10A block diagram 1000 of the inference phase according to some implementations of the present disclosure is shown. According to an example implementation of the present disclosure, the technical solutions described above can be applied in different situations. For example, a large amount of data (e.g., all training data in a dataset) can be used to perform pre-training and fine-tuning to obtain the action model 240. Figure 10 As shown, the reasoning process 1010 represents reasoning performed using the aforementioned motion model 240. Specifically, after inputting a language representation 1012 and the corresponding current state of the robot arm into the fine-tuned motion model 240, the motion model 240 outputs the result. Action 1016 represents the action that the robot arm will perform, and image prediction 1014 represents the image prediction of the robot arm when grasping the eggplant on the plate.

[0087] For example, distractions can be added to the robot arm's environment (e.g., changing the background of the table and / or adding a large amount of fruit and / or vegetables to the plate, etc.). The reasoning process 1020 represents the process in the presence of distractions. In this case, the action 1026 represents the action that the robot arm will perform corresponding to the language representation 1022, and the image prediction 1024 represents the image prediction when the robot arm grasps the green pepper on the plate in the presence of distractions.

[0088] For another example, a small amount of data (e.g., 10% or other proportion of the training data in the data set) can be used to perform pre-training and fine-tuning to obtain the action model 240. The reasoning process 1030 represents the reasoning performed using the above-mentioned action model 240. Specifically, after the language expression 1032 and the current state of the corresponding robot arm are input into the fine-tuned action model 240, the result is output by the action model 240. The action 1036 represents the action that the robot arm will perform, and the image prediction 1034 represents the image prediction of the robot arm when grabbing the green pepper in the plate. Figure 10 It can be seen that an exemplary implementation according to the present disclosure can achieve good results in many situations.

[0089] According to an exemplary implementation of the present disclosure, the positions of the steps for specifying an action to be performed by the robot arm (e.g., the parameter f described above) are received. Furthermore, based on the action model, action and image predictions that match the number of steps can be determined. For another example, the number of steps in the action to be performed by the robot arm can be specified, such as outputting action and image predictions corresponding to f and f+1. In this way, the existing knowledge in the action model 240 can be used to predict relevant information at different positions (i.e., time points).

[0090] According to one exemplary implementation of the present disclosure, the current state of the robot arm includes at least one of the following: an image and posture of the robot arm, and the state of its tool, and the action involves a change in the posture and the state of the tool. In this way, the accuracy of the prediction result can be improved based on multiple aspects of the current state.

[0091] According to an example implementation of the present disclosure, in the process of determining an action, a language encoder can be used to determine the language representation of the language description, a state encoder can be used to determine the state representation of the current state, and then an action decoder can be used to determine the action based on the language representation and the state representation. According to an example implementation of the present disclosure, in the process of determining an image prediction, an image decoder can be used to determine the image prediction based on the language representation and the state representation. In this way, the encoder-decoder architecture, which has been proven to be reliable, can be fully utilized to perform the desired prediction task.

[0092] It should be understood that while the above description of the process of operating a robotic arm uses the example of performing training and inference in a real-world application environment, the process described above can alternatively and / or additionally be applied in a virtual application environment. For example, the robotic arm can be operated in virtual manufacturing, virtual assembly, and other virtual simulation applications. In this case, the motion model matches the application environment of the robotic arm, and the application scenario includes at least one of the following: a virtual application environment and a real-world application environment.

[0093] In other words, if it is desired to implement the technical solution of the present disclosure in a real application environment (e.g., a real physical environment, such as a factory production line, etc.), the pre-training and fine-tuning process can be performed using video data collected in the real application environment. If it is desired to implement the technical solution of the present disclosure in a virtual application environment, the pre-training and fine-tuning process can be performed using video data collected in the virtual application environment. In this way, the accuracy of the motion model 240 can be improved, thereby improving the accuracy of the subsequent reasoning stage.

[0094] According to an exemplary implementation of the present disclosure, the motions from the motion model 240 can be used to directly drive the robot arm. Alternatively and / or additionally, the motions can be adjusted to determine the motion instructions for driving the robot arm, thereby obtaining more accurate motion instructions. For example, if it is found that directly using the obtained motions to drive the robot arm may result in errors, such as the robot arm being unable to accurately grasp an object, the posture of the robot arm and the state of the tool can be adjusted accordingly to obtain a more accurate robot arm trajectory, thereby determining a more accurate motion instruction.

[0095] According to an exemplary implementation of the present disclosure, a technical solution according to an exemplary implementation of the present disclosure can achieve better technical effects. Figure 11 A block diagram 1100 is shown comparing the results obtained using the motion model according to some implementations of the present disclosure with existing technical solutions. Figure 11 As shown, Table 1110 shows the results of multi-task learning in the ABCD->D scenario (that is, the training data involves the scenario ABCD, and the test data involves the scenario D), Table 1120 shows the processing results in the scenario of a small amount of training data (for example, 10% of the training data), and Table 1130 shows the test results in the case of a real robot experiment. Figure 11 It can be seen that a better technical effect can be achieved according to an example implementation of the present disclosure.

[0096] further, Figure 12 A block diagram 1200 is shown comparing the results obtained by using different methods to train the motion model according to some implementations of the present disclosure. Figure 12 As shown, in comparison diagram 1240, example 1210 shows the effect of the motion model obtained without video prediction and pre-training, example 1220 shows the effect of the motion model obtained without video pre-training, and example 1230 shows the effect of the motion model obtained with video prediction and video pre-training. Further, in comparison diagram 1242, example 1250 shows the effect of the motion model obtained without video prediction and pre-training, example 1260 shows the effect of the motion model obtained without video pre-training, and example 1270 shows the effect of the motion model obtained with video prediction and video pre-training. It can be seen that better results can be obtained when video prediction and video prediction are used.

[0097] Example Process

[0098] Figure 13 A flowchart of a method 1300 for operating a robotic arm according to some implementations of the present disclosure is shown. At block 1310, a verbal description specifying a goal to be achieved by the robotic arm is received. At block 1320, a current state of the robotic arm is obtained. At block 1330, an action to be performed by the robotic arm is determined based on the verbal description and the current state according to an action model, where the action model is pre-trained using reference data including data related to a human arm.

[0099] According to an example implementation of the present disclosure, the method 1300 further includes: determining, according to the action model, an image prediction of a scene in which the robot arm performs an action based on the language description and the current state.

[0100] According to an example implementation of the present disclosure, the method 1300 further includes: receiving a location and a number of steps for specifying an action to be performed by the robotic arm; and determining, based on the action model, an action and image prediction that matches the location and the number of steps.

[0101] According to an example implementation of the present disclosure, the current state of the robot arm includes at least any one of the following: an image and posture of the robot arm and a state of a tool of the robot arm, and the action involves changes in the posture and the state of the tool.

[0102] According to an example implementation of the present disclosure, an action model includes a language encoder, a state encoder, and an action decoder, wherein determining an action includes: using the language encoder to determine a language representation of a language description; using the state encoder to determine a state representation of a current state; and using the action decoder to determine an action based on the language representation and the state representation.

[0103] According to an example implementation of the present disclosure, the action model further includes an image decoder, and determining the image prediction includes: determining the image prediction based on the language representation and the state representation using the image decoder.

[0104] According to an example implementation of the present disclosure, the action model is obtained based on: pre-training the action model using a reference character video including reference character actions and a reference language description describing the character video to obtain a pre-trained action model; and fine-tuning the pre-trained action model using a reference robot arm video including reference robot arm actions and a reference action language description describing the robot arm video to obtain a fine-tuned action model.

[0105] According to an example implementation of the present disclosure, the pre-trained motion model includes: extracting a first group of reference frames and a second group of reference frames following the first group of reference frames from a reference character video; and using the motion model to determine a prediction of the second group of reference frames based on a reference language description and the first group of reference frames; and updating the motion model based on a first loss between the prediction of the second group of reference frames and the second group of reference frames.

[0106] According to an example implementation of the present disclosure, fine-tuning a pre-trained motion model includes: extracting a third group of reference frames and a fourth group of reference frames following the third group of reference frames from a reference robot arm video, respectively; and using the pre-trained motion model to determine a prediction of the fourth group of reference frames based on a reference motion language description and the third group of reference frames; and updating the motion model based on a second loss between the prediction of the fourth group of reference frames and the fourth group of reference frames.

[0107] According to an example implementation of the present disclosure, fine-tuning the pre-trained motion model further includes: obtaining a reference current state and a reference motion of a reference robot arm; and using the pre-trained motion model to determine a prediction of the reference motion based on the reference current state, the reference motion language description, and a third set of reference frames; and updating the motion model based on a third loss between the prediction of the reference motion and the reference motion.

[0108] According to an example implementation of the present disclosure, the reference current state includes at least any one of the following: a reference posture of the reference robot arm and a reference state of a reference tool of the reference robot arm, and the reference action involves at least any one of the following: a change in the reference posture and the reference state.

[0109] According to an example implementation of the present disclosure, the first number of the first set of reference frames is equal to the third number of the third set of reference frames, and the second number of the second set of reference frames is equal to the fourth number of the fourth set of reference frames.

[0110] According to an example implementation of the present disclosure, the motion model matches the application environment of the robot arm, and the application scenario includes at least any one of the following: a virtual application environment and a real application environment.

[0111] According to an example implementation of the present disclosure, the method 1300 further includes: adjusting the motion to determine a motion instruction for driving the robot arm.

[0112] According to an example implementation of the present disclosure, the motion model is pre-trained using reference data that does not include relevant data of the robot arm.

[0113] Example devices and equipment

[0114] Figure 14 A block diagram of an apparatus 1400 for operating a robotic arm according to some implementations of the present disclosure is shown. The apparatus includes: a receiving module 1410 configured to receive a language description specifying a goal to be achieved by the robotic arm; an acquiring module 1420 configured to acquire a current state of the robotic arm; and a determining module 1430 configured to determine an action to be performed by the robotic arm based on the language description and the current state according to an action model, wherein the action model is pre-trained using reference data including relevant data of a human arm.

[0115] According to an example implementation of the present disclosure, the apparatus further includes: a prediction module configured to determine an image prediction of a scene in which the robot arm performs an action based on a language description and a current state according to the action model.

[0116] According to an example implementation of the present disclosure, the apparatus further includes: a parameter receiving module configured to receive the position and number of steps for specifying an action to be performed by the robot arm; and a parameter-based determination module configured to determine, based on an action model, an action and image prediction that matches the position and number of steps.

[0117] According to an example implementation of the present disclosure, the current state of the robot arm includes at least any one of the following: an image and posture of the robot arm and a state of a tool of the robot arm, and the action involves changes in the posture and the state of the tool.

[0118] According to an example implementation of the present disclosure, an action model includes a language encoder, a state encoder, and an action decoder, wherein the determination module includes: a language encoding module configured to determine a language representation of a language description using the language encoder; a state encoding module configured to determine a state representation of a current state using the state encoder; and an action decoding module configured to determine an action based on the language representation and the state representation using the action decoder.

[0119] According to an example implementation of the present disclosure, the action model further includes an image decoder, and the determination module includes: an image decoding module configured to determine the image prediction based on the language representation and the state representation using the image decoder.

[0120] According to an example implementation of the present disclosure, the apparatus further includes: a pre-training module configured to pre-train an action model using a reference character video including reference character actions and a reference language description describing the character video to obtain a pre-trained action model; and a fine-tuning module configured to fine-tune the pre-trained action model using a reference robot arm video including reference robot arm actions and a reference action language description describing the robot arm video to obtain a fine-tuned action model.

[0121] According to an example implementation of the present disclosure, the pre-training module includes: a first extraction module, configured to extract a first group of reference frames and a second group of reference frames following the first group of reference frames from a reference character video; and a first prediction module, configured to utilize a motion model to determine a prediction of the second group of reference frames based on a reference language description and the first group of reference frames; and a first update module, configured to update the motion model based on a first loss between the prediction of the second group of reference frames and the second group of reference frames.

[0122] According to an example implementation of the present disclosure, the fine-tuning module includes: a second extraction module configured to extract a third group of reference frames and a fourth group of reference frames following the third group of reference frames from a reference robot arm video, respectively; and a second prediction module configured to utilize a pre-trained motion model to determine a prediction of the fourth group of reference frames based on a reference motion language description and the third group of reference frames; and a second update module configured to update the motion model based on a second loss between the prediction of the fourth group of reference frames and the fourth group of reference frames.

[0123] According to an example implementation of the present disclosure, the fine-tuning module further includes: an action acquisition module configured to acquire a reference current state and a reference action of a reference robot arm; and a third prediction module configured to utilize a pre-trained action model to determine a prediction of a reference action based on the reference current state, a reference action language description, and a third set of reference frames; and a third update module configured to update the action model based on a third loss between the prediction of the reference action and the reference action.

[0124] According to an example implementation of the present disclosure, the reference current state includes at least any one of the following: a reference posture of the reference robot arm and a reference state of a reference tool of the reference robot arm, and the reference action involves at least any one of the following: a change in the reference posture and the reference state.

[0125] According to an example implementation of the present disclosure, the first number of the first set of reference frames is equal to the third number of the third set of reference frames, and the second number of the second set of reference frames is equal to the fourth number of the fourth set of reference frames.

[0126] According to an example implementation of the present disclosure, the motion model matches the application environment of the robot arm, and the application scenario includes at least any one of the following: a virtual application environment and a real application environment.

[0127] According to an exemplary implementation of the present disclosure, the apparatus further includes: an adjustment module configured to adjust the action to determine an action instruction for driving the robot arm.

[0128] According to an example implementation of the present disclosure, the motion model is pre-trained using reference data that does not include relevant data of the robot arm.

[0129] Figure 15 FIG1 shows a block diagram of a device 1500 capable of implementing various implementations of the present disclosure. Figure 15 The illustrated computing device 1500 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. Figure 15 The illustrated computing device 1500 may be used to implement the methods described above.

[0130] like Figure 15 As shown, computing device 1500 is in the form of a general-purpose computing device. Components of computing device 1500 may include, but are not limited to, one or more processors or processing units 1510, memory 1520, storage device 1530, one or more communication units 1540, one or more input devices 1550, and one or more output devices 1560. Processing unit 1510 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1500.

[0131] The computing device 1500 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 1500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1500.

[0132] The computing device 1500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 15 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 1520 may include a computer program product 1525 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0133] Communication unit 1540 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of computing device 1500 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, computing device 1500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.

[0134] Input device 1550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1560 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 1500 may also communicate with one or more external devices (not shown) via communication unit 1540, as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 1500, or with any device that allows computing device 1500 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0135] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0136] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0137] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0138] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0139] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0140] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for operating a robotic arm, comprising: receiving a language description specifying a goal to be achieved by the robotic arm; Obtaining the current state of the robot arm; as well as determining, according to an action model, an action to be performed by the robot arm based on the language description and the current state, wherein the action model is pre-trained using reference data including relevant data of a human arm, the reference data including a reference human video of a reference human action and a reference language description describing the human video, and the action model is pre-trained based on: The action model is pre-trained using the reference character video including the reference character action and the reference language description describing the reference character video to obtain the pre-trained action model.

2. The method according to claim 1, further comprising: According to the action model, an image prediction of a scene in which the robot arm performs the action is determined based on the language description and the current state.

3. The method according to claim 2, further comprising: receiving a location and number of steps specifying an action to be performed by the robotic arm; as well as Based on the motion model, motions matching the positions and numbers of the steps and the image predictions are determined.

4. The method according to claim 1, wherein the current state of the robot arm includes at least any one of the following: an image and a posture of the robot arm and a state of a tool of the robot arm, and the action involves changes in the posture and the state of the tool.

5. The method of claim 2, wherein the action model comprises a language encoder, a state encoder, and an action decoder, wherein determining the action comprises: determining, using the language encoder, a language representation of the language description; determining a state representation of the current state using the state encoder; as well as The action is determined based on the language representation and the state representation using the action decoder.

6. The method of claim 2, wherein the motion model further comprises an image decoder, and determining the image prediction comprises: The image prediction is determined, with the image decoder, based on the language representation and the state representation.

7. The method according to claim 2, wherein the motion model is obtained based on: The pre-trained motion model is fine-tuned using a reference robotic arm video comprising a reference robotic arm motion and a reference motion language description describing the robotic arm video to obtain a fine-tuned motion model.

8. The method according to claim 7, wherein pre-training the motion model comprises: Extracting a first group of reference frames and a second group of reference frames subsequent to the first group of reference frames from the reference character video; as well as determining, using the motion model, a prediction for the second set of reference frames based on the reference language description and the first set of reference frames; as well as The motion model is updated based on a first loss between the prediction of the second set of reference frames and the second set of reference frames.

9. The method of claim 8, wherein fine-tuning the pre-trained motion model comprises: extracting a third group of reference frames and a fourth group of reference frames subsequent to the third group of reference frames from the reference robot arm video; as well as Determining, using the pre-trained motion model, a prediction of the fourth set of reference frames based on the reference motion language description and the third set of reference frames; as well as The motion model is updated based on a second loss between the prediction of the fourth set of reference frames and the fourth set of reference frames.

10. The method according to claim 9, wherein fine-tuning the pre-trained motion model further comprises: Obtaining a reference current state and a reference motion of the reference robot arm; as well as Determining, using the pre-trained motion model, a prediction of the reference motion based on the reference current state, the reference motion language description, and the third set of reference frames; as well as The motion model is updated based on a third loss between the prediction of the reference motion and the reference motion.

11. A method according to claim 10, wherein the reference current state includes at least any one of the following: a reference posture of the reference robot arm and a reference state of a reference tool of the reference robot arm, and the reference action involves at least any one of the following: a change of the reference posture and the reference state.

12. The method of claim 9, wherein a first number of the first set of reference frames is equal to a third number of the third set of reference frames, and a second number of the second set of reference frames is equal to a fourth number of the fourth set of reference frames.

13. The method according to claim 1, wherein the motion model matches an application environment of the robot arm, and the application environment includes at least any one of the following: a virtual application environment and a real application environment.

14. The method according to claim 1, further comprising: The motion is adjusted to determine motion instructions for driving the robot arm.

15. A device for operating a robot arm, comprising: a receiving module configured to receive a language description specifying a goal to be achieved by the robotic arm; an acquisition module, configured to acquire the current state of the robot arm; as well as a determination module configured to determine an action to be performed by the robot arm based on the language description and the current state according to an action model, wherein the action model is pre-trained using reference data including relevant data of a human arm, the reference data including a reference human video of a reference human action and a reference language description describing the human video, and the action model is pre-trained based on: The action model is pre-trained using the reference character video including the reference character action and the reference language description describing the reference character video to obtain the pre-trained action model.

16. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 14 when executed by the at least one processing unit. 17 . A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to claim 1 .

Citation Information

Patent Citations

  • Controlling robot based on free-form natural language input

    CN112136141A