Robot control method, device and equipment and storage medium

CN121816248APending Publication Date: 2026-04-07BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing machine learning models lack accuracy and robustness when facing complex tasks, mainly due to the limitations of single-modal instructions and the lack of fully labeled language and action data, which limits the improvement of model performance.

Method used

A multimodal fusion robot control method is adopted, which combines natural language description, environmental images and robot pose information. Action sequences are generated through causal relationship converter and decoder. The model is trained using incomplete labeled samples and its generalization ability is improved by fine-tuning.

Benefits of technology

It improves the robot's accuracy and adaptability in performing tasks in complex environments, enhances task understanding and completion efficiency, and significantly improves the model's robustness and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121816248A_ABST
    Figure CN121816248A_ABST
Patent Text Reader

Abstract

A control method of a robot (120) includes receiving description information of a target task executed by the robot (120), generating a first action image related to the target task based on the description information and a first environment image related to the robot (120), and generating a second action image related to the target task based on the first action image, the first environment image, the description information and pose information of the robot (120). At least one action to be performed by the robot (120) upon completion of the target task is determined, and control information indicative of the at least one action is sent to the robot (120). A control device (600) of a robot (120) includes a description information receiving module (601), a motion image generating module (602), a motion determining module (603), and a control information transmitting module (604). An electronic device (700) comprising at least one processing unit (710) and at least one memory (720) for executing a control method of the robot (120). A computer readable storage medium having a computer program stored thereon is executable by a processor to implement a control method of the robot (120). A computer program product, including computer programs / instructions, when executed by a processor, implements a control method of the robot (120). Through the arrangement, the task understanding ability in the robot control process can be improved based on multi-modal fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Robot control method, apparatus, device, and storage medium TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a robot control method, apparatus, device, and storage medium. BACKGROUND

[0002] In recent years, robot technology has been rapidly developed and has been widely used in multiple technical fields. For example, on a production line of a factory, a robot arm can be used to perform multiple tasks such as machining, grabbing, sorting, packaging, etc. For another example, a robot can be used to assist housework in family life. Further, machine learning technology has also been widely applied to multiple application scenarios. At this time, it is expected that robot technology and machine learning technology can be combined, and then the operation of the robot can be controlled in a simpler and more effective manner.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a robot control method is provided. The method can include receiving description information of a target task performed by a robot. A first action image related to the target task is generated based on the description information and a first environment image related to the robot. At least one action to be performed by the robot when completing the target task is determined based on the first action image, the first environment image, the description information, and pose information of the robot. Control information indicating the at least one action is sent to the robot.

[0005] In a second aspect of the present disclosure, a robot control apparatus is provided. The apparatus can include a description information receiving module configured to receive description information of a target task performed by a robot. An action image generating module configured to generate a first action image related to the target task based on the description information and a first environment image related to the robot. An action determining module configured to determine at least one action to be performed by the robot when completing the target task based on the first action image, the first environment image, the description information, and pose information of the robot. A control information sending module configured to send control information indicating the at least one action to the robot.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The medium has stored thereon a computer program which, when executed by a processor, implements the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional manners in the first aspect of the present embodiment. In other words, the computer instructions, when executed by the processor, implement the method provided in various optional manners in the first aspect of the present embodiment.

[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:

[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG. 2 shows a training and use schematic diagram of a control method of a robot according to some embodiments of the present disclosure;

[0013] FIG. 3 shows an example diagram of a control method process of a robot according to some embodiments of the present disclosure;

[0014] FIG. 4 shows a schematic diagram of a control method of a robot according to some embodiments of the present disclosure;

[0015] FIG. 5 shows a training schematic diagram of an action model according to some embodiments of the present disclosure;

[0016] FIG. 6 shows a schematic structural block diagram of a control device of a robot according to some embodiments of the present disclosure; and

[0017] FIG. 7 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meaning of "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.

[0020] In this document, unless explicitly stated, performing a step "in response to A" does not mean performing the step immediately after A, but can include one or more intermediate steps.

[0021] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained by appropriate means, wherein the relevant user can include any type of right subject, such as an individual, an enterprise or a group.

[0023] For example, in response to receiving the active request of the user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require the information of the relevant user to be obtained and used, so that the relevant user can voluntarily choose whether to provide the information to the software or hardware such as electronic device, application program, server or storage medium performing the operation of the technical solutions of the present disclosure according to the prompt information.

[0024] As an optional but non-limiting implementation manner, in response to receiving the active request of the relevant user, the prompt information can be sent to the relevant user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.

[0025] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0026] As used herein, the term “model” can learn the relationship between the corresponding input and output from the training data, so that the corresponding output can be generated for a given input after the training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. The neural network model is an example of a model based on deep learning. In this document, “model” can also be referred to as “machine learning model”, “learning model”, “machine learning network” or “learning network”, which are used interchangeably herein.

[0027] A “neural network” is a machine learning network based on deep learning. The neural network is capable of processing input and providing a corresponding output, which generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. The neural network used in deep learning applications usually includes many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.

[0028] Generally, machine learning can include three stages, namely a training stage, a testing stage and an application stage (also known as an inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are updated iteratively until the model can obtain consistent inference from the training data that meets the expected target. Through training, the model can be considered to learn the relationship between input and output (also known as the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained by training to determine the corresponding model output.

[0029] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, in an application environment 110, a robot 120 can be used to operate various objects 130 in the application environment 110. Here, the objects 130 can be food such as fruits and vegetables or items such as knives and forks. For example, the robot 120 can be used to pick up an object on a table and place it into a dish; for another example, the robot 120 can be used to take a building block out of a drawer and place it into a basket on the ground, and so on.

[0030] Machine learning models for controlling robot actions based on visual data and language data have been developed. However, most of the above machine learning models are guided by instructions in a single modality, such as language instructions or image instructions. As a result, the models often show insufficient accuracy and robustness when facing complex tasks. In addition, in the field of training machine learning models, data annotated with language instructions and actions is relatively scarce, and most existing technologies can only use part of the incomplete annotated data, such as data lacking language annotation or lacking action annotation. This limitation limits the performance improvement of the model.

[0031] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a control method of a robot is proposed. Referring to FIG. 2, which describes an overview of one example implementation of the present disclosure, FIG. 2 shows a schematic diagram of the principle 200 of the control method of the robot according to some implementations of the present disclosure. As shown in FIG. 2, in the training phase, the training of the machine learning model related to the robot can be completed by making full use of incomplete annotated samples 201. The incomplete annotated samples 201 can be a set of human activity samples, a set of action samples of different types of robots, and a set of robot sample pose information.

[0032] In the set of human activity samples and the set of action samples of different types of robots, only action videos and language annotations of the action videos are included, without annotations of the actions. The language annotation can indicate the purpose of the action, such as picking up a building block or closing a drawer, and so on. The action annotation can indicate the movement of the robot or the operation action of the mechanical arm. In the set of robot sample pose information, only the action annotation is included, without the language annotation. In the present disclosure, the training of the machine learning model related to the control of the robot can be completed by making full use of the incomplete annotated samples 201. The training process will be described in detail later. In addition, a small amount of complete annotated samples 202 can also be used to fine-tune the machine learning model.

[0033] In the control stage of the robot, the input to the machine learning model can include description information 211 of the target task, environment images 212 related to the robot, pose information of the robot, and other multi-dimensional information. Based on the multi-dimensional information, the machine learning model can determine the task completion progress 221 and determine at least one action 222. The task completion progress 221 can indicate the completion of the current task by the robot. Based on the at least one action 222, control information indicating the at least one action can be sent to the robot.

[0034] In the following, one example implementation according to the present disclosure will be described in a Chinese language environment. Alternatively and / or additionally, the technical solutions of one example implementation according to the present disclosure can be performed in other language environments. For example, the robot can be controlled in a Chinese, English, Japanese, French, etc. environment. For example, the robot can be controlled in different language application environments based on the multi-language capability provided by the machine learning technology. For ease of description, in the following, the process of controlling the robot will be described only by taking picking up a building block from a drawer as an example. Alternatively and / or additionally, the robot can perform other actions, for example, the robot arm can be used to process a part to a predetermined size, package various items, etc.

[0035] FIG. 3 shows an example flow 300 of a control method of a robot according to some embodiments of the present disclosure. For ease of discussion, the flow 300 will be described with reference to the environment of FIG. 1. The flow 300 can be implemented in a master control device of the robot 120. The master control device can be deployed on the robot 120 body. The master control device can also be deployed in the cloud with the function of communicating with the robot 120 body.

[0036] At block 301, the master control device receives description information of a target task to be performed by the robot 120. The description information is usually provided in the form of natural language, which describes in detail the specific task that the robot 120 needs to complete. For example, the description information can include “pick up the red building block from the drawer”, “put the bread slices into the toaster”, etc. The description information is the basis for the robot 120 to understand the task target, and the description information can be in the form of voice, text, etc.

[0037] At block 302, the master control device controls to generate a first action image related to the target task based on the description information and a first environment image related to the robot 120. The environment image can be obtained by a camera or other visual sensor built in the robot 120, which reflects the current state of the environment where the robot 120 is located. For example, the first environment image can be a plurality of images collected at the same time from different angles, or at least one panoramic image synthesized from a plurality of images from different angles, etc.

[0038] Based on the task description and the environment image, the host device can generate an expected target action image using an image generation model. The action image shows the ideal state of the robot 120 when performing a specific task. For example, the first environment image is an image including a drawer. The description information is "pick up the red block from the drawer", and the content of the action image can indicate a picture related to the robot 120's arm grabbing the red block. That is, the content indicated by the action image conforms to the content of the description information.

[0039] At block 303, the host device determines at least one action to be performed by the robot 120 in completing the target task based on the first action image, the first environment image, the description information, and the pose information of the robot 120.

[0040] Based on the target action image, the environment image, the description information, and the pose information of the robot 120, at least one action required to complete the target task can be determined. The pose information of the robot 120 can include the spatial coordinates of the robot 120 and the arm state of the robot 120. For example, the arm state can indicate that the arm is in an open state, a grasping state, and the like.

[0041] For the determination of the action, an action model can be used. Based on the input information, the action model can generate an action sequence containing one or more actions. These action sequences indicate the steps for the robot 120 to complete the task. For example, the action sequence can indicate that the robot 120 moves forward by 3 movement units, the arm moves down by 1 movement unit, and the like.

[0042] At block 304, the host device sends control information indicating the at least one action to the robot 120. After determining the action sequence, the host device converts the action sequence into specific control information and sends it to the action execution unit of the robot 120. For example, the control information can contain the instruction parameters required by the robot 120 to perform each action, such as joint angles, movement paths, and speeds, etc. The robot 120 receives and executes these control information, and gradually completes the target task.

[0043] Through the above process, the present disclosure effectively solves the problem of insufficient accuracy and robustness caused by single modality instruction in the related art. By comprehensively utilizing natural language description, environment image and pose information of the robot 120, multi-modal fusion is used to improve the accuracy of task understanding and execution. The whole process not only improves the task completion efficiency of the robot 120, but also significantly enhances the adaptability of the robot 120 in a diversified environment.

[0044] The control principle of the robot 120 is briefly introduced above. In order to improve the accuracy and robustness of the robot 120 in a complex task environment, the present disclosure introduces a multi-modal fusion action model. The action model generates an action sequence capable of guiding the robot 120 to perform a complex task by integrating task description information, environment images, target action images, and pose information of the robot 120. The working principle and implementation steps of the action model are described in detail below. In some embodiments, the feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and the first task label constitute a first input sequence of the action model, and the first task label indicates that the action model performs action prediction. At least one action is determined by processing the first input sequence using the action model.

[0045] FIG. 4 shows a control principle 400 schematic diagram of a master device according to some embodiments of the present disclosure. In the example of FIG. 4, the action model 402 can be composed of at least two parts, a causal transformer and a decoder. A feature encoder can be used as an auxiliary of the action model 402, responsible for converting different types of input data into unified feature representations. The causal transformer performs multi-modal fusion and processing using these feature representations. The decoder then converts the processing results into specific action instructions.

[0046] First, the description information 211 of the target task, the environment image 212 related to the robot 120, the target action image 431, and the pose information 432 of the robot 120 are input into the feature encoder. The feature encoder includes an image encoder 402-A1, a language encoder 402-A2, and a pose encoder 402-A3, which process different types of input data respectively. For example, the image encoder 402-A1 can use a masked auto encoder (MAE) to encode the environment image 212 related to the robot 120 and the target action image 431. The language encoder 402-A2 can use a contrastive language image pre-training (CLIP) technique to encode the description information 211 of the target task. The pose encoder 402-A3 can use a multilayer perceptron (MLP) to encode the pose information 432 of the robot 120.

[0047] The role of the task label includes increasing the functionality of the model. For example, the first task label can be represented as [ACT], which is used to indicate that the action model 402 can perform action prediction. The feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and the first task label together constitute the first input sequence of the causal relationship converter 402-B. The causal relationship converter 402-B performs processing and fusion related to determining actions on the first input sequence through a multi-layer self-attention mechanism to generate a high-level feature representation.

[0048] The high-level feature representation obtained by processing the causal relationship converter 402-B is input to the action decoder 402-C3 in the decoder. The action decoder 402-C3 converts the high-level feature representation into at least one action. These actions can guide the robot 120 to perform actions step by step until the target task is completed. In order to improve the accuracy of the actions determined by the action model 402, the live information 433 of the ground of the environment where the robot 120 is located can also be referred to in real time when predicting actions, and the feature representation of the ground live is obtained through the style encoder to indicate corresponding conditions such as slope conditions, obstacle conditions, and the like.

[0049] By introducing the multi-modal fusion action model 402, various information sources can be fully utilized to improve the task execution capability of the robot 120 in a complex environment. For example, the fusion of multi-modal information enables the model to more accurately understand the task and the environment, improving the accuracy of action prediction. It can still maintain a high task completion rate in the case of missing or noisy information.

[0050] The function of the first task label is described in detail above. In addition to the first task label, a second task label can also be included, which indicates that the action model 402 performs prediction of the completion progress of the task. That is, the action model 402 can determine the current completion progress of the robot 120 for the target task based on the first action image, the first environment image, the description information of the target task, and the pose information of the robot 120. For example, if it is the starting time, the completion progress of the target task is zero. With the determination and execution of actions, the completion progress of the target task will gradually change, such as completing 30%, completing 80%, and the like. Specifically, the feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and the second task label constitute the second input sequence of the action model 402, and the second task label indicates that the action model 402 performs task completion progress prediction. By processing the second input sequence using the action model 402, the completion progress of the target task can be determined.

[0051] The second task label can be denoted as [PRO] to indicate that the action model 402 can perform the prediction of the progress of the task completion. Still in connection with the illustration of FIG. 4, for example, if the feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and the second task label are taken as the input sequence, the action model 402 can predict the progress of the robot 120 for the target task based on the input sequence.

[0052] Similar to the foregoing principle, the processing and fusion related to determining the progress of the task are performed on the features in the input sequence in the causal relationship converter 402-B to generate a high-level feature representation. The high-level feature representation processed by the causal relationship converter 402-B is input to the progress decoder 402-C1 in the decoder to obtain the predicted value of the progress of the task completion. The predicted value of the progress of the task completion can indicate the percentage of the completion of the target task by the robot 120, such as 30%, 80%, and the like.

[0053] It is not difficult to understand that if the feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and the first task label and the second task label are taken as the input sequence, the action model 402 can output at least one action and the progress of the corresponding target task at the same time.

[0054] Based on the progress of the task completion, in some embodiments, the second action image can also be determined by processing a third input sequence by using the image model 401. The third input sequence is composed of the description information, the progress of the task completion, and a second environment image related to the robot 120, and the second environment image is collected at a time later than the first environment image.

[0055] During the execution of the task by the robot 120, the environment also changes with the change of the pose of the robot 120. During the execution of the task by the robot 120, the environment image needs to be collected in real time. Exemplarily, the second environment image is collected after the first environment image, reflecting the change of the environment during the execution of the task by the robot 120.

[0056] The progress of the target task determined by the action model 402 is taken as the new input of the image model 401. Exemplarily, the progress of the target task can be taken as additional information and appended to the description information to obtain new description information. Exemplarily, the new description information can be “pick up the red block from the drawer, and 30% of the action is completed”.

[0057] The new description information and the second environment image constitute a new input sequence of the image model 401. Based on the new input sequence, the image model 401 can generate a second action image. For example, in the case where the description information is “pick up the red block from the drawer, and 30% of the action has been completed”, the generated second action image can display that the arm of the robot 120 has approached the block. For another example, in the case where the description information is “pick up the red block from the drawer, and 80% of the action has been completed”, the generated second action image can display that the arm of the robot 120 has reached the block.

[0058] By integrating the task progress information into the description information, the image model 401 can generate more accurate target images to guide the actions determined by the robot 120. In this way, the target deviation caused by the lack of progress information is avoided, thereby improving the accuracy of task execution. In addition, since each step of action can be matched with the completion progress of the current task, the coherence of task execution is maintained and the accuracy of action determination is improved.

[0059] After introducing the control principle of the robot 120 in detail, the training of the machine learning model involved in the control process is introduced below. For example, the first action image is generated by the image model 401. In an embodiment, the image model 401 can be trained by the following method: selecting at least one sample image from a sample video corresponding to the execution process of a task sample. Determine the completion progress of the task sample corresponding to a given sample image in the at least one sample image. Obtain sample description information of the task sample. Based on the given sample image, the completion progress of the task sample, and the sample description information, train the image model 401.

[0060] The task sample video can be a video lacking action label annotation, and the trajectory τ am may be expressed as expression (1)

[0061] τ am = {l, o1, o2, o3, … o T} (1)

[0062] In expression (1), l can be expressed as language description information of the task sample video, o1, o2, …, o T may represent each video frame in the video, and T represents the number of video frames. The action label can indicate the movement trajectory of the robot and the arm state. The sample video lacking action label annotation can be a video of human activity, or a video corresponding to a cross-robot body. Thus, the number of available training materials can be significantly improved.

[0063] Taking a video of human activity as an example, it can be a video of a human taking a small ball out of a drawer, a video of a human putting a bread slice into a toaster, a video of a human cutting vegetables, and so on. The video corresponding to the robot body can be a video of picking up different objects such as a ball, a bread, and the like by any robot.

[0064] Each sample video corresponds to description information of a task sample. Taking a video of a human putting a bread slice into a toaster as an example, the description information can be “taking a bread slice out of a pocket and putting it into a groove of a toaster”. These description information provides the context information required for the image model 401 to generate a target image.

[0065] A plurality of sample images can be randomly extracted from the sample video, and each sample image in the sample video can correspond to a target image, which can be used as the annotation data of the sample image. For example, the i-th (i is a positive integer) frame image in the sample video can be used as a given sample image, and the i+n-th (n is a positive integer) frame image can be used as the target image corresponding to the given sample image. Based on the total length of the video and the time corresponding to the given frame image in the video, the task sample completion progress corresponding to the given sample image can be determined. If the sample video has a total of 20 frame images, the task sample completion progress corresponding to the 10th frame image can be 50%.

[0066] Based on the given sample image, the task sample completion progress, and the sample description information, the image model 401 can be trained. The specific process of training includes: combining the feature representation of the given sample image, the feature representation of the task sample completion progress, and the feature representation of the sample description information to form a first input sequence sample of the image model 401 to be trained. By processing the first input sequence sample by using the image model 401 to be trained, an action sample image is obtained. The training loss is determined based on the action sample image and the target image, and the target image is determined from the sample video corresponding to the given sample image.

[0067] The task sample completion progress and the sample description information can be spliced together, so that the feature representation of the task sample completion progress and the feature representation of the sample description information can be determined by using a language encoder. In addition, the feature representation of the given sample image can be determined by using an image encoder. The above feature representations can be spliced to form the first input sequence sample of the image model 401 to be trained.

[0068] By processing the first input sequence sample by using the image model 401 to be trained, an action sample image is obtained. The training loss is determined based on the difference between the action sample image and the target image. Based on the training loss, the parameters of the image model 401 can be adjusted.

[0069] To improve the generalization ability of the model, before generating the action sample image, random noise can be added to the feature representation corresponding to the action sample image to obtain a feature representation with added noise. The decoder is used to process the feature representation with added noise to obtain an action sample image containing noise. The training loss is determined based on the difference between the action sample image containing noise and the target image. Based on the training loss, the parameters of the image model 401 can be adjusted.

[0070] Through the above training process, the first aspect image model 401 can be free from the constraint of complete labeled samples and complete training using a large number of incomplete labeled samples. On the other hand, by introducing progress information, the image model 401 can more accurately predict the target image at different task stages, improving the accuracy of task execution. In addition, the image model 401 can dynamically adjust the generated target image according to the progress of the current task, ensuring that the robot can respond and adjust in real time during task execution, improving the efficiency and continuity of task execution.

[0071] As mentioned earlier, at least one action is determined by the action model 402. The training process of the action model 402 is described in detail below. The pre-trained action model 402 is subjected to a first fine-tuning using the sample video corresponding to the execution process of the task sample and the sample pose information of the robot. The action model 402 subjected to the first fine-tuning is subjected to a second fine-tuning using the sample video, the sample description information and the sample pose information of the task sample.

[0072] The pre-training process can be performed using the publicly available Ego4D dataset (or other datasets). The Ego4D dataset contains a large amount of first-person video data, which can be used to train the action model 402 to have the ability to understand the initial state and predict the action sequence.

[0073] The fine-tuning of the action model 402 can be divided into two stages. The first fine-tuning stage and the second fine-tuning stage. The first fine-tuning stage can be mainly based on samples lacking language annotation. In the first fine-tuning stage, the input of all description information of the action model 402 is regarded as an empty string. Such samples lacking language annotation can significantly improve the performance of the model when complete labeled samples are scarce. After the first fine-tuning stage, the action model 402 can perform the second fine-tuning stage based on complete labeled samples.

[0074] In the first fine-tuning stage, training can be completed based on samples lacking language annotation. These sample videos have frame-by-frame labeled actions and robot pose information, but no language description. The robot in these sample videos is not limited to manipulation tasks and can perform any meaningful action. These trajectories τ lm which can be expressed as expression (2):

[0075] τlm ={(o1, s1, a1), (o2, s2, a2), (o3, s3, a3), …, (o T , s T , a T )} (2)

[0076] In expression (2), o1, o2, …, o T may represent each video frame in the sample video, and T represents the number of video frames. s1, s2, …, s T may represent the pose information of the robot corresponding to the video frame. a1, a2, …, a T may represent the action of the robot corresponding to the video frame.

[0077] The action model 402 in the first fine-tuning stage can determine, based on the input, the action prediction and task progress prediction, etc. The training loss is determined based on the difference between the prediction result and the labeled result, so that the parameters of the action model 402 can be adjusted based on the training loss.

[0078] In the first fine-tuning stage, the training of the action model 402 does not depend on the language description, but is trained based only on the video and pose information, which improves the generalization ability of the model in different tasks. This means that the model can be more flexible and accurate when facing new tasks.

[0079] In the second fine-tuning stage, the action model 402 can be fine-tuned again using complete labeled samples. The complete labeled sample means that it includes the action with frame-by-frame labeling, the pose information of the robot, and the corresponding task description information in the sample video. The trajectory τ described by the sample video is completely labeled data, which can be expressed as expression (3):

[0080] τ={l,(o1, s1, a1), (o2, s2, a2), (o3, s3, a3), …, (o T , s T , a T )} (3)

[0081] In expression (3), l can be represented as the language description information of the sample video. o1, o2, …, o T may represent each video frame in the sample video, and T represents the number of video frames. s1, s2, …, s T may represent the pose information of the robot corresponding to the video frame. a1, a2, …, a T may represent the action of the robot corresponding to the video frame.

[0082] The lack of language-labeled samples in the first stage lays the foundation for the action model 402, which enhances its generalization ability by relying on visual and pose information, enabling the action model 402 to handle different tasks and environments. The fully labeled samples in the second stage can further fine-tune the parameters of the action model 402, optimizing the action model 402 by combining language descriptions. Through the two-stage fine-tuning process, the action model 402 not only has stronger generalization ability, but also improves the accuracy and robustness of task execution, ultimately performing well in practical applications.

[0083] The first fine-tuning stage is introduced below. The feature representation of the given sample image, the feature representation of the sample pose information, and the multiple task labels are combined to form a second input sequence sample. The given sample image is a first given sample image 512-A and a second given sample image 512-B with a specified time interval obtained from a sample video. The multiple task labels indicate that the action model 402 performs image prediction, action prediction, and task completion progress prediction. The predicted image indicates the prediction of the environment image during the execution of the task sample by the robot. By processing the second input sequence sample using the action model 402 that performs fine-tuning, image prediction results, at least one action prediction result, and task completion progress prediction results are obtained. Based on the image prediction results, at least one action prediction result, and task completion progress prediction results and their respective labeled results, a training loss is determined to adjust the parameters of the action model 402.

[0084] FIG. 5 shows a training diagram of the first fine-tuning stage 500 of the action model 402 according to some embodiments of the present disclosure. The first given sample image 512-A and the second given sample image 512-B with a specified time interval are obtained from a sample video. The first given sample image 512-A can simulate an environment image related to the robot, and the second given sample image 512-B can simulate an action image. Exemplarily, the first given sample image 512-A can simulate an environment image related to the robot from two dimensions of images of the environment where the robot is located and images of the robot’s arm. The image encoder 402-A1 encodes the first given sample image 512-A and the second given sample image 512-B to obtain the feature representations of the first given sample image 512-A and the second given sample image 512-B.

[0085] Since there is no sample description information of the task sample, the input of the speech encoder 402-A2 is empty. The pose encoder 402-A3 encodes the sample pose information 532 of the robot to obtain the feature representation of the sample pose information.

[0086] The plurality of task labels can include the first task label and the second task label mentioned above. The first task label can indicate the task completion progress prediction 521, and the second task label can indicate the action prediction 522. In addition, a third task label can also be included. The third task label can be represented as [OBS], and the third task label can indicate that the action model 402 performs image prediction 523. The predicted image indicates a prediction of an environment image during the execution of the task sample by the robot. For example, the current time is t1. Based on the third task label, the action model 402 can predict the environment image related to the robot at time t2 or time t3. By processing the feature representation output by the causal relationship converter 402-B through the image decoder 402-C3, the predicted image can be obtained.

[0087] The plurality of feature representations and the three task labels are used as inputs of the action model 402 in the first fine-tuning stage. Thus, the action model 402 can output the prediction result corresponding to the task completion progress, the prediction result corresponding to the image, and the prediction result corresponding to the action. During the prediction process, the live information 533 of the ground sample of the environment in which the robot is located can also be referred to in real time, and the live feature representation of the ground sample can be obtained through the style encoder to indicate corresponding conditions such as slope and obstacle.

[0088] The labeled result of the task completion progress is obtained, the prediction result corresponding to the task completion progress is used, and the labeled result of the task completion progress is used to determine the first training loss. The labeled result corresponding to the image is obtained, the prediction result corresponding to the image is used, and the labeled result corresponding to the image is used to determine the second training loss. The labeled result corresponding to the action is obtained, the prediction result corresponding to the action is used, and the labeled result corresponding to the action is used to determine the third training loss. The parameters of the action model 402 are adjusted by comprehensively considering the first training loss, the second training loss, and the third training loss.

[0089] The introduction of image prediction in the training process of the action model 402 can enable the action model 402 to have the ability to predict the future environment state. In this way, the action model 402 can complement the action prediction through video prediction to enhance the understanding of the environment, thereby improving the accuracy of the action prediction.

[0090] For the second fine-tuning stage, the training principle is the same as that of the first fine-tuning stage. The difference is that the second fine-tuning stage uses complete labeled samples, that is, the sample description information of the corresponding task sample is introduced on the basis of the first fine-tuning stage, so that the action model 402 has higher action prediction accuracy. In the example shown in FIG. 4, the output of the action model 402 includes the task completion progress and at least one action. Since the trained action model 402 also has the ability of image prediction. Thus, according to the actual scene requirements, the output of the predicted image can be realized by setting the image decoder.

[0091] For the aforementioned environment image, in some embodiments, the environment image can include an image of an environment where the robot 120 is located, and an image of an arm of the robot 120 being captured.

[0092] In a complex task environment, the robot 120 needs to have a comprehensive perception of the environment it is in, in order to accurately perform the task. The environment image can not only include the overall perspective of the environment where the robot 120 is located, but also include an image of the arm of the robot 120 being captured. This can provide more information about the interaction between the robot 120 and the environment, thereby improving the accuracy of action prediction and the effectiveness of task execution.

[0093] For the aforementioned at least one action, in some embodiments, the at least one action can correspond to a movement action of the robot and a manipulation action of the arm of the robot.

[0094] In the process of the robot 120 performing a task, the action can generally be divided into two categories: movement action of the robot 120 as a whole and manipulation action of the arm of the robot 120. Different types of actions play different roles in task execution. The movement action refers to the movement behavior of the robot 120 as a whole, including moving to a specified position in the environment, avoiding obstacles, etc. The movement action usually involves the movement mechanism of the robot 120, such as wheels, tracks, etc., and the displacement of the robot 120 is achieved by controlling these mechanisms. The arm manipulation action refers to the specific operation behavior of the arm of the robot 120, including grasping, carrying, placing objects, etc. The arm manipulation action usually involves multiple joints of the arm of the robot 120, and fine operations are achieved by controlling the movement of these joints.

[0095] FIG. 6 shows a schematic structural block diagram of a control device 600 of a robot according to some embodiments of the present disclosure. The device 600 may, for example, be implemented in or included in a master control device of a robot. Various modules / components in the device 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0096] As shown, the device 600 includes a description information receiving module 601 configured to receive description information of a target task performed by a robot. An action image generating module 602 is configured to generate a first action image related to the target task based on the description information and a first environment image related to the robot. An action determining module 603 is configured to determine at least one action to be performed by the robot in completing the target task based on the first action image, the first environment image, the description information, and pose information of the robot. A control information sending module 604 is configured to send control information indicating the at least one action to the robot.

[0097] In some embodiments, the action determination module 603 can be specifically configured to: compose the feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and a first task label into a first input sequence of the action model, the first task label indicating that the action model performs action prediction. By processing the first input sequence by using the action model, the at least one action is determined.

[0098] In some embodiments, further comprising a completion progress determination module configured to: determine a completion progress of the robot on the target task based on the first action image, the first environment image, the description information, and the pose information of the robot.

[0099] In some embodiments, the completion progress determination module is specifically configured to: compose the feature representation of the first action image, the feature representation of the first environment image, the feature representation of the description information, the feature representation of the pose information, and a second task label into a second input sequence of the action model, the second task label indicating that the action model performs task completion progress prediction. By processing the second input sequence by using the action model, the completion progress of the target task is determined.

[0100] In some embodiments, the action image generation module 602 can be further configured to: compose the description information, the completion progress of the target task, and a second environment image related to the robot into a third input sequence of the image model, the second environment image being captured at a time later than the first environment image. By processing the third input sequence by using the image model, the second action image is determined.

[0101] In some embodiments, further comprising an image model training module configured to: select at least one sample image from a sample video corresponding to an execution process of a task sample. Determine a task sample completion progress corresponding to a given sample image in the at least one sample image. Obtain sample description information of the task sample. Train the image model based on the given sample image, the task sample completion progress, and the sample description information.

[0102] In some embodiments, the image model training module can be specifically configured to: compose a feature representation of the given sample image, a feature representation of the task sample completion progress, and a feature representation of the sample description information into a first input sequence sample of the image model to be trained. By processing the first input sequence sample by using the image model to be trained, an action sample image is obtained. A training loss is determined based on the action sample image and a target image to adjust parameters of the image model, the target image being determined from the sample video and corresponding to the given sample image.

[0103] In some embodiments, an action model training module is further included. The module can be configured to perform fine-tuning on the pre-trained action model by using the sample video corresponding to the execution process of the task sample and the sample pose information of the robot.

[0104] In some embodiments, the action model training module can be specifically configured to: compose a second input sequence sample by the feature representation of a given sample image, the feature representation of the sample pose information, and a plurality of task labels, the given sample image being a first given sample image 512-A and a second given sample image 512-B with a given time interval acquired from the sample video, the plurality of task labels indicating image prediction, action prediction, and task completion progress prediction performed by the action model, the predicted image indicating a prediction of an environment image in the execution process of the robot on the task sample. By processing the second input sequence sample by using the action model performing fine-tuning, an image prediction result, at least one action prediction result, and a task completion progress prediction result are obtained. A training loss is determined based on the image prediction result, the at least one action prediction result, and the task completion progress prediction result and the respective corresponding labeled results, to adjust the parameters of the action model.

[0105] In some embodiments, the environment image includes an image of an environment where the robot is located, and an image of an arm of the robot.

[0106] In some embodiments, the at least one action corresponds to a movement action of the robot and a control action of the arm of the robot.

[0107] FIG. 7 illustrates a block diagram of an electronic device 700 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 700 illustrated in FIG. 7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 illustrated in FIG. 7 can include or be implemented as a master device of a robot, or the apparatus 600 of FIG. 6.

[0108] As shown in FIG. 7, the electronic device 700 is in the form of a general electronic device. The components of the electronic device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be an actual or virtual processor and is capable of performing various processing according to programs stored in the memory 720. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.

[0109] The electronic device 700 typically includes a plurality of computer storage media. Such media can be volatile and / or nonvolatile, removable and / or non-removable, and can be implemented in any method or technology for storage of information and / or data. The memory 720 can be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), EEPROM, flash memory, etc.), or some combination of the two. The storage device 730 can be a removable or non-removable media implemented in any method or technology for storage of information and / or data such as, for example, flash drives, magnetic disks, optical disks, or any other medium. The memory 720 and / or the storage device 730 can be used to store, among other things, the computer program product 725.

[0110] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive and / or a CD drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media such as a floppy disk, a CD-ROM, and so on. In such instances, each can be connected to the bus by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0111] The communication unit 740 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0112] The input device 750 can be one or more input devices such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices such as a display, a speaker, a printer, etc. The electronic device 700 can further communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 740, with one or more devices that enable a user to interact with the electronic device 700, or with any devices (e.g., a network card, a modem, etc.) that enables the electronic device 700 to communicate in a network environment. Such communication can be enabled by an input / output (I / O) interface (not shown).

[0113] According to an example implementation of the present disclosure, a computer readable storage medium is provided, having computer executable instructions stored thereon, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method described above.

[0114] According to an example implementation of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional manners in FIG. 3. Therefore, no further description will be given here.

[0115] Various aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0116] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can comprise a non-transitory computer readable medium. The instructions stored in the computer readable storage medium, which can comprise a non-transitory computer readable medium, can cause a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable medium having instructions stored therein comprises an article of manufacture including a manufacture of aspects of the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0117] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0118] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk drive, or any other suitable non-transitory computer readable medium can store the computer program product.

[0119] Having described several implementations of the present disclosure, it is to be appreciated various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be part of this disclosure. Accordingly, the foregoing description is by way of example only and is not intended to be limiting. The implementation described herein is implementations of the present disclosure. Other implementations of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure. Therefore, this disclosure is intended to cover all such modifications and variations as fall within the scope of the implementations. It is intended that the specification and depicted embodiments are to be considered exemplary only, with a true scope and spirit of the disclosure being indicated by the following claims.

Claims

1. A method for controlling a robot, comprising: receiving description information of a target task to be performed by a robot; generating a first action image related to the target task based on the description information and a first environment image related to the robot; determining at least one action to be performed by the robot when completing the target task based on the first action image, the first environment image, the description information, and pose information of the robot; and sending control information indicative of the at least one action to the robot. 2.The method of claim 1, wherein determining the at least one action comprises: composing a first input sequence of an action model with a feature representation of the first action image, a feature representation of the first environment image, a feature representation of the description information, a feature representation of the pose information, and a first task label, the first task label instructing the action model to perform action prediction; and determining the at least one action by processing the first input sequence with the action model. 3.The method of claim 1, further comprising: determining a completion progress of the target task by the robot based on the first action image, the first environment image, the description information, and the pose information of the robot. 4.The method of claim 3, wherein determining the completion progress of the target task by the robot comprises: composing a second input sequence of an action model with a feature representation of the first action image, a feature representation of the first environment image, a feature representation of the description information, a feature representation of the pose information, and a second task label, the second task label instructing the action model to perform task completion progress prediction; and determining the completion progress of the target task by processing the second input sequence with the action model. 5.The method of claim 3, further comprising: composing a third input sequence of an image model with the description information, the completion progress of the target task, and a second environment image related to the robot, the second environment image being captured at a time later than the first environment image; and determining a second action image by processing the third input sequence with the image model. 6.The method of claim 1, wherein the first action image is generated by an image model, and the image model is trained by: selecting at least one sample image from a sample video corresponding to an execution process of a task sample; determining a task sample completion progress corresponding to a given sample image in the at least one sample image; obtaining sample description information of the task sample; and training the image model based on the given sample image, the task sample completion progress, and the sample description information. 7.The method of claim 6, wherein training the image model comprises: composing a first input sequence sample of the image model to be trained with a feature representation of the given sample image, a feature representation of the task sample completion progress, and a feature representation of the sample description information. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ action sample image is obtained by processing the first input sequence sample using the image model to be trained; and a training loss is determined based on the action sample image and a target image, the target image being determined from the sample video and corresponding to the given sample image, to adjust parameters of the image model.

8. The method of claim 1, wherein the at least one action is determined by an action model trained by: performing fine-tuning of the pre-trained action model using a sample video corresponding to an execution process of a task sample and sample pose information of a robot.

9. The method of claim 8, wherein the action model is fine-tuned by: composing a second input sequence sample from a feature representation of a given sample image, a feature representation of the sample pose information, and a plurality of task labels, the given sample image being a first given sample image and a second given sample image having a time interval obtained from the sample video, the plurality of task labels indicating that the action model performs image prediction, action prediction, and task completion progress prediction, the predicted image indicating a prediction of an environment image during execution of the task sample on the robot; obtaining a predicted image, at least one action prediction, and a task completion progress prediction by processing the second input sequence sample using the fine-tuned action model; and determining a training loss based on the predicted image, the at least one action prediction, and the task completion progress prediction and respective ground truth results to adjust parameters of the action model.

10. The method of claim 1, wherein the environment image includes an image of an environment in which the robot is located and an image of an arm of the robot.

11. The method of claim 1, wherein the at least one action corresponds to a movement action of the robot and a manipulation action of the robot arm.

12. A control apparatus of a robot, comprising: a description information receiving module configured to receive description information of a target task to be executed by a robot; an action image generating module configured to generate a first action image related to the target task based on the description information and a first environment image related to the robot; an action determining module configured to determine at least one action to be executed by the robot when completing the target task based on the first action image, the first environment image, the description information, and pose information of the robot; and a control information sending module configured to send control information indicating the at least one action to the robot.

13. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit, cause the electronic device to perform the method of any one of claims 1-11. ​ ​ ​ 14. A computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method of any one of claims 1 to 11.

15. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the method of any one of claims 1 to 11.