Method, device, equipment and medium for determining action of robotic equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2024-09-10
- Publication Date
- 2026-05-12
AI Technical Summary
The robot equipment is not accurate enough when performing tasks, making it difficult to effectively acquire target items.
By receiving user tasks, the robot acquires image sequences of the environmental space and state sequences of the robot equipment. It then uses a motion model to predict future actions based on the images and state sequences, thereby controlling the robot equipment to acquire the target item.
It improves the accuracy and efficiency of robotic devices in acquiring target objects in complex environments, and enables them to flexibly handle complex situations in a variety of real-world environments.
Smart Images

Figure CN122029015A_ABST
Abstract
Description
Methods, apparatuses, devices, and media for determining actions of robotic devices TECHNICAL FIELD
[0001] Exemplary implementations of the present disclosure generally relate to the field of robotics, and particularly relate to methods, apparatuses, devices, and computer-readable storage media for determining actions of robotic devices using robots. BACKGROUND
[0002] Robotic technology has been rapidly developed and has been widely used in multiple technical fields. Currently, a variety of special robotic devices have been developed, for example, in an industrial environment, robots can be used to perform multiple tasks such as machining, grabbing, sorting, packaging, etc. For another example, in a home environment, a sweeping robot, a glass wiping robot, etc. have been developed. However, the accuracy of robotic devices performing various tasks is not satisfactory, and it is desirable to improve the accuracy of robotic devices to perform specified tasks in a more accurate manner.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a method for determining actions of a robotic device is provided. In the method, a user task is received, the user task indicating the robotic device to acquire a target item of a target type. A set of image sequences of an environment space in which the robotic device is located is acquired, an image sequence in the set of image sequences including a first plurality of images of the robotic device at a first plurality of time points, respectively. In response to determining that the set of image sequences includes the target item, a state sequence of the robotic device is acquired, the state sequence including a first plurality of states of the robotic device at the first plurality of time points, respectively. Using an action model, a second plurality of actions of the robotic device at a second plurality of time points is determined based on the set of image sequences and the state sequence, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robotic device to acquire the target item.
[0005] In a second aspect of the disclosure, there is provided an apparatus for determining actions of a robotic device. The apparatus comprises: a receiving module configured to receive a user task, the user task indicating the robotic device to obtain a target item of a target type; an image obtaining module configured to obtain a set of image sequences of an environment space where the robotic device is located, an image sequence in the set of image sequences comprising a first plurality of images of the robotic device at a first plurality of time points respectively; a state obtaining module configured to, in response to determining that the set of image sequences comprises the target item, obtain a state sequence of the robotic device, the state sequence comprising a first plurality of states of the robotic device at the first plurality of time points respectively; and a determining module configured to determine, based on the set of image sequences and the state sequence, a second plurality of actions of the robotic device at a second plurality of time points respectively using an action model, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robotic device to obtain the target item.
[0006] In a third aspect of the disclosure, there is provided an electronic device. The electronic device comprises: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the disclosure.
[0007] In a fourth aspect of the disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, causes the processor to implement the method according to the first aspect of the disclosure.
[0008] In a fifth aspect of the disclosure, there is provided a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.
[0009] It is to be understood that the details set forth herein are not intended to limit the key or critical features of the implementations of the disclosure, nor are they intended to limit the scope of the disclosure. Other features of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages and aspects of the implementations of the disclosure will become more apparent when described in conjunction with the following detailed description in conjunction with the accompanying drawings. In the drawings, like or similar reference numerals refer to like or similar elements, in which:
[0011] FIG. 1 shows a block diagram of an application environment according to one example implementation of the disclosure;
[0012] FIG. 2 shows a block diagram of obtaining an item by a robotic device according to some implementations of the disclosure;
[0013] FIG. 3 shows a block diagram of an image acquisition process according to some implementations of the present disclosure;
[0014] FIG. 4 shows a block diagram of an image sequence according to some implementations of the present disclosure;
[0015] FIG. 5 shows a block diagram of images from different acquisition units according to some implementations of the present disclosure;
[0016] FIG. 6 shows a block diagram of a structure of an action model according to some implementations of the present disclosure;
[0017] FIGS. 7A and 7B respectively show a block diagram of a process of acquiring an item according to some implementations of the present disclosure;
[0018] FIG. 8 shows a flowchart of a method for determining an action of a robotic device according to some implementations of the present disclosure;
[0019] FIG. 9 shows a block diagram of an apparatus for determining an action of a robotic device according to some implementations of the present disclosure; and
[0020] FIG. 10 shows a block diagram of a device capable of implementing a number of implementations of the present disclosure. DETAILED DESCRIPTION
[0021] Implementations of the present disclosure will be described in detail below with reference to the attached drawings. While certain implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the implementations set forth herein. Rather, these implementations are provided as exemplary embodiments for a more complete and thorough understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0022] In the description of the implementations of the present disclosure, the term "includes" and its conjugates are open-ended, meaning "including but not limited to". The term "based on" is intended to mean "based, at least in part, on" unless explicitly stated otherwise. The term "one implementation" or "an implementation" is understood to mean "at least one implementation". The term "some implementations" is understood to mean "at least some implementations". Other explicitly and implicitly recited definitions can be found in the description that follows. As used herein, the term "model" can represent a relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions that are currently known and / or will be developed in the future.
[0023] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0024] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0025] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0026] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above notification and obtaining of user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0028] The term "in response to" as used herein denotes a state in which a corresponding event occurs or a condition is met. It will be understood that the timing of the execution of the subsequent action performed in response to the event or condition is not necessarily strongly associated with the time at which the event occurs or the condition is met. For example, in some cases, the subsequent action can be performed immediately when the event occurs or the condition is met; in other cases, the subsequent action can be performed after a period of time after the event occurs or the condition is met.
[0029] Example Environment
[0030] Robot technology has been rapidly developed and has been widely used in multiple technical fields. However, the accuracy of robot devices performing various tasks is not satisfactory, and it is desirable to improve the accuracy of robot devices and thus perform specified tasks in a more accurate manner.
[0031] According to one example implementation of the present disclosure, a method for acquiring an item by a robotic device is proposed. Referring to FIG. 1, an application environment according to one example implementation of the present disclosure is described, which shows a block diagram 100 of an application environment according to one example implementation of the present disclosure. As shown in FIG. 1, a robotic device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robotic device 110 to perform a task of acquiring an item. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a warehouse environment, the physical space 160 can include one or more warehouses; in a home environment, the physical space 160 can include, but is not limited to, a living room, a bedroom, a study, a kitchen, a bathroom, etc., or a combination of one or more of the above.
[0032] As shown in FIG. 1, the robotic device 110 can include multiple parts. For example, a control unit 111 can serve as a control center of the robotic device 110, and an application program can be loaded into the control unit 111 in order to control various parts in the robotic device. The user 120 can use an interaction unit 112 to interact with the robotic device 110, for example, to input control instructions to the robotic device 110 in order to perform a desired task with the robotic device 110. The robotic device 110 can include an arm 113 for performing actions such as grabbing, releasing, etc. For example, the arm 113 can grab a certain item and move the item to a desired location, etc.
[0033] Alternatively and / or additionally, the robotic device 110 can also include a collection unit 114. Here, the collection unit 114 can include various types, such as an image collection unit, a sound collection unit, etc. Alternatively and / or additionally, the robotic device 110 can further include a sensing unit for detecting surrounding objects, for example, the distance between the robotic device and surrounding objects can be detected based on laser, etc. The robotic device 110 can also include a driving unit 115, for example, the robotic device 110 can be deployed on a movable base, and the driving unit 115 can drive the wheels of the base to move along a desired path.
[0034] The physical environment 160 can include one or more collection units 130, …, and 132, for example, one or more image collection units can be deployed in a room to collect images of the room from various angles. The physical environment 160 can include a control device 140, which can control one or more collection units 130, …, and 132, etc. via a network (not shown). Alternatively and / or additionally, in a warehouse, the control device 140 can control various devices in the physical space 160, for example, a conveying device for conveying items, a lifting device at a shelf, etc.
[0035] Alternatively and / or additionally, a machine learning model (e.g., model 150) can be provided to manage the robotic device 110. It should be appreciated that although FIG. 1 shows the model 150 located inside the physical space 160, alternatively and / or additionally, the model 150 can be located at a remote device outside the physical space 160, and the control device 140, the robotic device 110, or other devices can access the remote model 150 via a network.
[0036] The model 150 can include one or more models. If the model 150 includes multiple models, the multiple models can include multiple types of models. The model 150 can include at least a language model (LM) and an action model, for example. The language model can have a question-answering capability by learning from a large amount of corpus. The action model can control the robotic device 110 to perform various actions. The model 150 can also include an image recognition model, a text recognition model, and the like, for example.
[0037] As shown in FIG. 1, the user 120 can instruct the robotic device 110 to operate various items in the physical space 110. For example, the user 120 can instruct the robotic device 110 to find a certain item in the physical space 160; for another example, the user 120 can instruct the robotic device 110 to place the found item to a designated location, and the like.
[0038] Determining a summary of an action
[0039] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a method for determining an action of a robotic device is proposed. For ease of description, more details are described below only by way of example in a warehouse environment to control a robotic device to fetch an item.
[0040] A summary according to one example implementation of the present disclosure is described with reference to FIG. 2, which shows a block diagram 200 for determining an action of a robotic device according to some implementations of the present disclosure. As shown in FIG. 2, a user task can be received, which can instruct the robotic device 110 to fetch a target item of a target type. For example, the robotic device 110 can be instructed to fetch a “bottled water”.
[0041] In response to receiving the user task, the robotic device moves to a location (also referred to as a target location, e.g., a shelf in a warehouse) where the target item is located, the user task can instruct the robotic device to retrieve the target item from the target location. The robotic device can transition to a pose for retrieving the target item, e.g., the robotic device can rotate to face the shelf. The robotic device detects the target item based on attribute information of the target item (e.g., a name, an image, a feature, etc. of the target item) and environment information of the location (e.g., images acquired by the acquisition unit 114, 130, 132, etc.). In response to detecting the target item, the robotic device can place the target item into a target storage space (e.g., a tote, etc.).
[0042] Further, to retrieve the target item, a set of image sequences of an environment space where the robotic device is located can be received, the set of image sequences can include one or more image sequences, e.g., image sequence 210, …, image sequence 212, etc. Here, an image sequence in the set of image sequences includes a first plurality of images of the robotic device at a first plurality of time points (e.g., N time points), respectively. Here, the set of image sequences can be a historical set of image sequences before a current time point, each image sequence can relate to a plurality of time points.
[0043] It can be determined whether the target item is included in the set of image sequences. In response to determining that the target item is included in the set of image sequences (e.g., a bottle of water is included in the images of the shelf), a sequence of states of the robotic device is acquired. The sequence of states can include a first plurality of states of the robotic device at the first plurality of time points, respectively. Specifically, the sequence of states can include N states, the states of the robotic device can be stored in a variety of formats, e.g., the states can include a position, an orientation of the robotic device, positions and orientations of individual joints in the arm 113 of the robotic device, and a state of a tool at the arm (e.g., an open-close state of a gripper, etc.).
[0044] In turn, the action model 230 can be utilized to determine a second plurality of actions (e.g., action 240, …, and 242, etc.) of the robotic device at a second plurality of time points (e.g., future time points) based on the at least one image sequence 210, …, 212 and the sequence of states 220. Here, the second plurality of time points (e.g., future time points) are after the first plurality of time points, and the second plurality of actions are for controlling the robotic device to retrieve the target item. With some implementations of the present disclosure, the actions to be performed in the future can be determined with the historical states of the robotic device and the acquired historical images, thereby retrieving the target item in a more accurate and efficient manner.
[0045] According to some implementations of the present disclosure, the above-described method can be executed at any computing device with computing capability. For example, the above-described method can be executed with an application deployed at the robotic device 110. Alternatively and / or additionally, the application can be deployed at the control device 140 so as to execute the above-described method. Specifically, the powerful processing capability of the model 150 can be invoked so as to go to the target location. In turn, the robotic device 110 can grab the item and place it to the storage space.
[0046] Further, the robotic device 110 can return to the location of the user 120 and provide the user 120 with the retrieved item. Alternatively and / or additionally, the storage space can further include a storage space in a transportation tool such as a conveyor belt for transporting the item. At this time, the item can arrive at the designated location via the transportation tool, and so on.
[0047] Detailed procedure of determining action
[0048] Having described the outline of acquiring the item, in the following, more details according to some implementations of the present disclosure are described with reference to the accompanying drawings. According to some implementations of the present disclosure, the robotic device can go to the location where the target item is located and rotate to a direction facing the target item. The target location can be represented in a variety of functional ways. Specifically, in response to determining that the target location is represented with a location code, the target location coordinates corresponding to the location code are determined. For example, the location code is, for example, a code of a shelf in a physical space (i.e., a warehouse) (such as a number of the shelf, a bar code, a two-dimensional code, and so on). At this time, the robotic device can look up a warehouse map so as to locate the specific coordinates of the shelf (i.e., the target location coordinates). Further, a navigation path between the current location coordinates of the robotic device and the target location coordinates can be determined, and in turn, the robotic device is moved along the navigation path to the target location.
[0049] According to some implementations of the present disclosure, an image of the shelf can be taken so as to represent the target location. At this time, in response to determining that the target location is represented with a location image, the robotic device moves in the environmental space to capture an image of the environmental space. In response to determining that a location matching the location image is recognized in the image, the robotic device moves to the location.
[0050] The robot device can move to the location and transition to a pose for retrieving the target item, e.g., the robot device can rotate to face the shelf. The acquisition unit can continuously acquire images in front of the robot device to determine whether the pose of the robot device is suitable for grasping the item at the target location. The robot device can continuously adjust the pose based on the acquired images. Alternatively and / or additionally, in a warehouse environment, the shelf height can be high, in which case the robot device can reach a suitable pose for grasping the retrieval by means of a device such as an elevator, etc. For example, the robot device can enter the elevator and operate the elevator to reach a suitable height, etc. Alternatively and / or additionally, the robot device can also send a message to a central control system of the warehouse so that the central control system operates the elevator to reach a suitable height, etc.
[0051] According to some implementations of the present disclosure, images of the environment space where the robot device is located can be acquired by the acquisition unit. The images can be from at least any of the following: the acquisition unit at the robot device 110, and the acquisition unit in the physical space, etc. More details of the image acquisition are described with reference to FIG. 3, which shows a block diagram 300 of an image acquisition process according to some implementations of the present disclosure. As shown in FIG. 3, images (e.g., one or more images 310) can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images of various locations in the environment space, thereby facilitating the search for the target item.
[0052] Multiple acquisition units can be deployed at different locations of the robot device, e.g., one or more acquisition units can be deployed at the head of the robot device, one or more acquisition units can be deployed at the arm position, etc. In this way, the multiple acquisition units can work cooperatively and acquire images of the surrounding environment from different angles. For example, the acquisition unit at the arm position can acquire detailed images of the various joints of the arm, and the acquisition unit at the head can acquire an overall image of the arm, etc.
[0053] Alternatively and / or additionally, images of the environment space can be acquired from the acquisition units 130, …, and 132. Here, the acquisition units 130, …, and 132 can be pre-deployed at designated locations within the physical space, e.g., at the corners of the ceiling, etc. In this way, images of the environment space taken from a top-down perspective can be obtained, thereby facilitating an overall understanding of the layout of the environment space, thereby facilitating the localization of the target item.
[0054] FIG. 4 illustrates a block diagram 400 of image sequences according to some implementations of the present disclosure. As shown in FIG. 4, at each time point, multiple acquisition units can respectively acquire images. For example, at time point T1, acquisition unit 114 can acquire images 410, …, and acquisition unit 130 can acquire images 420; at time point TN, acquisition unit 114 can acquire images 420, …, and acquisition unit 130 can acquire images 422.
[0055] According to some implementations of the present disclosure, a set of image sequences of an environment space in which a robotic device is located can be acquired, an image sequence in the set of image sequences including a first plurality of images of the robotic device respectively at a first plurality of time points. For example, a first image sequence in the set of image sequences can be represented as and a second image sequence in the set of image sequences can be represented as Here, the subscript represents a time point at which an image is acquired, and the superscript represents an identity of an acquisition unit. At this time, the current time point is T i and the image sequence includes images at N-1 time points before the current time point.
[0056] FIG. 5 illustrates a block diagram 500 of images from different acquisition units according to some implementations of the present disclosure. As shown in FIG. 5, the arm of the robotic device in images 410 and 412 can respectively have poses 510 and 512. In this way, a plurality of images of the robotic device can be respectively acquired from multiple angles, thereby facilitating the robotic device to locate a target item in the images in a more accurate manner.
[0057] According to some implementations of the present disclosure, a variety of ways can be utilized to determine whether a target item is included in the set of image sequences. For example, the target item can be recognized from the images based on image recognition technology. Alternatively and / or additionally, a prompt can be constructed and input to a model in order to invoke the processing capability of the model to recognize the target item from the images. The prompt may, for example, be represented as: “Please identify the item represented by the second image from the first image”, and the acquired images (corresponding to the first image), the image of the bottled water (corresponding to the second image), and the prompt are submitted to the model.
[0058] The model can process the image, and in the case that the image includes the bottled water, the model can output a location where the bottled water is located (e.g., a region coordinate of the bottled water in the image, and / or directly output an image of the region where the bottled water is located, etc.). If the image does not include the bottled water, the robotic device can continue to move in the warehouse in order to find the bottled water. If the robotic device has traversed various locations of the warehouse and has not found the bottled water, the robotic device can feedback to the user that the bottled water cannot be found. With some implementations of the present disclosure, whether the image includes the bottled water can be detected based on a variety of manners, thereby improving the efficiency of the robotic device positioning.
[0059] According to some implementations of the present disclosure, in response to determining that the images in the sequence of images do not include the target item, the robotic device can be moved in order to find the target item within the warehouse. Alternatively and / or additionally, the robotic device can be instructed to return to the original position, and feedback to the user that the acquisition of the item is abnormal. In this way, a variety of complex situations in a variety of real environments can be flexibly handled.
[0060] According to some implementations of the present disclosure, in response to determining that the sequence of images includes the target item, a sequence of states of the robotic device is acquired, the sequence of states including a first plurality of states of the robotic device at a first plurality of time points, respectively. Here, the sequence of states can be represented as (sta i-N+1 ,sta i-N+2 ,…,sta i ), for example. The subscript of each state in the sequence indicates the time point at which the state is acquired, and each state can include at least any of the following: a pose of the robotic device, a pose of each joint in the arm, a state of a tool at the arm, etc. The pose can be described by a vector including 6 degrees of freedom, e.g., [pos1, pos2, pos3, rot1, rot2, rot3]. The first three dimensions in the vector represent the position, and the last three dimensions represent the orientation. The state of the tool can take different representation formats for different types of tools. Assuming that the tool is a gripper, the state can include an open state represented by 0, and a closed state represented by 1. Assuming that the tool is a drill bit, the state can include the rotational speed and model of the cutter, etc.
[0061] Further, a second plurality of actions of the robotic device at a second plurality of time points, respectively, can be determined based on the at least one sequence of images and the sequence of states, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robotic device to acquire the target item. Here, the second plurality of time points can include M time points after the current time point T i , e.g., T i+1 ,…,T i+M+1 .
[0062] More details are described with reference to FIG. 6, which shows a block diagram 600 of the structure of the action model according to some implementations of the present disclosure. As shown in FIG. 6, the input of the action model 230 can include at least one image sequence 610, a state sequence of the robot device 620. Alternatively and / or additionally, the input of the action model 230 can further include a rule 630 specifying for acquiring the target item. Specifically, the rule 630 can include at least any of the following: a quantity of the target item, an acquisition requirement of the target item, and a constraint condition for acquiring the target item.
[0063] Further, the output of the action model 230 can include actions 660 at a plurality of future time points. Here, the action can be represented by the difference Δ between the states of the robot device at two time points successively before and after. In the case that the state involves multiple dimensions (e.g., the pose of the robot device, the pose of each joint in the arm, the state of the reference tool of the arm), the action can accordingly include multiple dimensions. With some implementations of the present disclosure, the robot device can be controlled in a more accurate manner.
[0064] Alternatively and / or additionally, the output of the action model 230 can include image sequences 650, …, 654. Each image sequence can correspond to one acquisition unit, and each image sequence can include images associated with a plurality of time points, i.e., predicted images of the robot device at a plurality of future time points.
[0065] The action model can include an image encoder 612, a state encoder 622, an action decoder 662, and a backbone network 640. Alternatively and / or additionally, in the case that the input data involves the rule 630, the action model can further include a rule encoder 632. In the case that the output data involves the image sequences 650, …, 654, the action model 230 can further include image decoders 652. Here, each network branch in the action model 230 can share at least a part of the network layers in the backbone network 640.
[0066] According to some implementations of the present disclosure, reference data (also referred to as training data) can be acquired, and the action model can be trained. For example, during the operation of a reference robot device, a set of reference image sequences before a certain historical time point, a reference state sequence of the reference robot device, a set of reference state sequences after the historical time point can be acquired. Then, training samples are generated by using the acquired data, and the action model is trained accordingly.
[0067] Specifically, a first set of reference image sequences can be obtained in the reference space. This first set of reference image sequences includes first and a second set of reference images taken by the reference robot at first and a second set of reference time points, whereby the reference robot acquires the reference target item. The number of these first and a second set of reference time points can be represented as N, and the number of these second and a third set of reference time points can be represented as M. For time point T... i Specifically, the first reference image sequence in the first set of reference image sequences may include The second reference image sequence in the first set of reference image sequences may include Assuming only two image acquisition units are included, the two reference image sequences mentioned above can be concatenated: The image sequence at this point has a higher dimension.
[0068] Alternatively and / or additionally, the positions of individual images in an image sequence can be adjusted, for example, by ordering the images chronologically. For instance, multiple images acquired at earlier time points can be placed before multiple images acquired at later time points, i.e., at time point T... i-N Multiple images can be arranged at time point T i-N+1 Multiple images before: According to some implementations of this disclosure, an image sequence can be input to an image encoder 612 to determine image features. Alternatively and / or additionally, features related to the location of the acquisition unit can be inserted into the image features to describe the image from multiple perspectives.
[0069] According to some implementations of this disclosure, a first reference state sequence of the reference robot device can be obtained. The first reference state sequence includes a first plurality of reference states of the reference robot device at a first plurality of reference time points. For example, the reference state sequence can be represented as: (Rsta i-N+1 ,Rsta i-N+2 ,…,Rsta i A state sequence can be input into a state encoder to obtain state features.
[0070] Furthermore, a second reference state sequence of the reference robot device can be obtained. This second reference state sequence includes a second plurality of reference states of the reference robot device at a second plurality of reference time points, each following a first plurality of reference time points. Assuming the number of second plurality of reference time points is M, the second reference state sequence can be represented as: (Ract i+1 Ract i+2 Ract i+M+1 ).
[0071] Then, the action model can be updated using the set of reference image sequences, the first reference state sequence, and the second reference state sequence. Specifically, (Rsta (Rsta i-N+1 ,Rsta i-N+2 ,…,Rsta i ) can be used as the data part in the training sample, and (Ract i+1 ,Ract i+2 ,…,Ract i+M+1 ) can be used as the label part (ground truth) in the training sample.
[0072] Then, the data part can be input to the action model so that the predicted values of the plurality of actions can be determined by the action model. Further, the difference between the label part and the predicted values of the plurality of actions can be determined based on a predetermined loss function, and then the parameters of the action model can be updated in a direction that minimizes the difference. A large number of training samples can be constructed in a similar manner, and the action model can be updated in an iterative manner. With some implementations of the present disclosure, the trained action model can describe the relationship between the historical images, the historical states, and the future actions.
[0073] According to some implementations of the present disclosure, the action model can further process rules. At this time, the reference rules for obtaining the reference target items can be determined; and the action model can be updated using the reference rules. Here, the rules can include various contents, for example, including but not limited to the number of target items, the requirement for obtaining the target items, and the constraint conditions for obtaining the target items. For example, it can be specified to obtain a plurality of target items, for example, K. At this time, the training data further includes the "number" dimension. For another example, it can be specified to obtain the target items in the order from left to right (from near to far, etc.). At this time, the training data can further include the "obtaining requirement" dimension represented by text or other formats. For another example, it can be specified to keep the vertical direction of the items when obtaining the items. At this time, the training data can further include the "constraint condition" dimension represented by text or other formats.
[0074] The loss can be determined in a similar manner as described above, and then the action model 230 can be updated in a direction that minimizes the loss. With some implementations of the present disclosure, the action model with more functions can be generated in an end-to-end manner, so that the action model considers various rules when generating the robot actions.
[0075] According to some implementations of the present disclosure, the action model can further output a sequence of images of the robotic device at a plurality of subsequent time points. At this time, in the process of training the action model, a second set of reference image sequences of the reference space can be obtained, a reference image sequence in the second set of reference image sequences including a second plurality of reference images of the reference robotic device respectively at a second plurality of reference time points; and the action model is updated using the second plurality of reference images. For example, a first reference image sequence in the second set of reference image sequences can be represented as: A second reference image sequence in the second set of reference image sequences can be represented as: At this time, and may be used as the label part (ground truth) in the training sample.
[0076] At this time, the data part in the training data described above can be input to the action model so as to determine the predicted value of the sequence of images by the action model. Further, based on a predetermined loss function, the difference between the label part and the predicted value of the sequence of images can be determined, and then the parameters of the action model can be updated in the direction of minimizing the difference. A large number of training samples can be constructed in a similar manner, and the action model is updated in an iterative manner. Using some implementations of the present disclosure, the trained action model can describe the association between the historical images, the historical state and the future images.
[0077] It should be understood that although the above only schematically shows the case where there are two acquisition units, alternatively and / or additionally, there can be more or less acquisition units. In the above example, N and M are predetermined integers (i.e. hyperparameters), and generally N > M. In this way, the historical data can be fully utilized to predict the future action of the robotic device.
[0078] According to some implementations of the present disclosure, in the process of obtaining the first set of reference image sequences, for a reference image sequence in the set of reference image sequences, the reference robotic device is controlled using a reference action trajectory at the first plurality of reference time points; and a first plurality of reference images of the reference robotic device at the first plurality of reference time points are respectively obtained. Specifically, the reference action trajectory herein can be determined by, for example, a user manually controlling the robotic device. For example, the user can use a remote control device to control the robotic device to grasp an object. Alternatively and / or additionally, the reference action trajectory can be a trajectory pre-stored for controlling the robotic arm to grasp the object, which is a correct trajectory that has been verified to accurately grasp the object. In this way, accurate training data can be obtained based on various ways, thereby improving the accuracy of the action model.
[0079] According to some implementations of the present disclosure, the trained action model can be used to determine actions to control the robotic device. Specifically, the determined and (sta i-N+1 , sta i-N+2 ,..., sta i ) can be input to the action model to obtain a corresponding plurality of actions: (act i+1 , act i+2 ,..., act i+M+1 ). Further, the robotic device can be controlled using the plurality of actions to obtain the target object.
[0080] According to some implementations of the present disclosure, the captured image sequences and state sequences can be processed in a manner similar to that in the training process. For example, each image sequence can be input to an image encoder to generate a corresponding image feature, and a state sequence can be input to a state encoder to generate a corresponding state feature. For another example, two image sequences can be concatenated: The image sequences at this time have a higher dimension. Alternatively and / or additionally, the positions of the images in the image sequences can be adjusted, e.g., the images can be ordered in a time sequence. For example, the images captured at earlier time points can be arranged before the images captured at later time points, i.e., the images at time point T i-N may be arranged before the images at time point T i-N+1 :
[0081] According to some implementations of the present disclosure, the first set of image sequences includes a first image sequence and a second image sequence, the first image sequence is acquired by a first capturing unit deployed at a first position, and the second image sequence is acquired by a second capturing unit deployed at a second position. For example, the first image sequence can be acquired by a capturing unit deployed at an arm of the robotic device, and the second image sequence can be acquired by a capturing unit deployed at a head of the robotic device.
[0082] The importance of each image sequence can be determined. For example, the capturing unit at the arm is closer to the arm and the target object, and thus can provide more detailed information for action prediction. Thus, the first image sequence can be given a higher weight. For another example, the capturing device at the head is farther away from the arm and the target object, and thus can only provide effective details, and thus the second image sequence can be given a lower weight. At this time, the first image feature of the first image sequence and the second image feature of the second image sequence can be determined using the image encoder respectively. Further, the first image feature can be given a higher weight, and the second image feature can be given a lower weight.
[0083] According to some implementations of the present disclosure, directly applying the determined multiple actions can cause the robotic device to act stiffly and not smoothly. At this time, a second plurality of states of the robotic device at a second plurality of time points can be determined based on a second plurality of actions; and a smoothing process can be performed for the second plurality of states so as to update the second plurality of actions of the robotic device at the second plurality of time points. Specifically, a plurality of states after the robotic device performs actions (act i+1 , act i+2 ,…, act i+M+1 ) can be determined respectively. The smoothing process can be performed for the plurality of states so that the robotic device can acquire the target item in a more smooth manner.
[0084] According to some implementations of the present disclosure, the user task can further indicate a rule for acquiring the target item; and determining the second plurality of actions further comprises: determining the second plurality of actions based on the rule by using an action model. More details are described with reference to FIGS. 7A and 7B, which shows a block diagram 700A of a process of moving an item according to some implementations of the present disclosure. As shown in FIG. 7, in an image 710, the item 722 is not blocked by other items and can be directly grabbed. The historical image sequence and the historical state sequence can be input to the action model, and the action model can generate new actions to control the robotic device to grab the bottled water from the shelf. According to some implementations of the present disclosure, new instructions and states can be continuously input to the action model, and then subsequent actions can be determined.
[0085] According to some implementations of the present disclosure, a constraint condition 720 can be determined, i.e., a constraint condition that should be followed during the execution of the action. For example, it can be determined that during the movement of the bottled water, the original pose of the bottled water should be maintained (for example, the vertical direction is maintained and will not be tilted). The constraint condition 720 can be determined by using a language model. For example, a prompt word can be generated: “please determine the constraint condition that should be followed during the movement of the bottled water based on the following image”, or “please determine the precautions during the movement of the bottled water”, etc. At this time, the language model can generate the corresponding constraint condition. The historical image sequence, the historical state sequence and the constraint condition can be input to the action model, and a series of actions output by the action model will grab the bottled water while ensuring the constraint condition. According to some implementations of the present disclosure, the safety during the operation of the robotic device can be ensured, so as to avoid causing accidental damage to an item, etc.
[0086] According to some implementations of the present disclosure, if the bottled water is blocked by other items, the items can be removed first, and then the bottled water is retrieved. Specifically, in response to determining that the captured image indicates that the target item is blocked by another item, the other item can be moved to get the target item. More details are described with reference to FIG. 7B, which illustrates a block diagram 700B of a process of moving an item according to some implementations of the present disclosure. As shown in FIG. 7B, item 722 is the bottled water to be retrieved in image 730, and item 750 is located in front of item 722 and blocks item 722. At this time, the action generated by the action model can instruct the robotic device to move item 750 from position 760 to a position that does not hinder the retrieval of item 722 (e.g., position 760’ in image 740).
[0087] According to some implementations of the present disclosure, the action model can determine a placement position for placing item 750, and instruct the robotic device to move item 750 to the placement position. At this time, the action model will generate an action to control the robotic device to move item 750 from position 760 to position 760’. In this way, the robotic device can be supported to handle complex problems in a complex environment, and thus perform the user task in a more accurate manner.
[0088] According to some implementations of the present disclosure, during the process of moving item 750, the constraints during the movement of item 750 can be determined based on the pose of item 750, and the robotic device is instructed to move item 750 under the constraints. Similar to the process described above, a plurality of actions for moving item 750 can be generated under the constraints 720. In this way, it can be ensured that each action of the robotic device in a complex environment complies with safety specifications. Alternatively and / or additionally, after placing item 722 to the storage space, the robotic device can restore the position of item 750. Specifically, the arm of the robotic device can move item 750 from position 760’ back to position 760.
[0089] According to some implementations of the present disclosure, during the process of grabbing the item, as the robotic arm moves, images of the target item can be captured in real time, and other sensors at the robotic device can determine the distance between the arm and the target item in real time, and adjust the pose of the arm accordingly, so as to grab the item.
[0090] According to some implementations of the present disclosure, in response to detecting the plurality of target items at the target location, a rule for obtaining the target item can be received; and the target item is selected from the plurality of target items based on the rule. For example, the robotic device can feedback to the user that a plurality of bottled water is detected in the shelf, and ask the user how many bottled water is needed; alternatively and / or additionally, the user can be asked how much capacity of bottled water is needed, or which brand of bottled water is needed, etc. Then, the robotic device can grasp one or more bottled water according to the user’s answer.
[0091] According to some implementations of the present disclosure, the above-described method can be performed at a predetermined frequency. For example, the image sequence and the state sequence can be obtained in a sliding window manner. Assuming that the acquisition frequency of the acquisition unit is f1, and the control frequency of the robotic device is f2, the above-described method can be performed when a new image and / or the state of the robotic device is acquired. At this time, the image sequence and the state sequence can be updated in a sliding window manner. Alternatively and / or additionally, since the process of performing the above-described method requires time overhead, the above-described method can be performed at a frequency lower than f1 and f2. In this way, the workload of the control process can be reduced.
[0092] With some implementations of the present disclosure, the historical state of the robotic device and the historical images acquired can be utilized to determine the actions to be performed in the future, thereby obtaining the target item in a more accurate and efficient manner.
[0093] Example process
[0094] FIG. 8 illustrates a flowchart of a method 800 for determining actions of a robotic device according to some implementations of the present disclosure. At block 810, a user task is received, the user task indicating the robotic device to obtain a target item of a target type. At block 820, a set of image sequences of an environment space in which the robotic device is located is obtained, an image sequence in the set of image sequences including a first plurality of images of the robotic device at a first plurality of time points, respectively. At block 830, in response to determining that the set of image sequences includes the target item, a state sequence of the robotic device is obtained, the state sequence including a first plurality of states of the robotic device at the first plurality of time points, respectively. At block 840, a second plurality of actions of the robotic device at a second plurality of time points is determined based on the set of image sequences and the state sequence using an action model, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robotic device to obtain the target item.
[0095] According to some implementations of the present disclosure, the method further includes: determining a second plurality of states of the robotic device at a second plurality of time points respectively based on the second plurality of actions; and performing a smoothing process on the second plurality of states so as to update the second plurality of actions of the robotic device at the second plurality of time points.
[0096] According to some implementations of the present disclosure, the user task further indicates a rule for obtaining the target item; and determining the second plurality of actions further includes: determining the second plurality of actions based on the rule by using the action model.
[0097] According to some implementations of the present disclosure, the rule includes at least one of: a quantity of the target item, an obtaining requirement of the target item, and a constraint condition for obtaining the target item.
[0098] According to some implementations of the present disclosure, the action model is determined based on: obtaining a first set of reference image sequences of a reference environment space in which a reference robotic device is located, a reference image sequence in the first set of reference image sequences including a first plurality of reference images of the reference robotic device at a first plurality of reference time points respectively, the reference robotic device obtaining a reference target item; obtaining a first reference state sequence of the reference robotic device, the first reference state sequence including a first plurality of reference states of the reference robotic device at the first plurality of reference time points respectively; obtaining a second reference state sequence of the reference robotic device, the second reference state sequence including a second plurality of reference states of the reference robotic device at a second plurality of reference time points respectively, the second plurality of reference time points being after the first plurality of reference time points; and updating the action model by using the first set of reference image sequences, the first reference state sequence, and the second reference state sequence.
[0099] According to some implementations of the present disclosure, updating the action model further includes: obtaining a second set of reference image sequences of the reference environment space, a reference image sequence in the second set of reference image sequences including a second plurality of reference images of the reference robotic device at the second plurality of reference time points respectively; and updating the action model by using the second plurality of reference images.
[0100] According to some implementations of the present disclosure, updating the action model further includes: determining a reference rule for obtaining the reference target item; and updating the action model by using the reference rule.
[0101] According to some implementations of the present disclosure, obtaining the first set of reference image sequences includes: for a reference image sequence in the set of reference image sequences, controlling the reference robotic device by using a reference action trajectory at the first plurality of reference time points; and obtaining the first plurality of reference images of the reference robotic device at the first plurality of reference time points respectively.
[0102] According to some implementations of the present disclosure, the method further includes: in response to determining that the images in the set of image sequences do not include the target item, moving the robotic device.
[0103] According to some implementations of the present disclosure, the first set of image sequences includes a first image sequence and a second image sequence, the first image sequence is acquired by a first acquisition unit deployed at a first location, and the second image sequence is acquired by a second acquisition unit deployed at a second location.
[0104] Example apparatuses and devices
[0105] FIG. 9 illustrates a block diagram of an apparatus 900 for determining actions of a robotic device, according to some implementations of the present disclosure. The apparatus 900 includes: a receiving module 910 configured to receive a user task, the user task indicating the robotic device to acquire a target item of a target type; an image acquiring module 920 configured to acquire a set of image sequences of an environment space where the robotic device is located, the image sequences in the set of image sequences including a first plurality of images of the robotic device at a first plurality of time points, respectively; a state acquiring module 930 configured to, in response to determining that the set of image sequences includes the target item, acquire a state sequence of the robotic device, the state sequence including a first plurality of states of the robotic device at the first plurality of time points, respectively; and a determining module 940 configured to determine, based on the set of image sequences and the state sequence, a second plurality of actions of the robotic device at a second plurality of time points using an action model, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robotic device to acquire the target item.
[0106] According to some implementations of the present disclosure, the determining module is further configured to: determine, based on the second plurality of actions, a second plurality of states of the robotic device at the second plurality of time points, respectively; and perform smoothing processing on the second plurality of states so as to update the second plurality of actions of the robotic device at the second plurality of time points.
[0107] According to some implementations of the present disclosure, the user task further indicates a rule for acquiring the target item; and the determining module is further configured to determine, based on the rule, the second plurality of actions using the action model.
[0108] According to some implementations of the present disclosure, the rule includes at least any of: a quantity of the target item, an acquisition requirement of the target item, and a constraint condition for acquiring the target item.
[0109] According to some implementations of the present disclosure, the action model is determined based on: obtaining a first set of reference image sequences of a reference environment space in which a reference robot device is located, a reference image sequence in the first set of reference image sequences comprising a first plurality of reference images of the reference robot device at a first plurality of reference time points respectively, the reference robot device obtaining a reference target item; obtaining a first reference state sequence of the reference robot device, the first reference state sequence comprising a first plurality of reference states of the reference robot device at the first plurality of reference time points respectively; obtaining a second reference state sequence of the reference robot device, the second reference state sequence comprising a second plurality of reference states of the reference robot device at a second plurality of reference time points respectively, the second plurality of reference time points being after the first plurality of reference time points; and updating the action model using the first set of reference image sequences, the first reference state sequence, and the second reference state sequence.
[0110] According to some implementations of the present disclosure, updating the action model further comprises: obtaining a second set of reference image sequences of the reference environment space, a reference image sequence in the second set of reference image sequences comprising a second plurality of reference images of the reference robot device at the second plurality of reference time points respectively; and updating the action model using the second plurality of reference images.
[0111] According to some implementations of the present disclosure, updating the action model further comprises: determining a reference rule for obtaining the reference target item; and updating the action model using the reference rule.
[0112] According to some implementations of the present disclosure, the image obtaining module is further configured for: for a reference image sequence in a set of reference image sequences, controlling the reference robot device using the reference action trajectory at the first plurality of reference time points; and obtaining a first plurality of reference images of the reference robot device at the first plurality of reference time points respectively.
[0113] According to some implementations of the present disclosure, the apparatus further comprises: a moving module configured for moving the robot device in response to determining that an image in a set of image sequences does not comprise the target item.
[0114] According to some implementations of the present disclosure, the first set of image sequences comprises a first image sequence and a second image sequence, the first image sequence being obtained by a first collection unit deployed at a first location, and the second image sequence being obtained by a second collection unit deployed at a second location.
[0115] FIG. 10 illustrates a block diagram of a device 1000 that is capable of implementing the various implementations of the present disclosure. It should be appreciated that the computing device 1000 illustrated in FIG. 10 is merely an example and should not be construed as any limitation of the functionality and scope of the implementations described herein. The computing device 1000 illustrated in FIG. 10 can be used to implement the methods described above.
[0116] As illustrated in FIG. 10, the computing device 1000 is in the form of a general-purpose computing device. Components of the computing device 1000 can include, but are not limited to, one or more processors or processing units 1010, a memory 1020, a storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 1020. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of the computing device 1000.
[0117] The computing device 1000 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the computing device 1000 and includes both volatile and non-volatile media, removable and non-removable media. The memory 1020 can be volatile (such as register, cache, and random access memory (RAM)), non-volatile (such as read-only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 1030 can be a removable or non-removable media and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed by the computing device 1000.
[0118] The computing device 1000 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media such as a floppy disk, a ZIP® disk, a magnetic tape, or an optical disk. In these cases, each drive can be connected to the bus by one or more data media interfaces. The memory 1020 can include a computer program product 1025 having one or more program modules configured to carry out the various methods or actions of the various implementations of the present disclosure.
[0119] The communication unit 1040 enables communication with other computing devices over a communication medium. Additionally, the functionality of the components of the computing device 1000 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with each other through a communication connection. As such, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node in the networking environment.
[0120] The input device 1050 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 1060 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device 1000 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 1040, as needed, a device that enables a user to interact with the computing device 1000, or any device (e.g., a network card, a modem, etc.) that enables the computing device 1000 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0121] According to an example implementation of the present disclosure, a computer readable storage medium is provided, having stored thereon computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is provided, having stored thereon a computer program, which when executed by a processor implements the method described above.
[0122] Various aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0123] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0124] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0126] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for determining actions of a robotic device, comprising: receiving a user task, the user task indicating the robotic device to acquire a target item of a target type; acquiring a set of image sequences of an environment space where the robotic device is located, an image sequence in the set of image sequences comprising a first plurality of images of the robotic device at a first plurality of time points, respectively; in response to determining that the set of image sequences comprises the target item, acquiring a state sequence of the robotic device, the state sequence comprising a first plurality of states of the robotic device at the first plurality of time points, respectively; and determining, with an action model, a second plurality of actions of the robotic device at a second plurality of time points, respectively, based on the set of image sequences and the state sequence, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robotic device to acquire the target item. 2.The method of claim 1, further comprising: determining a second plurality of states of the robotic device at the second plurality of time points, respectively, based on the second plurality of actions; and performing smoothing processing for the second plurality of states so as to update the second plurality of actions of the robotic device at the second plurality of time points.
3. The method of claim 1, wherein the user task further indicates a rule for obtaining the target item; and determining the second plurality of actions further comprises: determining the second plurality of actions based on the rule with the action model. 4.The method of claim 3, wherein the rule comprises at least any one of: a quantity of the target item, an acquisition requirement of the target item, and a constraint condition for acquiring the target item. 5.The method of claim 1, wherein the action model is determined based on: acquiring a first set of reference image sequences of a reference environment space where a reference robotic device is located, a reference image sequence in the first set of reference image sequences comprising a first plurality of reference images of the reference robotic device at a first plurality of reference time points, respectively, the reference robotic device acquiring a reference target item; acquiring a first reference state sequence of the reference robotic device, the first reference state sequence comprising a first plurality of reference states of the reference robotic device at the first plurality of reference time points, respectively; acquiring a second reference state sequence of the reference robotic device, the second reference state sequence comprising a second plurality of reference states of the reference robotic device at a second plurality of reference time points, respectively, the second plurality of reference time points being after the first plurality of reference time points; and updating the action model with the first set of reference image sequences, the first reference state sequence, and the second reference state sequence. 6.The method of claim 5, wherein updating the action model further comprises: acquiring a second set of reference image sequences of the reference environment space, a reference image sequence in the second set of reference image sequences comprising a second plurality of reference images of the reference robotic device at the second plurality of reference time points, respectively; and updating the action model with the second plurality of reference images. 7.The method of claim 5, wherein updating the action model further comprises: determining a reference rule for acquiring the reference target object; and updating the action model using the reference rule.
8. The method of claim 5, wherein acquiring the first set of reference image sequences comprises: controlling, for a reference image sequence in the set of reference image sequences, the reference robot device using a reference action trajectory at the first plurality of reference time points; and acquiring a first plurality of reference images of the reference robot device at the first plurality of reference time points, respectively. moving the robot device in response to determining that an image in the set of image sequences does not include the target object.
9. The method of claim 1, further comprising:
10. The method of claim 1, wherein the first set of image sequences comprises a first image sequence and a second image sequence, the first image sequence being acquired by a first acquisition unit deployed at a first location, and the second image sequence being acquired by a second acquisition unit deployed at a second location.
11. An apparatus for determining actions of a robot device, comprising: a receiving module configured to receive a user task, the user task indicating a robot device to acquire a target object of a target type; an image acquiring module configured to acquire a set of image sequences of an environment space in which the robot device is located, an image sequence in the set of image sequences comprising a first plurality of images of the robot device at a first plurality of time points, respectively; a state acquiring module configured to acquire a state sequence of the robot device in response to determining that the set of image sequences includes the target object, the state sequence comprising a first plurality of states of the robot device at the first plurality of time points, respectively; and a determining module configured to determine, using an action model, a second plurality of actions of the robot device at a second plurality of time points, respectively, based on the set of image sequences and the state sequence, the second plurality of time points being after the first plurality of time points, and the second plurality of actions being for controlling the robot device to acquire the target object.
12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to perform the method of any of claims 1-10.
13. A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method of any of claims 1-10.
14. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method of any of claims 1-10.