A zero-sample robot control method, device, terminal and storage medium
By using the zero-sample robot control method in the robot system, using ordinary differential solution algorithms and action generation models to generate robot action sequences, the problem that the robot system cannot handle unknown tasks is solved, and the ability to efficiently complete tasks and adapt to new environments is achieved.
Patent Information
- Application Number
- CN202510237833.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing robot systems cannot handle unknown tasks, due to the dependence on data and the long training cycle.
A zero-sample robot control method is adopted. By obtaining the initial image and end image of the current manipulation task, the trained prediction model is input to generate an action mask sequence based on the ordinary differential solution algorithm, and input it to the action generation model to generate a robot action sequence to achieve the task completion.
Improves the ability of the robot to adapt to new environments, enables it to handle unknown tasks, and reduces dependence on data, and implements tasks with only the initial image and the end image without teaching.
Smart Images

Figure CN119748461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a zero-sample robot control method, device, terminal and storage medium. Background Art
[0002] Currently, a key goal in the field of robotic learning is to create general-purpose robotic systems that can perform a wide range of daily manipulation tasks in a variety of real-world scenarios that the robot may not have seen before. A straightforward approach to achieve this goal is to collect a huge dataset of robotic interactions for imitation learning. Although this approach is simple, it requires collecting data for different tasks, different objects, and different skills, and is also limited by physical access, so existing robotic systems cannot handle unknown tasks.
[0003] Therefore, the prior art has defects and needs to be improved and developed. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a zero-sample robot control method, device, terminal and storage medium in response to the above-mentioned defects of the prior art, aiming to solve the problem that the robot system in the prior art cannot handle unknown tasks.
[0005] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0006] A zero-sample robot control method, wherein the method comprises:
[0007] Acquire an initial image and an end image corresponding to the current manipulation task, wherein the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated;
[0008] Inputting the initial image and the end image into a trained prediction model, wherein an action mask sequence is generated based on an ordinary differential solution algorithm in the prediction model;
[0009] Inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence;
[0010] The current manipulation action is performed based on the robot action sequence. After the current manipulation action is performed, the current object state image of the robot is obtained, the initial image is updated based on the current object state image, and an updated action mask sequence is generated, and this cycle is repeated until the current manipulation task is completed.
[0011] In one embodiment of the present application, the initial image and the end image are input into a trained prediction model, and in the prediction model, an action mask sequence is generated based on an ordinary differential solution algorithm, including:
[0012] Inputting the initial image and the final image into a trained prediction model, wherein the initial image and the final image are respectively processed by a convolutional neural network in the prediction model to obtain a first latent feature and a second latent feature;
[0013] An action mask sequence is generated based on the first latent feature, the second latent feature and an ordinary differential solution algorithm, where the action mask sequence includes action mask graphs at several moments.
[0014] In one embodiment of the present application, generating an action mask sequence based on the first potential feature, the second potential feature and an ordinary differential solution algorithm includes:
[0015] Connecting the first latent feature with the second latent feature to obtain a hidden layer feature;
[0016] Integrate the hidden layer features based on an ordinary differential solution algorithm to obtain the hidden layer states at several moments;
[0017] The hidden layer states at several moments are input into the decoder in the prediction model to obtain an action mask sequence.
[0018] In one embodiment of the present application, a hand area and an object area are displayed on the action mask map.
[0019] In one embodiment of the present application, the action mask sequence, the initial image and the end image are input into a trained action generation model to generate a robot action sequence, including:
[0020] Inputting the action mask sequence, the initial image and the end image into a trained action generation model, and processing them through an encoder in the action generation model to obtain image features corresponding to each action mask image in the action mask sequence;
[0021] The image features at each moment are input into the decoder in the action generation model to obtain the robot action sequence.
[0022] In one embodiment of the present application, the robot motion sequence is a posture sequence of the robot end effector.
[0023] In one embodiment of the present application, the training steps of the prediction model and the action generation model include:
[0024] Acquire pre-collected action pairing data, wherein the action pairing data includes: one-to-one corresponding robot operation action images and human operation action images;
[0025] Extracting an initial training image and an ending training image corresponding to the manipulation task from the action pairing data;
[0026] Extracting action mask sequence training information from the human operation action image, using the action mask sequence training information as a true value label of a prediction model, and training the prediction model based on the initial training image and the end training image to obtain a trained prediction model;
[0027] The robot operation action training information is extracted from the robot operation action image, and the robot operation action training information is used as the true value label of the action generation model. The action generation model is trained based on the initial training image, the end training image and the action mask sequence training information to obtain a trained action generation model.
[0028] The present application also provides a zero-sample robot control device, wherein the device comprises:
[0029] An acquisition module, used to acquire an initial image and an end image corresponding to the current manipulation task, wherein the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated;
[0030] A prediction module, configured to input the initial image and the end image into a trained prediction model, wherein an action mask sequence is generated based on an ordinary differential solution algorithm in the prediction model;
[0031] An action generation module, used for inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence;
[0032] The execution module is used to execute the current manipulation action based on the robot action sequence. After executing the current manipulation action, the current object state image of the robot is obtained, the initial image is updated based on the current object state image, and an updated action mask sequence is generated, and this cycle is repeated until the current manipulation task is completed.
[0033] The present application also provides a terminal, which includes: a memory, a processor, and a zero-sample robot control program stored in the memory and executable on the processor, wherein the zero-sample robot control program implements the steps of the zero-sample robot control method as described above when executed by the processor.
[0034] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the zero-sample robot manipulation method as described above.
[0035] The present invention provides a zero-sample robot manipulation method, device, terminal and storage medium, the method comprising: obtaining an initial image and an end image corresponding to the current manipulation task, the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated; inputting the initial image and the end image into a trained prediction model, in which an action mask sequence is generated based on an ordinary differential solution algorithm; inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence; performing the current manipulation action based on the robot action sequence, after performing the current manipulation action, obtaining the current object state image of the robot, updating the initial image based on the current object state image, and generating an updated action mask sequence, and repeating this cycle until the current manipulation task is completed. The present application utilizes an ordinary differential solution algorithm to improve the robot's ability to adapt to new environments, so that the robot can handle unknown tasks, and the present application only requires an initial image and an end image to achieve the task without teaching. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flow chart of a preferred embodiment of the zero-sample robot control method of the present invention.
[0037] Figure 2 It is a structural principle block diagram of the prediction model in the present invention.
[0038] Figure 3 It is a structural principle block diagram of the action generation model in the present invention.
[0039] Figure 4 It is a functional principle block diagram of a preferred embodiment of the zero-sample robot control device in the present invention.
[0040] Figure 5 It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] In real-world environments, there are the following difficulties in training robots:
[0043] First, there are limitations to traditional learning methods, especially the reliance on data. Traditional machine learning methods usually require a large amount of labeled data, which is difficult to achieve in some fields (such as new product launches and rare object recognition).
[0044] Second, the collection cost of robot interaction datasets is high, and due to operational constraints, robot interaction data lacks diversity. In addition, the process of collecting and annotating robot data is time-consuming and expensive, which limits the robot's ability to be quickly deployed and updated.
[0045] Third, it lacks flexibility and cannot handle unknown task categories. Traditional machine learning methods perform poorly when faced with unseen categories or tasks, and cannot make effective inferences and decisions.
[0046] Fourth, the training cycle is long and the model is difficult to update. When new categories or tasks need to be introduced, traditional methods often require retraining the entire model, which consumes a lot of time and computing resources.
[0047] Fifth, knowledge transfer is limited, and it is difficult to use existing knowledge. Traditional methods often cannot effectively use existing knowledge for transfer learning, which limits their application capabilities in multiple tasks and multiple scenarios. Therefore, reducing the robot's dependence on data and speeding up the robot's training cycle are essential in robot control.
[0048] The following describes the zero-sample robot manipulation method, device, terminal and storage medium of the embodiment of the present application with reference to the accompanying drawings. In view of the problem that the robot system mentioned in the above background technology cannot handle unknown tasks, the present application provides a zero-sample robot manipulation method, in which the initial image and the end image corresponding to the current manipulation task are obtained, the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated; the initial image and the end image are input into the trained prediction model, in which an action mask sequence is generated based on the ordinary differential solution algorithm; the action mask sequence, the initial image and the end image are input into the trained action generation model to generate a robot action sequence; the current manipulation action is performed based on the robot action sequence, and after performing the current manipulation action, the current object state image of the robot is obtained, the initial image is updated based on the current object state image, and an updated action mask sequence is generated, and this cycle is repeated until the current manipulation task is completed. The present application uses the ordinary differential solution algorithm to improve the robot's ability to adapt to new environments, so that the robot can handle unknown tasks, and the present application only needs the initial image and the end image to achieve the task without teaching.
[0049] See also Figure 1 , Figure 1 is a flow chart of the zero-sample robot control method of the present invention. Figure 1 As shown, the zero-sample robot control method described in the embodiment of the present invention includes:
[0050] Step S100: Acquire an initial image and an end image corresponding to the current manipulation task, wherein the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated.
[0051] Specifically, the manipulation task of this application is divided into two parts: plan prediction and action generation. In robot learning, training a robot to perform a task requires imitation learning (teaching learning). Teaching learning means that a person manipulates the robot to complete a specific task and makes the robot reproduce its task. Traditional teaching can only allow the robot to repeat the teaching path and is not generalizable. The problem with learning-based teaching, that is, imitation learning, is that a large amount of video and robot motion data is collected for the same task in order for the model to learn the strategy to complete a task. This application only needs to obtain the initial image and the end image to perform any manipulation task without teaching.
[0052] In addition, before inputting the initial image and the final image into the prediction model, it is necessary to use image modification technology to remove the robotic arm on the initial image and the final image to ensure that the input data in the inference is consistent with the training data.
[0053] like Figure 1 As shown, the zero-sample robot control method described in this embodiment also includes:
[0054] Step S200: input the initial image and the end image into a trained prediction model, and generate an action mask sequence in the prediction model based on an ordinary differential solution algorithm.
[0055] The embodiment of the present application utilizes a prediction model based on neural differential equations to generate a human-object interaction plan, i.e., an action mask sequence, given an initial image and an end image, thereby improving the diversity of the robot dataset. The prediction model may be a visual transformer (ViT), a shifted window transformer (Swin Transformer), or other models.
[0056] In the embodiment of the present application, the step S200 specifically includes:
[0057] Step S210, inputting the initial image and the end image into a trained prediction model, wherein the initial image and the end image are processed by a convolutional neural network in the prediction model respectively to obtain a first latent feature and a second latent feature;
[0058] Step S220: Generate an action mask sequence based on the first latent feature, the second latent feature and the ordinary differential solution algorithm, wherein the action mask sequence includes action mask graphs at several moments.
[0059] Specifically, a prediction model (ODE-Plan) based on the Neural Ordinary Differential Equations (NODE) algorithm is used to generate an action mask sequence given an initial image and an end image. The ODE algorithm can model the dynamic changes of features, which can help the model understand the changes in the physical world, thereby generating a dynamic character interaction video plan. The action mask sequence is a form of expression of the character interaction video plan, which can be an RGB video. The solution to the prediction model can be Neural Ordinary Differential Equations, Euler method, etc.
[0060] The present application uses an ordinary differential solution algorithm to solve the first potential feature and the second potential feature to improve the robot's ability to adapt to new environments, so that the robot can handle unknown tasks.
[0061] In one embodiment of the present application, step S220 specifically includes:
[0062] Step S221, connecting the first latent feature and the second latent feature to obtain a hidden layer feature;
[0063] Step S222: Integrate the hidden layer features based on an ordinary differential solution algorithm to obtain hidden layer states at several moments;
[0064] Step S223: input the hidden layer states at several moments into the decoder in the prediction model to obtain an action mask sequence.
[0065] Specifically, the predictive model provides a new framework for designing and training flexible continuous-time neural network architectures, given an initial action The initial image and ending action of The ending image, represents the real number space, and are the height and width of the image, respectively, and a reasonable future mask is generated by the ordinary differential solution algorithm , to capture the expected trajectory of the hand and the object, as follows:
[0066] ;
[0067] ;
[0068] in, represents the initial state of the hidden layer, Conv is a convolutional neural network, is the concatenation operator, is the initial image corresponding to the initial state, is the ending image corresponding to the ending state, To model the hidden layer dynamics, it can be a linear layer, t represents the time, refers to the hidden state of the simulated time step interaction dynamics, where D represents the hidden feature dimension. The hidden state of the system is defined at an arbitrary time t and can be evaluated at any desired time using a numerical ODE solver, which will evaluate the dynamic To determine the solution, by integrating the above differential equation to different time steps, the output of the required time step can be obtained. Figure 2 As shown, after integration, the hidden layer state from the first moment to the kth moment can be obtained arrive ,Will arrive Input into the decoder (consisting of several convolutions), we can get the action mask sequence , including action mask map The formula is:
[0069] ;
[0070] in, represents a numerical ODE solver, T represents the time corresponding to the robot action sequence to be predicted, for example, if the desired output is the action in the future H seconds, then T = H. Use the image decoder to map the hidden state to the action mask ,Right now In the embodiment of the present application, a Runge-kunta 4th order solver is used to solve the equation.
[0071] In the embodiment of the present application, the hand region and the object region are displayed on the action mask map, that is, the hand region and the object region on the action mask map are segmented.
[0072] Specifically, the present application uses a predictive model to generate future feasible hand region masks and object region masks for interaction in the robot's physical scene. Figure 1 As shown, the zero-sample robot control method described in this embodiment also includes:
[0073] Step S300: input the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence.
[0074] The action mask does not directly tell the robot which actions should be performed to perform the desired interaction. In order to achieve robot operation in the context of predictive planning, this application uses an action generation model to convert the action mask into robot actions. The action generation model is built based on the Transformer model.
[0075] The action generation model of the embodiment of the present application is a sequence generation model, whose input is the initial image, the end image and the action sequence mask, and whose output is the robot action sequence. The action generation model converts the action sequence mask into robot actions in the physical world, so that the robot can imitate humans to complete the actions of picking up, placing, scooping, pouring, twisting, stacking and sliding. It also means that the robot can simulate human actions by watching a large number of human-computer interaction videos from the Internet, so as to achieve the purpose of controlling the robot. The action generation model can be a recurrent neural network (RNN), RWKV, Mamba, etc., or it can be an action generation model based on a diffusion model, such as a diffusion policy.
[0076] In the embodiment of the present application, the robot motion sequence is a posture sequence of the robot end effector.
[0077] Specifically, the output of the sequence generation model can be a posture sequence of the robot's end effector, a joint angle posture, velocity or torque sequence of the robot, or a velocity sequence of the robot's end effector.
[0078] In the embodiment of the present application, the step S300 specifically includes:
[0079] Inputting the action mask sequence, the initial image and the end image into a trained action generation model, and processing them through an encoder in the action generation model to obtain image features corresponding to each action mask image in the action mask sequence;
[0080] The image features at each moment are input into the decoder in the action generation model to obtain the robot action sequence.
[0081] This application uses an action generation model to convert action mask sequences into actual robot actions. In this way, robotic manipulation can benefit from the large amount of online human and object videos.
[0082] like Figure 1 As shown, the zero-sample robot control method described in this embodiment also includes:
[0083] Step S400, executing the current manipulation action based on the robot action sequence, after executing the current manipulation action, obtaining the current object state image of the robot, updating the initial image based on the current object state image, and generating an updated action mask sequence, and repeating this cycle until the current manipulation task is completed.
[0084] The action generation model predicts future actions based on the predicted action mask sequence and the actual actions of the robot currently observed, such as Figure 3 As shown. The initial image and the final image are used as the input of the action generation model, processed by the convolutional neural network and the fully connected network, and the image features are obtained using the Transformer encoder. , using the Transformer decoder to obtain the robot action sequence , the robot action sequence represents the H steps that the robot needs to perform in the future starting at time t, which is instantiated as a closed-loop strategy , where a represents an action; It refers to the initial image corresponding to time t, executes the first predicted action and obtains the current observation result of the next cycle. The first predicted action refers to the first action in the robot action sequence, and the current observation result refers to the current object state image of the robot, which can be captured by a camera. Figure 3 middle j represents any time step, represents the robot action sequence at time t+j, Indicates that the generation of the current action depends on the implementation of the previously generated action, that is, The generation of , , ,..., achieved.
[0085] Specifically, the initial image and the end image are input into the trained prediction model to generate an action mask sequence corresponding to the first moment; then the initial image, the end image and the action mask sequence at the first moment are input into the trained action generation model to obtain the robot action sequence at the first moment, and the robot executes the first action in the robot action sequence at the first moment. At this point, the action at the first moment is completed.
[0086] Since the actual action performed by the robot based on the first action in the robot action sequence may deviate, the present application will capture the current object state image of the robot and then use the current object state image as the initial image at the second moment in order to improve the accuracy and consistency of the action execution.
[0087] The end image and the initial image at the second moment are input into the trained prediction model to generate the action mask sequence corresponding to the second moment; the initial image at the second moment, the end image and the action mask sequence at the second moment are then input into the trained action generation model to obtain the robot action sequence at the second moment, and the robot performs the first action in the robot action sequence at the second moment, and the action at the second moment is completed. This cycle is repeated until the current manipulation task is completed.
[0088] In an embodiment of the present application, the training steps of the prediction model and the action generation model include:
[0089] Acquire pre-collected action pairing data, wherein the action pairing data includes: one-to-one corresponding robot operation action images and human operation action images;
[0090] Extracting an initial training image and an ending training image corresponding to the manipulation task from the action pairing data;
[0091] Extracting action mask sequence training information from the human operation action image, using the action mask sequence training information as a true value label of a prediction model, and training the prediction model based on the initial training image and the end training image to obtain a trained prediction model;
[0092] The robot operation action training information is extracted from the robot operation action image, and the robot operation action training information is used as the true value label of the action generation model. The action generation model is trained based on the initial training image, the end training image and the action mask sequence training information to obtain a trained action generation model.
[0093] It can be understood that the initial images of the paired robot operation action images and human operation action images are the same, and the ending images are also the same. Therefore, the initial images of the two are used as the initial training images, and the ending images of the two are used as the ending training images.
[0094] The present application is a framework for using robot zero-shot learning (ZSL) based on an ordinary differential solution algorithm to achieve robot manipulation, allowing the robot to use existing knowledge to reason in the real world without seeing a specific object or task, and to train the robot using an ordinary differential solution algorithm, thereby improving the robot's ability to adapt to new environments and being able to perform tasks in unknown environments, thereby solving the problems of the robot's dependence on data and long training cycles.
[0095] It should be noted that when training the model, the camera used to collect robot operation action images and human operation action images is in the same position as the robot base. When collecting specifically, a human operator can be used to remotely operate a robot in the scene, reset or in a parallel identical setting, the human manipulates a similar object with roughly similar movements as the robot arm. For example, collect trajectory pairs involving robots manipulating objects and humans manipulating similar objects in similar ways. Specifically, a paired demonstration collection method can be used, in which a human operator manipulates the robot in the scene by remote control. After resetting, or in the same setting, another human manipulates a similar object with roughly the same movements as the robot arm, so that collecting paired data is fast and convenient. In addition, training data can be generated with the help of a large visual language model to achieve data enhancement.
[0096] This application uses prediction models and action generation models to convert human-object interactions into robot operations to achieve zero-sample operations and reduce the high dependence on operation data. That is, the robot can learn human operation patterns through visual learning (allowing the robot to learn existing human actions, including human action videos from videos and various Internet platforms), thereby achieving machine learning without relying on large robot data sets. Specifically, when using traditional large robot data sets, if the robot needs to be able to perform any task, it is necessary to have a data set that includes any task, any camera external parameters, and any robot model, which is difficult to do. This application can perform any task given only the initial image and the end image.
[0097] In one embodiment, if Figure 4 As shown, based on the above zero-sample robot control method, the present invention also provides a zero-sample robot control device, the device comprising:
[0098] An acquisition module 100 is used to acquire an initial image and an end image corresponding to a current manipulation task, wherein the initial image is used to represent an initial state of the object to be manipulated, and the end image is used to represent an end state of the object to be manipulated;
[0099] A prediction module 200, configured to input the initial image and the end image into a trained prediction model, in which an action mask sequence is generated based on an ordinary differential solution algorithm;
[0100] An action generation module 300, for inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence;
[0101] The execution module 400 is used to execute the current manipulation action based on the robot action sequence. After executing the current manipulation action, the current object state image of the robot is obtained, the initial image is updated based on the current object state image, and an updated action mask sequence is generated, and this cycle is repeated until the current manipulation task is completed.
[0102] Figure 5 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:
[0103] A memory 501 , a processor 502 , and a computer program stored in the memory 501 and executable on the processor 502 .
[0104] When the processor 502 executes the program, the zero-sample robot control method provided in the above embodiment is implemented.
[0105] Furthermore, the terminal further includes:
[0106] The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0107] The memory 501 is used to store computer programs that can be executed on the processor 502 .
[0108] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0109] If the memory 501, the processor 502 and the communication interface 503 are implemented independently, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0110] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0111] The processor 502 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0112] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the zero-sample robot control method as described above is implemented.
[0113] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0114] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0115] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, in which the order shown or discussed may not be followed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0116] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can read instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary and then storing it in a computer memory.
[0117] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0118] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0119] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0120] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
[0121] In summary, the present invention discloses a zero-sample robot manipulation method, device, terminal and storage medium, the method comprising: obtaining an initial image and an end image corresponding to the current manipulation task, the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated; inputting the initial image and the end image into a trained prediction model, in which an action mask sequence is generated based on an ordinary differential solution algorithm; inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence; performing the current manipulation action based on the robot action sequence, after performing the current manipulation action, obtaining the current object state image of the robot, updating the initial image based on the current object state image, and generating an updated action mask sequence, and repeating this cycle until the current manipulation task is completed. The present application utilizes an ordinary differential solution algorithm to improve the robot's ability to adapt to new environments, so that the robot can handle unknown tasks, and the present application only requires an initial image and an end image to achieve the task without teaching.
[0122] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A zero-sample robot control method, characterized in that: The method comprises: Acquire an initial image and an end image corresponding to the current manipulation task, wherein the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated; Inputting the initial image and the end image into a trained prediction model, wherein an action mask sequence is generated based on an ordinary differential solution algorithm in the prediction model; Inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence; The current manipulation action is performed based on the robot action sequence. After the current manipulation action is performed, the current object state image of the robot is obtained, the initial image is updated based on the current object state image, and an updated action mask sequence is generated, and this cycle is repeated until the current manipulation task is completed.
2. The zero-sample robot control method according to claim 1, characterized in that: Inputting the initial image and the end image into a trained prediction model, wherein an action mask sequence is generated based on an ordinary differential solution algorithm in the prediction model, including: Inputting the initial image and the final image into a trained prediction model, wherein the initial image and the final image are respectively processed by a convolutional neural network in the prediction model to obtain a first latent feature and a second latent feature; An action mask sequence is generated based on the first latent feature, the second latent feature and an ordinary differential solution algorithm, where the action mask sequence includes action mask graphs at several moments.
3. The zero-sample robot control method according to claim 2, characterized in that: Generating an action mask sequence based on the first latent feature, the second latent feature, and an ordinary differential solution algorithm includes: Connecting the first latent feature with the second latent feature to obtain a hidden layer feature; Integrate the hidden layer features based on an ordinary differential solution algorithm to obtain the hidden layer states at several moments; The hidden layer states at several moments are input into the decoder in the prediction model to obtain an action mask sequence.
4. The zero-sample robot control method according to claim 2, characterized in that: The action mask image displays a hand area and an object area.
5. The zero-sample robot control method according to claim 2, characterized in that: Inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence, including: Inputting the action mask sequence, the initial image and the end image into a trained action generation model, and processing them through an encoder in the action generation model to obtain image features corresponding to each action mask image in the action mask sequence; The image features at each moment are input into the decoder in the action generation model to obtain the robot action sequence.
6. The zero-sample robot control method according to claim 1, characterized in that: The robot motion sequence is a position and posture sequence of the robot end effector.
7. The zero-sample robot control method according to claim 1, characterized in that: The training steps of the prediction model and the action generation model include: Acquire pre-collected action pairing data, wherein the action pairing data includes: one-to-one corresponding robot operation action images and human operation action images; Extracting an initial training image and an ending training image corresponding to the manipulation task from the action pairing data; Extracting action mask sequence training information from the human operation action image, using the action mask sequence training information as a true value label of a prediction model, and training the prediction model based on the initial training image and the end training image to obtain a trained prediction model; The robot operation action training information is extracted from the robot operation action image, and the robot operation action training information is used as the true value label of the action generation model. The action generation model is trained based on the initial training image, the end training image and the action mask sequence training information to obtain a trained action generation model.
8. A zero-sample robot control device, characterized in that: The device comprises: An acquisition module, used to acquire an initial image and an end image corresponding to the current manipulation task, wherein the initial image is used to represent the initial state of the object to be manipulated, and the end image is used to represent the end state of the object to be manipulated; A prediction module, configured to input the initial image and the end image into a trained prediction model, wherein an action mask sequence is generated based on an ordinary differential solution algorithm in the prediction model; An action generation module, used for inputting the action mask sequence, the initial image and the end image into a trained action generation model to generate a robot action sequence; The execution module is used to execute the current manipulation action based on the robot action sequence. After executing the current manipulation action, the current object state image of the robot is obtained, the initial image is updated based on the current object state image, and an updated action mask sequence is generated, and this cycle is repeated until the current manipulation task is completed.
9. A terminal, characterized in that: include: A memory, a processor, and a zero-sample robot control program stored in the memory and executable on the processor, wherein the zero-sample robot control program, when executed by the processor, implements the steps of the zero-sample robot control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the zero-sample robot manipulation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Instrument image recognition method based on inspection robot
CN114359552A
Depth neural network interpretable method, visualization method and related device
CN114419726A