Method and apparatus for determining a control policy model and method and apparatus for controlling an end effector
By adjusting the control strategy model of the end effector through reinforcement learning, the problems of adaptability and control accuracy of the end effector in complex tool operation tasks are solved, and the generalization ability and anti-interference ability are improved while reducing the training data.
Patent Information
- Application Number
- CN202411457051.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-10-17
AI Technical Summary
Existing technologies struggle to effectively address the adaptability and control accuracy issues of end effectors in complex tool operation tasks, especially in improving generalization and anti-interference capabilities with reduced training data.
By adjusting the control strategy model of the end effector through reinforcement learning and using specific reward scores to optimize the control strategy model, the model can be made to conform to physical constraints in the simulation environment, thereby adapting to the changing tool operation tasks.
It significantly reduces the need for human demonstration data and the training time for reinforcement learning, improves the adaptability and flexibility of the end effector in diverse tasks, and enhances control accuracy and generalization ability.
Smart Images

Figure CN119347753B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent robots, and in particular to a method for determining a control policy model, a method for controlling an end effector, an apparatus for controlling an end effector, an apparatus for determining a control policy model, an apparatus for controlling an end effector, an electronic device, a non-volatile computer-readable storage medium, and a computer program product. BACKGROUND
[0002] Tool operation using an end effector is a hot issue in the field of robotics, because it involves complex interactions among multiple objects, including the end effector, the tool, the target object, and the operating environment. For example, a robot needs to use a specific tool to grasp and move an object, such as taking an object out of or putting an object into a container. The complexity of this process increases significantly with different tool types, object characteristics, and environmental conditions. The industry is trying to develop end effectors that can adapt to these changes and effectively operate tools, but no effective solution has been proposed so far.
[0003] To improve the efficiency and adaptability of robot operation, end effectors equipped with artificial intelligence models have been proposed to complete the special and difficult task of tool operation using an end effector. However, such a solution often requires a large amount of training data. The industry is exploring how to reduce training data while improving the generalization ability of neural network models for end effectors. In addition, how to improve the control accuracy of end effectors, increase the number of operable objects, speed up the operation, and enhance the anti-interference ability are all problems that need to be solved urgently. By solving these problems, the end effector of the robot can better adapt to changing task scenarios, thus playing a greater role in practical applications. SUMMARY
[0004] To address the above problems, the present disclosure provides a method for determining a control policy model, a method for controlling an end effector, an apparatus for controlling an end effector, an apparatus for determining a control policy model, an apparatus for controlling an end effector, an electronic device, a non-volatile computer-readable storage medium, and a computer program product.
[0005] According to an aspect of the present disclosure, a method for determining a control policy model for controlling an end effector including a central link and a plurality of multi-joint operation assemblies is provided. The method includes obtaining a first control policy model loaded on the end effector, and adjusting the first control policy model loaded on the end effector using a plurality of action rounds to determine a second control policy model loaded on the end effector. In each action round of the plurality of action rounds, a learning task corresponding to the action round is determined based on information of the action round. A virtual body corresponding to the end effector simulates performing the learning task corresponding to the action round using the first control policy model in the adjusting, and determines a reward score corresponding to completing the learning task corresponding to the action round and a reward score related to a physical constraint in a process of simulating performing the learning task corresponding to the action round. The first control policy model is adjusted to increase the reward score corresponding to completing the learning task corresponding to the action round and the reward score related to the physical constraint in the process of simulating performing the learning task corresponding to the action round.
[0006] According to another aspect of the present disclosure, a method for controlling an end effector including a central link and a plurality of multi-joint operation assemblies is provided. The method includes obtaining, at a current time, a reference scene image when the end effector performs a target task and a pose corresponding to a contact point between the end effector and a second target tool, generating control information for controlling the end effector at a current time step using a control policy model loaded on the end effector, determining motor driving information of the central link and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step, and controlling the end effector based on the motor driving information of the central link and the plurality of multi-joint operation assemblies. The control policy model is determined by the above method.
[0007] According to another aspect of the present disclosure, an apparatus for controlling an end effector including a central connector and a plurality of multi-joint operation assemblies is provided, the apparatus including: a camera configured to acquire a reference scene image of the end effector performing a target task at a current time; a sensor configured to acquire a pose of a contact point of the end effector and a target tool at the current time; a processor configured to generate control information for controlling the end effector at a current time step using a control policy model carried on the end effector, and determine motor driving information of the central connector and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step; and a motor configured to control the end effector based on the motor driving information of the central connector and the plurality of multi-joint operation assemblies, wherein the control policy model is determined by the method described above.
[0008] According to another aspect of the present disclosure, an apparatus for determining a control policy model for controlling an end effector including a central connector and a plurality of multi-joint operation assemblies is provided, the apparatus including: a first module configured to acquire a first control policy model carried on the end effector; and a second module configured to adjust the first control policy model carried on the end effector using a plurality of action rounds to determine a second control policy model carried on the end effector, wherein in each action round of the plurality of action rounds, a learning task corresponding to the action round is determined based on information of the action round, a virtual body corresponding to the end effector simulates performing the learning task corresponding to the action round using the first control policy model in the adjusting, and a reward score corresponding to completion of the learning task corresponding to the action round and a reward score related to a physical constraint in a process of simulating performing the learning task corresponding to the action round are determined, and the first control policy model is adjusted to increase the reward score corresponding to completion of the learning task corresponding to the action round.
[0009] According to another aspect of the present disclosure, an apparatus for controlling an end effector is provided, the end effector comprising a central connector and a plurality of multi-joint operation assemblies, comprising: a first module configured to, at a current time, acquire a reference scene image when the end effector performs a target task and a pose corresponding to a contact point between the end effector and a target tool; a second module configured to generate control information for controlling the end effector at a current time step by using a control policy model carried on the end effector; a third module configured to determine motor driving information of the central connector and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step; and a fourth module configured to control the end effector based on the motor driving information of the central connector and the plurality of multi-joint operation assemblies, wherein the control policy model is determined by the method described above.
[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory, wherein the memory has computer executable code stored therein, the computer executable code, when executed by the processor, performs the method described above.
[0011] According to another aspect of the present disclosure, a non-volatile computer readable storage medium is provided, having executable code stored thereon, the executable code, when executed by a processor, causes the processor to perform the method described above.
[0012] According to another aspect of the present disclosure, a computer program product is provided, comprising computer executable instructions, the computer executable instructions, when executed by a processor, implement the method described above.
[0013] The present disclosure provides an improved method for determining a control policy model. Specifically, in order to obtain a second control policy model with higher generalization and robustness than a first control policy model, the present disclosure further adjusts the first control policy model through the reinforcement learning scheme to obtain the second control policy model, and uses specific reward scores in the adjustment process, i.e., both the "reward score corresponding to the learning task corresponding to the action round" and the "reward score related to the physical constraint in the process of simulating the execution of the learning task corresponding to the action round", which can ensure that the control policy model explores the control scheme as widely as possible while avoiding the trained control policy model in the simulation environment not meeting the physical constraints, thereby leading to the inability to be applied to the physical environment.
[0014] In addition, the method for determining the control policy model proposed in the present disclosure is particularly suitable for special tool operation tasks. Compared with the conventional method for determining a training scheme for determining a control policy model suitable for special tool operation tasks, the present disclosure adjusts the first control policy model using a reinforcement learning scheme on the basis of the pre-trained first control policy model. On the one hand, the demand for human demonstration data is reduced, and on the other hand, the training time required by reinforcement learning is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some exemplary embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings. The following drawings are not necessarily drawn in proportion to the actual size, and the emphasis is on showing the main idea of the present disclosure.
[0016] Figure 1 An example task scenario to which embodiments according to the present disclosure can be applied is schematically illustrated.
[0017] Figure 2 is a schematic view showing an end effector according to embodiments of the present disclosure.
[0018] Figure 3 is a flowchart showing a method for determining a control policy model according to embodiments of the present disclosure.
[0019] Figure 4 is a schematic view showing a method for determining a control policy model according to embodiments of the present disclosure.
[0020] Figure 5 is a schematic view showing a method for controlling an end effector according to embodiments of the present disclosure.
[0021] Figure 6 is a flowchart showing a method for controlling an end effector according to embodiments of the present disclosure.
[0022] Figure 7 is a schematic view of a device for controlling an end effector according to some embodiments of the present disclosure.
[0023] Figure 8 is an exemplary block diagram of an apparatus for determining a control policy model according to some embodiments of the present disclosure.
[0024] Figure 9 is an exemplary block diagram of an apparatus for controlling an end effector according to some embodiments of the present disclosure.
[0025] Figure 10 An example block diagram of a computing device is illustratively shown in accordance with some embodiments of the present disclosure. DETAILED DESCRIPTION
[0026] In order to make the objectives, technical solutions and advantages of the present disclosure more obvious, the following will describe example embodiments according to the present disclosure in detail with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the example embodiments described herein.
[0027] As shown in the present disclosure and claims, unless the context clearly suggests otherwise, the words “a”, “an”, “one”, and / or “the” do not necessarily refer to only one, but can include more than one. Generally, the terms “include” and “comprise” only indicate the inclusion of the explicitly identified steps and elements, and these steps and elements do not constitute an exclusive list of steps or elements.
[0028] Although the present disclosure makes various references to certain modules in the apparatuses according to embodiments of the present disclosure, however, any number of different modules can be used and run on the user terminal and / or server. The modules are merely illustrative, and different aspects of the apparatuses and methods can use different modules.
[0029] Flowcharts are used in the present disclosure to illustrate the operations performed by the methods and apparatuses according to embodiments of the present disclosure. It should be understood that the preceding or following operations are not necessarily performed in sequence. Instead, various steps can be processed in reverse order or simultaneously, as needed. Meanwhile, other operations can also be added to these processes, or one or more steps of the operations can be removed from these processes.
[0030] For the convenience of describing the present disclosure, the following introduces concepts related to the present disclosure.
[0031] Embodiments of the present disclosure relate to the design and control of “dexterous hands” in the field of intelligent robots. Dexterous hands are used to simulate the complex movements and functions of human hands, aiming to endow robots with similar flexibility and operating ability as human hands. Through precise mechanical structures and control systems, dexterous hands can perform tasks such as grasping, carrying and operating tools. Key technologies of dexterous hands include multi-degree-of-freedom finger design, integrated sensor systems and advanced control algorithms, enabling dexterous hands to perform precise operations in complex environments.
[0032] The mechanical structure of a "dexterous hand" typically consists of fingers, a palm, and connectors with multiple degrees of freedom. Each finger contains multiple joints, which are controlled by servo motors to achieve precise position control and dynamic response. To enhance the dexterous hand's sensing capabilities, it integrates various sensors, such as IMU sensors and joint angle encoders. These sensors provide real-time data on the actuator's acceleration, posture, joint angles, and angular velocities to achieve precise motion control.
[0033] With technological advancements, the applications of "dexterous hands" are constantly expanding. Beyond precision assembly and quality inspection in industrial automation, dexterous hands are demonstrating broad application potential in medical rehabilitation, service robotics, and education and research. For example, in medical rehabilitation, dexterous hands can assist in surgical procedures or be used to develop highly realistic prostheses, helping people with disabilities regain hand function. In the field of service robotics, dexterous hands enable robots to perform more complex household tasks, such as tidying clothes and preparing food. Furthermore, with advancements in artificial intelligence and machine learning technologies, the control strategies of dexterous hands are continuously being optimized to adapt to increasingly complex and dynamic task requirements.
[0034] The solutions provided in this application mainly involve artificial intelligence technology and apply it to the control field of robot end effectors, as illustrated in the following embodiments.
[0035] In summary, the solutions provided by the embodiments of this disclosure involve technologies such as artificial intelligence and machine learning. The embodiments of this disclosure will be further described below with reference to the accompanying drawings.
[0036] Figure 1 The illustration shows example scenarios 100 that can be applied to some embodiments of this disclosure. For example... Figure 1 As shown, scenario 100 includes a robot 110 with an end effector 111, which can be various types of execution structures, such as a dexterous hand or an end gripper. Scenario 100 also includes a tool 120 and a container 130 containing multiple objects 131. Exemplarily, a control strategy model or tool control model determined according to some embodiments of this disclosure can be deployed in the robot 110, such as in a controller network or processor structure for decision-making or control functions within the robot 110, to control the end effector 111 of the robot 110 to use the tool 120 to move at least a portion of the multiple objects 131 in the container 130 to a target location, such as to a specified height above the container 130.
[0037] exist Figure 1In the example shown in FIG. 1, the robot 110 is shown as an anthropomorphic robot, the tool 120 is shown as a spoon-like tool, the container 130 is shown as a bowl-like container, and the objects 131 are shown as spherical objects. However, depending on the specific application scenario, the robot 110 can also be other types of robots, such as a robotic arm, etc., the tool 120 can also be other forms of tools, such as other forms of scooping tools or other types of tools such as a pinching tool, the container 130 can also be other shapes or types of containers, and the objects 131 can also be other shapes, other sizes of objects. In addition, optionally, the container 130 can also not exist, i.e., the multiple objects 131 can be placed directly at the corresponding positions. Also, optionally, the robot can also be controlled to move at least part of the multiple objects 131 to another container.
[0038] Figure 2 FIG. 2 is a schematic diagram showing an end effector according to an embodiment of the present disclosure.
[0039] Figure 2 The end effector 111 is further described below with an example of a “dexterous hand”. Figure 2 The “dexterous hand” (i.e., the end effector 111) in FIG. 2 is composed of a central link and multiple multi-joint operating components, each finger is composed of multiple segments, each segment is equipped with at least one active bending joint. These joints are usually controlled by servo motors, which can provide precise position control and dynamic response, enabling the dexterous hand to perform delicate operations. The design of the joints allows for a wide range of motion, including extension and flexion, enabling diverse grasping modes and operational capabilities. In addition, the mechanical structure of the dexterous hand also includes lightweight but strong materials, such as aluminum alloys or carbon fibers, to ensure strength and durability while reducing weight, improving operational flexibility and response speed.
[0040] To enhance the perception and feedback capabilities of the “dexterous hand”, multiple sensors are integrated on its central link. Among them are inertial measurement unit (IMU) sensors, which measure the linear acceleration and angular velocity of the effector by integrating accelerometers and gyroscopes, providing data on the end effector’s attitude and motion. These data are helpful for performing precise motion control, especially when complex operations are required or when working in dynamic environments. In addition, optionally, each joint is equipped with a joint angle encoder to provide feedback on joint position and velocity, enabling precise motion control. The integration of these sensors enables adaptive control of the dexterous hand to cope with unpredictable external disturbances and changes.
[0041] The control system of the "dexterous hand" is responsible for coordinating the actions of the motors to achieve complex manipulation tasks. Through advanced algorithms, the "dexterous hand" can simulate the actions of a human hand, such as grasping, handling objects, and using tools. This ability makes the dexterous hand have application prospects in automated production lines, robotic surgery assistance, service robots, and scientific research fields. The control system usually includes a real-time operating system and advanced control algorithms, such as PID control, model predictive control, or machine learning algorithms, to achieve highly precise and adaptive operations.
[0042] Further, Figure 2 The "dexterous hand" in the above can be used to simulate human hands to use / operate tools. In daily life, the use of tools is essential to extend the physical capabilities of humans. For example, tools such as spoons and screwdrivers enable people to perform tasks such as scooping granular objects or liquids, tightening screws, etc., which would be difficult to complete without tools. The design of the "dexterous hand" enables it to mimic these basic human actions, and through integrated sensors and precise control systems, it can operate tools with high flexibility and precision.
[0043] However, the end effector often has difficulty operating tools to control target objects. Specifically, the interaction between the end effector, the tool, and the target object or environment is very complex. This complex interaction relationship means that the method designed for the end effector to directly manipulate objects cannot be simply applied to tasks that use tools to manipulate objects (i.e., tool manipulation tasks). Tool manipulation tasks require the end effector to understand and adapt to the dynamics of the tool and the interaction with the target object, which requires improvements and adjustments to the overall control strategy of the end effector.
[0044] In addition, the end effector is usually not in direct contact with the target object or environment during tool manipulation, which leads to insufficient observation and feedback on the contact state between the tool and the target object. The lack of such direct feedback makes it more difficult to achieve effective feedback control. Tool manipulation tasks often involve multiple target objects and complex physical constraints, making it difficult to deploy control methods based on traditional dynamics models. Therefore, in order to achieve effective tool manipulation in the end effector, improved control methods need to be developed to adapt to the complex relationship between the end effector, the tool, and the target object or environment, and to address the challenges posed by the lack of direct contact and the involvement of multiple objects and physical constraints.
[0045] Currently, to overcome the challenges in tool manipulation, data-driven approaches have become a hot research topic. Imitation learning methods show potential in this regard. The imitation learning approach guides the robot learning by analyzing the behavior of human operators (i.e., human demonstrations). However, the imitation learning approach usually relies on humans directly manipulating or remotely controlling the end effector of the robot during data collection, which is not only time-consuming but also inefficient.
[0046] To reduce the dependence on human demonstrations, reinforcement learning provides an alternative approach that allows the robot to learn skills autonomously through interaction with the environment with little or no direct guidance from humans. Although the reinforcement learning approach has achieved some success, the diversity and complexity of tool manipulation tasks make the training process of reinforcement learning challenging. For example, in a scooping task, the motion strategy of the robot needs to be adjusted according to factors such as the shape of the spoon and the bowl, the quantity and material of the target content, etc. In order to train a general strategy that can adapt to various scenarios, a large amount of training data is often needed to ensure that the neural network model carried by the end effector can effectively learn from the training data.
[0047] Therefore, in view of the above problems, according to an aspect of the present disclosure, a method for determining a control strategy model for controlling an end effector is provided, the end effector comprising a central connector and a plurality of multi-joint operation components, the method comprising: obtaining a first control strategy model carried on the end effector, and adjusting the first control strategy model carried on the end effector using a plurality of action rounds to determine a second control strategy model carried on the end effector, wherein in each action round of the plurality of action rounds, a learning task corresponding to the action round is determined based on information of the action round; a virtual body corresponding to the end effector simulates the learning task corresponding to the action round using the first control strategy model in the adjustment, and determines a reward score corresponding to completion of the learning task corresponding to the action round and a reward score related to a physical constraint in the process of simulating the learning task corresponding to the action round; and the first control strategy model is adjusted so that the reward score corresponding to completion of the learning task corresponding to the action round and the reward score related to the physical constraint in the process of simulating the learning task corresponding to the action round increase.
[0048] According to an aspect of the present disclosure, a method for controlling an end effector is also provided, the end effector comprising a central connector and a plurality of multi-joint operation assemblies, comprising: at a current time, obtaining a reference scene image when the end effector performs a target task and a pose corresponding to a contact point between the end effector and a second target tool; using a control policy model carried on the end effector, generating control information for controlling the end effector at a current time step; based on the control information for controlling the end effector at the current time step, determining motor driving information of the central connector and the plurality of multi-joint operation assemblies; and based on the motor driving information of the central connector and the plurality of multi-joint operation assemblies, controlling the end effector; wherein the control policy model is determined by the above method.
[0049] In order to obtain a second control policy model with higher generalization and robustness than the first control policy model, the present disclosure further adjusts the first control policy model through a reinforcement learning scheme to obtain the second control policy model, and uses specific reward scores in the adjustment process, i.e., both the "reward score corresponding to the learning task corresponding to the action round" and the "reward score related to the physical constraint in the process of simulating the learning task corresponding to the action round". This can ensure that the control policy model explores the control scheme as widely as possible while avoiding the trained control policy model in the simulation environment not meeting the physical constraint, thereby causing the control policy model to be unable to be applied to the physical environment.
[0050] The method for determining the control policy model provided by the present disclosure is particularly suitable for special tool operation tasks. Compared with the traditional training scheme for determining a control policy model suitable for special tool operation tasks, the present disclosure adjusts the first control policy model using a reinforcement learning scheme on the basis of the pre-trained first control policy model. This reduces the demand for human demonstration data and the training time required for reinforcement learning.
[0051] Specifically, in the method provided by the present disclosure, only training data for training the first control policy model needs to be collected, which significantly reduces the amount of data required for training compared with the traditional scheme. At the same time, since the first control policy model is adjusted to obtain a second control policy model in the method provided by the present disclosure to adapt to a second target task different from the first target task, the generalization of the control policy model is significantly improved compared with the traditional scheme. At the same time, since reinforcement learning is performed on the basis of the first control policy model, the control policy model does not need to perform excessive and invalid exploration in this process, which significantly reduces the time required for reinforcement learning.
[0052] In some embodiments of the present disclosure, imitation learning and reinforcement learning can also be combined to enable the neural network model onboard the end effector to learn from imitation with less human demonstrations and to obtain the ability to complete tool operation tasks through reinforcement learning based on imitation learning. Thus, the present disclosure can significantly reduce the need for collecting human demonstration data, enabling the neural network model onboard the end effector to quickly learn from imitation with a small amount of training data and effectively work in unseen scenarios through reinforcement learning. Through the embodiments of the present disclosure, the neural network model onboard the end effector can extract key information from limited training data and generalize it to new and unknown environments. Embodiments of the present disclosure not only improve learning efficiency, but also enhance the adaptability and flexibility of robots in diverse tasks.
[0053] Next, reference is made to Figure 3 The embodiments of the present disclosure are described in general.
[0054] Figure 3 is a flowchart showing a method 30 of determining a control policy model according to an embodiment of the present disclosure.
[0055] The method 30 can be used to determine a control policy model (e.g., a first control policy model and a second control policy model as detailed below). The control policy model is used to control the end effector to perform a target task (e.g., a first target task and a second target task as detailed below) with a target tool (e.g., a first target tool and a second target tool as detailed below). The target task can include moving at least one target object (e.g., a first target object and a second target object as detailed below) from an initial position (e.g., a first initial position and a second initial position as detailed below) to a target position (e.g., a first target position and a second target position as detailed below) with the target tool (e.g., a first target tool and a second target tool as detailed below). Of course, the present disclosure is not limited thereto.
[0056] Exemplarily, the method 30 can be used to control Figure 1 and Figure 2 the end effector 111 of the robot 110 as shown to control the end effector 111 of the robot 110 to move the object 131 in the container 130 using the tool 120. Optionally, the method 30 can be implemented in a real machine environment or in a virtual environment (or simulation environment). Optionally, the target tool can be any type of tool, such as a digging tool, a clamping tool, a sticking tool, a sucking tool, a fork tool, etc., and the target object can be any form of object, such as a spherical object, a cuboid object, an irregularly shaped object, etc. Of course, the present disclosure is not limited thereto.
[0057] The initial position should be understood as the location of multiple target objects before the execution of the target task begins. For example, multiple target objects can be placed in a container at the initial position, or they can be scattered or piled up at the initial position. Similarly, the target position should be understood as the position where the target objects are expected to be moved. For example, it may be expected that at least one of the multiple target objects will be moved to a container at the target position, or it may be expected that only the object will be moved or placed at the target position. It should be understood that the initial position or target position can refer to a range of positions, and not necessarily a precise point. Of course, this disclosure is not limited to this.
[0058] The method 30 for determining a control strategy model according to embodiments of the present disclosure may include, for example: Figure 3 The steps S301 to S302 are shown. Of course, method 30 may include more or fewer operations, and this disclosure is not limited thereto. Optionally, step S301 may be an operation in the pre-training phase. Operation S302 may be an operation in the reinforcement learning phase.
[0059] As described above, the end effector includes a central connector and multiple joint-operated components (e.g., a robotic finger). The central connector and multiple joint-operated components can be referred to as... Figure 1 and Figure 2 The relevant descriptions will not be repeated here.
[0060] In step S301, the first control strategy model mounted on the end effector is obtained.
[0061] Optionally, the first control strategy model is used to control the end effector to perform a first target task, wherein, during the execution of the first target task, the end effector uses a first target tool to move a first target object from a first initial position to a first target position. Optionally, during the movement of the first target object from the first initial position to the first target position, the first target tool provides a force to the first target object, and the end effector does not contact the first target object. This disclosure is not limited thereto.
[0062] Optionally, the first control strategy model can be a neural network model with a preset network structure, which can be constructed based on existing neural network structures in related technologies, or it can be any new network structure. Furthermore, the first control strategy model can also be other types of models. This disclosure does not specifically limit the specific model type and structure of the first control strategy model.
[0063] Optionally, the first control policy model is a pre-trained neural network model. The first control policy model can be pre-trained in various ways, including but not limited to transfer learning, imitation learning, self-supervised learning. The following will briefly introduce the process of "obtaining the first control policy model carried on the end effector" by taking imitation learning as an example. Those skilled in the art should understand that the present disclosure is not limited thereto.
[0064] Optionally, step S301 comprises: obtaining a reference scene image when the end effector executes the first target task and a reference trajectory corresponding to the contact point between the end effector and the first target tool; and determining the first control policy model carried on the end effector based on the reference scene image when the end effector executes the first target task and the reference trajectory corresponding to the contact point between the end effector and the first target tool, the first control policy model being used to control the end effector to execute the first target task. The present disclosure is not limited thereto.
[0065] Optionally, the reference scene image when the end effector executes the first target task is not limited to one image, but can include multiple images or even a video. These reference scene images can show the scene at different times during the execution of the first target task by the end effector, providing more comprehensive information to help the end effector better understand and identify the progress of the target task. For example, for the scene of scooping particles with a spoon in Figure 1 , the reference scene images at different times show different relative positions and states of the bowl and the spoon, and these images continuously show the progress of the task.
[0066] Optionally, the reference scene image when the end effector executes the first target task is a depth image captured by a depth camera. The depth camera calculates the distance of an object by emitting infrared light and measuring the time of the reflected light, generating a depth image. Optionally, the reference scene image includes three-dimensional information of the scene in which the end effector executes the first target task. In this way, the first strategy model carried on the end effector can perform accurate spatial positioning and object recognition based on the reference scene image, thereby enabling the robot to better understand its environment and determine the target task to be executed. Of course, the present disclosure is not limited thereto.
[0067] Optionally, the reference trajectory corresponding to the contact point between the end effector and the first target tool is a motion trajectory of the contact point when the end effector performs the first target task. The reference trajectory can include at least one of pose information, position information, velocity information, acceleration information, angular velocity information, and angular acceleration information of the contact point at each time step. Optionally, the reference trajectory can be collected by a sensor on the end effector. The position information can be three-dimensional coordinates in a specified coordinate system, such as three-dimensional coordinates of the contact point, and the pose information can be described by a quaternion. Alternatively, the pose of the target tool can be described in other ways, or the number of parameters can be adjusted according to the actual allowed degrees of freedom. Optionally, the position change or the pose change can be described in the same way as the pose of the target tool. Of course, the present disclosure is not limited in this way.
[0068] In one example, the reference trajectory can be represented by a time sequence of position information of the contact point at each time step. Optionally, each element in the time sequence can have multiple dimensions, which respectively represent the position of the contact point in the x-axis direction, the position of the contact point in the y-axis direction, the position of the contact point in the z-axis direction (the direction of gravity), and the pose information represented by a quaternion. Of course, the reference trajectory can also be represented by other data structures, and the present disclosure is not limited in this way.
[0069] In one example, the reference scene image and the reference trajectory can be obtained in the following way. In a real physical space, a human operator controls the end effector to operate a first target tool to perform a first target task. During this process, the human operator controls the end effector to interact with the selected first target tool (such as a spoon, a screwdriver, a pliers, etc.) to complete a specific target task (such as scooping granular or liquid, tightening a screw, clamping an object, etc.) through direct or indirect control. The reference scene image and the reference trajectory corresponding to the contact point between the end effector and the first target tool can be captured in real time by deploying a depth camera in the real physical space and a sensor on the end effector. Of course, the present disclosure is not limited in this way.
[0070] In another example, the reference scene image and the reference trajectory can also be obtained in the following way. In a simulation environment, an animation of a human operator controlling a virtual end effector to operate a virtual first target tool to perform a first target task is drawn. In the animation, the end effector is controlled to interact with the virtual first target tool (such as a virtual spoon, a virtual screwdriver, a virtual pliers, etc.) to complete a specific target task (such as scooping virtual granular matter or liquid, fastening a virtual screw, clamping a virtual object, etc.). The reference scene image and the reference trajectory corresponding to the contact point between the end effector and the first target tool can be captured in real time by deploying a perspective camera in the simulation environment. Of course, the present disclosure is not limited thereto.
[0071] Optionally, the process of determining the first control policy model mounted on the end effector is the process of training the first control policy model. In this training process, the reference scene image and the reference trajectory corresponding to the contact point between the end effector and the first target tool when the end effector performs the first target task are used as training data, and the model parameters of the first control policy model are constantly adjusted until a predetermined number of iterations or model parameter convergence is reached. Thus, the first control policy model attempts to enable the end effector to autonomously move through its own motor without the control of a human operator, so that the end effector can control the first target tool to move along the reference trajectory to move the first target object from the first initial position to the first target position or the vicinity of the first target position.
[0072] Optionally, as mentioned above, the reference scene image and the reference trajectory can be obtained in both the physical space and the simulation environment. The reference scene image and the reference trajectory provide reference information for the first control policy model, which can gradually adjust the control information for controlling the end effector. The control information for controlling the end effector indicates the action that the end effector should perform at the current time step. By performing the action, the end effector can make the contact point between the end effector and the first target tool at the next time step substantially consistent with the expected pose of the reference trajectory.
[0073] Although the training process of the first control policy model can adopt the method of imitation learning to ensure that the first control policy model learns a rough control policy capable of achieving the first target task in limited training data. However, due to the inherent limitations of imitation learning, once the real scene changes or any parameter in the target task changes (for example, at least one of the target tool, the target object, the initial position of the target object, or the target position changes), the performance of the first control policy model trained by imitation learning will be poor. Therefore, the first control policy model will be adjusted (or fine-tuned) through step S302 to obtain a control policy model (i.e., a second control policy model) with higher generalization and robustness.
[0074] In step S302, the first control policy model carried on the end effector is adjusted using multiple action rounds to determine a second control policy model carried on the end effector. In each action round of the multiple action rounds, a learning task corresponding to the action round is determined based on information of the action round; a virtual body corresponding to the end effector simulates the learning task corresponding to the action round using the first control policy model in the adjustment, and determines a reward score corresponding to completion of the learning task corresponding to the action round and a reward score related to a physical constraint in a process of simulating the learning task corresponding to the action round; and the first control policy model is adjusted to increase the reward score corresponding to completion of the learning task corresponding to the action round and the reward score related to the physical constraint in the process of simulating the learning task corresponding to the action round.
[0075] Optionally, the second control policy model carried on the end effector is used to control the end effector to perform a second target task, wherein in the process of the end effector performing the second target task, the end effector moves the second target object from the second initial position to the second target position using the second target tool, and the first target task and the second target task are different.
[0076] Optionally, in the process of moving the second target object from the second initial position to the second target position, the second target tool provides a force to the second target object, and the end effector is not in contact with the second target object. Of course, the present disclosure is not limited thereto.
[0077] Optionally, the first target task and the second target task satisfy at least one of the following conditions: the first target tool is different from the second target tool; the first target object is different from the second target object; the number of the first target objects is different from the number of the second target objects; the first initial position is different from the second initial position; the first target position is different from the second target position; or the first target object is in a first container and the second target object is in a second container, and the first container and the second container are different. Of course, the present disclosure is not limited thereto.
[0078] Optionally, to obtain a second control policy model with higher generalization and robustness than the first control policy model, the second control policy model can be the first control policy model fine-tuned by reinforcement learning. The second control policy model has the same architecture as the first control policy model, but the generalization of the second control policy model is higher than that of the first control policy model. The training process of reinforcement learning enables the second control policy model to learn more environmental dynamics and potential strategies, so that it can make more reasonable decisions when facing new situations.
[0079] Specifically, in the reinforcement learning process, by learning the same or different learning tasks in each action round, the obtained second control policy model can explore as many control policies as possible, so that the control policy model is more "intelligent". Optionally, in the process of simulating the learning task corresponding to the action round by the virtual body corresponding to the end effector, the end effector simulates the movement of the target object from the initial position specific to the action round to the target position corresponding to the action round by using the target tool specific to the action round.
[0080] Optionally, in the process of reinforcement learning, the idea of human learning courses can be used for reference. Specifically, in the process of learning a certain skill in school, humans often need to consolidate learned knowledge and improve performance through reasonably set courses, and then learn higher-order skills. Generally, humans need to learn courses from easy to difficult to gradually adapt to more complex environments, where a course can be a developing concept, can be composed of certain educational goals, specific knowledge and experience, and expected learning activity methods, and contains a rich, basic, and creative set of plans and settings. Based on this, the present disclosure proposes that a control policy model can be adapted to different task difficulty scenarios from easy to difficult through a set of designed processes, and finally converges to a model capable of completing high difficulty tasks. Inspired by the human course learning process, the present disclosure introduces the idea of course learning into the training process of reinforcement learning to solve the problem that the control policy model lacks exploratory after learning basic strategies or is difficult to converge when directly facing high difficulty tasks.
[0081] Optionally, in S302, the information of the action round includes course difficulty information corresponding to the action round and course progress information, and the target object specific to the action round or the target tool specific to the action round is determined.
[0082] Optionally, in order to enhance the robustness of reinforcement learning, the initial position specific to the action round and the target position corresponding to the action round can also be randomly determined. Of course, the present disclosure is not limited thereto.
[0083] The course difficulty information corresponding to the current action round is associated with the course progress information of the current action round, and the course progress information indicates that the higher the course difficulty information corresponding to the current action round is, the later the order in which the action round is performed is. Of course, the present disclosure is not limited thereto.
[0084] Optionally, in some embodiments of the present disclosure, the course difficulty information can also be dynamically adjusted. For example, the course difficulty information corresponding to the current action round can be associated with the reward score of the learning task corresponding to the previous action round, and the higher the reward score of the learning task corresponding to the previous action round is, the higher the course difficulty information corresponding to the current action round is. Of course, the present disclosure is not limited thereto.
[0085] Optionally, in some embodiments of the present disclosure, the same course difficulty information corresponds to multiple action rounds, and the learning tasks corresponding to different action rounds are different, wherein the learning tasks corresponding to different action rounds are different at least one of the following conditions: the target tool specific to different action rounds is different; the target object specific to different action rounds is different; the number of target objects specific to different action rounds is different; or the target objects specific to different action rounds are located in different containers. In addition, in the learning tasks corresponding to different action rounds, the initial position specific to different action rounds can also be different, or the target position specific to different action rounds can also be different. Of course, the present disclosure is not limited thereto.
[0086] Thus, in the process of reinforcement learning, the model parameters of the preset model can be updated based on the execution of the learning task, so as to converge towards the direction of successfully completing the learning task. When the course progress information indicates that the number of times of successfully executing the learning task in multiple consecutive action rounds satisfies a threshold condition, the course difficulty information can be updated, and subsequently, the next stage of model training operation can be performed based on the updated course difficulty information. Thus, the number of targets can be gradually increased to gradually increase the course difficulty, thereby gradually guiding the virtual body corresponding to the end effector to complete a higher difficulty target task through exploration, avoiding the problem that the control policy model converges to only complete a low difficulty task or the exploration space is too large to converge.
[0087] In addition, since the reinforcement learning process is implemented in a simulation environment, this solution can avoid real machine loss as much as possible. Alternatively, the "reward score corresponding to the completion of the learning task corresponding to the action round" refers to the reward score obtained by the virtual body corresponding to the end effector after completing a specific learning task in a virtual environment, and the "reward score related to physical constraints in the process of simulating the execution of the learning task corresponding to the action round" refers to the reward score obtained by the virtual body corresponding to the end effector when executing the task in a simulation environment, taking into account the physical constraint conditions. Both of these reward scores are sparse rewards that can encourage the control policy model to explore more extensively in the simulation environment to find behaviors that may lead to rewards, thereby possibly discovering new and more effective strategies.
[0088] The "reward score corresponding to the completion of the learning task corresponding to the action round" can make the virtual body corresponding to the end effector have to learn from limited feedback, which forces the algorithm to use samples more efficiently, thereby improving the sample efficiency of the learning process. At the same time, since the virtual body corresponding to the end effector is exposed to more diverse situations during training, it is more likely to develop a strategy that can generalize to new situations, rather than being limited to specific behaviors that often result in rewards.
[0089] For example, for the completion of the scooping action, the following conditions can be used to evaluate: first, whether the target tool specific to the action round (i.e., the virtual body corresponding to the spoon) is located directly above the virtual body corresponding to the container; second, whether the amount of objects scooped into the virtual body corresponding to the spoon (i.e., the number of target objects specific to the action round) meets a predetermined standard. Specifically, the sufficiency of the amount of objects needs to consider factors such as the capacity of the spoon, the volume of the objects, and the height of the contents in the spoon to determine. Through these specific conditions, it can be effectively determined whether the virtual body corresponding to the end effector has completed the corresponding learning task.
[0090] Alternatively, the "reward score corresponding to the completion of the learning task corresponding to the action round" can be set to indicate at least one or a combination of multiple items, such as whether the virtual body corresponding to the end effector successfully moves the target object specific to the action round from the initial position specific to the action round to the target position corresponding to the action round using the target tool specific to the action round, or whether the target object spills during the process of simulating the execution of the learning task corresponding to the action round by the virtual body corresponding to the end effector. Of course, the present disclosure is not limited thereto.
[0091] By "setting the reward score related to the physical constraint in the process of simulating the learning task corresponding to the action round", the virtual body corresponding to the end effector can only perform actions that meet the environmental physical constraints and prevent undesirable training results. For example, for the scooping action, due to the difference between the simulation environment and the physical space, the actions simulated by the virtual body corresponding to the end effector in the simulation environment may not meet the physical constraints. These actions are not possible in the real space, so it is necessary to determine whether the virtual body corresponding to the spoon is correctly inserted into the virtual body corresponding to the container, whether the virtual body corresponding to the spoon is in contact with the virtual body corresponding to the target object, and so on. These all need to be designed to avoid corresponding reward scores to ensure the effectiveness and safety of training.
[0092] Optionally, the "reward score related to the physical constraint in the process of simulating the learning task corresponding to the action round" can be set to indicate a combination of at least one or more of the following: whether the target tool specific to the action round collides with the target object specific to the action round in the process of simulating the learning task corresponding to the action round by the virtual body corresponding to the end effector; whether the target tool specific to the action round collides with the container in which the target object specific to the action round is located in the process of simulating the learning task corresponding to the action round by the virtual body corresponding to the end effector; whether the end effector is separated from the target tool in the process of simulating the learning task corresponding to the action round by the virtual body corresponding to the end effector; and whether the target tool is out of the observation range in the process of simulating the learning task corresponding to the action round by the virtual body corresponding to the end effector. Of course, the present disclosure is not limited thereto.
[0093] Specifically, the training process of the reinforcement learning can be briefly described as follows: based on the scene image at the current time step in the process of the end effector performing the second target task and the pose corresponding to the contact point between the end effector and the second target tool, the control information of the end effector at the current time step is determined by using the first control policy model; the reward score for the control information of the end effector at the current time step is determined based at least in part on the control information of the end effector at the current time step; and the model parameters of the first control policy model are updated based at least in part on the reward score for the control information of the end effector at the current time step to determine the second control policy model carried on the end effector. Then reference will be made to Figure 4 Further description of the process, the present disclosure will not be repeated here. Of course, the present disclosure is not limited thereto.
[0094] Through the training process of reinforcement learning, the second control policy model determines the control information corresponding to each time step at each time step, and then learns according to the feedback (reward or punishment). As the process of reinforcement learning deepens, the second control policy model learns how to transition between different states to achieve long-term goals, improving its generalization ability. Through reinforcement learning, the second control policy model not only inherits the knowledge learned by the first control policy model, but also enhances flexibility and adaptability.
[0095] Therefore, the present disclosure proposes an improved method for determining a control policy model. Specifically, in order to obtain a second control policy model with higher generalization and robustness than the first control policy model, the present disclosure further adjusts the first control policy model through the reinforcement learning scheme to obtain the second control policy model, and uses specific sparse reward scores in the adjustment process, that is, both "reward score corresponding to completing the learning task corresponding to the action round" and "reward score related to physical constraints in the process of simulating the execution of the learning task corresponding to the action round". It can ensure that the control policy model explores the control scheme as widely as possible while avoiding the trained control policy model in the simulation environment not meeting the physical constraints, thereby leading to the inability to be applied to the physical environment.
[0096] In addition, the method for determining a control policy model proposed by the present disclosure is particularly suitable for special tool operation tasks. Compared with the traditional training scheme for determining a control policy model suitable for special tool operation tasks, the present disclosure adjusts the first control policy model using the reinforcement learning scheme based on the pre-trained first control policy model. On the one hand, it reduces the demand for human demonstration data, and on the other hand, it reduces the training time required for reinforcement learning.
[0097] Further, the present disclosure combines reinforcement learning with the idea of curriculum learning, so that the control policy model carried by the end effector can adapt to changing operating environments and task requirements during the reinforcement learning process, achieving more flexible and efficient operating performance.
[0098] Notably, in some embodiments of the present disclosure, reinforcement learning is combined with imitation learning, thereby significantly reducing the need for collecting human demonstration data, enabling the control policy model (e.g., any one of the first control policy model and the second control policy model mentioned above) carried by the end effector to quickly perform imitation learning through a small amount of training data and effectively work in unseen scenarios using reinforcement learning. Through embodiments of the present disclosure, the control policy model carried by the end effector can extract key information from limited training data and generalize it to new and unknown environments. Embodiments of the present disclosure not only improve learning efficiency, but also enhance the adaptability and flexibility of robots in diverse tasks.
[0099] Next, reference is made to Figure 4 Details of embodiments of the present disclosure are further described.
[0100] Figure 4 is a schematic diagram illustrating a method 30 of determining a control policy model according to an embodiment of the present disclosure.
[0101] Optionally, embodiments of the present disclosure improve the architecture of the control policy model. Compared with the structure of a conventional policy model, the improved control policy model (i.e., the first control policy model and the second control policy model) is composed of two decoupled modules: a visual encoder and a controller network. In the process of adjusting the first control policy model to obtain the second control policy model, a value network decoupled from both the visual encoder and the controller network can also be introduced to assist the adjustment of the first control policy model. In addition, the visual encoder can also be combined with an encoder for encoding sensory information such as a haptic information encoder to better capture scene information and progress information of target task execution. Of course, the present disclosure is not limited thereto.
[0102] Optionally, the visual encoder processes the input reference scene image to extract a visual encoding feature vector. The visual encoding feature vector fuses relative pose information between the end effector and the target tool and progress information of the target task. The visual encoder converts the reference scene image into a form understandable by machines, providing the controller network with the information needed for decision-making. The controller network calculates the pose change amount of the target tool corresponding to each time step according to the visual encoding feature and the reference trajectory corresponding to the contact point between the end effector and the target tool.
[0103] Reference is made to Figure 4In the imitation learning stage, in step S301, the first control policy model can be obtained by training the first control policy model. Optionally, as a specific example, in the imitation learning stage, reference scene images and reference trajectories corresponding to the contact points between the end effector and the first target tool when the end effector performs the first target task in the physical space can be collected first. Then, in the simulation environment, the data collected in the physical space is projected, and the data projected into the simulation environment is augmented, so as to collect relatively more training data while reducing human demonstration as much as possible. Of course, the training data can also be augmented in the simulation environment, and the present disclosure is not limited thereto.
[0104] Optionally, assuming that in the physical space, the reference scene images and the reference trajectories corresponding to the contact points between the end effector and the first target tool when the end effector performs the first target task are obtained. Then, the reference scene images and the reference trajectories corresponding to the contact points between the end effector and the first target tool when the end effector performs the first target task in the physical space can be projected into the simulation environment first to generate the projection data of the reference scene images in the simulation environment and the projection data of the reference trajectories projected in the simulation environment. The first control policy model carried on the end effector is trained to enable the first control policy model to generate control information for controlling the end effector at a current time step based on the pose data of the first target tool in the projection data of the reference scene images in the simulation environment and the projection data of the reference trajectories projected in the simulation environment. The control information for controlling the end effector indicates the action performed by the end effector at the current time step.
[0105] Specifically, the reference scene images and the reference trajectories captured in the physical space can be obtained by sensors and cameras. These data are preprocessed, such as scaling, cropping and normalization, to adapt to the input requirements of the simulation environment. Then, a simulation environment that simulates the original physical scene as much as possible is constructed in MuJoCo. MuJoCo is a high-level physics engine for simulating multi-joint characters and rigid body dynamics, including contact and collision. On this basis, trajectory models in the simulation environment are designed according to the actually captured reference trajectories, which can be used as both the target of agent training and dynamic elements in the environment. Subsequently, the processed reference scene images and reference trajectories can be projected into the simulation environment, which involves the combination of image rendering and physical simulation. In this process, OpenAI Gym serves as a toolkit for developing and comparing reinforcement learning algorithms, providing a unified interface to describe the interaction between the environment and the agent. Gym supports a variety of predefined environments, including simulation environments that use MuJoCo as the backend, providing convenience for imitation training.
[0106] More specifically, a MuJoCo XML file can be written to model the physics simulation. For example, a corresponding model can be loaded in the MuJoCo environment through the XML file, such as a scenario of scooping a target object (e.g., a small ball) from a target container (e.g., a bowl) by a digging tool (e.g., a spoon). The corresponding virtual bodies of the end effector, the spoon, and the bowl (referred to as virtual end effector, virtual spoon, and virtual bowl, respectively) can be loaded in the MuJoCo environment, and the virtual end effector and the virtual spoon can be rigidly connected together, where the virtual end effector can be regarded as an actuator to control the virtual spoon. The welding point of the virtual end effector and the virtual spoon can be regarded as a contact point of the end effector and the spoon in the real robot environment. Since the spoon and the bowl are typical non-convex bodies with complex geometries, they cannot be directly implemented in the virtual environment. An embodiment of the present disclosure can slice the scanned real spoon and bowl models into multiple convex hulls and then recombine them into complete objects.
[0107] In addition, a number of small balls can also be added to the environment through the XML file. The MuJoCo environment written can be accessed using the MujocoEnv provided by OpenAI Gym. The MujocoEnv can be created, and the environment description variables such as action_space, observation_space, etc. can be defined, and the function functions such as init (initialization), step, reset, etc. can be implemented. The MujocoEnv can correspond to the written MuJoCo environment. The action_space can be used to define the action information, which can include the pose change amount of the target tool, which can be represented by 7 bits of information, for example. The observation_space can be used to define the model input information, which can include the reference trajectory composed of the pose information of the contact point, the relative position relationship between the target object and the target tool, etc. In order to determine the relative pose between the target tool and the target object, an embodiment of the present disclosure can also set a plurality of position points on the target tool to determine whether the target object is accurately located on the target tool.
[0108] For example, for the scenario of scooping up small balls, 11-bit input parameters can be defined, in which 7 bits represent the parameters of the spoon pose, 3 bits of one-hot code represent the position relationship between the small ball and the spoon, and 1 bit represents the height of the uppermost small ball (i.e., the liquid level) in the initial state, wherein in the one-hot code design, 100 is used to represent that there is no small ball in the spoon, 010 represents that there is a small ball in the spoon, and 001 represents that there is a small ball in the spoon and the centroid of the small ball is higher than the centroid of the spoon, or 18-bit input parameters can be defined, in addition to the aforementioned parameters, 7 bits represent the pose of the contact point between the virtual spoon and the virtual end effector. For example, for the scenario of clamping a block with a clamp, 11-bit input parameters can be defined, in which 7 bits represent the parameters of the clamp pose, 1 bit represents the opening angle of the clamp, 2 bits of one-hot code represent the position relationship between the target object and the clamp, and 1 bit represents the task scenario, wherein the position relationship between the target object and the clamp has two types, i.e., clamped and not clamped.
[0109] As described above, the first control strategy model carried on the end effector includes a visual encoder and a controller network. The visual encoder and the controller network can be trained separately. The input of the visual encoder is a reference scene image corresponding to a single moment, and the output is a corresponding visual encoding vector, which can capture the key features of the reference scene image and provide the controller network with the execution progress and current state of the current target task.
[0110] To facilitate training in a simulation environment and subsequent deployment on a real machine, the visual encoder should be able to encode both reference scene images in a physical space and reference scene images in a simulation environment, and output the same visual encoding vector for the same scene. The visual encoder will serve as a bridge between the simulation environment and the physical space. To this end, the visual encoder can be trained using the following scheme: using the visual encoder in training, encoding the reference scene image when the end effector in the physical space performs a first target task to generate a visual encoding vector; using the visual encoder in training, encoding the projection data of the reference scene image in the simulation environment to generate visual label data; adjusting the model parameters of the visual encoder to make the difference between the visual encoding vector and the visual label data converge. In this way, the visual encoder can output the same visual encoding vector for the same scene in the physical space or in the simulation environment, thereby providing accurate input data to the controller network. Of course, the present disclosure is not limited thereto.
[0111] Next, in step S301, the controller network can be trained separately after the training of the visual encoder is completed. Optionally, the training process comprises: generating, by the trained visual encoder, a visual encoding vector corresponding to the current time step based on the projection data of the reference scene image corresponding to the current time step in the simulation environment; predicting, by the controller network under training, a predicted value of the pose data of the contact point at the next time step based on the visual encoding vector corresponding to the current time step and a ground truth value of the pose data of the contact point at the current time step in the projection data of the projected reference trajectory in the simulation environment; and adjusting the model parameters of the controller network under training so that the difference between the predicted value of the pose data of the contact point at the next time step and the ground truth value of the pose data of the contact point at the next time step in the projection data of the projected reference trajectory in the simulation environment converges. Of course, the present disclosure is not limited thereto.
[0112] It is worth noting that the pose data of the contact point at the next time step can be used to indicate the control information for controlling the end effector at the current time step. Specifically, the difference between the pose of the contact point at the next time step and the pose at the current time step will determine how the individual motors of the end effector should be driven at the current time step. That is, if the pose data of the contact point between the end effector and the target tool at the next time step can be accurately predicted, it is known how to control the end effector.
[0113] As an example, assume that the visual encoding vector corresponding to the current time step can be characterized as an 8-dimensional feature vector O visual . Subsequently, this visual encoding vector is concatenated with a 7-dimensional pose vector O pose of the end effector (including position and quaternion information) to form a 15-dimensional feature vector (O pose , O visual ). This composite feature vector is then input into the controller network to generate an action vector a output that guides the end effector to perform the next operation. The action vector a output is the predicted value of the pose data of the contact point at the next time step, and its training label a data is the ground truth value of the pose data of the contact point at the next time step in the projection data of the projected reference trajectory in the simulation environment.
[0114] Optionally, a mean square error loss function (MSE) can be used to measure the output action vector a output and its training label a dataThe difference between the two. This loss function can quantify the accuracy of action prediction, providing guidance for controller network training. At the same time, the Adam optimizer is used to optimize the parameters of the visual encoder and the controller network. The Adam optimizer is widely used due to its adaptive learning rate feature, which can automatically adjust the learning rate during training based on historical gradient information, thereby speeding up convergence and improving training results.
[0115] After completing the imitation training (where the visual encoder and the controller are trained separately) in step S301, a first control strategy model that can only barely complete the basic goal task has been obtained. According to the method 30 described in the foregoing, the reinforcement learning training process in step S302 can be started to obtain a second control strategy model with stronger generalization ability.
[0116] Optionally, in operation S302, the process of determining the reward score of the control information of the end effector at the current time step includes: determining the cumulative reward score of the control information of the end effector at the current time step using a value network. Specifically, the training process of reinforcement learning needs to introduce a value network. The value network receives the same input as the controller network and outputs a value for evaluating the action output generated by the controller network (i.e., the reward score related to the physical constraint in the process of completing the learning task corresponding to the action round and simulating the learning task corresponding to the action round).
[0117] Thus, the reinforcement learning training process can be modeled as a Markov decision process defined by the tuple wherein, represents the state space, represents the action space, p(s ′ |s,a) is the state transition function, representing the probability of reaching state s ′ after performing action a in state s. The function defines the cumulative reward score in the target task, the value of which is determined by the value network. γ ∈ [0, 1) is the discount factor, used to evaluate the importance of future reward scores.
[0118] As mentioned above, some embodiments of the present disclosure can also combine the training process of reinforcement learning with curriculum learning, so that the first control strategy model not only continues to explore the environment of the tool operation task, but also constantly adjusts and optimizes the predicted action according to the exploration result. In this way, the first control strategy model can gradually improve the accuracy of the action to better adapt to and complete various tool operation tasks.
[0119] In the process of reinforcement learning, sparse reward scores are used to describe task requirements, including "reward scores for completing the learning task corresponding to the action round" and "reward scores related to physical constraints in the process of simulating the execution of the learning task corresponding to the action round". The "reward scores for completing the learning task corresponding to the action round" are only given when the task is completed. In addition, the embodiments of the present disclosure also design "reward scores related to physical constraints in the process of simulating the execution of the learning task corresponding to the action round" to constrain actions to meet environmental physical constraints and prevent undesirable training results.
[0120] Specifically, the reward scores can be defined in the following way to guide the end effector to perform the learning task in the desired way.
[0121] For "reward scores for completing the learning task corresponding to the action round", if the virtual body corresponding to the end effector successfully completes the task, i.e., successfully scoops up a target object, the reward score will increase by 150 points; and for each overflow of a target object, the reward score will decrease by 250 points. Thus, such reward scores will guide the end effector to learn how to effectively scoop up objects while avoiding excessive object scattering, ensuring that its behavior meets the requirements for accuracy and efficiency in actual operations.
[0122] For "reward scores related to physical constraints in the process of simulating the execution of the learning task corresponding to the action round", if the target tool (e.g., a spoon) collides with a bowl or an object, the target tool falls off the target effector, or the target tool exceeds the preset position limit, the reward score will be deducted by 100 points. These reward scores are designed to ensure that the end effector maintains the stability of the target tool during operation, avoids collisions, and limits its actions within the workspace of the end effector. Thus, the end effector can learn how to efficiently complete the task while complying with physical constraints.
[0123] In the curriculum design of tool operation tasks, the focus is on the attributes of the virtual body corresponding to the target tool, the attributes of the virtual body corresponding to the target object, and the attributes of the task environment (e.g., the virtual body of the container where the target object is located). For example, in the scooping task, the difficulty of the curriculum can be achieved by adjusting three key factors: the height of the target object in the container, the size of the target object, and the size of the target tool - the spoon.
[0124] Optionally, when the height of the target object in the container is lower, the spoon needs to be inserted deeper to contact the object, which not only increases the risk of collision between the spoon and the wall and the bottom of the container, but also makes the entire scooping task more challenging. Similarly, the larger the size of the target object, the less amount of object the spoon can scoop in a single operation, which also correspondingly increases the difficulty of the course. In addition, using a smaller spoon will limit the range of visual perception, making it more difficult for the visual encoder to capture object features, further increasing the complexity of the task. Through these meticulous adjustments, the course can gradually improve the operation skills and adaptability of the trainee.
[0125] The present disclosure designs multiple structured courses that mix different target object heights, target object sizes, and target tool sizes as multi-dimensional difficulty variables of the control policy model to learn various scooping tasks in the reinforcement learning training process. The difficulty information of these structured courses is evaluated and sorted in order from easy to difficult, allowing the control policy model to steadily improve performance and acquire new skills during learning. By continuously evaluating the performance of the control policy model in the current reinforcement learning process, it can be decided whether to switch to a more difficult or easier course. In addition, previously learned courses are interspersed at certain stages to prevent potential problems of catastrophic forgetting of the control policy network. This course design strategy aims to promote the robustness of the control policy model by gradually increasing the difficulty and regularly reviewing.
[0126] For example, in one embodiment of the present disclosure, the course can be set through two stages, and then the learning task corresponding to each action round is designed. Of course, the present disclosure is not limited thereto.
[0127] Specifically, in the first stage, the difficulty variable types of each course can be determined first, such as the size of the target tool, the size of the target object, the number of target objects, the initial position of the target tool, the target position of the target tool, the container where the target object is located, and the like. In the scooping task, the difficulty variable types are, for example, the height of the target object in the container, the size of the target object, and the size of the target tool-spoon, and the difficulty coefficients corresponding to each specific difficulty variable are evaluated, such as 1 centimeter or small size. Then, different learning tasks are obtained by combining different types of difficulty variables, and the combined difficulty coefficients of these learning tasks are calculated. Next, these courses are sorted in ascending order according to the course difficulty information.
[0128] In the second stage, the control policy model is trained according to this structured course training strategy, and the performance of the control policy model in the current course is continuously evaluated. If the performance of the control policy model in the current course reaches a preset passing threshold δ passThen we will move on to the next, more difficult course. Train the model; if the model is in the current course The performance is lower than the preset failure threshold δ failed If the threshold is not reached within the maximum number of training epochs n, then the program will switch to the previous, simpler course. Training will be conducted. Furthermore, if the model can continuously upgrade to reach the review threshold e times, review sessions will be interspersed throughout the training. This design aims to ensure that the control strategy model gradually transitions to more difficult courses after mastering the skills at the current level, while consolidating learned knowledge through review to prevent forgetting.
[0129] Specifically, you can first set up a course set. For course sets Each course Determine the corresponding course difficulty coefficient V i Course difficulty level V i It can be composed of multiple dimensions, such as {v1, v2, ..., v m Then, based on the equation... Calculate course difficulty information C i It is the difficulty coefficient measurement function d(v) k The integral is calculated across all dimensions. Then, all courses can be sorted according to the course difficulty information C using the function D = Sort(D, C).
[0130] Then, the control policy model is trained starting with the simplest lesson, iterating through all the learning tasks in the sorted lessons. The control policy model is then used in the lessons during the training process. The model is trained and its parameters and reward score are adjusted accordingly. If the performance measurement function P(M,reward) is greater than or equal to the threshold δ, then... pass If you pass e courses consecutively, then you need to review the previous t courses. If you do not pass e courses consecutively, you will begin learning the next course.
[0131] If the control strategy model performs below the preset failure threshold δ in the current course. failed It will revert to the previous lesson and continue training until the upgrade conditions are met. If the pass threshold δ is not reached within the maximum number of training cycles n, it will be reset. pass Then we will switch to the previous, simpler course. Conduct training.
[0132] Finally, the control policy model outputs the adjusted first control policy model as the second control policy model after passing all the courses. This step-by-step increasing difficulty training method helps the model gradually master more complex skills and eventually achieve optimal performance. Through regular review and timely adjustment of training difficulty, this course learning strategy can effectively improve the generalization ability and adaptability of the control policy model.
[0133] After completing the reinforcement learning training process, the adjusted first control policy model can also be tested. The control policy model is tested on each course, and the evaluation result according to the performance measurement function P is compared with the passing threshold δ pass If the performance exceeds the threshold, the control policy model decides to enter the next more difficult course, or if the performance is below the threshold, it returns to the previous simpler course for additional training. This process continues until the model can successfully pass all environment tests, gradually improving its performance at different difficulty levels.
[0134] In addition, since the controller network has some ability after the pre-training stage, directly optimizing the mismatched controller network and value network may cause the untrained value network to fail to accurately evaluate the output of the controller network at the beginning of training or result in longer training time. Therefore, in the initial stage of reinforcement learning training, the parameters of the controller network will be frozen first, and the value network will be trained. Once this value network has sufficient evaluation ability, the parameters of both networks will be unfrozen and they will be trained simultaneously. The training process can be briefly described as follows.
[0135] Optionally, in the initial stage of reinforcement learning training, only the value network can be trained. The training of the value network includes: performing a plurality of action rounds for the first target task, in each time step in each action round: based on the scene image at the current time step in the process of simulating the execution of the first target task by the end effector and the pose corresponding to the contact point between the end effector and the first target tool, determining the control information of the end effector at the current time step by using the controller network in the first control policy model; and based on the control information of the end effector at the current time step, predicting a predicted value of the cumulative reward score of the control information of the end effector at the current time step; after completing each action round, calculating the predicted value of the cumulative reward score of the control information of the last time step in the action round as the predicted value of the reward corresponding to the action round; determining the true value of the reward score corresponding to the action round based on the execution result of the first target task in the action round; and adjusting the model parameters in the value network so that the difference between the predicted value and the predicted value of the reward corresponding to the action round converges. Of course, the present disclosure is not limited thereto.
[0136] Optionally, in a subsequent stage of reinforcement learning training, both the value network and the controller network can be jointly trained. The joint training of the value network and the controller network comprises: performing a plurality of action rounds for the second target task, in each time step in each action round: in each time step in each action round: based on the scene image at the current time step in the process of the end effector performing the learning task and the pose corresponding to the contact point between the end effector and the second target tool, determining, by using the controller network of the first control policy model in training, a predicted value of the control information of the end effector at the current time step; based at least in part on the predicted value of the control information of the end effector at the current time step, determining, by using the value network in training, a predicted value of the cumulative reward score for the control information of the end effector at the current time step; after completing each action round, calculating the predicted value of the cumulative reward score for the control information of the last time step in the action round as the predicted value of the reward score corresponding to the action round; determining the true value of the reward score corresponding to the action round based on the execution result of the learning task in the action round; adjusting the model parameters in the value network so that the difference between the predicted value and the true value of the reward score corresponding to the action round converges; and adjusting the model parameters of the controller network in the first control policy model so that the predicted value of the reward score corresponding to the action round increases.
[0137] Optionally, in step S302, reinforcement learning training can be performed in a simulation environment. Optionally, a depth image of the simulation environment is captured by using a built-in camera of MuJoCo and is input into the visual encoder trained in step S302. The pose of the virtual spoon is directly obtained from MuJoCo. The training environment for reinforcement learning is constructed by using OpenAI Gym. The observation space is defined as a 15-dimensional vector ranging in (-∞, +∞), and the action space is a 6-dimensional vector ranging in [-1, 1]. The control information of the end effector at the current time step corresponds to a fixed execution time, for example, 0.002 seconds, until an action round is completed. In each time step, the simulation environment receives a 6-dimensional vector a output , which represents the control information output by the controller network in training, for controlling the motion of the virtual spoon. Subsequently, the simulation environment returns the updated depth image and the latest pose information of the virtual contact point. Thus, this setting can accurately simulate the dynamic process of the scooping task, while providing a structured observation and action space for the reinforcement learning algorithm. By updating the simulation environment state and the pose information of the contact point in real time, continuous scooping actions can be effectively simulated, providing rich feedback information for adjusting the first control policy model.
[0138] Next, as Figure 5As shown, after step S302 is completed, the second control strategy model can be directly deployed on the real machine of the end effector for performing one or more target tasks, especially the first target task or the second target task.
[0139] At this time, the input of the second control strategy model can be a scene image captured in real time in the physical space, and the output of the second control strategy model can be control information in the physical space for controlling the end effector.
[0140] Optionally, the control information is pose information of a contact point between the end effector and the second target tool expected at a next time step. Optionally, after the pose information of the contact point between the end effector and the second target tool expected at the next time step is determined, control information of each motor of the end effector can be determined by an inverse kinematics method, which can be acceleration of each motor or torque of the motor. Although there is no big difference between the two physical quantities as control information for controlling the motor to rotate in a mathematical sense, in an actual physical system, not all of the two physical quantities can be accurately measured. Therefore, in the actual use of the end effector, a person skilled in the art can select a physical quantity with better data test effect and more in line with the model for subsequent calculation according to the specific circumstances. Of course, the present disclosure is not limited thereto.
[0141] For example, inverse kinematics is a method of solving the rotation angle and speed of each component of the end effector according to the expected pose information of the end effector (such as the expected pose information of the contact point between the end effector and the target tool). By using inverse kinematics, target action control of the end effector can be achieved.
[0142] The process of determining the control information required by the end effector to perform the second target task by the inverse kinematics method can be mathematically described by the following formula (1).
[0143] q = IK(p end -p base ) (1)
[0144] Where q is a vector of angle values of each joint of the end effector in the physical space, p end -p base is a position vector of the contact point between the end effector and the second target tool. Formula (1) establishes a mapping from the space corresponding to each joint to the Cartesian space. For example, by using a numerical method or an analytical method, the joint angle satisfying the target position and attitude can be solved based on the corresponding dynamics model of the end effector, and then the control information for controlling each joint is determined by using the joint angle information, so that the end effector can perform the target action.
[0145] Optionally, in the case of inverse kinematics mapping without analytical solution, a numerical approximation method can be adopted, given the position vector p of the end effector in the physical space end -p base , the joint angle solution q corresponding to the constraint is searched by an iterative algorithm. After obtaining the inverse kinematics solution, the kinematics solution q in the joint space corresponds to the joint control quantity that makes the end effector reach the target pose p end . Thus, the dynamics analysis can be continued to calculate the joint driving force / torque, complete the motion control, and obtain the control information in the physical space. Of course, the present disclosure is not limited thereto.
[0146] Thus, the embodiments of the present disclosure also provide an end effector, comprising: a central connecting member and a plurality of multi-joint operation assemblies; and a control strategy model carried on the end effector and determined by the method 30 according to the embodiments of the present disclosure. Wherein, the control strategy model can be the first control strategy model or the second control strategy model as described above, and the present disclosure is not limited thereto.
[0147] Optionally, the end effector is configured to perform a second target task, and the performing the second target task comprises: at a current time, acquiring a reference scene image when the end effector performs the second target task and a pose corresponding to a contact point between the end effector and a second target tool, generating control information for controlling the end effector at a current time step by using a second control strategy model carried on the end effector, determining motor driving information of the central connecting member and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step, and controlling the end effector based on the motor driving information of the central connecting member and the plurality of multi-joint operation assemblies.
[0148] Optionally, the end effector can also be configured to perform a first target task. At this time, although the first control strategy model can also be used to complete the task, generally speaking, the control effect of the second control strategy model will be better than that of the first control strategy model. Therefore, the process of performing the first target task can comprise: at a current time, acquiring a reference scene image when the end effector performs the first target task and a pose corresponding to a contact point between the end effector and a second target tool, generating control information for controlling the end effector at a current time step by using a second control strategy model carried on the end effector, determining motor driving information of the central connecting member and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step, and controlling the end effector based on the motor driving information of the central connecting member and the plurality of multi-joint operation assemblies.
[0149] Next, reference will be made toFigure 6 The method 60 for controlling an end effector according to embodiments of the present disclosure is further illustrated. Wherein, Figure 6 is a flow chart illustrating the method 60 for controlling an end effector according to embodiments of the present disclosure.
[0150] The method 60 for controlling an end effector according to embodiments of the present disclosure can include steps S601-S603 as shown. Figure 5 Of course, the method 60 can also include more or less operations, and the present disclosure is not limited thereto.
[0151] In step S601, at the current time, a reference scene image when the end effector performs a target task and a pose corresponding to a contact point between the end effector and a target tool are obtained.
[0152] In step S602, a control policy model carried on the end effector is used to generate control information for controlling the end effector at the current time step.
[0153] In step S603, based on the control information for controlling the end effector at the current time step, motor driving information of the central connecting member and the plurality of multi-joint operation assemblies is determined.
[0154] In step S604, based on the motor driving information of the central connecting member and the plurality of multi-joint operation assemblies, the end effector is controlled.
[0155] Wherein, the control policy model is the second control policy model described above, determined by the method described in reference Figure 3 to 4 .
[0156] Figure 7 The schematic diagram of the device 70 for controlling an end effector according to some embodiments of the present disclosure is schematically shown. The end effector includes a central connecting member and a plurality of multi-joint operation assemblies.
[0157] As Figure 7As shown, the device 70 includes a camera 71 configured to acquire a reference scene image of the end effector performing a target task at a current time; a sensor 72 configured to acquire a pose of a contact point between the end effector and a second target tool at the current time; a processor 73 configured to generate control information for controlling the end effector at a current time step by using a control policy model carried on the end effector, and determine motor driving information of the central connecting member and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step; and a motor 74 configured to control the end effector based on the motor driving information of the central connecting member and the plurality of multi-joint operation assemblies. The control policy model is determined by a method as described above. Figure 3 to 4 The method determines.
[0158] Figure 8 An exemplary block diagram of an apparatus 80 for determining a control policy model according to some embodiments of the present disclosure is schematically shown. The control policy model is used to control an end effector, which includes a central connecting member and a plurality of multi-joint operation assemblies.
[0159] Specifically, the apparatus 80 includes a first module 81 configured to acquire a first control policy model carried on the end effector, and a second module 82 configured to adjust the first control policy model carried on the end effector by using a plurality of action rounds to determine a second control policy model carried on the end effector, wherein in each action round of the plurality of action rounds, a learning task corresponding to the action round is determined based on information of the action round, a virtual body corresponding to the end effector simulates performing the learning task corresponding to the action round by using the first control policy model in the adjustment, and a reward score corresponding to completing the learning task corresponding to the action round and a reward score related to a physical constraint in a process of simulating performing the learning task corresponding to the action round are determined, and the first control policy model is adjusted so as to increase the reward score corresponding to completing the learning task corresponding to the action round.
[0160] Figure 9 An exemplary block diagram of an apparatus 90 for controlling an end effector according to some embodiments of the present disclosure is schematically shown. The end effector includes a central connecting member and a plurality of multi-joint operation assemblies. The apparatus 90 includes a first module 91, a second module 92, a third module 93, and a fourth module 94.
[0161] Specifically, the first module 91 is configured to acquire, at a current time, a reference scene image when the end effector performs a target task and a pose corresponding to a contact point between the end effector and a target tool. The second module 92 is configured to generate control information for controlling the end effector at a current time step by using a control strategy model carried on the end effector. The third module 93 is configured to determine motor driving information of the central connecting member and the plurality of multi-joint operation assemblies based on the control information for controlling the end effector at the current time step. The fourth module 94 is configured to control the end effector based on the motor driving information of the central connecting member and the plurality of multi-joint operation assemblies.
[0162] The control strategy model is the second control strategy model described above. Figure 3 to 4 The method determines.
[0163] It should be understood that the apparatuses 80 and 90 can be implemented in software, hardware, or a combination of software and hardware. Different modules can be implemented in the same software or hardware structure, or one module can be implemented by different software or hardware structures.
[0164] In addition, the apparatus 80 can be used to implement the method 30 described above, and the apparatus 90 can be used to implement the method 60 described above. Details related thereto have been described in detail above, and for the sake of brevity, will not be repeated here. The apparatuses 80 and 90 can have the same features and advantages as described with respect to the foregoing methods.
[0165] Figure 10 An example block diagram of a computing device 1900 according to some embodiments of the present disclosure is schematically illustrated. For example, it can represent a computing device that can be used to deploy the apparatus 80 and the apparatus 90 provided by the present disclosure.
[0166] As shown, the example computing device 1900 includes a processing system 1901, one or more computer readable medium 1902, and one or more I / O interface 1903, which are communicatively coupled to each other. Although not shown, the computing device 1900 can also include a system bus or other data and command transfer system that couples the various components to each other in communication. The system bus can include any one or combination of different bus structures, e.g., a memory bus or memory controller, a peripheral bus, a serial bus, a parallel bus, and / or a processor or local bus with various bus architectures that can be utilized in and / or by various embodiments of the present disclosure, or can also include control and data lines.
[0167] The processing system 1901 is representative of the functionality performed by hardware as an example. As such, the processing system 1901 is illustrated as including hardware elements 1904 that can be configured to perform a variety of functions as an example. These hardware elements 1904 can be collectively configured to perform one or more of the various functionalities as described herein. For example, the hardware elements 1904 can be configured to execute instructions stored in the computer-readable medium 1902 to perform a variety of functionalities as described herein. Accordingly, the computer-readable medium 1902 can be used in implementing the processes, algorithms, etc. described herein.
[0168] The computer-readable medium 1902 is illustrated as including memory / storage 1905. The memory / storage 1905 represents the memory / storage associated with one or more computer-readable media. The memory / storage 1905 can include volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, optical disks, magnetic disks, and so forth). The memory / storage 1905 can include fixed and removable storage devices. Exemplary, the memory / storage 1905 can be used to store various poses, location data, etc. mentioned in the embodiments above. The computer-readable medium 1902 can be configured in a variety of other ways as further described below.
[0169] The one or more input / output interfaces 1903 are representative of functionality to allow a user to enter commands and information to computing device 1900, and also allow information to be presented to the user and / or other components or devices using various input / output devices. Examples of input devices include a keyboard, cursor control device (e.g., a mouse), microphone (e.g., for voice inputs), a scanner, touch functionality (e.g., capacitive or other sensors that are configured to detect physical touch), a camera (e.g., which can employ visible or non-visible wavelengths such as infrared frequencies to detect movement that does not involve touch as gestures), a
[0170] The computing device 1900 also includes application program 1906. The application program 1906 can be stored in the memory / storage 1905, for example, as an example. The application program 1906 can implement various processes, algorithms, etc. as described herein, along with the processing system 1901, etc. Figure 7 or Figure 8 all of the functionality of the various modules of the apparatus 70 or 80 described.
[0171] Various techniques can be described in the general context of software, hardware, elements, or program modules. Generally, such modules include routines, programs, objects, elements, components, data structures, and the like that perform particular tasks or implement particular abstract data types. The terms "module," "functionality," and the like as used herein generally represent software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that the techniques can be implemented on a variety of computing platforms having a variety of processors.
[0172] Implementations of the described modules and techniques can be stored or transmitted across some form of computer-readable media. Computer-readable media can include various media that can be accessed by the computing device 1900. By way of example, and not limitation, computer-readable media can include "computer-readable storage media" and "computer-readable signal media."
[0173] In contrast to signal-bearing media, "computer-readable storage media" refers to media or means configured to store information in a form that is suitable for storage and / or access by a computing device, and / or transmission between computing devices. Thus, computer-readable storage media refers to non-signal bearing media. Computer-readable storage media includes volatile and non-volatile, removable and non-removable media implemented in a method or technology for storage and / or transmission of information such as computer-executable instructions, data structures, program modules, logical elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, hard disks, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture that are appropriate for storage and / or access of desired information and that can be accessed by a computing device.
[0174] "Computer-readable signal media" refers to a signal-bearing medium that is configured to transmit information that is accessible to the computing device 1900. The computer-readable signal media can be implemented in various forms, by various mechanisms. The computer-readable signal media includes, but is not limited to, a propagating signal, a carrier wave, a wave equation, and the like. The computer-readable signal media can also be referred to as a signal source or a communication medium. The signal bearing medium can include a computer-readable storage medium.
[0175] As previously described, hardware elements 1904 and computer-readable media 1902 are representative of instructions, modules, programmable device logic and / or fixed device logic implemented in a hardware form that can be employed in some embodiments to implement at least portions of the techniques described herein. Hardware elements can include components of an integrated circuit or
[0176] The foregoing combination of software and hardware modules also can be used to implement various techniques and modules described herein. Thus, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage media and / or by one or more hardware elements 1904. The computing device 1900 can be configured to implement particular instructions and / or functions corresponding to the software and / or hardware modules. Accordingly, implementation of a module as a software module and / or hardware module, either in whole or in part, can be determined as a matter of design choice by one skilled in the art. For example, software modules can be implemented as software stored in memory or storage devices and executed on general purpose computer hardware, but can also be implemented as executing in hardware elements 1904, e.g., as a central processing unit of a processing system 1901.
[0177] The technology described herein can be supported by various configurations of the computing device 1900 and is not limited to the specific examples of technology described herein.
[0178] It will be appreciated that, for clarity, the embodiments of the disclosure have been described hereinafter with reference to different functional elements. It will be apparent, however, that the functionality of each of the functional elements can be implemented in a single element, multiple elements or as part of other functional elements without detracting from the disclosure. For example, the functionality of an element can be split between elements, multiple elements or combined with functionality of other elements. Hence, the references to specific functional elements are only to be seen as references to suitable means for providing the described functionality rather than indicative of a strict logical or physical structure or organization. As such, the disclosure can be implemented in a single element, or can be physically and functionally distributed between different elements and circuitry.
[0179] The disclosure provides a computer-readable storage medium having stored thereon computer-executable instructions that, when executed, implement the method for determining a control policy model, the method for determining a tool control model, or the control method described above.
[0180] The present disclosure provides a computer program product or computer program comprising computer executable instructions stored in a computer readable storage medium. A processor of a computing device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions to cause the computing device to perform the method for determining a control policy model, the method for determining a tool control model, or the control method provided in various embodiments described above.
[0181] Variations to the disclosed embodiments can become apparent to those of ordinary skill in the art from the foregoing description and accompanying claims. In the claims, the word "comprising" does not exclude other elements or steps, and the word only "one" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage.
Claims
1. A method for determining a control strategy model, the control strategy model being used to control an end effector, the end effector including a central connector and multiple multi-joint operating components, the method comprising: Obtain the first control strategy model mounted on the end effector, and By utilizing multiple action cycles, the first control strategy model mounted on the end effector is adjusted to determine the second control strategy model mounted on the end effector. In each of the plurality of action rounds, Based on the information of the action rounds, the learning task corresponding to the action rounds is determined; The virtual body corresponding to the end effector uses the first control strategy model under adjustment to simulate the learning task corresponding to the action round, and determines the reward score corresponding to the completion of the learning task corresponding to the action round, as well as the reward score related to physical constraints during the simulation of the learning task corresponding to the action round; and The first control strategy model is adjusted to increase the reward score for completing the learning task corresponding to the action round and the reward score related to physical constraints during the simulation of the learning task corresponding to the action round. in, The first control strategy model is used to control the end effector to perform a first target task, wherein, during the process of the end effector performing the first target task, the end effector uses a first target tool to move a first target object from a first initial position to a first target position; The second control strategy model mounted on the end effector is used to control the end effector to perform a second target task. During the execution of the second target task, the end effector uses a second target tool to move the second target object from a second initial position to a second target position. The first target task and the second target task are different.
2. The method as described in claim 1, wherein, During the process of the virtual body corresponding to the end effector simulating the execution of the learning task corresponding to the action round, the end effector simulates the execution of: using the target tool specific to the action round to move the target object specific to the action round from the initial position specific to the action round to the target position corresponding to the action round.
3. The method as described in claim 2, wherein, The reward score for completing the learning task corresponding to the action round indicates a combination of at least one or more of the following: Does the end effector successfully move the target object specific to the action cycle from the initial position specific to the action cycle to the target position corresponding to the action cycle using the target tool specific to the action cycle? or During the process of the virtual body corresponding to the end effector simulating the learning task corresponding to the action round, whether the target object overflows.
4. The method of claim 2, wherein, The reward score related to physical constraints during the simulated execution of the learning task corresponding to the action round indicates a combination of at least one or more of the following: During the process of the virtual body corresponding to the end effector simulating the execution of the learning task corresponding to the action round, whether the target tool specific to the action round collides with the target object specific to the action round; During the process of the virtual body corresponding to the end effector simulating the execution of the learning task corresponding to the action round, whether the target tool specific to the action round collides with the container where the target object specific to the action round is located; During the process of the virtual body corresponding to the end effector simulating the execution of the learning task corresponding to the action round, whether the end effector is separated from the target tool; as well as During the process of the virtual body corresponding to the end effector simulating the learning task corresponding to the action round, whether the target tool is outside the observation range.
5. The method of claim 2, wherein, The step of determining the learning task corresponding to the action round based on the information of the action round includes: Based on the information of the action round, including the course difficulty information and course progress information corresponding to the action round, a specific target tool or a specific target object of the action round is determined.
6. The method of claim 5, wherein, The course difficulty information corresponding to the current action round is associated with the course progress information of the current action round. The course progress information indicates that the later the action round is executed, the higher the course difficulty information corresponding to the current action round.
7. The method of claim 6, wherein, The same course difficulty information corresponds to multiple action rounds, and different action rounds correspond to different learning tasks. The different learning tasks corresponding to the different action rounds satisfy at least one of the following conditions: Different action rounds have different specific target tools; Different action rounds target different specific objects; The number of specific target objects varies in different action rounds; or The specific target objects in different action rounds are located in different containers.
8. The method of claim 2, wherein, The step of adjusting the first control strategy model mounted on the end effector using multiple action cycles to determine the second control strategy model mounted on the end effector includes: In each of the plurality of action rounds, During the process of the end effector simulating the execution of the learning task, based on the scene image at the current time step and the pose information corresponding to the contact point between the end effector and the target tool specific to the action round, the control information of the end effector at the current time step is determined using the first control strategy model under adjustment. At least in part, a reward score is determined for the control information of the end effector at the current time step, based on the control information of the end effector at the current time step; and After completing the action rounds, the first control strategy model mounted on the end effector is adjusted, at least in part, based on the reward scores for the control information of the end effector at each time step.
9. The method of claim 8, wherein, The determination of the reward score for the control information of the end effector at the current time step includes: using a value network to determine the cumulative reward score for the control information of the end effector at the current time step. The training of the value network includes at least: executing multiple action rounds for the first target task, and at each time step in each action round: Based on the scene image at the current time step during the simulated execution of the first target task by the end effector, and the pose corresponding to the contact point between the end effector and the first target tool, the control information of the end effector at the current time step is determined using the controller network in the first control strategy model; and Based on the control information of the end effector at the current time step, predict the cumulative reward score of the control information of the end effector at the current time step; After completing each round of actions The predicted value of the cumulative reward score of the control information in the last time step of the action round is calculated as the predicted value of the reward corresponding to the action round; Based on the execution result of the first objective task corresponding to the action round, determine the true value of the reward score corresponding to the action round; and Adjust the model parameters in the value network so that the predicted value of the reward corresponding to the action round and the difference between the predicted values converge.
10. The method of claim 9, wherein, The step of adjusting the first control strategy model mounted on the end effector using multiple action cycles to determine the second control strategy model mounted on the end effector includes: In each time step of each action round: Based on the scene image at the current time step during the execution of the learning task by the end effector and the pose corresponding to the contact point between the end effector and the second target tool, the predicted value of the control information of the end effector at the current time step is determined using the controller network of the first control strategy model in training. Based at least in part on the predicted value of the control information of the end effector at the current time step, the predicted value of the cumulative reward score for the control information of the end effector at the current time step is determined using the value network in training. After completing each round of actions The predicted value of the cumulative reward score of the control information in the last time step of the action round is calculated as the predicted value of the reward score corresponding to the action round; Based on the execution results of the learning tasks corresponding to the action rounds, the true value of the reward score corresponding to the action rounds is determined; Adjust the model parameters in the value network so that the difference between the predicted and actual reward scores for the action rounds converges; and Adjust the model parameters of the controller network in the first control strategy model so that the predicted value of the reward score corresponding to the action round increases.
11. The method of claim 2, wherein, During the process of the first target object moving from the first initial position to the first target position, the first target tool applies a force to the first target object, and the end effector does not contact the first target object. During the process of the second target object moving from the second initial position to the second target position, the second target tool provides force to the second target object, and the end effector does not contact the second target object.
12. The method of claim 1, wherein, The first objective task and the second objective task do not satisfy at least one of the following conditions: The first target tool is different from the second target tool; The first target object is different from the second target object; The number of the first target objects is different from the number of the second target objects; The first initial position is different from the second initial position; The first target location is different from the second target location; or The first target object is located in a first container, and the second target object is located in a second container. The first container and the second container are different.
13. A method for controlling an end effector, the end effector comprising a central connector and a plurality of multi-joint operating components, including: At the current moment, acquire a reference scene image of the end effector when it is performing the target task, as well as the pose of the contact point between the end effector and the target tool; Using the control strategy model mounted on the end effector, control information for controlling the end effector at the current time step is generated; Based on the control information used to control the end effector at the current time step, determine the motor drive information of the central connector and the plurality of multi-joint operating components; as well as The end effector is controlled based on the motor drive information of the central connector and the plurality of multi-joint operating components; The control strategy model is determined by the method described in any one of claims 1-11.
14. An apparatus for controlling an end effector, the end effector including a central connector and a plurality of multi-joint operating components, the apparatus comprising: A camera is used to acquire a reference scene image at the current moment when the end effector is performing the target task; A sensor is used to determine the pose of the end effector at the point of contact with the target tool at the current moment. The processor is configured to use a control strategy model mounted on the end effector to generate control information for controlling the end effector at the current time step; and to determine motor drive information for the central connector and the plurality of multi-joint operating components based on the control information for controlling the end effector at the current time step. as well as A motor is used to control the end effector based on motor drive information from the central connector and the plurality of multi-joint operating components; The control strategy model is determined by the method described in any one of claims 1-11.
15. An apparatus for determining a control strategy model, the control strategy model being used to control an end effector, the end effector including a central connector and a plurality of multi-joint operating components, the apparatus comprising: The first module is used to obtain the first control strategy model mounted on the end effector; The second module is used to adjust the first control strategy model mounted on the end effector using multiple action cycles, in order to determine the second control strategy model mounted on the end effector. In each of the plurality of action rounds, Based on the information of the action rounds, the learning task corresponding to the action rounds is determined; The virtual body corresponding to the end effector uses the first control strategy model under adjustment to simulate the learning task corresponding to the action round, and determines the reward score corresponding to the completion of the learning task corresponding to the action round, as well as the reward score related to physical constraints during the simulation of the learning task corresponding to the action round; and The first control strategy model is adjusted to increase the reward score for completing the learning task corresponding to the action round. in, The first control strategy model is used to control the end effector to perform a first target task, wherein, during the process of the end effector performing the first target task, the end effector uses a first target tool to move a first target object from a first initial position to a first target position; The second control strategy model mounted on the end effector is used to control the end effector to perform a second target task. During the execution of the second target task, the end effector uses a second target tool to move the second target object from a second initial position to a second target position. The first target task and the second target task are different.
16. An apparatus for controlling an end effector, the end effector comprising a central connector and a plurality of multi-joint operating components, including: The first module is used to acquire, at the current moment, a reference scene image of the end effector performing the target task and the pose corresponding to the contact point between the end effector and the target tool; The second module is used to generate control information for controlling the end effector at the current time step by utilizing the control strategy model mounted on the end effector; The third module is used to determine the motor drive information of the central connector and the plurality of multi-joint operating components based on the control information used to control the end effector at the current time step. as well as The fourth module is used to control the end effector based on the motor drive information of the central connector and the plurality of multi-joint operating components; The control strategy model is determined by the method described in any one of claims 1-11.
17. An electronic device comprising: processor; and A memory, wherein the memory stores computer-executable code, which, when run by the processor, performs the method of any one of claims 1-13.
18. A non-volatile computer-readable storage medium having executable code stored thereon, the executable code, when executed by a processor, causing the processor to perform the method of any one of claims 1-13.
19. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1 to 13.
Citation Information
Patent Citations
Efficient adaption of robot control policy for new task using meta-learning based on meta-imitation learning and meta-reinforcement learning
CN113677485A
Dynamic interactive representation-based dexterous manipulator grabbing method
CN117798919A
Cited By
Manipulation method learning apparatus, manipulation method learning system, manipulation method learning method, and program
US12617083B2