Method and apparatus for determining control policy model, and method and apparatus for controlling end effector
By adjusting the control strategy model on the end effector and combining reinforcement learning and reward score optimization in a virtual environment, the adaptability and control accuracy issues of the end effector in complex tool operation tasks are solved, achieving efficient generalization and robustness improvement.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-08-08
- Publication Date
- 2026-04-23
AI Technical Summary
Existing technologies struggle to effectively address the adaptability and control accuracy issues of end effectors in complex tool operation tasks, especially in improving generalization and anti-interference capabilities with reduced training data.
By acquiring the first control policy model on the end effector and using reinforcement learning across multiple action rounds, the model is adjusted to determine the second control policy model. The control policy is then optimized by incorporating reward scores and physical constraints in the virtual environment to improve generalization and robustness.
It significantly reduces the need for training data, improves the adaptability and flexibility of the end effector in diverse tasks, and enhances control accuracy and anti-interference capabilities.
Smart Images

Figure CN2025113520_23042026_PF_FP_ABST
Abstract
Description
Methods and apparatus for determining control strategy models, and methods and apparatus for controlling end effectors.
[0001] Related applications
[0002] This application claims priority to Chinese patent application filed on October 17, 2024, application number 202411457051.7, entitled “Method and apparatus for determining a control strategy model and method and apparatus for controlling an end effector”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of intelligent robot technology, specifically to a method for determining a control strategy model, a method for controlling an end effector, a device for controlling an end effector, an apparatus for determining a control strategy model, an apparatus for controlling an end effector, an electronic device, a non-volatile computer-readable storage medium, and a computer program product. Background Technology
[0004] Tool manipulation using end effectors is a hot topic in robotics because it involves complex interactions between multiple objects, including the end effector, the tool, the object, and the operating environment. For example, robots need to use specific tools to grasp and move objects, such as removing or placing objects from a container. The complexity of this process increases significantly with the type of tool, the characteristics of the object, and the environmental conditions. Industry is attempting to develop end effectors that can adapt to these variations and operate tools effectively, but no truly effective solution has yet been proposed.
[0005] To improve the efficiency and adaptability of robot operations, end effectors equipped with artificial intelligence models have been proposed to perform the unique and challenging task of tool manipulation. However, such solutions often require large amounts of training data. Industry is exploring ways to improve the generalization ability of neural network models used in end effectors while reducing the amount of training data. Furthermore, improving the control precision of end effectors, increasing the number of operable objects, accelerating operation speed, and enhancing anti-interference capabilities are all pressing issues that need to be addressed. Solving these problems will enable robot end effectors to better adapt to diverse task scenarios, thereby playing a greater role in practical applications. Summary of the Invention
[0006] To address the above problems, this disclosure provides a method for determining a control strategy model, a method for controlling an end effector, a device for controlling an end effector, an apparatus for determining a control strategy model, an apparatus for controlling an end effector, an electronic device, a non-volatile computer-readable storage medium, and a computer program product.
[0007] According to one aspect of this disclosure, a method for determining a control strategy model is proposed. The control strategy model is used to control an end effector, which includes a central connector and multiple multi-joint operating components. The method includes: acquiring a first control strategy model mounted on the end effector; and adjusting the first control strategy model mounted on the end effector using multiple action rounds to determine a second control strategy model mounted on the end effector. In each of the multiple action rounds, a learning task corresponding to the action round is determined based on information from the action round. A virtual entity corresponding to the end effector uses the adjusted first control strategy model to virtually execute the learning task corresponding to the action round, and determines a reward score for completing the learning task and a reward score related to physical constraints during the virtual execution of the learning task. The method also involves adjusting the first control strategy model to increase the reward score for completing the learning task and the reward score related to physical constraints during the virtual execution of the learning task.
[0008] According to another aspect of this disclosure, a method for controlling an end effector is proposed, the end effector including a central connector and multiple joint operating components, comprising: at a current moment, acquiring a reference scene image of the end effector performing a task and the pose corresponding to the contact point between the end effector and a second task tool; generating control information for controlling the end effector at the current time step using a control strategy model mounted on the end effector; determining motor drive information of the central connector and the multiple joint operating components based on the control information for controlling the end effector at the current time step; and controlling the end effector based on the motor drive information of the central connector and the multiple joint operating components; wherein the control strategy model is determined by the above method.
[0009] According to another aspect of this disclosure, a device for controlling an end effector is proposed. The end effector includes a central connector and multiple joint operating components. The device includes: a camera for acquiring a reference scene image of the end effector performing a task at a current time; a sensor for determining the pose of the end effector corresponding to the contact point between the end effector and the task tool at the current time; a processor for generating control information for controlling the end effector at a current time step using a control strategy model mounted on the end effector; and determining motor drive information of the central connector and the multiple joint operating components based on the control information for controlling the end effector at the current time step; and a motor for controlling the end effector based on the motor drive information of the central connector and the multiple joint operating components; wherein the control strategy model is determined by the above method.
[0010] According to another aspect of this disclosure, an apparatus for determining a control strategy model is provided. The control strategy model is used to control an end effector, which includes a central connector and multiple multi-joint operating components. The apparatus includes: a first module for acquiring a first control strategy model mounted on the end effector; a second module for adjusting the first control strategy model mounted on the end effector using multiple action rounds to determine a second control strategy model mounted on the end effector, wherein in each of the multiple action rounds, a learning task corresponding to the action round is determined based on information from the action round; a virtual entity corresponding to the end effector uses the adjusted first control strategy model to virtually execute the learning task corresponding to the action round, and determines a reward score corresponding to completing the learning task and a reward score related to physical constraints during the virtual execution of the learning task; and adjusts the first control strategy model such that the reward score for completing the learning task corresponding to the action round increases.
[0011] According to another aspect of this disclosure, an apparatus for controlling an end effector is proposed. The end effector includes a central connector and multiple joint operating components, comprising: a first module for acquiring, at a current moment, a reference scene image of the end effector performing a task and the pose corresponding to the contact point between the end effector and the task tool; a second module for generating control information for controlling the end effector at the current time step using a control strategy model mounted on the end effector; a third module for determining motor drive information of the central connector and the multiple joint operating components based on the control information for controlling the end effector at the current time step; and a fourth module for controlling the end effector based on the motor drive information of the central connector and the multiple joint operating components; wherein the control strategy model is determined by the method described above.
[0012] According to another aspect of this disclosure, an electronic device is proposed, comprising: a processor; and a memory, wherein computer-executable code is stored in the memory, the computer-executable code performing the above-described method when run by the processor.
[0013] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is proposed, on which executable code is stored, which, when executed by a processor, causes the processor to perform the above-described method.
[0014] According to another aspect of this disclosure, a computer program product is proposed, comprising computer executable instructions that, when executed by a processor, implement the method described above.
[0015] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.
[0017] Figure 1 schematically illustrates an example task scenario that can be applied to embodiments of the present disclosure.
[0018] Figure 2 is a schematic diagram illustrating an end effector according to an embodiment of the present disclosure.
[0019] Figure 3 is a flowchart illustrating a method for determining a control strategy model according to an embodiment of the present disclosure.
[0020] Figure 4 is a schematic diagram illustrating a method for determining a control strategy model according to an embodiment of the present disclosure.
[0021] Figure 5 is a schematic diagram illustrating a method for controlling an end effector according to an embodiment of the present disclosure.
[0022] Figure 6 is a flowchart illustrating a method for controlling an end effector according to an embodiment of the present disclosure.
[0023] Figure 7 schematically illustrates a device for controlling an end effector according to some embodiments of the present disclosure.
[0024] Figure 8 schematically illustrates an exemplary block diagram of an apparatus for determining a control strategy model according to some embodiments of the present disclosure.
[0025] Figure 9 schematically illustrates an exemplary block diagram of a device for controlling an end effector according to some embodiments of the present disclosure.
[0026] Figure 10 schematically illustrates an example block diagram of a computing device according to some embodiments of the present disclosure. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0028] As shown in this disclosure and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0029] While this disclosure makes various references to certain modules in the apparatus according to embodiments of this disclosure, any number of different modules may be used and run on user terminals and / or servers. The modules are merely illustrative, and different aspects of the apparatus and methods may use different modules.
[0030] Flowcharts are used in this disclosure to illustrate the operations performed by the methods and apparatus according to embodiments of this disclosure. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0031] To facilitate the description of this disclosure, the following concepts related to this disclosure are introduced.
[0032] The various embodiments disclosed herein relate to the design and control of a "dexterous hand" in the field of intelligent robotics. A dexterous hand is used to simulate the complex movements and functions of a human hand, aiming to endow robots with dexterity and manipulative capabilities similar to a human hand. Through sophisticated mechanical structures and control systems, the "dexterous hand" is capable of performing tasks such as grasping, carrying, and manipulating tools. Key technologies of the "dexterous hand" include multi-degree-of-freedom finger design, integrated sensor systems, and advanced control algorithms, enabling the "dexterous hand" to perform precise operations in complex environments.
[0033] The mechanical structure of a "dexterous hand" typically consists of fingers, a palm, and connectors with multiple degrees of freedom. Each finger contains multiple joints, which are controlled by servo motors to achieve precise position control and dynamic response. To enhance the dexterous hand's sensing capabilities, it integrates various sensors, such as IMU sensors and joint angle encoders. These sensors provide real-time data on the actuator's acceleration, posture, joint angles, and angular velocities to achieve precise motion control.
[0034] With technological advancements, the applications of "dexterous hands" are constantly expanding. Beyond precision assembly and quality inspection in industrial automation, dexterous hands are demonstrating broad application potential in medical rehabilitation, service robotics, and education and research. For example, in medical rehabilitation, dexterous hands can assist in surgical procedures or be used to develop highly realistic prostheses, helping people with disabilities regain hand function. In the field of service robotics, dexterous hands enable robots to perform more complex household tasks, such as tidying clothes and preparing food. Furthermore, with advancements in artificial intelligence and machine learning technologies, the control strategies of dexterous hands are continuously being optimized to adapt to increasingly complex and dynamic task requirements.
[0035] The solutions provided in this application mainly involve artificial intelligence technology and apply it to the control field of robot end effectors, as illustrated in the following embodiments.
[0036] In summary, the solutions provided by the embodiments of this disclosure involve technologies such as artificial intelligence and machine learning. The embodiments of this disclosure will be further described below with reference to the accompanying drawings.
[0037] Figure 1 schematically illustrates an example scenario 100 to which some embodiments of the present disclosure may be applied. As shown in Figure 1, scenario 100 includes a robot 110 having an end effector 111, which can be various types of execution structures, such as a dexterous hand, an end gripper, etc. Furthermore, scenario 100 also includes a tool 120 and a container 130 containing a plurality of objects 131. Exemplarily, a control strategy model or tool control model determined by the scheme provided according to some embodiments of the present disclosure can be deployed in the robot 110, for example, in a controller network or processor structure for decision-making or control functions within the robot 110, so as to control the end effector 111 of the robot 110 to use the tool 120 to move at least a portion of the plurality of objects 131 in the container 130 to a target location, such as to a specified height above the container 130.
[0038] In Figure 1, robot 110 is shown as a humanoid robot, tool 120 is shown as a spoon-shaped tool, container 130 is shown as a bowl-shaped container, and object 131 is shown as a spherical object. However, depending on the specific application scenario, robot 110 can also be other types of robots, such as a robotic arm, tool 120 can also be other forms of tools, such as other forms of digging tools or other types of tools such as gripping tools, container 130 can also be other shapes or types of containers, and object 131 can also be other shapes and sizes of objects. Furthermore, optionally, container 130 may not be present, meaning that multiple objects 131 can be directly placed at their respective locations. Also optionally, the robot's end effector can be controlled to move at least a portion of the multiple objects 131 into another container.
[0039] Figure 2 is a schematic diagram illustrating an end effector according to an embodiment of the present disclosure.
[0040] Figure 2 further illustrates the end effector 111 using a "dexterous hand" as an example. The "dexterous hand" (i.e., end effector 111) in Figure 2 consists of a central connector and multiple multi-joint manipulation components. Each finger is composed of multiple segments, each equipped with at least one active flexion joint. These joints are typically controlled by servo motors, which provide precise positional control and dynamic response, enabling the dexterous hand to perform fine manipulations. The joint design allows for a wide range of movements, including extension and flexion, thus enabling diverse grip patterns and manipulation capabilities. Furthermore, the mechanical structure of the dexterous hand incorporates lightweight yet robust materials, such as aluminum alloys or carbon fiber, to ensure strength and durability while reducing weight and improving operational flexibility and responsiveness.
[0041] To enhance the dexterous hand's perception and feedback capabilities, multiple sensors are integrated into its central connector. These include an inertial measurement unit (IMU) sensor, which integrates accelerometers and gyroscopes to measure the linear acceleration and angular velocity of the actuator, providing data on the end effector's attitude and motion. This data is helpful for performing precise motion control, especially when performing complex operations or working in dynamic environments. Additionally, each joint is optionally equipped with a joint angle encoder to provide feedback on joint position and velocity, further enabling precise motion control. The integration of these sensors allows the dexterous hand to perform adaptive control in response to unpredictable external disturbances and changes.
[0042] The control system of a "dexterous hand" coordinates the movement of motors to accomplish complex tasks. Through advanced algorithms, the "dexterous hand" can simulate the movements of a human hand, such as grasping, moving objects, and using tools. This capability makes dexterous hands promising for applications in automated production lines, robotic surgery, service robots, and scientific research. The control system typically includes a real-time operating system and advanced control algorithms, such as PID control, model predictive control, or machine learning algorithms, to achieve highly precise and adaptable operation.
[0043] Furthermore, the "dexterous hand" in Figure 2 can be used to simulate the human hand in using / manipulating tools. In daily life, the use of tools is crucial for extending human physical capabilities. For example, tools such as spoons and screwdrivers enable people to perform tasks that would be difficult to complete without tools, such as scooping particles or liquids or tightening screws. The "dexterous hand" is designed to mimic these basic human movements, and through integrated sensors and a sophisticated control system, it can manipulate tools with a high degree of dexterity and precision.
[0044] However, end effectors often struggle to manipulate tools to control objects. Specifically, the interactions between the end effector, the tool, and the task object or environment are highly complex. This complex interaction means that methods designed for end effectors to directly manipulate objects cannot be simply applied to tasks that use tools to manipulate objects (i.e., tool-operated tasks). Tool-operated tasks require end effectors to understand and adapt to the dynamics of the tool and its interaction with the object, necessitating improvements and adjustments to the overall control strategy of the end effector.
[0045] Furthermore, end effectors typically do not directly contact objects or the environment during tool manipulation, leading to insufficient observation and feedback on the contact state between the tool and the object. This lack of direct feedback makes effective feedback control more difficult. Tool manipulation tasks often involve multiple task objects and complex physical constraints, making it difficult to deploy control methods based on traditional dynamic models. Therefore, to achieve effective tool manipulation in end effectors, improved control methods need to be developed to adapt to the complex relationships between the end effector, the tool, and the object or environment, and to address the challenges arising from the lack of direct contact and the involvement of multiple objects and physical constraints.
[0046] Currently, data-driven approaches have become a research hotspot to overcome the challenges of tool operation. Imitation learning methods have shown potential in this regard. Imitation learning schemes guide robot learning by analyzing the behavior of human operators (i.e., human demonstrations). However, imitation learning schemes typically rely on humans directly operating or remotely controlling the robot's end effector during data collection, which is not only time-consuming but also inefficient.
[0047] To reduce reliance on human demonstrations, reinforcement learning offers an alternative, allowing robots to autonomously learn skills through interaction with their environment with little or no direct human guidance. While reinforcement learning has achieved some success, the diversity and complexity of tool manipulation tasks make training challenging. For example, in a scooping task, the robot's motion strategy needs to be adjusted based on factors such as the shape of the spoon and bowl, the quantity and material of the target contents. Training a general policy adaptable to various scenarios often requires a large amount of training data to ensure that the neural network model on the end effector can learn effectively from the training data.
[0048] Therefore, to address the aforementioned problems, according to one aspect of this disclosure, a method for determining a control strategy model is proposed. The control strategy model is used to control an end effector, which includes a central connector and multiple multi-joint operating components. The method includes: acquiring a first control strategy model mounted on the end effector; and adjusting the first control strategy model mounted on the end effector using multiple action rounds to determine a second control strategy model mounted on the end effector. In each of the multiple action rounds, a learning task corresponding to the action round is determined based on information from the action round. A virtual entity corresponding to the end effector uses the adjusted first control strategy model to virtually execute the learning task corresponding to the action round, and determines a reward score for completing the learning task and a reward score related to physical constraints during the virtual execution of the learning task. The method also involves adjusting the first control strategy model to increase the reward score for completing the learning task and the reward score related to physical constraints during the virtual execution of the learning task.
[0049] According to one aspect of this disclosure, a method for controlling an end effector is also proposed. The end effector includes a central connector and multiple joint operating components. The method includes: acquiring, at a current moment, a reference scene image of the end effector performing a task and the pose corresponding to the contact point between the end effector and a second task tool; generating control information for controlling the end effector at the current time step using a control strategy model mounted on the end effector; determining motor drive information of the central connector and the multiple joint operating components based on the control information for controlling the end effector at the current time step; and controlling the end effector based on the motor drive information of the central connector and the multiple joint operating components; wherein the control strategy model is determined by the above method.
[0050] Optionally, based on the control information used to control the end effector at the current time step, the motor drive information of the central connector and the plurality of multi-joint operating components is determined by inverse kinematics. The inverse kinematics method calculates the rotation angle and speed of each joint of the end effector based on the expected pose information of the contact point between the end effector and the task tool at the next time step, thereby determining the control information of each motor. The specific formula is q = IK(p,θ), where q is the angle value vector of each joint, p is the position vector of the contact point between the end effector and the task tool, θ is the attitude information of the end effector, and IK represents the inverse kinematics solution function.
[0051] To obtain a second control strategy model with higher generalization and robustness than the first control strategy model, this disclosure further adjusts the first control strategy model using a reinforcement learning scheme to obtain the second control strategy model. In the adjustment process, specific reward scores are used, namely, "reward scores corresponding to the completion of the learning task corresponding to the action round" and "reward scores related to physical constraints during the virtual execution of the learning task corresponding to the action round". This ensures that the control strategy model explores control schemes as widely as possible, while avoiding the control strategy model trained in the virtual environment from not conforming to physical constraints, thus preventing it from being applied in the physical environment.
[0052] The method for determining the control policy model proposed in this disclosure is particularly suitable for specific tool operation tasks. Compared to traditional methods for determining the training scheme of the control policy model suitable for specific tool operation tasks, this disclosure uses a reinforcement learning approach to adjust the first control policy model based on a pre-trained first control policy model. This reduces the need for human demonstration data and also reduces the training time required for reinforcement learning.
[0053] Specifically, the method provided in this disclosure only requires collecting training data for training the first control policy model, significantly reducing the amount of data required for training compared to traditional methods. Furthermore, since the method also adjusts the first control policy model to obtain a second control policy model to adapt to a second task different from the first task, the generalization ability of the control policy model is significantly improved compared to traditional methods. Additionally, because reinforcement learning is performed based on the first control policy model, the control policy model does not need to engage in excessive ineffective exploration during this process, greatly reducing the time required for the reinforcement learning process.
[0054] In some embodiments of this disclosure, imitation learning and reinforcement learning can be combined to enable the neural network model on the end effector to learn by imitation from limited human demonstrations and acquire the ability to perform tool operation tasks through reinforcement learning based on this imitation learning. Therefore, this disclosure significantly reduces the need to collect human demonstration data, allowing the neural network model on the end effector to quickly learn by imitation with minimal training data and effectively operate in unseen scenarios using reinforcement learning. Through embodiments of this disclosure, the neural network model on the end effector can extract key information from limited training data and generalize it to new, unknown environments. Embodiments of this disclosure not only improve learning efficiency but also enhance the robot's adaptability and flexibility in diverse tasks.
[0055] Next, the various embodiments of this disclosure will be described in general with reference to FIG3.
[0056] Figure 3 is a flowchart illustrating a method 30 for determining a control strategy model according to an embodiment of the present disclosure.
[0057] Method 30 can be used to determine a control strategy model (e.g., a first control strategy model and a second control strategy model detailed below). This control strategy model is used to control the end effector to perform a task (e.g., a first task and a second task detailed below) using a task tool (e.g., a first task tool and a second task tool detailed below). The task may include moving at least one task object (e.g., a first task object and a second task object detailed below) from an initial position (e.g., a first initial position and a second initial position detailed below) to a target position (e.g., a first target position and a second target position detailed below) using the task tool (e.g., a first task tool and a second task tool detailed below). This disclosure is not limited thereto.
[0058] Exemplarily, method 30 can be used to control the end effector 111 of the robot 110 shown in Figures 1 and 2, so as to control the end effector 111 of the robot 110 to use tool 120 to move object 131 in container 130. Optionally, method 30 can be implemented in a real machine environment or in a virtual environment. Optionally, the task tool can be any type of tool, such as a digging tool, gripping tool, sticking tool, suction tool, fork tool, etc., and the task object can be an object of any shape, such as a spherical object, a cube-shaped object, an irregularly shaped object, etc. Of course, this disclosure is not limited thereto.
[0059] The initial position should be understood as the location of multiple task objects before the task begins execution. For example, multiple task objects can be placed in a container at the initial position, or they can be scattered or piled up at the initial position. Similarly, the target position should be understood as the location where the task objects are expected to be moved. For example, it may be expected that at least one of the multiple task objects will be moved to a container at the target position, or it may be expected that only they will be moved or placed at the target position. It should be understood that the initial position or target position can refer to a range of positions, and not necessarily a precise point. Of course, this disclosure is not limited to this.
[0060] The method 30 for determining a control policy model according to embodiments of the present disclosure may include steps S301 to S302 as shown in FIG3. Of course, method 30 may also include more or fewer operations, and the present disclosure is not limited thereto. Optionally, step S301 may be an operation in the pre-training phase. Operation S302 may be an operation in the reinforcement learning phase.
[0061] As described above, the end effector includes a central connector and multiple joint-operated components (e.g., a mechanical finger). The central connector and multiple joint-operated components can be described with reference to the relevant descriptions in Figures 1 and 2, and will not be repeated here.
[0062] In step S301, the first control strategy model mounted on the end effector is obtained.
[0063] Optionally, the first control strategy model is used to control the end effector to perform a first task, wherein, during the execution of the first task, the end effector uses a first task tool to move a first task object from the first initial position to the first target position. Optionally, during the movement of the first task object from the first initial position to the first target position, the first task tool provides a force to the first task object, and the end effector does not contact the first task object. This disclosure is not limited thereto.
[0064] Optionally, the first control strategy model can be a neural network model with a preset network structure, which can be constructed based on existing neural network structures in related technologies, or it can be any new network structure. Furthermore, the first control strategy model can also be other types of models. This disclosure does not specifically limit the specific model type and structure of the first control strategy model.
[0065] Optionally, the first control strategy model is a pre-trained neural network model. The first control strategy model can be pre-trained in various ways, including but not limited to transfer learning, imitation learning, and self-supervised learning. The following uses imitation learning as an example to briefly describe the process of "obtaining the first control strategy model mounted on the end effector," and those skilled in the art should understand that this disclosure is not limited thereto.
[0066] Optionally, step S301 includes: acquiring a reference scene image of the end effector performing the first task and a reference trajectory corresponding to the contact point between the end effector and the first task tool; and determining a first control strategy model mounted on the end effector based on the reference scene image of the end effector performing the first task and the reference trajectory corresponding to the contact point between the end effector and the first task tool, wherein the first control strategy model is used to control the end effector to perform the first task. This disclosure is not limited thereto.
[0067] Specifically, reference scene images and reference trajectories corresponding to the contact points between the end effector and the first task tool can be collected in physical space for training the end effector to perform the first task. Then, in a virtual environment, the data collected in the physical space is projected, and the projected data in the virtual environment is amplified. Next, in the virtual environment, using the reference scene images and reference trajectories corresponding to the contact points between the end effector and the first task tool as training data, the model parameters of the first control strategy model are continuously adjusted until a predetermined number of iterations is reached or the model parameters converge. Thus, the first control strategy model attempts to enable the end effector to move autonomously via its own motor without human operator control, allowing the end effector to control the first task tool to move approximately along the reference trajectory to move the first task object from a first initial position to a first target position or its vicinity.
[0068] Optionally, the reference scene images during the execution of the first task by the end effector are not limited to a single image, but may include multiple images, or even a video. These reference scene images may depict the scene at different moments during the execution of the first task by the end effector, providing more comprehensive information to help the end effector better understand and identify the progress of the task. For example, regarding the scene of using a spoon to scoop up particles in Figure 1, the reference scene images at different moments present the different relative positions and states of the bowl and spoon, and these images continuously demonstrate the progress of the task.
[0069] Optionally, the reference scene image when the end effector performs the first task is a depth image captured by a depth camera. The depth camera calculates the distance to objects and generates a depth image by emitting infrared light and measuring the time it takes for the reflected light to return. Optionally, the reference scene image includes three-dimensional information of the scene in which the end effector performs the first task. Thus, the first strategy model mounted on the end effector can perform accurate spatial localization and object recognition based on the reference scene image, thereby enabling the robot to better understand its environment and determine the task to be performed. Of course, this disclosure is not limited thereto.
[0070] Optionally, the reference trajectory corresponding to the contact point between the end effector and the first task tool is the motion trajectory of the contact point when the end effector performs the first task. The reference trajectory may include at least one of the following: attitude information, position information, velocity information, acceleration information, angular velocity information, and angular acceleration information of the contact point at each time step. Optionally, the reference trajectory can be acquired by sensors on the end effector. The position information can be three-dimensional coordinates in a specified coordinate system, such as the three-dimensional coordinates of the contact point, and the attitude information can be described by quaternions. Alternatively, the pose of the task tool can be described in other ways, or the number of parameters can be adjusted according to the actual allowed degrees of freedom. Optionally, the change in position or attitude can be described in the same way as the pose of the task tool. Of course, this disclosure is not limited to this.
[0071] In one example, the reference trajectory can be represented as a temporal numerical sequence consisting of the position information of the contact points at each time step. Optionally, each element in the temporal numerical sequence can have multiple dimensions, representing the position of the contact point in the x-axis direction, the position of the contact point in the y-axis direction, the position of the contact point in the z-axis direction (gravity direction), and attitude information represented by quaternions at a certain time step. Of course, the reference trajectory can also be represented by other data structures, and this disclosure is not limited thereto.
[0072] In one example, a reference scene image and reference trajectory can be acquired as follows: In a real physical space, a human operator controls the end effector to operate a first task tool to perform a first task. During this process, the human operator directly or indirectly controls the end effector to interact with the selected first task tool (such as a spoon, screwdriver, pliers, etc.) to complete a specific task (such as scooping up particles or liquid, tightening screws, clamping objects, etc.). A depth camera deployed in the real physical space and sensors deployed on the end effector can be used to capture, in real time, images of the reference scene and the reference trajectory corresponding to the contact point between the end effector and the first task tool. Of course, this disclosure is not limited to this.
[0073] In another example, reference scene images and reference trajectories can also be obtained in the following way: In a virtual environment, a human operator creates an animation of controlling a virtual end effector to operate a virtual first task tool to perform a first task. In this animation, the end effector is controlled to interact with the virtual first task tool (such as a virtual spoon, virtual screwdriver, virtual pliers, etc.) to complete a specific task (such as scooping virtual particles or liquid, tightening virtual screws, clamping virtual objects, etc.). Reference scene images and reference trajectories corresponding to the contact points between the end effector and the first task tool can be captured in real time using a viewpoint camera deployed in the virtual environment. Of course, this disclosure is not limited to this.
[0074] Optionally, determining the first control strategy model mounted on the end effector is equivalent to training the first control strategy model. During this training process, the reference scene image when the end effector performs the first task and the reference trajectory corresponding to the contact point between the end effector and the first task tool are used as training data. The model parameters of the first control strategy model are continuously adjusted until a predetermined number of iterations are reached or the model parameters converge. Thus, the first control strategy model attempts to enable the end effector to move autonomously via its own motor without human operator control, allowing the end effector to control the first task tool to move approximately along the reference trajectory to move the first task object from a first initial position to a first target position or its vicinity.
[0075] Optionally, as described above, the reference scene image and reference trajectory can be acquired either in physical space or in a virtual environment. The reference scene image and reference trajectory provide reference information for the first control strategy model, enabling it to progressively adjust the control information used to control the end effector. The control information for controlling the end effector indicates the action the end effector should perform at the current time step. By performing this action, the end effector can ensure that the contact point between the end effector and the first task tool approximately matches the desired pose of the reference trajectory at the next time step.
[0076] While the training process of the first control policy model can employ imitation learning to ensure that it learns a coarse control policy capable of achieving the first task from limited training data, the inherent limitations of imitation learning mean that changes in the real-world scenario or any parameter in the task (e.g., changes in at least one of the task tool, task object, initial position of the task object, or target position) will lead to poor performance of the first control policy model trained through imitation learning. Therefore, step S302 will adjust (or fine-tune) the first control policy model to obtain a control policy model (i.e., the second control policy model) with higher generalization and robustness.
[0077] In step S302, the first control strategy model mounted on the end effector is adjusted using multiple action rounds to determine the second control strategy model mounted on the end effector. Specifically, in each of the multiple action rounds, a learning task corresponding to that action round is determined based on the information of that action round; the virtual entity corresponding to the end effector uses the adjusted first control strategy model to virtually execute the learning task corresponding to that action round, and determines the reward score for completing the learning task and the reward score related to physical constraints during the virtual execution of the learning task; and the first control strategy model is adjusted to increase the reward score for completing the learning task and the reward score related to physical constraints during the virtual execution of the learning task.
[0078] Physical constraints refer to the physical rules and limitations that the virtual entity corresponding to the end effector must follow when performing learning tasks in a virtual environment. These constraints ensure that the virtual entity's actions conform to real-world physical laws, avoiding unrealistic operations such as collisions between the task tool and the task object or container, separation of the end effector from the task tool, or the task tool going out of the observation range. In reinforcement learning, by setting reward scores related to physical constraints, the control policy model is guided to learn control policies that conform to physical constraints, thus enabling the trained model to be applied in real-world physical environments.
[0079] Reward scores are quantitative metrics used in reinforcement learning to evaluate the performance and effectiveness of an end effector when performing a learning task. They include "the reward score for completing the learning task corresponding to the action round" and "the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round." The former indicates whether the end effector successfully completes the learning task, such as whether the round-adapted task object is moved from its initial position to the target position, or whether the task object overflows; the latter is used to constrain the actions of the virtual object, ensuring they conform to physical constraints, such as whether the task tool collides with the task object or container, or whether the end effector separates from the task tool. The control policy model learns and adjusts through reward scores to increase them, thereby improving the model's performance and generalization ability.
[0080] Optionally, in each of the plurality of action rounds, based on the course difficulty information and course progress information corresponding to the action round, a round-adapted task tool and a round-adapted task object are determined, thereby determining the learning task corresponding to the action round. The course difficulty information is associated with the course progress information; the later the action round is executed, the higher the course difficulty. For example, when the course difficulty is low, the round-adapted task tool can be a larger spoon, and the round-adapted task object can be a larger object with a smaller number of objects; as the course difficulty increases, the size of the round-adapted task tool decreases, and the size and number of the round-adapted task objects decrease.
[0081] Optionally, the second control strategy model mounted on the end effector is used to control the end effector to perform a second task. During the execution of the second task, the end effector uses the second task tool to move the second task object from the second initial position to the second target position, and the first task and the second task are different.
[0082] Optionally, during the process of the second task object moving from the second initial position to the second target position, the second task tool provides a force to the second task object, and the end effector does not contact the second task object. Of course, this disclosure is not limited thereto.
[0083] Optionally, the first task and the second task satisfy at least one of the following conditions: the first task tool is different from the second task tool; the first task object is different from the second task object; the number of the first task objects is different from the number of the second task objects; the first initial position is different from the second initial position; the first target position is different from the second target position; or the first task object is located in a first container, the second task object is located in a second container, and the first container and the second container are different. Of course, this disclosure is not limited thereto.
[0084] Optionally, to obtain a second control policy model with higher generalization and robustness than the first control policy model, the second control policy model can be a fine-tuned version of the first control policy model using reinforcement learning. The second control policy model has the same architecture as the first control policy model, but its generalization is superior. The reinforcement learning training process allows the second control policy model to learn more about environmental dynamics and potential policies, enabling it to make more rational decisions when facing new situations.
[0085] Specifically, in the reinforcement learning process, by learning the same or different learning tasks in each action round, the obtained second control policy model can explore as many control policies as possible, making the control policy model more "intelligent". Optionally, during the process of the virtual object corresponding to the end effector virtually executing the learning task corresponding to the action round, the action round adaptation task tool is used to move the action round adaptation task object from the initial position of the action round to the target position corresponding to the action round.
[0086] Optionally, the concept of human learning curricula can be drawn upon in reinforcement learning. Specifically, when humans learn a skill in school, they often need a well-designed curriculum to consolidate their knowledge and improve their performance, thereby learning higher-level skills. Generally, humans need to learn curricula from easy to difficult to gradually adapt to more complex environments. Here, curriculum can be an evolving concept, consisting of specific educational goals, particular knowledge and experience, and anticipated learning activities, containing a rich, fundamental, yet creative and potential set of plans and settings. Based on this, this disclosure proposes that a designed process can be used to adapt control policy models from easy to difficult to different task difficulty scenarios, ultimately converging into a model capable of completing high-difficulty tasks. Inspired by the human curriculum learning process, this disclosure introduces this curriculum learning concept into the reinforcement learning training process to address the problems of control policy models lacking exploratory potential after learning basic strategies or failing to converge when directly facing high-difficulty tasks.
[0087] Optionally, the information of the action round includes the course difficulty information and course progress information corresponding to the action round. In step S302, based on the course difficulty information and course progress information corresponding to the action round, the round adaptation task tool or the round adaptation task object of the action round is determined.
[0088] Optionally, to enhance the robustness of reinforcement learning, the initial position of the action round and the target position corresponding to the action round can be determined randomly. Of course, this disclosure is not limited thereto.
[0089] The course difficulty information corresponding to the current action round is associated with the course progress information of the current action round. The course progress information indicates that the later the action round is executed, the higher the course difficulty information corresponding to the current action round. However, this disclosure is not limited to this. The course difficulty information corresponding to the current action round and the course progress information of the current action round can be positively correlated, with the course progress information indicating that the later the action round is executed, the higher the course difficulty information corresponding to the current action round.
[0090] Optionally, in some embodiments of this disclosure, the course difficulty information can also be dynamically adjusted. For example, the course difficulty information corresponding to the current action round can be associated with the reward score of the learning task corresponding to the previous action round; the higher the reward score of the learning task corresponding to the previous action round, the higher the course difficulty information corresponding to the current action round. Of course, this disclosure is not limited thereto.
[0091] Optionally, in some embodiments of this disclosure, the same course difficulty information corresponds to multiple action rounds, and different action rounds correspond to different learning tasks. The difference in learning tasks corresponding to different action rounds satisfies at least one of the following conditions: the task tools adapted to different action rounds are different; the task objects adapted to different action rounds are different; the number of task objects adapted to different action rounds is different; or the task objects adapted to different action rounds are located in different containers. Furthermore, in the learning tasks corresponding to different action rounds, the initial positions of different action rounds may also be different, or the specific target positions of different action rounds may also be different. Of course, this disclosure is not limited to these limitations.
[0092] Therefore, during reinforcement learning, the model parameters of the first control strategy model can be updated based on the execution status of the learning task, guiding it towards successful completion of the learning task. When the course progress information indicates that the number of times the learning task is successfully executed within multiple consecutive action rounds meets a threshold condition, the course difficulty information can be updated. Subsequently, the next stage of model training can be performed based on the updated course difficulty information. This allows for a gradual increase in the number of objectives, thereby gradually increasing the course difficulty and guiding the virtual entity corresponding to the end effector to complete more challenging tasks through exploration. This avoids the control strategy model converging only to low-difficulty tasks or experiencing convergence issues due to an excessively large exploration space. The threshold condition can be a preset threshold number of times the learning task is successfully executed within multiple consecutive action rounds, as indicated by the course progress information.
[0093] Furthermore, since the reinforcement learning process is implemented in a virtual environment, this approach can minimize the overhead of using a physical device. Optionally, "the reward score corresponding to completing the learning task for the action round" refers to the reward score obtained by the virtual agent corresponding to the end effector after completing a specific learning task in the virtual environment, while "the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round" refers to the reward score obtained by the virtual agent corresponding to the end effector when performing the task in the simulated environment, taking into account physical constraints. Both types of reward scores are sparse rewards, which can encourage the control policy model to explore more extensively in the virtual environment to find behaviors that may lead to rewards, thereby potentially discovering new and more effective strategies.
[0094] Sparse reward refers to a reward method in reinforcement learning where the reward signal does not appear at every time step, but is only given when a specific event or condition is met. In this application, the "reward score corresponding to the learning task for completing the action round" is given only when the task is completed, which is a sparse reward. Sparse rewards can encourage the control policy model to explore more extensively in the virtual environment to find behaviors that may lead to rewards, thereby potentially discovering new and more effective strategies. However, it also increases the difficulty of learning because the model needs to learn an effective control policy with limited feedback.
[0095] By assigning a reward score to the learning task corresponding to the completed action round, the virtual agent corresponding to the end effector must learn from limited feedback. This forces the algorithm to use samples more efficiently, thereby improving the sample efficiency of the learning process. Simultaneously, because the virtual agent corresponding to the end effector is exposed to a wider variety of situations during training, it is more likely to develop strategies that can generalize to new situations, rather than being limited to specific behaviors that frequently yield rewards.
[0096] For example, the completion status of the scooping action can be evaluated based on the following conditions: First, whether the task tool (i.e., the virtual object corresponding to the spoon) in the action round is directly above the virtual object corresponding to the container; second, whether the amount of object scooped up by the virtual object corresponding to the spoon (i.e., the number of task objects in the action round) meets the predetermined standard. Specifically, the sufficiency of the object amount needs to be determined by comprehensively considering factors such as the capacity of the spoon, the volume of the object, and the height of the contents within the spoon. Through these specific conditions, it is possible to effectively determine whether the virtual object corresponding to the end effector has completed the corresponding learning task.
[0097] Optionally, the "reward score for completing the learning task corresponding to the action round" can be set to indicate a combination of at least one or more of the following: whether the end effector successfully moves the action round adaptation task object from the initial position of the action round to the target position corresponding to the action round using the action round adaptation task tool; or whether the action round adaptation task object overflows during the virtual execution of the learning task corresponding to the action round by the virtual body corresponding to the end effector in a virtual manner. Of course, this disclosure is not limited thereto.
[0098] By assigning reward scores related to physical constraints during the virtual execution of the learning task corresponding to each action round, the virtual object corresponding to the end effector can be ensured to perform actions that satisfy the physical constraints of the environment, preventing undesirable training results. For example, for a scooping action, due to the difference between the virtual environment and physical space, the actions performed virtually by the virtual object corresponding to the end effector in the virtual environment may not conform to physical constraints. These actions are impossible to achieve in real space. Therefore, it is necessary to determine whether the virtual object corresponding to the spoon is correctly inserted into the virtual object corresponding to the container, whether the virtual object corresponding to the spoon and the virtual object corresponding to the task object are clipping through each other, etc. Corresponding reward scores need to be designed to avoid these issues, ensuring the effectiveness and safety of training.
[0099] Environmental physical constraints refer to the physical rules and limitations that an end effector must follow when performing tasks in a virtual or real physical environment. These constraints include collisions between the task tool and the task object or container, the task tool detaching from the end effector, and the task tool exceeding preset positional limits. In reinforcement learning, the actions of the end effector are constrained by setting a "reward score related to physical constraints during the virtual execution of the learning task corresponding to the action rounds." This ensures that the end effector can only perform actions that satisfy the environmental physical constraints, preventing undesirable training results and ensuring the feasibility and effectiveness of the model in practical applications.
[0100] Optionally, the "reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round" can be set to indicate a combination of at least one or more of the following: whether the round-adaptation task tool of the action round collides with the round-adaptation task object of the action round during the virtual execution of the learning task corresponding to the action round by the virtual body corresponding to the end effector; whether the round-adaptation task tool of the action round collides with the container where the round-adaptation task object of the action round is located during the virtual execution of the learning task corresponding to the action round by the virtual body corresponding to the end effector; whether the end effector separates from the round-adaptation task tool during the virtual execution of the learning task corresponding to the action round by the virtual body corresponding to the end effector; and whether the round-adaptation task tool goes out of the observation range during the virtual execution of the learning task corresponding to the action round by the virtual body corresponding to the end effector. This disclosure is not limited thereto.
[0101] Specifically, the training process of this reinforcement learning can be summarized as follows: Based on the scene image at the current time step during the execution of the second task by the end effector and the pose corresponding to the contact point between the end effector and the second task tool, the control information of the end effector at the current time step is determined using a first control policy model; a reward score for the control information of the end effector at the current time step is determined, at least in part, based on the control information of the end effector at the current time step; and the model parameters of the first control policy model are updated, at least in part, based on the reward score for the control information of the end effector at the current time step, to determine the second control policy model mounted on the end effector. This process will be further described with reference to Figure 4, and will not be repeated here. Of course, this disclosure is not limited thereto.
[0102] Through reinforcement learning training, the second control policy model determines the corresponding control information at each time step and then learns based on feedback (reward or punishment). As the reinforcement learning process progresses, the second control policy model learns how to transition between different states to achieve long-term goals, improving its generalization ability. Through reinforcement learning, the second control policy model not only inherits the knowledge learned by the first control policy model but also enhances its flexibility and adaptability.
[0103] Therefore, this disclosure proposes an improved method for determining a control strategy model. Specifically, in order to obtain a second control strategy model with higher generalization and robustness than the first control strategy model, this disclosure further adjusts the first control strategy model using a reinforcement learning scheme to obtain the second control strategy model. During the adjustment process, specific sparse reward scores are used, namely, "the reward score corresponding to completing the learning task corresponding to the action round" and "the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round." This ensures that the control strategy model explores control schemes as broadly as possible while avoiding the situation where the control strategy model trained in the virtual environment does not conform to physical constraints, thus preventing its application in the physical environment.
[0104] Furthermore, the method for determining the control policy model proposed in this disclosure is particularly suitable for specific tool operation tasks. Compared to traditional methods for determining the training scheme of the control policy model suitable for specific tool operation tasks, this disclosure uses a reinforcement learning approach to adjust the first control policy model based on a pre-trained first control policy model. This reduces the need for human demonstration data and also reduces the training time required for reinforcement learning.
[0105] Furthermore, this disclosure combines the ideas of reinforcement learning and curriculum learning, so that the control strategy model carried by the end effector can adapt to the changing operating environment and task requirements during the reinforcement learning process, and achieve more flexible and efficient operation performance.
[0106] Notably, in some embodiments of this disclosure, reinforcement learning is combined with imitation learning, thereby significantly reducing the need for collecting human demonstration data. This allows the control strategy model on the end effector (such as either the first or second control strategy model mentioned above) to quickly learn through imitation with limited training data and effectively operate in unseen scenarios using reinforcement learning. Through the embodiments of this disclosure, the control strategy model on the end effector can extract key information from limited training data and generalize it to new, unknown environments. The embodiments of this disclosure not only improve learning efficiency but also enhance the robot's adaptability and flexibility in diverse tasks.
[0107] Next, the details of the embodiments of this disclosure will be further described with reference to FIG4.
[0108] Figure 4 is a schematic diagram illustrating a method 30 for determining a control strategy model according to an embodiment of the present disclosure.
[0109] Optionally, embodiments of this disclosure improve the architecture of the control strategy model. Compared to the structure of traditional strategy models, the improved control strategy model (i.e., the first control strategy model and the second control strategy model) consists of two decoupled modules: a visual encoder and a controller network. During the adjustment of the first control strategy model to obtain the second control strategy model, a value network decoupled from both the visual encoder and the controller network can be introduced to assist in the adjustment of the first control strategy model. Furthermore, the visual encoder may be combined with encoders for encoding sensory information, such as tactile encoders, to better capture scene information and task execution progress information. Of course, this disclosure is not limited to this. A tactile encoder is a device or module for encoding tactile sensory information. In this application, it may be combined with encoders for encoding sensory information, such as visual encoders, to better capture scene information and task execution progress information. By encoding tactile information, the control strategy model can be provided with more comprehensive environmental perception, thereby enabling the model to make more accurate decisions.
[0110] Optionally, the visual encoder processes the input reference scene image and extracts a visually encoded feature vector. The visually encoded feature vector is a vector extracted by the visual encoder after processing the input reference scene image. This vector integrates the relative pose information between the end effector and the task tool, as well as the task's progress information. It converts the reference scene image into a machine-understandable form, providing the controller network with the information needed for decision-making and helping the controller network calculate the pose change of the task tool at each time step. The controller network calculates the pose change of the task tool at each time step based on the visually encoded features and the reference trajectory corresponding to the contact point between the end effector and the task tool. Specifically, the controller network can employ a neural network structure such as a multilayer perceptron to calculate the pose change of the task tool at each time step based on the visually encoded features and the reference trajectory corresponding to the contact point between the end effector and the task tool. Specifically, the visually encoded features and the reference trajectory are used as input, undergoing feature transformation and nonlinear mapping through the hidden layer of the multilayer perceptron, and finally the pose change of the task tool is obtained through the output layer. The pose change refers to the numerical change in the position and posture of the task tool at each time step. The controller network calculates the pose change of the task tool at each time step based on visually encoded features and a reference trajectory corresponding to the contact point between the end effector and the task tool. This pose change guides the movement of the end effector to achieve precise control of the task tool, enabling it to complete the learning task.
[0111] Referring to Figure 4, in the imitation learning stage, in step S301, the first control strategy model can be obtained by training the first control strategy model. Optionally, as a specific example, in the imitation learning stage, reference scene images and reference trajectories corresponding to the contact points between the end effector and the first task tool can be collected in physical space when training the end effector to perform the first task. Then, in a virtual environment, the data collected in the physical space is projected, and the data projected into the virtual environment is amplified to collect a relatively large amount of training data while minimizing human demonstration. Of course, it is also possible not to amplify the training data in the virtual environment, and this disclosure is not limited thereto.
[0112] Optionally, assuming that in physical space, a reference scene image of the end effector performing the first task and a reference trajectory corresponding to the contact point between the end effector and the first task tool are acquired, then the reference scene image of the end effector performing the first task and the reference trajectory corresponding to the contact point between the end effector and the first task tool in physical space can be projected into a virtual environment to generate projection data of the reference scene image and the reference trajectory in the virtual environment. The first control strategy model mounted on the end effector is trained so that the first control strategy model can generate control information for controlling the end effector at the current time step based on the pose data of the first task tool at the current time step in the projection data of the reference scene image and the projection data of the reference trajectory in the virtual environment. The control information for controlling the end effector instructs the action performed by the end effector at the current time step.
[0113] Specifically, reference scene images and trajectories captured in physical space are acquired via sensors and cameras. This data undergoes preprocessing, such as scaling, cropping, and normalization, to adapt to the input requirements of the virtual environment. Next, a virtual environment is built in MuJoCo to simulate the original physical scene as closely as possible. MuJoCo is a high-level physics engine used to simulate multi-joint characters and rigid body dynamics, including contact and collision. Based on this, trajectory models in the virtual environment are designed according to the actually captured reference trajectories. These models can serve as both training targets for the agent and dynamic elements within the environment. Subsequently, the processed reference scene images and trajectories can be projected onto the virtual environment, involving a combination of image rendering and physical simulation. In this process, OpenAI Gym, as a toolkit for developing and comparing reinforcement learning algorithms, provides a unified interface for describing the interaction between the environment and the agent. Gym supports various predefined environments, including virtual environments using MuJoCo as a backend, facilitating imitation training.
[0114] More specifically, MuJoCo XML files can be written for physical simulation modeling. For example, corresponding models can be loaded into the MuJoCo environment via XML files. For instance, in a scenario where a scooping tool, such as a spoon, scoops a task object (e.g., a ball) from a target container (e.g., a bowl), virtual bodies corresponding to the end effector, spoon, and bowl (referred to as virtual end effector, virtual spoon, and virtual bowl, respectively) can be loaded, and the virtual end effector and virtual spoon can be fixed together. The virtual end effector can be considered as the actuator controlling the virtual spoon. The welding point between the virtual end effector and the virtual spoon can be considered as the contact point between the end effector and the spoon in the real machine environment. Since spoons and bowls are typical non-convex bodies with complex geometries, they cannot be directly implemented in the virtual environment. One embodiment of this disclosure can slice the scanned real spoon and bowl models into multiple convex hulls and then recombine them into a complete object.
[0115] In addition, several small balls can be added to the environment via an XML file. The MujocoEnv provided by OpenAI Gym can be used to access the written MuJoCo environment. A MujocoEnv can be created, defining environment description variables such as action_space and observation_space, and implementing functionalities such as init (initialization), step, and reset. The MujocoEnv corresponds to the written MuJoCo environment. The action_space can be used to define action information, which may include the pose changes of the task tool, for example, represented by 7 bits. The observation_space can be used to define model input information, which may include a reference trajectory composed of the pose information of contact points, the relative positional relationships between multiple task objects and the task tool, etc. To determine the relative pose between the task tool and the task object, embodiments of this disclosure can also set multiple position points on the task tool to determine whether the task object is accurately located on the task tool.
[0116] For example, in a scenario where a spoon scoops up a ball, an 11-bit input parameter can be defined. Seven bits represent the spoon's pose, three bits represent the positional relationship between the ball and the spoon, and one bit represents the initial height of the uppermost ball (i.e., the liquid level). In the one-hotspot design, 100 indicates no ball inside the spoon, 010 indicates a ball inside, and 001 indicates a ball inside with its center of mass higher than the spoon's center of mass. Alternatively, an 18-bit input parameter can be defined, including seven bits representing the pose of the contact point between the virtual spoon and the virtual end effector, in addition to the aforementioned parameters. As another example, in a scenario where a clamp picks up a block, an 11-bit input parameter can be defined. Seven bits represent the clamp's pose, one bit represents the clamp's opening angle, two bits represent the positional relationship between the object and the clamp, and one bit represents the task scenario. The positional relationship between the object and the clamp can be either being picked up or not.
[0117] As described above, the first control strategy model mounted on the end effector includes a visual encoder and a controller network. The visual encoder and controller network can be trained separately. The input to the visual encoder is a reference scene image corresponding to a single time step, and the output is the corresponding visual encoding vector. This vector captures key features of the reference scene image, providing the controller network with information such as the execution progress and current state of the current task.
[0118] To facilitate training in a virtual environment and subsequent deployment on a real device, the visual encoder should be able to encode reference scene images in both physical space and virtual environment, and output the same visual encoding vector for the same scene. The visual encoder will act as a bridge for information conversion between the virtual environment and physical space. Optionally, the visual encoder can be trained using the following method: Using the trained visual encoder, the reference scene image in physical space when the end effector performs the first task is encoded to generate a visual encoding vector; the projection data of the reference scene image in the virtual environment is encoded using the trained visual encoder to generate visual label data; the model parameters of the visual encoder are adjusted so that the difference between the visual encoding vector and the visual label data converges. Specifically, a loss function, such as the mean squared error loss function, can be calculated between the visual encoding vector and the visual label data. Then, the gradient of the model parameters is calculated based on the loss function, and the model parameters are updated in the opposite direction of the gradient, iterating continuously until the difference converges.
[0119] This allows the visual encoder to output the same visual encoding vector for the same scene, whether in physical space or a virtual environment, thus providing accurate input data to the controller network. Of course, this disclosure is not limited thereto.
[0120] Next, in step S301, the controller network can be trained separately after the visual encoder training is completed. Optionally, the training process includes: generating a visual encoding vector corresponding to the current time step using the trained visual encoder, based on the projection data of the reference scene image in the virtual environment corresponding to the current time step; predicting the pose data of the contact point in the next time step using the trained controller network, based on the visual encoding vector corresponding to the current time step and the true value of the pose data of the contact point in the projection data of the reference trajectory in the virtual environment at the current time step; and adjusting the model parameters of the trained controller network so that the difference between the predicted value of the pose data of the contact point in the next time step and the true value of the pose data of the contact point in the projection data of the reference trajectory in the virtual environment at the next time step converges. Of course, this disclosure is not limited thereto.
[0121] Optionally, the mean squared error (MSE) loss function can be used to measure the difference between the predicted pose data of the contact point at the next time step and the true pose data of the contact point at the next time step in the projection data of the reference trajectory projected onto the virtual environment. The Adam optimizer is used to adjust the model parameters of the controller network during training so that the difference between the two converges. The Adam optimizer automatically adjusts the learning rate based on the historical gradient information, updates the model parameters in the opposite direction of the gradient, and iterates until the difference converges.
[0122] It is worth noting that the pose data of the contact point in the next time step can be used to indicate the control information for controlling the end effector in the current time step. Specifically, the difference between the pose of the contact point in the next time step and the pose in the current time step will determine how to drive the various motors of the end effector in the current time step. That is, if the pose data of the contact point between the end effector and the task tool in the next time step can be accurately predicted, it is possible to know how to control the end effector.
[0123] As an example, suppose the visual encoding vector corresponding to the current time step can be represented as an 8-dimensional feature vector O. visual Subsequently, this visual encoding vector is combined with the 7-dimensional pose vector O of the end effector. pose (Including position and quaternion information) are concatenated to form a 15-dimensional feature vector (O). pose O visual This composite feature vector is then fed into the controller network, which generates the action vector a. output This guides the end effector to perform the next operation. Action vector a output That is, the controller network predicts the pose data of the contact point in the next time step, and its training label adata The true value of the pose data of the contact point at the next time step in the projection data of the reference trajectory projected into the virtual environment.
[0124] Alternatively, the mean squared error loss function (MSE) can be used to measure the output action vector a. output and its training label a data The difference between them. This loss function can quantify the accuracy of action prediction, providing guidance for the training of the controller network. Simultaneously, the Adam optimizer is used to optimize the parameters of the visual encoder and controller network. The Adam optimizer is widely used due to its adaptive learning rate characteristic, which can automatically adjust the learning rate based on historical gradient information during training, thereby accelerating convergence and improving training performance.
[0125] After completing the imitation training in step S301 (where the visual encoder and controller are trained separately), a first control policy model that can barely complete basic tasks is obtained. According to the previously described method 30, the reinforcement learning training process in step S302 can be started to obtain a second control policy model with stronger generalization ability.
[0126] Optionally, in operation S302, the process of determining the reward score for the control information of the end effector at the current time step includes: using a value network to determine the cumulative reward score for the control information of the end effector at the current time step.
[0127] The cumulative reward score refers to the accumulated reward score obtained by the end effector in executing control information over a period of time during reinforcement learning. The value network receives the same input as the controller network and determines the cumulative reward score through predefined reward calculation rules. It comprehensively considers the "reward score corresponding to the completion of the learning task corresponding to the action round" and the "reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round". It is used to evaluate the value of the action output generated by the controller network and assist in the adjustment of the first control strategy model.
[0128] Optionally, determining the reward score for the control information of the end effector at the current time step includes: using a value network to determine the cumulative reward score for the control information of the end effector at the current time step. The value network receives the same input as the controller network, including the scene image at the current time step and the pose information corresponding to the contact point between the end effector and the task tool, and determines the cumulative reward score R through a predefined reward calculation rule R = αR1 + βR2. Here, R1 represents the reward score corresponding to completing the learning task corresponding to the action round, R2 represents the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round, α and β are preset weight coefficients, and α + β = 1.
[0129] Specifically, the training process of reinforcement learning requires the introduction of a value network. This value network receives the same input as the controller network and outputs a value used to evaluate the action output generated by the controller network (i.e., the reward score corresponding to the completion of the learning task corresponding to the action round and the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round).
[0130] Therefore, the reinforcement learning training process can be modeled as a process consisting of tuples. The defined Markov decision process, wherein, Representing the state space, Let p(s′|s,a) represent the action space, and p(s′|s,a) be the state transition function, representing the probability of reaching state s′ after performing action a in state s. Function r: The cumulative reward score in the task is defined, and its value is determined by the value network. γ∈[0,1) is a discount factor used to assess the importance of future reward scores.
[0131] A Markov Decision Process (MDP) is a mathematical model used to describe an agent's decision-making and interactions in an environment. It consists of elements such as a state space, an action space, a state transition function, and a reward function. In the reinforcement learning training process of this application, it is modeled as a Markov Decision Process. The state space represents all possible states of the environment, the action space represents all actions the agent can take, the state transition function defines the probability of transitioning to the next state after performing an action in a given state, and the reward function evaluates the reward obtained by the agent for taking an action in each state. The control policy model continuously learns and adjusts within this Markov Decision Process to achieve long-term goals and improve its generalization ability.
[0132] The state transition function is a crucial element in Markov decision processes, defining the probability of transitioning to the next state after performing an action in a given state. In the reinforcement learning training process of this application, the state transition function describes the change in the environmental state after the end effector takes control information in the current state. The control policy model uses the state transition function to predict the impact of different actions on the environmental state, thereby making more rational decisions.
[0133] The discount factor is a parameter in Markov decision processes used to assess the importance of future reward scores. Its value ranges from 0 to 1; a discount factor closer to 1 indicates that the agent values future rewards more, while a discount factor closer to 0 indicates that the agent focuses more on current rewards. In the reinforcement learning training process of this application, the discount factor is used to balance the weights of current and future rewards, influencing the decision-making and learning process of the control policy model.
[0134] As described above, some embodiments of this disclosure can also combine the reinforcement learning training process with curriculum learning, so that the first control policy model can not only continuously explore the environment of the tool operation task, but also continuously adjust and optimize the predicted actions based on the exploration results. In this way, the first control policy model can gradually improve the accuracy of the actions to better adapt to and complete various tool operation tasks.
[0135] In reinforcement learning, sparse reward scores are used to describe task requirements: "the reward score for completing the learning task corresponding to the action round and the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round." The "reward score for completing the learning task corresponding to the action round" is given only upon task completion. Furthermore, embodiments of this disclosure also design "reward scores related to physical constraints during the virtual execution of the learning task corresponding to the action round" to constrain actions, thereby satisfying environmental physical constraints and preventing undesirable training results.
[0136] Specifically, reward scores can be defined in the following way to guide the end effector to perform the learning task in the desired manner.
[0137] Regarding the "reward score corresponding to the learning task for completing the specified action round," if the virtual object corresponding to the end effector successfully completes the task, i.e., successfully scoops up a task object, the reward score will increase by 150 points; and for each task object overflowing, the reward score will decrease by 250 points. Thus, this reward score system guides the end effector to learn how to effectively scoop up objects while avoiding excessive object spillage, ensuring its behavior meets the accuracy and efficiency requirements of actual operation.
[0138] Regarding the "reward score related to physical constraints during the virtual execution of the learning task corresponding to the action rounds," 100 points are deducted if the task tool (e.g., a spoon) collides with a bowl or object, detaches from the target actuator, or exceeds a preset positional limit. These reward scores are designed to ensure the end effector maintains the stability of the task tool during operation, avoids collisions, and restricts its movements within the end effector's workspace. Thus, the end effector can learn how to efficiently complete tasks while adhering to physical constraints.
[0139] In the course design for tool operation tasks, the focus is on the attributes of the virtual entity corresponding to the task tool, the attributes of the virtual entity corresponding to the task object, and the attributes of the task environment (such as the virtual entity of the container where the task object is located). Taking the scooping task as an example, the course difficulty can be adjusted by adjusting three key factors: the height of the task object inside the container, the size of the task object, and the size of the task tool - the spoon.
[0140] Optionally, when the object inside the container is relatively low, the spoon needs to reach deeper to touch it. This not only increases the risk of the spoon colliding with the container walls and bottom but also makes the entire scooping task more challenging. Similarly, the larger the object, the less the spoon can scoop up in a single operation, which correspondingly increases the difficulty of the course. Furthermore, using a smaller spoon limits the range of visual perception, making it more difficult for the visual encoder to capture object features, further increasing the task's complexity. Through these subtle adjustments, the course can gradually improve learners' operational skills and adaptability.
[0141] This disclosure designs multiple structured courses that use a mix of different task object heights, task object sizes, and task tool sizes as multidimensional difficulty variables for the control policy model to learn various retrieval tasks during reinforcement learning training. The difficulty information of these structured courses is evaluated and ordered from easiest to hardest, allowing the control policy model to steadily improve performance and acquire new skills during learning. By continuously evaluating the control policy model's performance in the current reinforcement learning process, it can be determined whether to switch to more difficult or easier courses. Furthermore, reviews of previously learned courses are interspersed at certain stages to prevent the potential problem of catastrophic forgetting in the control policy network. This course design strategy aims to promote the robustness of the control policy model through progressively increasing difficulty and regular review.
[0142] The structured courses are a set of learning tasks with specific difficulty and order designed in this application for training the control policy model. These courses use various task object heights, task object sizes, and task tool sizes as multidimensional difficulty variables for the control policy model to learn various retrieval tasks during reinforcement learning training. Evaluating the difficulty information of these structured courses and ordering them from easy to difficult allows the control policy model to steadily improve its performance and acquire new skills during learning. By continuously evaluating the control policy model's performance in the current course, it can be determined whether to switch to a more difficult or easier course, while interspersing reviews of previously learned courses at certain stages to prevent the potential problem of catastrophic forgetting in the control policy network.
[0143] For example, in one embodiment of this disclosure, the course can be set up in two stages, thereby designing learning tasks corresponding to each round of actions. Of course, this disclosure is not limited thereto.
[0144] Specifically, in the first stage, the various courses can be determined first. The difficulty variables are categorized into types such as the size of the task tool, the size of the task object, the number of task objects, the initial position of the task tool, the target position of the task tool, and the container in which the task object is located. In the scooping task, difficulty variables include the height of the task object within the container, the size of the task object, and the size of the task tool—the spoon. A difficulty coefficient corresponding to each specific difficulty variable is evaluated, such as 1 centimeter or a small size. Then, different types of difficulty variables are combined to obtain different learning tasks, and the combined difficulty coefficient of these learning tasks is calculated. Finally, these courses are sorted in ascending order of difficulty information.
[0145] In the second phase, the training strategy model is implemented according to this structured curriculum, and the control strategy model is continuously evaluated in the current curriculum. The performance. If the control strategy model is in the current course The performance reached the preset pass threshold δ pass Then we will move on to the next, more difficult course. Train the model; if the model is in the current course The performance is lower than the preset failure threshold δ failed If the threshold is not reached within the maximum number of training epochs n, then the program will switch to the previous, simpler course. Training will be conducted. Furthermore, if the model can continuously upgrade to reach the review threshold e times, review sessions will be interspersed throughout the training. This design aims to ensure that the control strategy model gradually transitions to more difficult courses after mastering the skills at the current level, while consolidating learned knowledge through review to prevent forgetting.
[0146] The threshold is a preset standard used to determine whether the performance of the control policy model in the current course meets the requirements for moving on to a more difficult course for training. When the value of the performance measurement function of the control policy model in the current course is greater than or equal to the threshold, it indicates that the model has mastered the skills of the current course and can switch to a more difficult course to continue learning, so as to gradually improve the model's ability and adaptability.
[0147] The review threshold is a preset number of times used to determine whether the control policy model needs to undergo course review. When the control policy model can continuously upgrade to reach the review threshold, it indicates that the model has made good progress in the learning process. However, in order to prevent the potential problem of catastrophic forgetting in the control policy network, it is necessary to intersperse reviews of previously learned courses to consolidate the learned knowledge and improve the robustness of the model.
[0148] Specifically, you can first set up a course set. For course sets Each course Determine the corresponding course difficulty coefficient V i Course difficulty level V i It can be composed of multiple dimensions, such as {v1, v2, ..., v m Then, based on the equation... Calculate course difficulty information C i It is the difficulty coefficient measurement function d(v) k The integral is calculated across all dimensions. Then, all courses can be sorted according to the course difficulty information C using the function D = Sort(D, C).
[0149] A course set is a collection of multiple structured courses with varying levels of difficulty and task settings. When training the control policy model, it starts with the simplest course in the set and learns courses sequentially according to their difficulty. Through continuous learning and adjustment within the course set, the control policy model gradually improves its performance and adaptability, ultimately achieving the ability to complete more challenging tasks. The course difficulty coefficient is a quantitative indicator used to measure the difficulty of each structured course. It consists of multiple dimensions, such as the height of the task object within the container, the size of the task object, and the size of the task tool. By comprehensively evaluating and calculating these dimensions, the course difficulty coefficient for each course can be obtained. Based on the course difficulty coefficient, course difficulty information can be calculated, and the structured courses can be ranked, allowing the control policy model to adapt to different task difficulty scenarios from easy to difficult.
[0150] Then, the control policy model is trained starting with the simplest lesson, iterating through all the learning tasks in all sorted lessons. The control policy model is then used in the lessons during the training process. Training is performed on the model, and the model parameters and reward score of the control policy model are adjusted accordingly. If the value of the performance measurement function P(M,reward) is greater than or equal to the threshold δ, the system will be considered successful. pass If you pass e courses consecutively, then you need to review the previous t courses. If you do not pass e courses consecutively, you will begin learning the next course.
[0151] If the control strategy model performs below the preset failure threshold δ in the current course. failed It will revert to the previous lesson and continue training until the upgrade conditions are met. If the passing threshold δ is not reached within the maximum number of training cycles n, it will be considered a failure. pass Then we will switch to the previous, simpler course. Conduct training.
[0152] The failure threshold is a preset standard used to determine whether the control policy model performs poorly in the current lesson and needs to return to a previous, easier lesson for training. When the performance measurement function of the control policy model in the current lesson falls below the failure threshold, or fails to reach the pass threshold within the maximum number of training epochs, it indicates that the model is encountering difficulties in the current lesson. The lesson difficulty needs to be reduced, and training should return to a previous, easier lesson to consolidate learned knowledge and improve the model's learning effectiveness.
[0153] The maximum number of training epochs is a preset upper limit on the number of training rounds used to limit the number of times the control policy model is trained in each course. If the control policy model still fails to reach the passing threshold within the maximum number of training epochs, it indicates that the model is struggling to achieve the desired learning effect in the current course. In this case, the model will be switched to a previous, simpler course for training to adjust the training strategy and improve learning efficiency.
[0154] Finally, after passing all the challenges in the courses, the adjusted first control policy model is output as the second control policy model. This progressively increasing difficulty training method helps the model gradually master more complex skills and ultimately achieve optimal performance. Through regular review and timely adjustment of training difficulty, this course-based learning strategy can effectively improve the generalization ability and adaptability of the control policy model.
[0155] After completing the reinforcement learning training process, the adjusted first control policy model can be tested. The control policy model is tested in each lesson, and the results are evaluated based on the performance measurement function P and the pass threshold δ. passThe comparison determines whether the control policy model should move on to the next, more difficult lesson (if performance exceeds a threshold) or revert to the previous, easier lesson for additional training (if performance falls below a threshold). This process continues until the model successfully passes tests in all environments, progressively improving its performance at different difficulty levels.
[0156] Specifically, it can be based on the performance measurement function. The evaluation results are compared with a threshold to determine whether the control policy model should move to the next more difficult course (if performance exceeds the threshold) or return to the previous easier course for additional training (if performance falls below the threshold). Where N... s N represents the number of times the learning task is successfully completed in one action round. t This indicates the total number of attempts in the round for that action.
[0157] Furthermore, since the controller network already possesses certain capabilities after the pre-training phase, directly optimizing a controller network and a value network with mismatched capabilities may lead to failures or longer training times in the early stages of training due to the untrained value network's inability to accurately evaluate the controller network's output. Therefore, in the initial stage of reinforcement learning training, the parameters of the controller network are first frozen, focusing on training the value network. Once this value network has sufficient evaluation capabilities, the parameters of both networks are unfrozen, and they are trained simultaneously. This training process can be summarized as follows.
[0158] Optionally, in the initial stage of reinforcement learning training, only the value network can be trained. The training of the value network includes: executing multiple action rounds for the first task; at each time step in each action round: based on the scene image at the current time step during the virtual execution of the first task by the end effector and the pose corresponding to the contact point between the end effector and the first task tool, using the controller network in the first control policy model, determining the control information of the end effector at the current time step; and based on the control information of the end effector at the current time step, predicting the predicted value of the cumulative reward score of the control information at the current time step; after completing each action round, calculating the predicted value of the cumulative reward score of the control information at the last time step in the action round as the predicted reward value corresponding to the action round; based on the execution result of the first task corresponding to the action round, determining the true value of the reward score corresponding to the action round; and adjusting the model parameters in the value network so that the difference between the predicted value and the true value of the reward corresponding to the action round converges. Of course, this disclosure is not limited to this.
[0159] Optionally, in subsequent stages of reinforcement learning training, both the value network and the controller network can be trained jointly. The joint training of the value network and the controller network includes: executing multiple action rounds for the second task; at each time step within each action round: based on the scene image at the current time step during the execution of the learning task by the end effector and the pose corresponding to the contact point between the end effector and the second task tool, using the controller network of the trained first control policy model, determining the predicted value of the control information of the end effector at the current time step; and at least partially based on the predicted value of the control information of the end effector at the current time step, using the trained value network, determining the control information for the end effector at the current time step. The system calculates the predicted cumulative reward score of the control information at the current time step; after completing each action round, it calculates the predicted cumulative reward score of the control information at the last time step in the action round as the predicted reward score for that action round; based on the execution result of the learning task corresponding to the action round, it determines the true reward score for that action round; it adjusts the model parameters in the value network to converge the difference between the predicted and true reward scores for that action round; and it adjusts the model parameters of the controller network in the first control strategy model to increase the predicted reward score for that action round.
[0160] Optionally, in step S302, reinforcement learning training can be performed in the virtual environment. Optionally, depth images of the virtual environment are captured using the built-in camera of MuJoCo and input into the visual encoder trained in step S302. The pose of the virtual spoon is obtained directly from MuJoCo. The reinforcement learning training environment is constructed using OpenAI Gym. The observation space is defined as a 15-dimensional vector ranging from (-∞, +∞), and the action space is a 6-dimensional vector ranging from [-1, 1]. The control information of the end effector at the current time step corresponds to a fixed execution time, such as 0.002 seconds, until one action round is completed. At each time step, the virtual environment receives a 6-dimensional vector a. output This vector represents the control information output by the controller network during training, used to control the movement of the virtual spoon. Subsequently, the virtual environment returns an updated depth image and the latest pose information of the virtual contact points. Thus, this setup can accurately simulate the dynamic process of the scooping task, while providing a structured observation and action space for the reinforcement learning algorithm. By updating the virtual environment state and the pose information of the contact points in real time, continuous scooping actions can be effectively simulated, providing rich feedback information for adjusting the first control strategy model.
[0161] Next, as shown in Figure 5, after completing step S302, the second control strategy model can be directly deployed on the actual end effector to perform one or more tasks, especially the first task or the second task.
[0162] At this time, the input to the second control strategy model can be a scene image captured in real time in the physical space, and the output of the second control strategy model can be control information used to control the end effector in the physical space.
[0163] Optionally, the control information is the expected pose information of the contact point between the end effector and the second tooling in the next time step. Optionally, after determining the expected pose information of the contact point between the end effector and the second tooling in the next time step, the control information of each motor of the end effector can be determined by inverse kinematics, which can be either the acceleration or the torque of each motor. Although mathematically these two physical quantities are not significantly different as control information for controlling motor rotation, in actual physical systems, not both physical quantities can be accurately measured. Therefore, those skilled in the art can, in actual use of the end effector, select the physical quantity with better data testing results and more consistent with the model for subsequent calculations, depending on the specific circumstances. Of course, this disclosure is not limited thereto.
[0164] For example, inverse kinematics is a method that calculates the rotation angles and velocities of the components controlling the end effector based on the desired pose information of the end effector (e.g., the desired pose information of the contact point between the end effector and the task tool). Using inverse kinematics, target motion control of the end effector can be achieved.
[0165] The process of determining the control information required for the end effector to perform the second task using inverse kinematics can be mathematically described by the following formula (1): q = IK(p end -p base (1)
[0166] Where q is the angle vector of each joint of the end effector in physical space, p end -p base It is the position vector of the contact point where the end effector contacts the second task tool. Formula (1) establishes the mapping from the space corresponding to each joint to the Cartesian space. For example, the joint angles that satisfy the target position and attitude can be solved based on the dynamic model corresponding to the end effector by using numerical or analytical methods. Then, the control information for each joint can be determined using the joint angle information so that the end effector can perform the target action.
[0167] Alternatively, when the inverse kinematics mapping lacks an analytical solution, a numerical approximation method can be used, given the position vector p of the end effector in physical space. end -p base An iterative algorithm is used to search for the joint angle solution q that satisfies the constraints. After obtaining the inverse kinematics solution, the kinematic solution q in the joint space corresponds to the solution that enables the end effector to achieve the target pose p. end The joint control parameters are then obtained. From this, dynamic analysis can be performed to calculate the joint driving force / torque, complete motion control, and obtain control information in physical space. Of course, this disclosure is not limited to this.
[0168] Therefore, embodiments of this disclosure also provide an end effector, including: a central connector and a plurality of multi-joint operating components; and a control strategy model mounted on the end effector and determined by method 30 according to embodiments of this disclosure. The control strategy model can be either the first control strategy model or the second control strategy model described above, and this disclosure is not limited thereto.
[0169] Optionally, the end effector is configured to perform a second task, which includes: at the current moment, acquiring a reference scene image of the end effector performing the second task and the pose corresponding to the contact point between the end effector and the second task tool; using a second control strategy model mounted on the end effector, generating control information for controlling the end effector at the current time step; determining the motor drive information of the central connector and the plurality of multi-joint operating components based on the control information for controlling the end effector at the current time step; and controlling the end effector based on the motor drive information of the central connector and the plurality of multi-joint operating components.
[0170] Optionally, the end effector can also be configured to perform a first task. In this case, although a first control strategy model can also be used to complete the task, the control effect of a second control strategy model is generally better than that of the first control strategy model. Therefore, the process of performing the first task may include: at the current moment, acquiring a reference scene image of the end effector performing the first task and the pose corresponding to the contact point between the end effector and the second task tool; using the second control strategy model mounted on the end effector, generating control information for controlling the end effector at the current time step; determining the motor drive information of the central connector and the plurality of multi-joint operating components based on the control information for controlling the end effector at the current time step; and controlling the end effector based on the motor drive information of the central connector and the plurality of multi-joint operating components.
[0171] Next, a method 60 for controlling an end effector according to an embodiment of the present disclosure will be further described with reference to FIG6. FIG6 is a flowchart illustrating a method 60 for controlling an end effector according to an embodiment of the present disclosure.
[0172] The method 60 for controlling an end effector according to embodiments of the present disclosure may include steps S601 to S603 as shown in FIG6. Of course, the method 60 may also include more or fewer operations, and the present disclosure is not limited thereto.
[0173] In step S601, at the current moment, a reference scene image of the end effector performing the task and the pose corresponding to the contact point between the end effector and the task tool are acquired.
[0174] In step S602, control information for controlling the end effector at the current time step is generated using the control strategy model mounted on the end effector.
[0175] In step S603, based on the control information used to control the end effector at the current time step, the motor drive information of the central connector and the plurality of multi-joint operating components is determined.
[0176] In step S604, the end effector is controlled based on the motor drive information of the central connector and the plurality of multi-joint operating components.
[0177] The control strategy model is the second control strategy model described above, which is determined by the method described in Figures 3 to 4.
[0178] Figure 7 schematically illustrates a device 70 for controlling an end effector according to some embodiments of the present disclosure. The end effector includes a central connector and multiple multi-joint operating components.
[0179] As shown in Figure 7, the device 70 includes: a camera 71 for acquiring a reference scene image of the end effector performing a task at the current moment; a sensor 72 for determining the pose of the contact point between the end effector and the second task tool at the current moment; a processor 73 for generating control information for controlling the end effector at the current time step using a control strategy model mounted on the end effector; and determining motor drive information for the central connector and the plurality of multi-joint operating components based on the control information for controlling the end effector at the current time step; and a motor 74 for controlling the end effector based on the motor drive information for the central connector and the plurality of multi-joint operating components. The control strategy model is determined by the method described in Figures 3 to 4.
[0180] Figure 8 schematically illustrates an exemplary block diagram of an apparatus 80 for determining a control strategy model according to some embodiments of the present disclosure. The control strategy model is used to control an end effector, which includes a central connector and multiple multi-joint operating components.
[0181] Specifically, the device 80 includes: a first module 81, configured to acquire a first control strategy model mounted on the end effector; and a second module 82, configured to adjust the first control strategy model mounted on the end effector using multiple action rounds to determine a second control strategy model mounted on the end effector, wherein, in each of the multiple action rounds, a learning task corresponding to the action round is determined based on the information of the action round; the virtual entity corresponding to the end effector uses the adjusted first control strategy model to virtually execute the learning task corresponding to the action round, and determines the reward score corresponding to completing the learning task corresponding to the action round and the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round; and adjusts the first control strategy model to increase the reward score for completing the learning task corresponding to the action round.
[0182] Figure 9 schematically illustrates an exemplary block diagram of an apparatus 90 for controlling an end effector according to some embodiments of the present disclosure. The end effector includes a central connector and a plurality of multi-joint operating components. The apparatus 90 includes a first module 91, a second module 92, a third module 93, and a fourth module 94.
[0183] Specifically, the first module 91 is used to acquire, at the current moment, a reference scene image of the end effector performing a task and the pose corresponding to the contact point between the end effector and the task tool. The second module 92 is used to generate control information for controlling the end effector at the current time step using the control strategy model mounted on the end effector. The third module 93 is used to determine the motor drive information of the central connector and the plurality of multi-joint operating components based on the control information for controlling the end effector at the current time step. The fourth module 94 is used to control the end effector based on the motor drive information of the central connector and the plurality of multi-joint operating components.
[0184] The control strategy model is the second control strategy model described above, which is determined by the method described in Figures 3 to 4.
[0185] It should be understood that devices 80 and 90 can be implemented in software, hardware, or a combination of both. Multiple different modules can be implemented in the same software or hardware architecture, or a single module can be implemented by multiple different software or hardware architectures.
[0186] Furthermore, device 80 can be used to implement method 30 as described above, and device 90 can be used to implement method 60 as described above. The relevant details have already been described in detail above and will not be repeated here for the sake of brevity. Devices 80 and 90 can have the same features and advantages as described with respect to the foregoing methods.
[0187] Figure 10 schematically illustrates an example block diagram of a computing device 1900 according to some embodiments of the present disclosure. For example, it may represent a computing device that can be used to deploy the apparatus 80 and apparatus 90 provided in the present disclosure.
[0188] As shown in the figure, the example computing device 1900 includes a processing system 1901 communicatively coupled to each other, one or more computer-readable media 1902, and one or more I / O interfaces 1903. Although not shown, the computing device 1900 may also include a system bus or other data and command transmission system that couples the various components to each other. The system bus may include any or a combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any of a variety of bus architectures, or may include control and data lines.
[0189] Processing system 1901 represents the functionality of performing one or more operations using hardware. Therefore, processing system 1901 is illustrated as including hardware elements 1904 that can be configured as processors, function blocks, etc. This may include application-specific integrated circuits (ASICs) or other logic devices formed using one or more semiconductors implemented in the hardware. Hardware element 1904 is not limited by its forming material or the processing mechanism employed therein. For example, a processor may consist of semiconductors and / or transistors (e.g., integrated circuits (ICs)). In such a context, processor-executable instructions may be electronically executable instructions.
[0190] Computer-readable medium 1902 is illustrated as including memory / storage device 1905. Memory / storage device 1905 represents a memory / storage device associated with one or more computer-readable media. Memory / storage device 1905 may include volatile storage media (such as random access memory (RAM)) and / or non-volatile storage media (such as read-only memory (ROM), flash memory, optical disk, magnetic disk, etc.). Memory / storage device 1905 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disk, etc.). Exemplarily, memory / storage device 1905 may be used to store various pose, location data, etc., mentioned in the above embodiments. Computer-readable medium 1902 may be configured in various other ways as further described below.
[0191] One or more input / output interfaces 1903 represent the functionality that allows a user to type commands and information into the computing device 1900 and also allows information to be presented to the user and / or sent to other components or devices using various input / output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones (e.g., for voice input), scanners, touch functionality (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., capable of detecting non-touch-related movements as gestures using visible or invisible wavelengths (such as infrared frequencies), network interface cards (NICs), receivers, and so on. Examples of output devices include display devices (e.g., monitors or projectors), speakers, printers, haptic-responsive devices, network interface cards (NICs), transmitters, and so on.
[0192] The computing device 1900 also includes an application program 1906. The application program 1906 can be stored as computing program instructions in the memory / storage device 1905. The application program 1906, together with the processing system 1901, etc., can implement all the functions of the various modules of the device 70 or 80 described with respect to FIG. 7 or FIG. 8.
[0193] This document describes various technologies in the general context of software, hardware, components, or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc., that perform specific tasks or implement specific abstract data types. The terms "module," "function," etc., as used herein generally refer to software, firmware, hardware, or a combination thereof. The technologies described herein are platform-independent, meaning that these technologies can be implemented on a variety of computing platforms with various processors.
[0194] Implementations of the described modules and technologies may be stored on or transmitted across some form of computer-readable medium. Computer-readable medium may include a variety of media accessible by the computing device 1900. By way of example and not limitation, computer-readable medium may include "computer-readable storage medium" and "computer-readable signal medium".
[0195] In contrast to simple signal transmission, carrier waves, or signals themselves, a "computer-readable storage medium" refers to a medium and / or device capable of persistently storing information, and / or a tangible storage device. Therefore, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented using methods or techniques suitable for storing information (such as computer-executable instructions, data structures, program modules, logic elements / circuits, or other data). Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD or other optical storage devices, hard disks, magnetic tape cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of art suitable for storing desired information and accessible by a computer.
[0196] "Computer-readable signal medium" refers to a signal-bearing medium configured to transmit instructions, such as via a network, to computing device 1900. Signal media typically embody computer-executable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves, data signals, or other transmission mechanisms. Signal media also includes any information transmission medium. By way of example and not limitation, signal media includes wired media such as wired networks or direct connections, and wireless media such as acoustic, RF, infrared, and other wireless media.
[0197] As previously described, hardware element 1904 and computer-readable medium 1902 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements may include components of integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implementations or other hardware devices in silicon. In this context, hardware elements can serve as processing devices for executing program tasks defined by instructions, modules, and / or logic embodied by the hardware element, and as hardware devices for storing instructions for execution, such as the previously described computer-readable storage medium.
[0198] The foregoing combinations can also be used to implement the various techniques and modules described herein. Therefore, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage medium and / or by one or more hardware elements 1904. The computing device 1900 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, modules can be implemented at least partially in hardware as modules executable as software by the computing device 1900, for example, by using the computer-readable storage medium and / or hardware element 1904 of the processing system. Instructions and / or functions can be executed / operated by, for example, one or more computing devices 1900 and / or processing system 1901 to implement the techniques, modules, and examples described herein.
[0199] The techniques described herein can be supported by these various configurations of the computing device 1900, and are not limited to specific examples of the techniques described herein.
[0200] It should be understood that, for clarity, embodiments of this disclosure have been described with reference to different functional units. However, it will be apparent that, without departing from this disclosure, the functionality of each functional unit may be implemented in a single unit, in multiple units, or as part of other functional units. For example, functionality described as being performed by a single unit may be performed by multiple different units. Therefore, references to a particular functional unit are considered merely as references to the appropriate unit used to provide the described functionality, and not as indicating a strict logical or physical structure or organization. Thus, this disclosure may be implemented in a single unit, or may be physically and functionally distributed among different units and circuits.
[0201] This disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed, implement the above-described method for determining a control strategy model, method for determining a tool control model, or control method.
[0202] This disclosure provides a computer program product or computer program including computer-executable instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the computing device to perform the methods for determining a control strategy model, the methods for determining a tool control model, or the control methods provided in the various embodiments described above.
[0203] In summary, this application provides a method for determining a control strategy model, a method for controlling an end effector, a device for controlling an end effector, an apparatus for determining a control strategy model, an apparatus for controlling an end effector, an electronic device, a non-volatile computer-readable storage medium, and a computer program product. By acquiring a first control strategy model mounted on the end effector, a second control strategy model is determined by adjusting it using multiple action rounds. In each action round, a learning task is determined based on the action round information. The virtual end effector executes the task virtually, determining the completion score and the reward score related to physical constraints, and adjusting the model to increase these two scores. From a technical perspective, extensively exploring control schemes in a virtual environment allows the model to traverse more possible control paths, avoiding getting trapped in local optima, much like exploring more paths in a maze to find the optimal exit. Simultaneously, considering the reward score related to physical constraints ensures that the model's learned strategy conforms to real-world physical laws, avoiding training strategies feasible in a virtual environment but unapplicable in a real physical environment. This improves the model's generalization and robustness, enabling stable operation in different physical environments, reducing debugging costs and failure risks in practical applications, and improving resource utilization.
[0204] Furthermore, the first control strategy model is used to control the end effector to execute the first task, that is, to use the first task tool to move the first task object from the first initial position to the first target position; the second control strategy model is used to control the execution of a second task different from the first task, that is, to use the second task tool to move the second task object from the second initial position to the second target position. During the virtual execution of the learning task, a round-adapted task tool is used to move the round-adapted task object. This allows the control strategy model to adapt to different task scenarios, much like a multi-functional tool that can function in different work environments. Different tasks require different control strategies. By learning different tasks, the model can master multiple strategies, improving resource reusability, avoiding separate model training for each task, saving significant training time and computational resources, and improving the overall efficiency of the system.
[0205] Furthermore, the reward score for completing the learning task corresponding to each action round indicates whether the end effector successfully moved the task object or whether the task object overflowed. This reward score setting provides a clear goal orientation for the end effector's learning, much like setting a clear competition goal for an athlete. When the end effector successfully completes the task and receives a reward, and when the task object overflows and is penalized, it will prompt it to continuously optimize its control strategy, improving the accuracy and efficiency of task execution. Accurate task completion can reduce task execution time and resource consumption, improve the stability and reliability of robot operation, reduce the cost increase caused by improper task execution, and make the robot more economical in practical applications.
[0206] Furthermore, when performing learning tasks corresponding to action rounds virtually, reward scores related to physical constraints indicate whether the task tool collides, separates from the end effector, or goes beyond the observation range. In actual operation, violating these physical constraints may lead to task failure, equipment damage, or safety accidents. By setting relevant reward scores, the end effector automatically avoids violating constraints during the learning process, ensuring the safety and stability of task execution, much like adding a safety net to robot operation. This reduces equipment maintenance costs and task delays caused by collisions, separations, and other issues, improves the reliability and operational efficiency of the entire system, and enables robots to work more safely and efficiently in industrial production and other scenarios.
[0207] Furthermore, the action round information includes course difficulty and progress information. Based on this, the appropriate rounds are determined to fit the task tools or objects when identifying the learning task. Introducing the concept of course-based learning into training, much like humans gradually learn from simple to complex courses, the model starts with easy tasks. As the course progresses, the task difficulty increases, allowing it to gradually adapt to different difficulty scenarios. This approach avoids the model struggling to learn or failing to converge when initially faced with overly complex tasks, improving learning efficiency and enabling the model to master effective control strategies more quickly. It enhances the model's adaptability, allowing it to handle a wider range of task difficulty, much like a student acquiring more advanced knowledge and skills through gradual learning.
[0208] Furthermore, the difficulty information of the current action round is linked to the course progress information; the later the course progress, the higher the difficulty. This association aligns with human learning patterns, allowing the model to gradually accumulate experience and skills before tackling more challenging tasks. As the difficulty increases, the model continuously explores new control strategies, improving its performance and capabilities, much like an athlete gradually increases the difficulty of training to improve their competitive level. This helps the model make more rational decisions when facing complex tasks, improving its generalization ability and its ability to cope with complex scenarios, making the robot more intelligent and efficient in practical applications.
[0209] Furthermore, the same course difficulty information corresponds to multiple action rounds, with different learning tasks in each round, such as different tools, objects, quantities, or containers. This diverse task setting allows the model to encounter more different scenarios at the same difficulty level, increasing the diversity of learning samples, much like exposing students to various types of exercises. The model learns different control strategies from different tasks and integrates and optimizes them, improving its generalization ability. Even when faced with new and unseen task scenarios, the model can make reasonable decisions based on existing experience, enhancing the robot's adaptability and flexibility in diverse environments, enabling the robot to play a role in more diverse scenarios.
[0210] Furthermore, when adjusting the first control strategy model to determine the second control strategy model, in each action round, based on the scene image and contact point pose information at the current time step, the control information is determined using the first control strategy model under adjustment, and then the reward score is determined. After completing the action round, the model is adjusted based on the reward scores at each time step. This adjustment method based on real-time scene information and reward scores allows the model to continuously optimize the control strategy according to the actual situation, just like a navigation system continuously adjusts its route according to real-time traffic conditions. The scene image and contact point pose information at each time step provide rich environmental information, which the model uses to make more accurate decisions. The reward score provides feedback for the adjustment, causing the model to optimize in the direction of obtaining higher rewards, improving the model's learning efficiency and accuracy, enabling it to converge to the optimal strategy faster, and reducing training time and computational resource consumption.
[0211] Furthermore, when determining the reward score for control information, a value network is used to determine the cumulative reward score. The value network is trained by performing multiple action rounds on the first task, determining the control information and predicting the cumulative reward score at each time step. After completing the action rounds, the parameters are adjusted to bring the difference between the predicted and actual values closer. The introduction of the value network can more accurately evaluate the value of the action output generated by the controller network. In reinforcement learning, accurately evaluating the value of actions is crucial for model learning, much like providing investors with accurate investment assessment reports. Through multiple rounds of learning and adjustment, the value network can more accurately predict the cumulative reward score, providing reliable feedback for controller network adjustments, improving the stability and effectiveness of model training, reducing uncertainty during training, and enabling the model to learn effective control strategies more quickly.
[0212] Furthermore, when adjusting the first control strategy model to determine the second control strategy model, at each time step of each action round, the controller network is used to determine the predicted value of the control information, and the value network is used to determine the predicted value of the cumulative reward score. After completing the action round, the parameters of the value network and the controller network are adjusted to increase the predicted value of the reward score. This joint training method allows the controller network and the value network to cooperate with each other to jointly optimize the control strategy model, much like two experts working together to solve a complex problem. The controller network is responsible for generating control information, and the value network evaluates its value. Through continuous interaction and adjustment, the two enable the model to better adapt to environmental and task requirements, improving the model's performance and generalization ability. This allows the model to make more reasonable decisions in different task scenarios, improving the efficiency and accuracy of task execution.
[0213] Furthermore, in both the first and second tasks, the task tool applies force to the task object, while the end effector does not contact the object. This setup aligns with real-world tool operation scenarios, avoiding potential problems from direct contact, such as damage to the task object. It also reduces friction and other interference factors caused by contact, resulting in more stable and precise task execution, much like how using a tool to indirectly manipulate an object allows for more precise control of force and direction. Simultaneously, it simplifies the design of the control strategy model, eliminating the need to consider complex contact mechanics between the end effector and the task object, thus improving the safety and reliability of task execution and reducing the difficulty of model design and debugging.
[0214] Furthermore, the first and second tasks differ in terms of task tools, objects, quantity, initial position, target position, or container. This diverse task setting allows the control strategy model to adapt to more different scenarios, enhancing the model's generalization ability and adaptability, much like an all-around athlete excelling in different competitions. Different tasks bring different requirements and challenges. By learning from multiple tasks, the model can quickly adjust its strategy based on existing experience when facing new task scenarios, improving the robot's practicality and application range in various complex environments, enabling the robot to play a role in more fields.
[0215] Furthermore, the method for controlling the end effector involves acquiring a reference scene image and contact point pose at the current moment, generating control information using a control strategy model, determining motor drive information, and then controlling the end effector. This method precisely controls the end effector's movements based on real-time information from the actual scene and an optimized control strategy model. The reference scene image and contact point pose information provide detailed information about the current task, allowing the control strategy model to generate the most suitable control information. The accurate determination of motor drive information ensures that the end effector moves precisely as required, improving the control accuracy and task execution efficiency of the end effector. This enables the robot to complete various tasks more accurately, reduces task failures and resource waste caused by inaccurate control, and increases the success rate of the robot in practical applications.
[0216] Furthermore, when training the first control strategy model, training data is first collected from physical space and then projected and augmented in a virtual environment. Collecting data from physical space ensures the authenticity and reliability of the data, much like collecting real samples for research. Projecting and augmenting in a virtual environment increases the diversity and quantity of data. More training data allows the model to learn more features and patterns, improving its generalization ability and accuracy, much like a student mastering more knowledge and problem-solving methods by doing more different types of exercises. At the same time, reducing human demonstrations lowers data collection costs and time, improves training efficiency, and makes model training more efficient and economical.
[0217] Furthermore, by combining imitation learning and reinforcement learning, the neural network model on the end effector learns from limited human demonstrations and then acquires the ability to perform tool operation tasks through reinforcement learning. Imitation learning allows the model to quickly learn basic control strategies from human demonstrations, much like a novice learning basic skills from an experienced person. Reinforcement learning allows the model to further explore and optimize strategies in unknown scenarios, enabling it to work effectively in unseen environments. This combined approach significantly reduces the need for collecting human demonstration data, improves learning efficiency, and enhances the robot's adaptability and flexibility in diverse tasks. It allows the robot to better cope with various complex real-world scenarios, much like a learner who possesses basic skills and can continuously improve themselves in new environments.
[0218] Furthermore, the control strategy model is architecturally improved, consisting of a visual encoder, a controller network, and a decoupled value network. The visual encoder processes the reference scene image, extracts visually encoded feature vectors, and provides decision-making information to the controller network, much like providing accurate intelligence to a decision-maker. The controller network calculates the pose changes of the task tool based on the visually encoded features and the reference trajectory, guiding the movement of the end effector. The value network assists in adjusting the first control strategy model and evaluates the value of the controller network's action output. This architectural improvement enables the model to better capture scene information and task execution progress information, improving the model's decision-making accuracy and control precision. The decoupled modular design makes model training and adjustment more flexible and efficient, reduces mutual interference between modules, improves the model's maintainability and scalability, and makes the model more resilient in the ever-evolving technological environment.
[0219] By studying the accompanying drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps, and "a" or "an" does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not imply that a combination of these measures cannot be used for profit.
Claims
1. A method for determining a control strategy model, the control strategy model being used to control an end effector, the end effector including a central connector and multiple multi-joint operating components, the method comprising: Obtain the first control strategy model mounted on the end effector, and By utilizing multiple action cycles, the first control strategy model mounted on the end effector is adjusted to determine the second control strategy model mounted on the end effector. In each of the plurality of action rounds, Based on the information of the action rounds, the learning task corresponding to the action rounds is determined; The virtual body corresponding to the end effector uses the first control strategy model under adjustment to virtually execute the learning task corresponding to the action round, and determines the reward score corresponding to the completion of the learning task corresponding to the action round and the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round; and The first control strategy model is adjusted to increase the reward score for completing the learning task corresponding to the action round and the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round.
2. The method as described in claim 1, wherein the first control strategy model is used to control the end effector to perform a first task, and during the process of the end effector performing the first task, the end effector uses a first task tool to move a first task object from the first initial position to the first target position; The second control strategy model mounted on the end effector is used to control the end effector to perform a second task. During the execution of the second task, the end effector uses the second task tool to move the second task object from the second initial position to the second target position. The first task and the second task are different. During the process of the virtual object corresponding to the end effector performing the learning task corresponding to the action round in a virtual manner, the action round adaptation task tool is used to move the action round adaptation task object from the initial position of the action round to the target position corresponding to the action round.
3. The method of claim 2, wherein the reward score for completing the learning task corresponding to the action round indicates a combination of at least one or more of the following: Does the end effector successfully move the action cycle adaptation task object from the initial position of the action cycle to the target position corresponding to the action cycle using the action cycle adaptation task tool? During the process of the virtual body corresponding to the end effector performing the learning task corresponding to the action round in a virtual manner, whether the round-adapted task object overflows.
4. The method of claim 2 or 3, wherein the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round indicates a combination of at least one or more of the following: During the process of the virtual object corresponding to the end effector performing the learning task corresponding to the action round in a virtual manner, whether the round adaptation task tool of the action round collides with the round adaptation task object of the action round; During the process of the virtual body corresponding to the end effector performing the learning task corresponding to the action round in a virtual manner, whether the round adaptation task tool of the action round collides with the container where the round adaptation task object of the action round is located; During the process of the virtual entity corresponding to the end effector performing the learning task corresponding to the action round in a virtual manner, is the end effector separated from the round-adaptive task tool? as well as During the process of the virtual entity corresponding to the end effector performing the learning task corresponding to the action round in a virtual manner, whether the round adaptation task tool is outside the observation range.
5. The method according to any one of claims 2 to 4, wherein the information of the action round includes the course difficulty information and course progress information corresponding to the action round; The step of determining the learning task corresponding to the action round based on the information of the action round includes: Based on the course difficulty information and course progress information corresponding to the action round, determine the round-adapted task tool or the round-adapted task object of the action round.
6. The method as described in claim 5, wherein the course difficulty information corresponding to the current action round is associated with the course progress information of the current action round, and the course progress information indicates that the later the action round is executed, the higher the course difficulty information corresponding to the current action round.
7. As described in claim 6, the same course difficulty information corresponds to multiple action rounds, and different action rounds correspond to different learning tasks. The learning tasks corresponding to different action rounds are different, satisfying at least one of the following conditions: Different action rounds are adapted to different task tools; Different action rounds are suitable for different task objects; Different action rounds are suitable for a different number of task objects; or The task objects for different action rounds are located in different containers.
8. The method according to any one of claims 2 to 7, wherein adjusting the first control strategy model mounted on the end effector using multiple action cycles to determine the second control strategy model mounted on the end effector comprises: In each of the plurality of action rounds, During the process of the virtual body corresponding to the end effector performing the learning task corresponding to the current action round in a virtual manner, based on the scene image of the current time step in the action round and the pose information corresponding to the contact point between the end effector and the round adaptation task tool of the action round, the control information of the end effector at the current time step in the current action round is determined by using the first control strategy model under adjustment. A reward score for the control information of the end effector at the current time step is determined, at least in part, based on the control information of the end effector at the current time step. as well as After completing the action rounds, the first control strategy model mounted on the end effector is adjusted, at least in part, based on the reward scores for the control information of the end effector at each time step.
9. The method of claim 8, the determining a reward score for the control information for the end effector at a current time step comprises: The cumulative reward score for the control information of the end effector at the current time step is determined using a value network. The training of the value network includes at least: performing multiple action rounds for the first task, and at each time step in each action round: Based on the scene image at the current time step during the virtual execution of the first task by the end effector and the pose corresponding to the contact point between the end effector and the first task tool, the control information of the end effector at the current time step is determined using the controller network in the first control strategy model; and Based on the control information of the end effector at the current time step, predict the cumulative reward score of the control information of the end effector at the current time step; After completing each round of actions The predicted value of the cumulative reward score of the control information in the last time step of the action round is calculated as the predicted value of the reward corresponding to the action round; Based on the execution result of the first task corresponding to the action round, determine the true value of the reward score corresponding to the action round; and Adjust the model parameters in the value network so that the difference between the predicted and actual values of the reward corresponding to the action round converges.
10. The method of claim 9, wherein adjusting the first control strategy model mounted on the end effector using multiple action cycles to determine the second control strategy model mounted on the end effector comprises: In each time step of each action round: Based on the scene image at the current time step during the execution of the learning task by the end effector and the pose corresponding to the contact point between the end effector and the second task tool, the predicted value of the control information of the end effector at the current time step is determined using the controller network of the first control strategy model in training. Based at least in part on the predicted value of the control information of the end effector at the current time step, the predicted value of the cumulative reward score for the control information of the end effector at the current time step is determined using the value network in training. After completing each round of actions The predicted value of the cumulative reward score of the control information in the last time step of the action round is calculated as the predicted value of the reward score corresponding to the action round; Based on the execution results of the learning tasks corresponding to the action rounds, the true value of the reward score corresponding to the action rounds is determined; Adjust the model parameters in the value network so that the difference between the predicted and actual reward scores for the action rounds converges; as well as Adjust the model parameters of the controller network in the first control strategy model so that the predicted value of the reward score corresponding to the action round increases.
11. The method of any one of claims 2 to 19, wherein during the movement of the first task object from the first initial position to the first target position, the first task tool provides a force to the first task object, and the end effector does not contact the first task object, and During the process of the second task object moving from the second initial position to the second target position, the second task tool provides force to the second task object, and the end effector does not contact the second task object.
12. The method according to any one of claims 1 to 11, wherein the first task and the second task do not satisfy at least one of the following conditions: The first task tool is different from the second task tool; The first task object is different from the second task object; The number of the first task objects is different from the number of the second task objects; The first initial position is different from the second initial position; The first target location is different from the second target location; or The first task object is located in a first container, and the second task object is located in a second container. The first container and the second container are different.
13. A method for controlling an end effector, the end effector comprising a central connector and a plurality of multi-joint operating components, including: At the current moment, acquire the reference scene image when the end effector is performing the task and the pose corresponding to the contact point between the end effector and the task tool; Using the control strategy model mounted on the end effector, control information for controlling the end effector at the current time step is generated; Based on the control information used to control the end effector at the current time step, determine the motor drive information of the central connector and the plurality of multi-joint operating components; as well as The end effector is controlled based on the motor drive information of the central connector and the plurality of multi-joint operating components; The control strategy model is determined by the method described in any one of claims 1-11.
14. An apparatus for controlling an end effector, the end effector including a central connector and a plurality of multi-joint operating components, the apparatus comprising: A camera is used to acquire a reference scene image of the end effector performing a task at the current moment; A sensor is used to determine the pose of the end effector at the point of contact with the task tool at the current moment. The processor is configured to use a control strategy model mounted on the end effector to generate control information for controlling the end effector at the current time step; and based on the control information for controlling the end effector at the current time step, to determine the motor drive information of the central connector and the plurality of multi-joint operating components. as well as A motor is used to control the end effector based on motor drive information from the central connector and the plurality of multi-joint operating components; The control strategy model is determined by the method described in any one of claims 1-11.
15. An apparatus for determining a control strategy model, the control strategy model being used to control an end effector, the end effector including a central connector and a plurality of multi-joint operating components, the apparatus comprising: The first module is used to obtain the first control strategy model mounted on the end effector; The second module is used to adjust the first control strategy model mounted on the end effector using multiple action cycles, in order to determine the second control strategy model mounted on the end effector. In each of the plurality of action rounds, Based on the information of the action rounds, the learning task corresponding to the action rounds is determined; The virtual entity corresponding to the end effector uses the first control strategy model under adjustment to virtually execute the learning task corresponding to the action round, and determines the reward score corresponding to the completion of the learning task corresponding to the action round and the reward score related to physical constraints during the virtual execution of the learning task corresponding to the action round; and Adjust the first control strategy model so that the reward score for completing the learning task corresponding to the action round increases.
16. An apparatus for controlling an end effector, the end effector comprising a central connector and a plurality of multi-joint operating components, including: The first module is used to acquire, at the current moment, a reference scene image of the end effector performing a task and the pose corresponding to the contact point between the end effector and the task tool; The second module is used to generate control information for controlling the end effector at the current time step by utilizing the control strategy model mounted on the end effector; The third module is used to determine the motor drive information of the central connector and the plurality of multi-joint operating components based on the control information used to control the end effector at the current time step. as well as The fourth module is used to control the end effector based on the motor drive information of the central connector and the plurality of multi-joint operating components; The control strategy model is determined by the method described in any one of claims 1-11.
17. An electronic device comprising: processor; and A memory, wherein the memory stores computer-executable code, which, when run by the processor, performs the method of any one of claims 1-13.
18. A non-volatile computer-readable storage medium having executable code stored thereon, the executable code, when executed by a processor, causing the processor to perform the method of any one of claims 1-13.
19. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1 to 13.
Citation Information
Patent Citations
Autonomous operation decision-making method for picking manipulator
CN117621046A
Robot assembly learning method based on teaching reward state machine and residual reinforcement learning
CN117863152A
Humanoid robot object grabbing method and device based on reinforcement learning control
CN117961888A
Quadruped robot motion control method and system based on reinforcement learning action simulation
CN118012077A
Mechanical arm control method based on simulation and variable parameter two-stage reinforcement learning
CN118357922A
Cited By
Transition control methods and systems to assist robots in switching control strategies
CN122308109A