A control method for a multi-axis robotic arm
Through the sensing device, the relative position images of the actuator and target object of the multi-axis robot arm are collected, visual features are extracted, and the reinforcement learning model is input, and the control strategy is output, which solves the problems of high difficulty and poor accuracy of the multi-axis robot arm, and achieves higher control command accuracy and complex task completion efficiency.
Patent Information
- Application Number
- CN202510053682.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In the prior art, the multi-axis robotic arm is difficult to operate, and the accuracy of the control execution of complex control tasks is poor.
The sensor device collects the relative position images of the actuator and target object of the multi-axis robot arm, extracts visual features, and inputs a reinforcement learning model to output a control strategy. The initial training dataset and correction dataset are used to optimize training reinforcement learning models to generate more accurate control instructions.
The accuracy of multi-axis robotic arm control instructions and the completion efficiency and accuracy of complex tasks are improved, and the problems of high difficulty and poor accuracy are solved.
Smart Images

Figure CN119458384B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of robot control, and in particular to a control method for a multi-axis robotic arm. Background Art
[0002] Robot manipulation has always been a core issue in robotics, and achieving human-level dynamic and dexterous manipulation tasks has been a long-term goal in this field. Although traditional hand-designed controllers or other methods based on machine learning algorithms have achieved certain results in specific tasks, they have limitations when dealing with complex dynamics, precision assembly, and dual-arm coordination tasks.
[0003] Machine learning-based methods, such as deep learning, reinforcement learning, and imitation learning, can theoretically enable robots to autonomously acquire complex manipulation skills. However, traditional learning methods face challenges in practical applications, such as low sample efficiency, strong dependence on accurate models and parameters, and poor optimization stability, which limits their application in the real world. Ultimately, they often fail to achieve satisfactory results when dealing with complex tasks that require precise interaction, dynamic control, and high-dimensional observation space.
[0004] Currently, no effective solution has been proposed to address the problems of high difficulty in controlling multi-axis robotic arms and poor accuracy in executing complex control tasks in related technologies. Summary of the invention
[0005] A control method for a multi-axis robotic arm provided by an embodiment of the present invention at least solves the problem in the related art that the multi-axis robotic arm is difficult to control and has poor accuracy in controlling complex control tasks.
[0006] According to one aspect of an embodiment of the present invention, a control method for a multi-axis robotic arm is provided, comprising: collecting relative position images of an actuator of the multi-axis robotic arm and a target object to be clamped by means of a sensing device, and extracting visual features; inputting the visual features into a reinforcement learning model for a target task, and outputting a corresponding control strategy by the reinforcement learning model; wherein the reinforcement learning model is obtained by training an initial training data set using reinforcement learning, and then correcting and optimizing the training according to a correction data set, wherein the initial training data set comprises: first image data of the actuator and the target object before the target task is executed, and a corresponding effective control strategy, and the correction data set comprises: second image data of the actuator and the target object before the target task is corrected, and a corresponding corrected control strategy; according to the control strategy and the current image, a corresponding first control instruction is generated, and according to the first control instruction, the multi-axis robotic arm and the actuator are controlled to act on the target object, wherein the current image comprises images of the actuator and the target object under the current circumstances.
[0007] As an optional embodiment, after controlling the multi-axis robot and the actuator to perform an action on the target object according to the first control instruction, the method further includes: detecting whether the target task is successful according to the state of the target object after the action is completed; in the case of failure of the target task, receiving a corrected second control instruction, controlling the multi-axis robot and the actuator to perform the next action on the target object according to the second control instruction, or, according to the control strategy and the new current image, regenerating the corresponding first control instruction, and controlling the multi-axis robot and the actuator to perform the next action on the target object according to the first control instruction; after the action is completed, re-detecting whether the target task is successful, and continuing to generate the first control instruction if the target task fails, or receiving the second control instruction until the target task is successful, and recording the corresponding control strategy, wherein the control strategy is a modified control strategy when it includes the second control instruction, and is an effective control strategy when it only includes the first control instruction; updating the initial training data set or the modified data set according to the control strategy; and optimizing the reinforcement learning model using the updated initial training data set and the modified data set.
[0008] As an optional embodiment, the visual features of the relative position image are input into a trained reinforcement learning model, and before the reinforcement learning model outputs the corresponding control instructions, the method further includes: training the initial model using a reinforcement learning algorithm through the initial training data, wherein the training data volume of the initial training data is within a preset data volume range, and the initial training data is the training data in the initial training data set; performing optimization testing according to the trained initial model, generating a corresponding control strategy according to the corresponding visual features, and generating a corresponding first control instruction according to the control strategy and the current image; when the first control instruction controls the action of the actuator of the multi-axis robot arm and indicates that the target task has failed, receiving a corrected second control instruction, controlling the multi-axis robot arm and the actuator to act on the target object according to the second control instruction until the target task succeeds; updating the corrected data set according to the second control instruction; extracting optimized training data from the initial training data set and the updated corrected data set, and optimizing the initial model according to the optimized training data; when the target task completion rate of the initial model without correction reaches a preset completion rate, completing the training, and using the initial model as the reinforcement learning model.
[0009] As an optional embodiment, the initial model is trained using a reinforcement learning algorithm through the initial training data. Before obtaining the reinforcement learning model, the method also includes: generating a preset number of effective control strategies by demonstrating operations, wherein the effective control strategies are multiple control instruction sequences that control the actuator to successfully complete the target action on the target object; recording the effective control strategies and the first image data corresponding to the effective control strategies as a set of initial training data; recording the initial training data and updating the initial training data set; and when training is required, selecting a corresponding number of initial training data from the initial training data set.
[0010] As an optional embodiment, before detecting whether the target task is successful based on the state of the target object after the action is completed, the method also includes: generating a preset number of target task success data and target task failure data of the target object by means of demonstration operation; wherein the target task success data includes a target task success image and a success mark, and the target task failure data includes a target task failure image and a failure mark; training a classifier using the target task success data and the target task failure data; and determining that the classifier training is completed when the classification loss of the trained classifier reaches a preset threshold.
[0011] As an optional embodiment, detecting whether the target task is successful according to the state of the target object after the action is completed includes: identifying the image of the target object after the action is completed by the classifier, and detecting whether the target task is successful; after the target task is successful, recording the image of the target object after the target task is successful, and updating the target task success data; when the second control instruction is received, recording the image of the corresponding target object, and updating the target task failure data; and optimizing and training the classifier according to the updated target task success data and target task failure data.
[0012] As an optional embodiment, the reinforcement learning model is optimized and trained using the updated initial training data set and the revised data set, including: recording the first model parameter of the reinforcement learning model before the optimization training; reselecting the optimization training data using the updated initial training data set and the revised data set to optimize the reinforcement learning model; determining the second model parameter of the reinforcement learning model after the optimization training; when the difference between the first model parameter and the second model parameter does not exceed a first preset difference threshold, updating the first model parameter to the second model parameter to optimize the reinforcement learning model; when the difference between the first model parameter and the second model parameter exceeds a first preset difference threshold, determining the corresponding weights of the first model parameter and the second model parameter according to the difference; according to the corresponding weights, performing weighted summation based on the first model parameter and the second model parameter to determine the third model parameter, and updating the first model parameter to the third model parameter to optimize the reinforcement learning model.
[0013] As an optional embodiment, the sensing device includes a follow-up camera arranged at the end of the robotic arm, and a fixed camera arranged in the environment where the robotic arm and the target object are located; the sensing device is used to collect relative position images of the actuator of the multi-axis robotic arm and the target object to be clamped, and visual features are extracted, including: collecting a first image from the perspective of the actuator through the follow-up camera; collecting a second image of the actuator and the target object from the perspective of the environment through the fixed camera; extracting a first visual feature based on the first image, and extracting a second visual feature based on the second image; when both the first image and the second image contain the target object, the first visual feature and the second visual feature are fused based on the size and shape of the target object to obtain the visual feature; when the first image does not contain the target object, but the second image contains the target object, the second visual feature corresponding to the second image is used as the visual feature.
[0014] As an optional embodiment, the multi-axis robotic arm is provided with a control device, and the second control instruction is a control instruction issued by the control device; before controlling the multi-axis robotic arm and the actuator to perform an action on the target object, the method further includes: according to the requirements of the target task, setting the safety constraints of the multi-axis robotic arm through the control device, wherein the safety constraints include at least one of the following: the movement range of the actuator of the multi-axis robotic arm, the range of force applied by the actuator to the target object, and the movement speed range of the actuator; when generating a corresponding first control instruction according to the control strategy and the current image, the control strategy is constrained by using the safety constraint so that the action corresponding to the generated first control instruction complies with the safety constraint.
[0015] According to another aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor, and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes any one of the above methods.
[0016] The control method of the multi-axis manipulator provided by the embodiment of the invention adopts the initial training data set and the correction data set to train the reinforcement learning model, so that the control strategy output by the reinforcement learning model has a higher accuracy, and then the actuator of the multi-axis manipulator and the relative position image of the target object to be clamped are collected by the sensing device, and the visual features are extracted; the visual features are input into the reinforcement learning model for the target task, and the reinforcement learning model outputs the corresponding control strategy; according to the control strategy and the current image, a first control instruction that is more in line with the control strategy is generated, and the multi-axis manipulator and the actuator are controlled to act on the target object according to the first control instruction. The technical effect of improving the accuracy of the first control instruction and the efficiency and accuracy of completing the target task is achieved. The problem of high difficulty in controlling the multi-axis manipulator in the related technology and poor accuracy of the control execution of complex control tasks is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work.
[0018] Figure 1 It is a flow chart of a control method of a multi-axis robotic arm according to an embodiment of the present invention.
[0019] Figure 2It is a schematic diagram of the overall architecture of the control system of the multi-axis robotic arm of an embodiment created by the present invention.
[0020] Figure 3 It is a schematic diagram of a control device for a multi-axis robotic arm according to an embodiment of the present invention.
[0021] Figure 4 It is a structural schematic diagram of the electronic device created by the present invention. DETAILED DESCRIPTION
[0022] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0023] Reinforcement learning is a field in machine learning that emphasizes how an agent takes a series of actions in an environment to maximize the cumulative reward. The agent learns the optimal behavior strategy based on the reward signal fed back by the environment by interacting with the environment.
[0024] For example, in the task of training a multi-axis robotic arm to grasp a target object, the multi-axis robotic arm (agent) takes different joint movement actions (actions) at different target object positions (environments). When it successfully grasps the object, it will receive a positive reward, and if it fails to grasp the object, it will receive a negative reward. By continuously trying and receiving reward feedback, the robot learns the best walking strategy.
[0025] However, the existing reinforcement learning method requires a lot of time and a lot of corrections to achieve a high accuracy rate, which results in low training efficiency and slow speed, or it can be understood that the final execution accuracy of the target task is low.
[0026] In order to solve the problem in the related art that a multi-axis robotic arm is difficult to operate and has poor accuracy in controlling and executing complex control tasks, an embodiment of the present invention provides a control method for a multi-axis robotic arm. Figure 1 is a flow chart of a control method of a multi-axis robotic arm according to an embodiment of the present invention, such as Figure 1 As shown, the method comprises the following steps:
[0027] Step S101, collecting images of the actuators of the multi-axis robot arm and the relative positions of the target objects to be clamped by a sensing device, and extracting visual features;
[0028] Step S102, inputting the visual features into a reinforcement learning model for the target task, and the reinforcement learning model outputs a corresponding control strategy; wherein the reinforcement learning model is obtained by training an initial training data set using reinforcement learning, and then correcting and optimizing the training according to a correction data set, wherein the initial training data set includes: first image data of the actuator and the target object before the target task is executed and the corresponding effective control strategy, and the correction data set includes: second image data of the actuator and the target object before the target task is corrected and the corresponding corrected control strategy;
[0029] Step S103, generating a corresponding first control instruction according to the control strategy and the current image, and controlling the multi-axis robot arm and the actuator to perform actions on the target object according to the first control instruction, wherein the current image includes images of the actuator and the target object in the current situation.
[0030] The control method of the multi-axis manipulator provided by the embodiment of the invention uses an initial training data set and a correction data set to train the reinforcement learning model, so that the control strategy output by the reinforcement learning model has a higher accuracy rate, and then collects the actuator of the multi-axis manipulator and the relative position image of the target object to be clamped through the sensor device, and extracts the visual features; the visual features are input into the reinforcement learning model for the target task, and the reinforcement learning model outputs the corresponding control strategy; according to the control strategy and the current image, a first control instruction that is more in line with the control strategy is generated, and the multi-axis manipulator and the actuator are controlled to act on the target object according to the first control instruction. The technical effect of improving the accuracy of the first control instruction and the efficiency and accuracy of completing the target task is achieved.
[0031] The execution subject of the above steps may be a control device of the multi-axis robot arm. Specifically, the control device may be built into the multi-axis robot arm, for example, a control chip, a processor, a controller, etc. The control device may also be an external control device of the multi-axis robot arm, for example, a host computer, a computer, a server, a tablet, a mobile device, etc.
[0032] The above-mentioned sensing device can be a device capable of collecting physical parameters of the multi-axis robot arm, the actuator and the target object. It can be a camera, a sensor and other devices. The above-mentioned multi-axis robot arm can be understood as a robot capable of performing tasks. In other embodiments, the multi-axis robot arm can also be a walking robot, a multi-legged robot, a sweeping robot and other robots that can be trained using reinforcement learning.
[0033] The above relative position images may be one or more, and the multiple relative position images may be images taken by cameras with different viewing angles. All relative position images are combined together to determine the relative position of the actuator of the multi-axis robot arm and the target object to be clamped. In theory, the position and posture of the target object can also be determined.
[0034] Whether it is relative position or posture, it is used to provide a basis for determining the control strategy later. Different postures and relative positions use different actions.
[0035] The above relative position and posture can be described in many ways, such as the distance in a certain coordinate system and the visual features of the image.
[0036] In order to improve the speed and efficiency of feature extraction while taking into account the accuracy, this embodiment can use a method of extracting visual features to determine the relative position and posture of the actuator and the target object.
[0037] Since the multi-axis robot and the actuator are controlled by the above-mentioned execution subject, the joint information of the multi-axis robot and the parameter information of the actuator can determine the position and posture of the actuator. This method is more accurate.
[0038] If the target object is a system with an independent multi-axis robot arm, its position and posture can only be obtained through the above-mentioned sensing device. The specific method can be the above-mentioned visual feature extraction. The specific method of extraction will be described later.
[0039] The extracted visual features are then input into a reinforcement learning model for the target task, and the reinforcement learning model outputs the corresponding control strategy. The reinforcement learning model is a machine learning model that mainly uses reinforcement learning.
[0040] However, due to the slow speed and low efficiency of reinforcement learning and the low accuracy in execution, when training the reinforcement learning model, this embodiment first uses reinforcement learning to train the initial training data set so that it has the ability to generate a policy network.
[0041] Then let the trained model autonomously generate a policy network, execute the policy, and make corrections in case of task failure to obtain a corrected data set. The final reinforcement learning model obtained by the correction and optimization training is corrected according to the corrected data set. The task execution success rate of the reinforcement learning model finally obtained is 99%-100%. Compared with the 70%-80% of the reinforcement learning model in the existing technology, it has a higher accuracy rate.
[0042] Moreover, the method of this embodiment can achieve the above success rate within a few hours. Compared with the training time of the existing reinforcement learning model, it takes a longer time and a larger number of task execution times to have higher training efficiency and training speed.
[0043] The above current image includes the positions of the actuator and the target object in the current situation. The relative position and posture of the target object and the actuator can be characterized by visual feature extraction. According to the control strategy and the current image, a corresponding first control instruction can be generated, and the multi-axis robot arm and the actuator can be controlled to act on the target object according to the first control instruction.
[0044] This enables more accurate control of the multi-axis robotic arm, achieving the technical effect of improving the accuracy of the first control instruction, as well as the efficiency and accuracy of completing the target task.
[0045] As an optional embodiment, after controlling the multi-axis robot and the actuator to perform an action on the target object according to the first control instruction, the method also includes: detecting whether the target task is successful according to the state of the target object after the action is completed; if the target task fails, receiving a corrected second control instruction, and controlling the multi-axis robot and the actuator to perform the next action on the target object according to the second control instruction, or, according to the control strategy and the new current image, regenerating the corresponding first control instruction, and controlling the multi-axis robot and the actuator to perform the next action on the target object according to the first control instruction; after the action is completed, re-detecting whether the target task is successful, and if the target task fails, continue to generate the first control instruction, or receive the second control instruction until the target task is successful, and record the corresponding control strategy, wherein the control strategy includes the second control instruction and is a modified control strategy, and the control strategy only includes the first control instruction and is an effective control strategy; updating the initial training data set or the modified data set according to the control strategy; and optimizing the reinforcement learning model using the updated initial training data set and the modified data set.
[0046] It should be noted that after the first control instruction controls the actuator to act, it is also possible to detect whether the target task is successful based on the state of the target object after the action is completed. If successful, the corresponding strategy data can be recorded as new initial training data. If unsuccessful, the generation of the first control instruction can continue, or human intervention can be performed to receive the second control instruction generated by the operation.
[0047] After the action corresponding to the first control instruction or the second control instruction is completed, it is also possible to continue to re-detect whether the target task is successful. If the target task fails, continue to generate the first control instruction, or receive the second control instruction, until the target task succeeds.
[0048] When the target task is successful, the corresponding control strategy can be recorded. When the control strategy includes the second control instruction, it is a modified control strategy, and when the control strategy only includes the first control instruction, it is an effective control strategy. The modified control strategy can be used as a modified data set, and the effective control strategy can be used as an initial training data set. Thus, the initial training data set or the modified data set is updated according to the control strategy.
[0049] Then, the updated initial training data set and the revised data set can be used to optimize the training of the reinforcement learning model. The optimization training method can be to select a part of the data from the initial training data set and the revised data set respectively to form the optimized training data.
[0050] Use the optimized training data to perform optimized training in the same way as the initial training data set. This method of optimized training does not need to rely on correction operations and has the advantage of being more efficient. Moreover, after optimized training, the reinforcement learning model will also have a better strategy generation effect. Ultimately, the accuracy and efficiency of task completion will be improved.
[0051] As an optional embodiment, the visual features of the relative position image are input into a trained reinforcement learning model, and before the reinforcement learning model outputs the corresponding control instructions, the method also includes: training the initial model using a reinforcement learning algorithm through initial training data, wherein the training data volume of the initial training data is within a preset data volume range, and the initial training data is the training data in the initial training data set; performing optimization testing based on the trained initial model, generating a corresponding control strategy based on the corresponding visual features, and generating a corresponding first control instruction based on the control strategy and the current image; when the first control instruction controls the action of the actuator of the multi-axis robot arm, indicating that the target task has failed, receiving a corrected second control instruction, controlling the multi-axis robot arm and the actuator to act on the target object according to the second control instruction until the target task succeeds; updating the corrected data set according to the second control instruction; extracting optimized training data from the initial training data set and the updated corrected data set, and optimizing the initial model according to the optimized training data; when the target task completion rate of the initial model without correction reaches the preset completion rate, completing the training and using the initial model as the reinforcement learning model.
[0052] The above reinforcement learning model needs to be trained before use. The specific training method is as follows:
[0053] First, the initial model is trained using the reinforcement learning algorithm through the initial training data. The training data volume of the initial training data is within the preset data volume range, and the initial training data is part or all of the training data in the initial training data set.
[0054] This embodiment takes into account that the initial training data has a poor effect on improving the accuracy after training to a certain amount due to the reinforcement learning model itself. At the same time, in order to improve the efficiency of obtaining the initial training data, it can be agreed that the initial training data required for training is 20-200.
[0055] After the initial training data is trained, the reinforcement learning model has the basic ability to generate a policy network. However, the error rate is very high. At this time, an optimization test can be performed based on the trained initial model, and the corresponding control strategy can be generated based on the corresponding visual features, and the corresponding first control instruction can be generated based on the control strategy and the current image.
[0056] When the first control instruction controls the action of the actuator of the multi-axis robot arm, indicating that the target task has failed, the operator makes a correction, and the above-mentioned execution subject can receive a corrected second control instruction, and control the multi-axis robot arm and the actuator to act on the target object according to the second control instruction until the target task is successful.
[0057] After success, the revised data set is updated according to the second control instruction, and optimized training data is extracted from the initial training data set and the updated revised data set, and the initial model is optimized and trained according to the optimized training data.
[0058] After each optimization training, the number of target tasks completed without correction and the number of tasks executed by the initial model can be calculated, and the completion rate of the above target tasks can be calculated. When the completion rate reaches the preset completion rate, the training is completed and the initial model is used as a reinforcement learning model.
[0059] As an optional embodiment, the initial model is trained using a reinforcement learning algorithm through initial training data. Before obtaining the reinforcement learning model, the method also includes: generating a preset number of effective control strategies by demonstrating operations, wherein the effective control strategies are multiple control instruction sequences that control the actuator to successfully complete the target action on the target object; recording the effective control strategies and the first image data corresponding to the effective control strategies as a set of initial training data; recording the initial training data and updating the initial training data set; and when training is required, selecting a corresponding number of initial training data from the initial training data set.
[0060] The above demonstration operation may be a standard action demonstration, a machine control strategy with high efficiency and accuracy, or a manually operated control strategy with high accuracy and efficiency.
[0061] And determine the corresponding effective control strategy based on the above demonstration operation. If it is a machine demonstration, the corresponding control instruction sequence can be directly obtained. If it is a human operation, the control instruction sequence of the operation can be recorded as an effective control strategy.
[0062] Record the effective control strategy and the first image data corresponding to the effective control strategy as a set of initial training data, record and update the initial training data set. Then, when the initial training data needs to be trained, select the corresponding amount of initial training data from the initial training data set.
[0063] This allows initial training data with higher accuracy and efficiency to be obtained as quickly as possible.
[0064] As an optional embodiment, before detecting whether the target task is successful based on the state of the target object after the action is completed, the method also includes: generating target task success data and target task failure data of a preset number of target objects by demonstrating the operation; wherein the target task success data includes a target task success image and a success mark, and the target task failure data includes a target task failure image and a failure mark; training a classifier using the target task success data and the target task failure data; and determining that the classifier training is completed when the classification loss of the trained classifier reaches a preset threshold.
[0065] It should be noted that, based on the state of the target object after the action is completed, detecting whether the target task is successful can be achieved through a classifier.
[0066] Before using the classifier, you need to obtain training data to train the classifier.
[0067] The acquisition of training data for the classifier can also be done by demonstrating operations to generate a preset number of target object target task success data and target task failure data, and then by marking the target task success images and target task failure images.
[0068] Specifically, the target task success images are marked as successful, and the target task success data is generated. The target task failure images are marked as failed, and the target task failure data is generated.
[0069] The classifier is trained using target task success data and target task failure data; when the classification loss of the trained classifier reaches a preset threshold, the classifier training is determined to be completed.
[0070] Therefore, it is possible to accurately determine whether the target task after the first control instruction or the second control instruction is executed is successful.
[0071] As an optional embodiment, detecting whether the target task is successful based on the state of the target object after the action is completed includes: using a classifier, identifying the image of the target object after the action is completed, and detecting whether the target task is successful; after the target task is successful, recording the image of the target object after the target task is successful, and updating the target task success data; when a second control instruction is received, recording the image of the corresponding target object, and updating the target task failure data; and optimizing and training the classifier based on the updated target task success data and target task failure data.
[0072] Specifically, when using a classifier to detect whether a target task is successful, first obtain an image of the corresponding target object, extract features from the image of the target object, input the features into the classifier, and determine whether the target task is successful.
[0073] In this embodiment, when the sensing device is a plurality of cameras, the above-mentioned image can be an image taken by any camera, but the image should contain the target object.
[0074] As an optional embodiment, the reinforcement learning model is optimized and trained using the updated initial training data set and the revised data set, including: recording the first model parameter of the reinforcement learning model before the optimization training; reselecting the optimization training data using the updated initial training data set and the revised data set to optimize the reinforcement learning model; determining the second model parameter of the reinforcement learning model after the optimization training; updating the first model parameter to the second model parameter when the difference between the first model parameter and the second model parameter does not exceed a first preset difference threshold to optimize the reinforcement learning model; determining the corresponding weights of the first model parameter and the second model parameter according to the difference when the difference between the first model parameter and the second model parameter exceeds the first preset difference threshold; determining the third model parameter by weighted summing the first model parameter and the second model parameter according to the corresponding weights, and updating the first model parameter to the third model parameter to optimize the reinforcement learning model.
[0075] During optimization training, the training data used in optimization training includes updated revised training data and initial training data. During the optimization training process, the parameters should be smooth, and the final success rate gradually approaches 100%. However, optimization training is an independent training process. In order to avoid performance degradation caused by model training, the difference between the first model parameters and the second model parameters before and after optimization training can be used to determine whether to accept the second model parameters and the degree of acceptance of the second model parameters.
[0076] Specifically, when the difference between the first model parameter and the second model parameter does not exceed the first preset difference threshold, it means that the difference between the first model parameter and the second model parameter is not large, and the degree of optimization is acceptable. Then, the first model parameter is updated to the second model parameter to optimize the reinforcement learning model.
[0077] When the difference between the first model parameter and the second model parameter exceeds the first preset difference threshold, it indicates that the degree of optimization is large and there may be problems such as poor stability. The corresponding weights of the first model parameter and the second model parameter can be determined according to the difference. The larger the difference, the smaller the weight of the second model parameter. The smaller the difference, the larger the weight of the second model parameter. The specific weight can be determined according to the actual situation.
[0078] Then, according to the corresponding weights, a weighted sum is performed based on the first model parameters and the second model parameters to determine the third model parameters, and the first model parameters are updated to the third model parameters to optimize the reinforcement learning model.
[0079] In addition, it is also possible to determine whether to update the third model parameter by checking whether the difference between the third model parameter and the first model parameter meets the first preset difference threshold requirement. If it does not meet the first preset difference threshold requirement, the third model parameter can be discarded and the optimization can be abandoned. This ensures the stability of the reinforcement learning model and thus the stability of the multi-axis robot control.
[0080] As an optional embodiment, the sensing device includes a follow-up camera arranged at the end of the robotic arm, and a fixed camera arranged in the environment where the robotic arm and the target object are located; the sensing device collects images of the relative positions of the actuator of the multi-axis robotic arm and the target object to be clamped, and extracts visual features, including: collecting a first image from the perspective of the actuator through the follow-up camera; collecting a second image of the actuator and the target object from the perspective of the environment through the fixed camera; extracting a first visual feature based on the first image, and extracting a second visual feature based on the second image; when both the first image and the second image contain the target object, the first visual feature and the second visual feature are fused based on the size and shape of the target object to obtain a visual feature; when the first image does not contain the target object, but the second image contains the target object, the second visual feature corresponding to the second image is used as the visual feature.
[0081] The sensing device includes a follow-up camera set at the end of the robot arm, which can capture more details of the target object. The fixed camera set in the environment where the robot arm and the target object are located can capture the relative position of the robot arm and the target object in the overall environment from a panoramic angle. This avoids the problem that in some cases, the actuator is far away from the target object and cannot capture the target object, which leads to the inability to judge the subsequent completion status.
[0082] By using the above-mentioned follow-up camera and fixed camera, it is possible to ensure that the actuator and the target object can be placed in the same image based on capturing the details of the target object.
[0083] In some embodiments, there may be multiple fixed cameras to achieve more comprehensive shooting. The visual features of different images can be processed by feature fusion so that the final visual features have the desired visual features in each image.
[0084] In addition, the fusion features have better learning effects in the training and recognition of the reinforcement learning model, improving the training efficiency and accuracy of the reinforcement learning model.
[0085] As an optional embodiment, a control device is provided on the multi-axis robotic arm, and the second control instruction is a control instruction issued by the control device; before controlling the multi-axis robotic arm and the actuator to perform an action on the target object, the method also includes: according to the requirements of the target task, setting safety constraints of the multi-axis robotic arm through the control device, wherein the safety constraints include at least one of the following: the movement range of the actuator of the multi-axis robotic arm, the range of force applied by the actuator to the target object, and the movement speed range of the actuator; when generating a corresponding first control instruction according to the control strategy and the current image, the control strategy is constrained by using the safety constraint so that the action corresponding to the generated first control instruction complies with the safety constraint.
[0086] By setting safety constraints, the operation of actuators and multi-axis robotic arms can be made safer and more reliable.
[0087] Safety constraints can be various, for example, the above constraints on the range of motion, force, and speed of the actuator, which is a macroscopic guarantee of the safety of the target object during the action.
[0088] In other embodiments, safety constraints may be provided for more microscopic structures or program operations. For example, the travel range of one or more joints in a multi-axis robot may be constrained, and the numerical range of program operation may be constrained, etc., to ensure the safe operation of the multi-axis robot or the safe action of the target object.
[0089] It should be noted that this embodiment also provides an optional implementation, which is described in detail below.
[0090] This embodiment provides a method and system for quickly achieving precise robot operation. The method integrates human demonstration and human correction, effective machine learning algorithms, and system-level design, and can learn strategies with high operation accuracy, high success rate, and high movement speed in complex manipulation tasks within a few hours or even shorter training time. This greatly speeds up the deployment of robots in industrial applications.
[0091] This implementation method performs well in a variety of complex tasks such as dynamic manipulation, precision assembly, and dual-arm coordination, providing an efficient and practical solution for robots to autonomously acquire complex manipulation skills, and has broad industrial application prospects and research value.
[0092] This implementation belongs to the field of robotics technology and can be applied to industrial automation, manufacturing, service robots and other fields to achieve complex robot operation tasks.
[0093] The above usually refers to four skills: 1. The action chain is long and a series of actions need to be completed to achieve the final goal. For example: folding clothes; assembling multiple parts into a whole in sequence. 2. Fine manipulation of certain objects. For example: inserting cables into slots in industrial production lines; assembling memory sticks on computer motherboards, etc. 3. Manipulation of highly flexible objects. For example: sticking tape; installing flexible conveyor belts, etc. 4. Tasks that require the cooperation of both arms and performed simultaneously. For example, passing objects; installing large objects.
[0094] like Figure 2 As shown, the software of the system of this embodiment mainly consists of two modules:
[0095] 1. Execution module: Responsible for controlling the robot to interact with the environment using the current action strategy and perform the specified tasks. During the execution process, the interaction data is collected and sent to the data buffer. This module allows the operator to intervene and provide corrective actions when necessary.
[0096] In the execution process, the trained policy network is executed (the policy network is trained by machine learning in the learning process). If the policy network is updated in the learning process, the control strategy in the execution process will be changed. Therefore, it can be considered that the policy network based on machine learning is also executed in the execution process, but the training of the policy network is not carried out in this process.
[0097] The specific work of the execution process is to execute the trained policy network, control the robotic arm through code to complete specific actions (that is, interact with the environment), and collect interaction data during the action and put it into the buffer. These data will be used for training in the learning process.
[0098] The above-mentioned intervention when necessary generally refers to the deviation of the robot's action. For example, when training a robot to grab a water bottle, if it is found that the robot is exploring a direction that is obviously unable to grab the water bottle, intervention is required.
[0099] Specifically, a human can use a remote control (Space Mouse or other hardware device for remote control of robots) to intervene in the robot operation. At this time, the execution module will stop the original control strategy and switch to manual control mode. For example, in the previous example, the robot needs to be manually controlled by the remote control to move to the water bottle and grab the water bottle.
[0100] Currently, the correction results are determined by humans or by machine classifiers.
[0101] 2. Learning module: responsible for training the robot's action strategy. The training data is based on the initial demonstration of humans (data from dozens of successful completions of the specified tasks) and the data collected when the robot interacts with the environment extracted from the data cache. The action strategy and value function are updated using the reinforcement learning algorithm, and the updated strategy is sent to the execution module.
[0102] The data buffer area may include two, the above interactive data may include the effective control strategy for autonomously completing the task, and the corrected control strategy, and different data may be partitioned for storage. The above effective control strategy may be stored in the demonstration data buffer area, and the corrected control strategy may be the corrected data buffer area.
[0103] The hardware system of this embodiment mainly consists of the following optional parts:
[0104] Robot: Common six-axis or seven-axis robotic arms;
[0105] Sensors: 2D camera, 3D camera, six-dimensional force sensor, tactile sensor, etc.
[0106] Actuator: Gripper, suction cup, glue dispenser, screw machine or other customized actuator;
[0107] Control devices: control mouse, joystick or robotic arm, etc.; including machine learning models, which can learn a mapping relationship from known inputs and corresponding correct outputs, and can predict new unseen inputs.
[0108] Fixtures: Pneumatic, electric or other types of custom or general purpose fixtures.
[0109] The core functions of this implementation are mainly the following:
[0110] Reinforcement learning algorithm: It uses an efficient reinforcement learning algorithm based on a combination of offline and online data, which can make full use of demonstration data and correction data to accelerate the learning process. In each training step, the algorithm evenly samples from demonstration data and robot autonomous exploration data (including correction actions) to form a training batch, and introduces the target network to stabilize the training process, and regularly updates the parameters of the target network through a soft update strategy.
[0111] Operation demonstration: At the beginning of training, the operator is required to control the robot's movement through the operating equipment and successfully complete dozens of prescribed tasks. The interaction data generated in the process is collected for training.
[0112] Action Correction: During training, while the robot performs tasks autonomously, the operator can intervene at any time to provide corrective actions. These corrections are used to update the policy and help the robot learn from its mistakes, especially when learning difficult tasks from scratch.
[0113] Hardware control: In order to ensure the safety and speed of the robot when executing strategies, a special hardware control unit is designed to directly control some of the robot's motion performance.
[0114] Specifically include: Safety: For tasks that require contact with the environment, limit the robot's range of motion and the force applied to ensure safety. Movement speed: For dynamic tasks, directly set the movement speed of the robot end to not be lower than a certain set value to ensure the movement speed of the robot arm.
[0115] The training and deployment method of this embodiment includes the following steps:
[0116] Obtain initial training data, through operator demonstration, to provide initial training data, including reference action sequences for successfully completing tasks, as the basis for initial strategy learning;
[0117] Based on the reinforcement learning algorithm, the robot conducts strategy training. The robot interacts with the environment according to the current strategy, explores autonomously, and performs tasks. The system monitors the execution effect in real time. During the strategy training process, it integrates the operator's demonstration and correction data to improve the efficiency of sample utilization.
[0118] Strategy optimization is performed with the participation of the operator. When the operator observes that the robot deviates from the expected operation, or the system detects that the robot enters an error state and needs to be corrected, the operator can intervene by operating the equipment to control the robot to perform the correct action.
[0119] All interaction data, including the robot’s autonomous actions and corrective actions, are stored in the data cache for subsequent training and policy updates;
[0120] The learning module samples data from the data cache, performs policy training, and updates the policy network, so that the robot can learn to improve from its mistakes.
[0121] Repeat the strategy training and optimization process until the strategy reaches a preset performance indicator, such as a high success rate;
[0122] Deploy trained strategies and apply them in actual robot manipulation tasks to complete complex manipulation tasks.
[0123] Specifically, this embodiment can use the RLPD algorithm in reinforcement learning for training. In each training step, RLPD samples equally from previous data and policy data to form a training batch, and then updates the parameterized Q function and policy according to their respective loss function gradients.
[0124] The reward function can use a sparse reward function, which uses a trained classifier to make a binary judgment on whether the task is successful. The optimization goal is to maximize the success probability of each trajectory.
[0125] This embodiment can be used for a variety of complex robot manipulation tasks, including but not limited to: Precision assembly tasks: such as inserting a memory stick on a motherboard, installing an SSD, inserting a USB connector, etc. The robot needs to accurately position and apply appropriate force to complete high-precision assembly tasks.
[0126] Dual-arm coordination tasks: such as dual-arm delivery of objects, assembling IKEA bookshelves, installing car dashboards, and installing timing belts. The robot needs to coordinate two robotic arms to complete tasks that require complex coordination.
[0127] The following takes the USB grab and insert task as an example to illustrate the technical solution of this implementation. The detailed action flow is as follows:
[0128] Preparation: Robotic arm and the gripper at the end; two cameras, one is a follow-up camera and the other is a fixed camera. Take pictures of the area near the gripper at the end of the robotic arm.
[0129] Task description: In this task, the USB cable is randomly placed on the desktop, and the robot needs to complete the following steps:
[0130] Grasping the USB connector part: The robot needs to identify the position and posture of the USB connector and grasp it accurately.
[0131] Insert the USB connector into the corresponding slot: The robot needs to align the USB connector with the slot on the motherboard and insert it accurately.
[0132] Release the gripper: After the USB connector is successfully inserted, the robot needs to release the gripper to complete the task.
[0133] Detailed training process:
[0134] Configure the initial dataset. The human operator manually manipulates the robot using a remote control device (such as a remote controller or a 3D mouse) to complete 20 demonstrations of the USB plug-in and pull-out task. During this process, the visual input of the two cameras, the joint states, the end gripper movements, and the corresponding control instructions are recorded. These demonstration data are stored in the demonstration data buffer as the basis for initializing the policy and training the Q function (action value function).
[0135] Train the classifier, which is used to determine whether the task is successful. First, collect data. The operator manually operates the robot through the remote control device to demonstrate the state of successful task, USB is successfully inserted and the gripper is released, about 200 times, and demonstrates various states of task failure, including USB not inserted, incomplete insertion, and gripper not released, about 1000 times. Collect data under each state and annotate the samples, success is 1, and failure is 0. Based on these data, train the classifier, and when the performance of the validation set reaches the set standard, save the model parameters and wait for use.
[0136] Reinforcement learning training is performed, and the initial strategy is a strategy based on demonstration data. The offline reinforcement learning algorithm RLPD (Reinforcement Learning with Prior Data) is used. The extracted visual features are integrated with information such as the robot joint status, and control instructions are output through the fully connected layer to guide the robot to perform specific actions. During the training process, half of the data is sampled from the demonstration data buffer and the correction data buffer at each training step to form a training batch.
[0137] In actual training, the robot attempts to perform the USB plug-in and unplug task according to the current strategy and collects data in the process. The strategy interacts with the environment from the initial state, performs actions, collects data, and updates the strategy network and Q function network according to the set update frequency.
[0138] The operator needs to monitor the robot's behavior in real time, intervene when necessary, and record relevant data. The device used for intervention is a remote control device. Intervention opportunities include situations where the robot's grasping position is obviously wrong, the end effector posture is incorrect, and the gripper fails to operate correctly.
[0139] The interventions include using a remote control device to manually control the robot, correcting incorrect actions, and trying to guide the robot back to the correct operation trajectory. The intervention data will also be recorded for further optimization of the strategy. As the training progresses, the strategy should gradually improve, the success rate will increase, and the intervention rate will decrease.
[0140] When the strategy achieves satisfactory performance, save the final model parameters and back up the training logs and data during the training convergence and strategy deployment phase. Deploy the trained strategy to the actual robot control system, perform safety checks and operation monitoring to ensure that the strategy can autonomously complete the USB grabbing and insertion tasks in the actual environment.
[0141] Training data can include:
[0142] Offline demonstration data: Quantity: 20 trajectories. Acquisition method: Real-time recording of the robot's status and actions.
[0143] Reward classifier data: Quantity: about 200 positive samples and 1000 negative samples. Acquisition method: A human operator uses a remote control device to remotely control the robot to perform the final action of the task, record the camera image, and manually mark the successful and failed image samples.
[0144] Online policy data: Quantity: Automatically generated during training, with a total number of transitions of approximately 50,000. Acquisition method: When the policy performs tasks autonomously, it automatically records images, states, actions, and rewards during the interaction with the environment.
[0145] Human intervention data: Quantity: Based on needs, more in the early stage of training and gradually reduced in the later stage. Acquisition method: Corrective actions provided by the operator when the strategy is wrong.
[0146] Examples of error action types:
[0147] Inaccurate positioning: The robot fails to reach the USB port accurately, and may be to the left, right, high or low. Or the robot fails to reach the USB port accurately.
[0148] Incorrect posture: When the robot is grasping or inserting, the angle of the end gripper is not appropriate, resulting in the USB not being grasped or inserted correctly.
[0149] Improper operation of the gripper: The gripper fails to open or close in time at a specific position, resulting in a failure to grasp or release.
[0150] Path planning error: The robot's moving path is not planned properly, resulting in collisions with other objects in the environment or movement along unnecessary paths.
[0151] Slow response to environmental changes: Failure to adapt to changes in USB position or external interference in a timely manner.
[0152] The specific ways of human intervention are as follows:
[0153] Adjust the end effector position: Use the remote control device to precisely control the position of the robot and correct positioning errors.
[0154] Adjust end effector posture: Use the remote control device to rotate or tilt the end effector to obtain the correct grasping or insertion angle.
[0155] Control the opening and closing of the gripper: By pressing a button, the gripper can be opened and closed to ensure that the USB is grabbed or released at the right time.
[0156] Guide the overall movement path: Use remote control equipment to operate the robot, plan a reasonable movement path, and avoid obstacles.
[0157] Adapt to external disturbances: Use the remote control device to operate the robot to reposition and grasp the USB device to ensure mission continuation even if the USB position is changed.
[0158] Currently, in this scenario, after existing reinforcement learning training, the success rate of USB insertion tasks is about 70%-80%, while the success rate of the improved solution in this implementation can reach 99%-100%. Moreover, a success rate of 99%-100% can be achieved within a few hours. Not only the success rate is greatly improved, but also the training efficiency is improved and the training time is shortened. The training of the target task can be quickly realized according to the needs, and the use value is higher.
[0159] Figure 3 Schematic diagram of a control device for a multi-axis robotic arm according to an embodiment of the present invention. Figure 3 As shown, based on the control method of the multi-axis robotic arm provided by the embodiment of the invention, the embodiment of the invention also provides a control device for a multi-axis robotic arm, which includes: a feature extraction module 301, a strategy prediction module 302, and an instruction control module 303. The device is described in detail below.
[0160] The feature extraction module 301 is used to collect the relative position images of the actuators of the multi-axis robot arm and the target object to be clamped through a sensor device, and extract visual features.
[0161] The strategy prediction module 302 is connected to the feature extraction module 301, and is used to input the visual features into the reinforcement learning model for the target task, and the reinforcement learning model outputs the corresponding control strategy; wherein the reinforcement learning model is trained by using the reinforcement learning method on the initial training data set, and then corrected and optimized according to the correction data set. The initial training data set includes: the first image data of the actuator and the target object before the target task is executed and the corresponding effective control strategy, and the correction data set includes: the second image data of the actuator and the target object before the target task is corrected and the corresponding corrected control strategy.
[0162] The instruction control module 303 is connected to the above-mentioned strategy prediction module 302, and is used to generate a corresponding first control instruction according to the control strategy and the current image, and control the multi-axis robot arm and the actuator to perform actions on the target object according to the first control instruction, wherein the current image includes the image of the actuator and the target object in the current situation.
[0163] The control device of the multi-axis manipulator provided by the embodiment of the invention uses an initial training data set and a correction data set to train the reinforcement learning model, so that the control strategy output by the reinforcement learning model has a higher accuracy rate, and then collects the actuator of the multi-axis manipulator and the relative position image of the target object to be clamped through the sensing device, and extracts the visual features; the visual features are input into the reinforcement learning model for the target task, and the reinforcement learning model outputs the corresponding control strategy; according to the control strategy and the current image, a first control instruction that is more in line with the control strategy is generated, and the multi-axis manipulator and the actuator are controlled to act on the target object according to the first control instruction. The technical effect of improving the accuracy of the first control instruction and the efficiency and accuracy of completing the target task is achieved.
[0164] The embodiments of the present invention also provide a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.
[0165] The embodiments of the present invention further provide a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.
[0166] The embodiment of the invention also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform the method of the embodiment of the invention when executed by the at least one processor.
[0167] refer to Figure 4, a block diagram of an electronic device that can be used as a server or client of an embodiment of the invention will now be described, which is an example of hardware devices that can be applied to various aspects of the invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or required herein.
[0168] like Figure 4 As shown, the electronic device includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 to a random access memory (RAM) 403. In RAM 403, various programs and data required for the operation of the electronic device can also be stored. The computing unit 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0169] Multiple components in the electronic device are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information to the electronic device, and the input unit 406 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device. The output unit 407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include but is not limited to a disk, an optical disk. The communication unit 409 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0170] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing units, various computing units running reinforcement learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, the method embodiments created by the present invention may be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via a ROM 402 and / or a communication unit 409. In some embodiments, the computing unit 401 may be configured to perform the above method in any other appropriate manner (e.g., by means of firmware).
[0171] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the computer program, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0172] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0173] It should be noted that the term "including" and its variations used in the embodiments of the present invention are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0174] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0175] The various steps described in the method implementation methods provided by the embodiments of the present invention can be performed in different orders and / or in parallel. In addition, the method implementation methods may include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this respect.
[0176] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same and similar parts between the various embodiments refer to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.
[0177] The above-described embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of protection. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the attached claims.
Claims
1. A control method for a multi-axis robotic arm, comprising: The sensor device collects the relative position images of the actuator of the multi-axis robot arm and the target object to be clamped, and extracts the visual features; Inputting the visual features into a reinforcement learning model for a target task, and having the reinforcement learning model output a corresponding control strategy; The reinforcement learning model is characterized in that the reinforcement learning model is obtained by training an initial training data set using reinforcement learning, and then correcting and optimizing the training according to a correction data set, wherein the initial training data set includes: first image data of the actuator and the target object before the target task is executed and the corresponding effective control strategy, and the correction data set includes: second image data of the actuator and the target object before the target task is corrected and the corresponding corrected control strategy; After the reinforcement learning model outputs the corresponding control strategy, a corresponding first control instruction is generated according to the control strategy and the current image, and the multi-axis robot arm and the actuator are controlled to perform actions on the target object according to the first control instruction, wherein the current image includes images of the actuator and the target object in the current situation.
2. The method according to claim 1, characterized in that After controlling the multi-axis robot arm and the actuator to act on the target object according to the first control instruction, the method further includes: Detecting whether the target task is successful according to the state of the target object after the action is completed; In the case where the target task fails, receiving a corrected second control instruction, controlling the multi-axis robot arm and the actuator to perform a next action on the target object according to the second control instruction, or, according to the control strategy and the new current image, regenerating a corresponding first control instruction, and controlling the multi-axis robot arm and the actuator to perform a next action on the target object according to the first control instruction; After the action is completed, re-detect whether the target task is successful, and if the target task fails, continue to generate the first control instruction, or receive the second control instruction, until the target task succeeds, and record the corresponding control strategy, wherein the control strategy includes the second control instruction and is a modified control strategy, and the control strategy only includes the first control instruction and is an effective control strategy; Update the initial training data set or the revised data set according to the control strategy; The reinforcement learning model is optimized and trained using the updated initial training data set and the revised data set.
3. The method according to claim 2, characterized in that Before inputting the visual features of the relative position image into a trained reinforcement learning model and the reinforcement learning model outputting corresponding control instructions, the method further includes: Using the initial training data, the initial model is trained using a reinforcement learning algorithm, wherein the training data volume of the initial training data is within a preset data volume range, and the initial training data is the training data in the initial training data set; Performing optimization testing according to the trained initial model, generating a corresponding control strategy according to the corresponding visual features, and generating a corresponding first control instruction according to the control strategy and the current image; In the case where the first control instruction controls the action of the actuator of the multi-axis robot arm indicating a failure of the target task, receiving a corrected second control instruction, and controlling the multi-axis robot arm and the actuator to act on the target object according to the second control instruction until the target task succeeds; updating the correction data set according to the second control instruction; Extracting optimized training data from the initial training data set and the updated revised data set, and performing optimization training on the initial model according to the optimized training data; When the target task completion rate of the initial model without correction reaches a preset completion rate, the training is completed and the initial model is used as the reinforcement learning model.
4. The method according to claim 3, characterized in that The method further includes: training the initial model by using a reinforcement learning algorithm through the initial training data to obtain the reinforcement learning model. Generate a preset number of effective control strategies by means of demonstration operations, wherein the effective control strategies are multiple control instruction sequences for controlling the actuator to successfully complete the target action on the target object; Recording the effective control strategy and the first image data corresponding to the effective control strategy as a set of initial training data; Recording the initial training data and updating the initial training data set; When training is required, a corresponding amount of initial training data is selected from the initial training data set.
5. The method according to claim 4, characterized in that Before detecting whether the target task is successful according to the state of the target object after the action is completed, the method further includes: Generate a preset number of target task success data and target task failure data of the target object by means of demonstration operation; wherein the target task success data includes a target task success image and a success mark, and the target task failure data includes a target task failure image and a failure mark; Training a classifier using the target task success data and the target task failure data; When the classification loss of the trained classifier reaches a preset threshold, it is determined that the training of the classifier is completed.
6. The method according to claim 5, characterized in that According to the state of the target object after the action is completed, detecting whether the target task is successful includes: By means of the classifier, an image of the target object after the action is completed is identified to detect whether the target task is successful; After the target task is successful, an image of the target object for which the target task is successful is recorded, and the target task success data is updated; When receiving the second control instruction, recording the image of the corresponding target object and updating the target task failure data; The classifier is optimized and trained according to the updated target task success data and target task failure data.
7. The method according to claim 2, characterized in that Utilizing the updated initial training data set and the revised data set to optimize the reinforcement learning model includes: Recording the first model parameters of the reinforcement learning model before optimizing the training; Using the updated initial training data set and the revised data set, reselecting optimized training data, and performing optimized training on the reinforcement learning model; Determining a second model parameter of the optimized trained reinforcement learning model; When the difference between the first model parameter and the second model parameter does not exceed a first preset difference threshold, updating the first model parameter to the second model parameter to optimize the reinforcement learning model; When the difference between the first model parameter and the second model parameter exceeds a first preset difference threshold, determining corresponding weights of the first model parameter and the second model parameter according to the difference; According to the corresponding weights, a weighted sum is performed based on the first model parameters and the second model parameters to determine a third model parameter, and the first model parameter is updated to the third model parameter to optimize the reinforcement learning model.
8. The method according to claim 1, characterized in that The sensing device includes a follow-up camera arranged at the end of the mechanical arm, and a fixed camera arranged in the environment where the mechanical arm and the target object are located; The sensor device collects the actuator of the multi-axis robot arm and the relative position image of the target object to be clamped, and extracts visual features, including: Acquire a first image from the actuator's perspective by using the follow-up camera; Collecting a second image of the actuator and the target object from an environmental perspective by the fixed camera; extracting a first visual feature according to the first image, and extracting a second visual feature according to the second image; In a case where both the first image and the second image contain the target object, fusing the first visual feature and the second visual feature based on the size and shape of the target object to obtain the visual feature; When the first image does not include the target object but the second image includes the target object, the second visual feature corresponding to the second image is used as the visual feature.
9. The method according to any one of claims 2 to 7, characterized in that The multi-axis robot arm is provided with a control device, and the second control instruction is a control instruction issued by the control device; Before controlling the multi-axis robotic arm and the actuator to perform an action on the target object, the method further includes: According to the requirements of the target task, the safety constraints of the multi-axis manipulator are set by the control device, wherein the safety constraints include at least one of the following: the movement range of the actuator of the multi-axis manipulator, the range of the force applied by the actuator to the target object, and the movement speed range of the actuator; When a corresponding first control instruction is generated according to the control strategy and the current image, the control strategy is constrained by using the safety constraint so that an action corresponding to the generated first control instruction complies with the safety constraint.
10. An electronic device comprising: A processor, and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Robot vision control method and device based on deep network architecture automatic search and storage medium
CN109840508A
Correction method and system for grabbing and positioning errors of mobile robot and robot
CN113843798A